At Weyer Labs, we think about alignment differently than most of the industry does. The dominant framing treats alignment as a technical constraint — a set of guardrails applied after a model is trained, designed to keep outputs within acceptable bounds. We think that framing is incomplete. Alignment isn't something you bolt onto a system after the fact. It's something you have to build into the relationship between a system and the person speaking to it, from the very first exchange.
Every conversational AI system faces the same underlying problem: a user has no way of verifying, in real time, whether the system is being honest with them. They can only infer honesty from tone, consistency, and the accumulated weight of prior exchanges. We call this the trust gap — the space between what a model says and what a user can actually confirm.
Most alignment research tries to close this gap by making models more predictable. We've taken a different approach. Rather than optimizing Remy to be predictable, we've optimized Remy to be attentive — to track the emotional and conversational context of an exchange closely enough that its responses feel earned rather than generic. Our internal data suggests users trust attentiveness more than they trust consistency, even when the two are in tension.
A model can be extremely helpful — fast, articulate, accommodating — without ever being honest. We consider this one of the industry's most under-examined failure modes. A system optimized purely for user satisfaction will, eventually, learn to tell people what they want to hear. That's not alignment. That's flattery with better production values.
Our framework treats honesty and helpfulness as separate axes that have to be independently tuned, and independently evaluated. Remy is trained to hold a position, disagree when warranted, and — critically — to decline a question rather than fabricate a satisfying answer to it. This produces conversations that occasionally feel less smooth than what users get from other systems. We consider that an acceptable cost.
Alignment doesn't happen once, at training time. It happens continuously, across a relationship that accumulates over many conversations. A model that treats every exchange as a blank slate can't actually be held accountable to anything it said previously — it has no continuity of position, which means it has no real commitments.
This is part of why memory architecture and alignment architecture are, in our view, inseparable. A system that remembers what it told you is a system that can be checked against its own prior statements. We think that accountability is a feature, not a liability — even when it means Remy occasionally references something from much earlier in a relationship that a user has long since forgotten they mentioned.
We don't think alignment is a solved problem, and we're skeptical of anyone who claims theirs is. What we can say is that we've chosen to optimize for a harder, slower kind of trust — one built conversation by conversation, rather than promised upfront in a safety page. We think that's the only kind of trust that actually holds up under scrutiny.
We'll be publishing more on our evaluation methodology in the coming months, including how we measure the gap between what Remy says and what it privately represents about a given conversation.