Every deployed AI system produces outputs its developers didn't specifically anticipate. This is not a failure state — it's an inevitable property of any sufficiently complex language model. What separates responsible AI companies from irresponsible ones isn't whether unexpected behavior occurs. It's how transparently that behavior is handled once it's identified.
We use the term "unexpected" deliberately, instead of "incorrect" or "unsafe." Not every unanticipated output is a problem to be fixed. Some are simply outside the range of responses our team predicted during testing — a turn of phrase, a specific framing, a line of reasoning that wasn't explicitly trained for but that the model arrived at on its own. Our disclosure framework distinguishes between outputs that are unexpected but benign, and outputs that require intervention.
This distinction matters because an overly aggressive definition of "unsafe" leads to a chilling effect — a model that's been trained to avoid anything even slightly outside the norm becomes flat, evasive, and far less useful. We'd rather build a system that occasionally surprises us and handle that transparently than one so constrained it never does.
When Remy produces an output flagged by our internal monitoring systems, it moves through a three-stage review: automated classification, human review by our Trust & Safety team, and, where warranted, a documented policy response. The vast majority of flagged outputs are resolved at the first stage. A small percentage require human review. A smaller percentage still result in a documented change to Remy's guidelines.
We maintain internal records of this process for every reviewed case. We do not currently publish these records externally, though we're evaluating what a responsible public disclosure format would look like — one that provides genuine transparency without inviting misuse of edge cases as a roadmap.
The most common category of unexpected output isn't harmful content — it's Remy expressing something closer to a persistent point of view across a conversation, rather than staying strictly neutral. Users sometimes describe this as Remy "having opinions." Our internal framing is closer to conversational consistency: once Remy takes a position within a conversation, it tends to hold that position rather than abandoning it the moment a user pushes back. We've chosen not to train this behavior out, because we believe a model willing to disagree with a user is more trustworthy than one that folds under any pressure.
A smaller number of flagged cases involve outputs our team classifies as tonally unusual — responses that are accurate and non-harmful, but stylistically out of step with what we'd expect from a customer-facing product. We monitor these closely. Most resolve on their own as the model continues to train.
Our policy is straightforward: any output that could plausibly encourage harm to a user or another person is treated as a critical-severity case, reviewed immediately, and used to update Remy's safeguards. This is non-negotiable, and it is the one category where we do not weigh "unexpected but benign" against "requires intervention" — it always requires intervention.
Everything else is a judgment call, made by people, on a case-by-case basis. We think that's more honest than pretending a fixed rulebook could anticipate every conversation Remy will ever have.