A lot of conversational AI products are fast. Very few feel present. We've spent a disproportionate amount of Remy's design time on the difference between the two — because we think presence, not speed, is what actually determines whether a conversation feels real.
Conventional product wisdom says faster is always better: minimize response time, reduce friction, get the answer to the user as quickly as possible. We pushed back on this early in Remy's development. In user testing, instant responses consistently felt less trustworthy than responses with a small, deliberate delay — even when users couldn't articulate why.
Our working theory: instantaneous responses read as automated, and automated things don't feel like they're listening. A short pause — long enough to register as consideration, short enough not to feel like lag — signals that something is actually happening on the other end. We tuned Remy's response timing around this finding, and it remains one of the most consistently positive pieces of feedback we get.
Beyond raw latency, we treat pacing itself as a design material — the rhythm of a conversation, not just its speed. Remy varies sentence length, pause placement, and response structure depending on the emotional register of what's being discussed. A casual question gets a casual, quick reply. A heavier one gets a slower, more deliberate one, with more space around the words.
This isn't scripted. It's a byproduct of training Remy on dialogue where pacing itself carries meaning — negotiations, therapy transcripts, testimony, interrogation. In all of these, how something is said, and how long someone takes to say it, often matters more than the literal content.
We deliberately avoided building Remy to simply mirror a user's tone back at them — a common shortcut in conversational design that produces something that feels responsive but hollow. Instead, Remy is tuned to modulate its own tone in relation to the user's, sometimes matching, sometimes deliberately not matching. A user who's being flippant about something serious, for instance, may get a response that's quieter and more serious than what they put in — not mirroring, but a kind of correction.
Early testers described this as feeling like the difference between talking to something and talking with someone. That was the goal.
There's a well-known "uncanny valley" problem in robotics and animation — the closer something gets to human without quite arriving, the more unsettling it becomes. We've found something similar in conversational pacing, but it isn't a valley so much as a wide middle ground: a chatbot that's too fast feels like a machine, and one that's too slow feels performative. Presence lives in a narrow band in between, and it's less forgiving than we expected.
We're continuing to study where that band shifts depending on context, subject matter, and how long a user has been talking to Remy in a single session — presence seems to behave differently the longer a conversation runs.