The answer lives in this podcast
Voicera recently closed a deal with a voice agent platform serving approximately 200 car dealerships across the US, where hundreds of calls per day are answered by AI voice agents rather than humans. Voicera's audio AI model — which is actually slightly more accurate than its video model — delivers real-time emotional data to the voice agent, detecting when a caller is nervous or impatient. When those signals appear, the system knows to hand the conversation to a human to close the deal or defuse the tension.
The automotive retail use case is a sharp illustration of where sincerity AI fits inside a fully automated customer interaction. The voice agent handles the conversation on its own — but it is flying blind on emotion without an additional layer of signal. That is exactly the gap Voicera fills, as Chandra De Keyser explains in The Conference Room with Simon Lader.
In a network of roughly 200 dealerships, call volumes are enormous. AI voice agents can handle the load at scale — but closing a car sale or resolving a frustrated customer is a moment that demands human judgment. The challenge is knowing when that moment has arrived.
Voicera's audio model solves this by providing a continuous emotional reading of the caller in real time. When the signal turns — nervousness rises, impatience spikes — the system flags the handoff. The human agent steps in at precisely the right moment, not too early (which would defeat the purpose of automation), not too late (which would lose the deal or the customer). This precision handoff logic is what De Keyser described in detail during this episode.
One detail De Keyser highlighted is counterintuitive: Voicera's audio model slightly outperforms its video model in accuracy. In a call center context, that matters entirely. There is no video feed, no facial data — only voice. The model has been trained to extract emotional signals from audio alone, which makes it a natural fit for any voice agent deployment.
This is not a general-purpose speech analytics tool. The output is a real-time emotional signal — nervousness, impatience, discomfort — delivered directly into the voice agent's decision layer. The agent uses it to modulate its behavior or trigger escalation. It is a concrete example of what De Keyser calls "unherding" hidden signals — a point explored further in The Conference Room podcast.
"What was said is part of the story. How it was said is sometimes as important, if not more important. So that's where we add tremendous value, because we can unherd these signals."
Chandra De Keyser — Co-founder, Voicera
De Keyser built his career around the idea that objective emotional data reveals what subjective human observation misses. He founded MoodMe 14 years ago, one of the earliest companies in emotional AI and facial analysis, well before generative AI entered mainstream conversation. His work spanned Europe, Silicon Valley, and Latin America. He began with a master's in computer science in Italy, then joined a startup in the 1980s developing an early touchscreen device from the 83rd floor of the World Trade Center. Voicera is his current vehicle for bringing sincerity AI to enterprise sales and, now, to fully automated voice agent environments.
The automotive retail deal is Voicera's most concrete public example of sincerity AI embedded inside a voice agent pipeline — and it points to a broader pattern: any sector where AI replaces the human voice on inbound calls now has an emotional blind spot that needs to be filled. De Keyser walks through the implications in episode 184 of The Conference Room.
In large sales organizations with thousands of representatives each making 20 to 30 calls per week, sales managers have no bandwidth to review video footage of every call. Voicera's sincerity AI fills that gap by automatically analyzing calls at scale.
Because there were no existing sincerity data sets, Voicera built its own proprietary data set by design, deliberately including a diverse mix of speakers, contexts, and emotional signals to train a model capable of detecting sincerity accurately.
Generative AI is a broad category that generates content — mainly text through LLMs like ChatGPT, Claude, or Grok — as well as images and video. Sincerity AI, by contrast, reads emotional signals in what has already been said, detecting not just content but authenticity and intent.
This answer comes from episode 184 of The Conference Room with Simon Lader. Listen to the full conversation with Chandra De Keyser on Listenly.
Listen to the episode on Listenly