Voice AI is attracting billions of dollars in investment and a steady flow of products that promise humanlike conversation, but industry executives say the technology has not yet reached a defining mass-market breakthrough. In a TechCrunch report from the HumanX conference, leaders from enterprise voice platform PolyAI and meeting assistant Otter described a gap between sounding natural and being dependable enough to handle consequential tasks.
PolyAI chief technology officer Shawn Wen told TechCrunch reporter Ivan Mehta that full-duplex models are an important milestone because they can listen while speaking. His view, however, is that the next challenge is faster reasoning: a system must retrieve and formulate an answer quickly enough for the exchange to feel natural. On that account, fluid turn-taking is necessary but does not by itself create a capable conversational agent.

Customer service makes that distinction especially visible. Wen argued that an automated agent should avoid a robotic delivery, but he tied user confidence to whether it can actually resolve a problem. He suggested that callers may become willing to continue with an AI after the first few turns if the system demonstrates that it can complete the task. That is an executive’s expectation, not evidence that callers universally accept voice agents today.
The deeper weakness begins before reasoning. Wen said automatic speech recognition systems can miss important keywords, causing the system to lose the context of an exchange. A voice agent can therefore sound polished while working from an incomplete or incorrect interpretation of what the person said. The more actions connected to that interpretation, the greater the potential downstream impact.
Otter chief marketing officer Alex Gay described transcription as a foundation rather than the final product. According to TechCrunch, Otter wants transcripts to support follow-up automation and other productivity features. Gay warned that when the initial transcript is inaccurate, subsequent actions inherit the error; once a system acts incorrectly, users can lose trust in the platform.

Otter is also working on digital twins that could represent people in meetings. Gay said such systems would need to reproduce emotional expression well enough to support debate, strategy and the sense of an existing relationship. Without those qualities, he argued, a meeting avatar amounts to little more than a question-and-answer bot. TechCrunch did not report a release date or performance results for the digital-twin work.
Transparency adds another requirement. Both executives said people should know when they are speaking with AI or being recorded. Otter is considering notifications in meeting chats even when its bot is not visibly present, while Wen said enterprise callers should be told that the other side is an AI system. Those practices address consent and trust, but the report does not claim that disclosure standards are consistent across the industry.
The picture that emerges is less a race to create the most lifelike voice than a test of the entire operating chain. Recognition must capture the right words, reasoning must be fast and grounded in context, automation must avoid compounding mistakes, and the system must identify itself clearly. TechCrunch’s reporting suggests that voice AI will not have its breakthrough merely when it sounds human; it will need to become predictably useful without hiding what it is.

Comments
Loading comments…