
These execs think voice AI hasn’t reached its ChatGPT moment yet
Quick Answer
Voice AI has not yet reached its 'ChatGPT moment,' according to PolyAI's CTO Shawn Wen, who emphasizes the need for faster reasoning in full-duplex models.
Quick Take
Otter's CMO Alex Gay highlights the importance of accurate transcription and emotional expression in AI communication, as current models struggle with understanding context and user intent.
Key Points
- PolyAI's full-duplex models can speak and listen simultaneously but lack fast reasoning.
- Otter focuses on improving transcription accuracy to enhance productivity and trust.
- Voice AI tools must ensure users know they are interacting with AI.
- Emotional expression in AI is crucial for effective communication in meetings.
- Investors are heavily funding voice AI startups across various applications.
DeepSignal Analysis
What happened
Voice AI is seen as a promising interface, yet experts like Shawn Wen from PolyAI argue it hasn't achieved a significant breakthrough akin to ChatGPT. Current models still struggle with reasoning speed and contextual understanding, as highlighted by Otter's Alex Gay, who emphasizes the need for accurate transcription and emotional expression in AI communication.
Key evidence
- Shawn Wen, CTO of PolyAI, stated that while full-duplex models exist, voice AI has not yet reached its 'ChatGPT moment' due to slow reasoning capabilities.
- Alex Gay, CMO of Otter, noted that accurate transcription and emotional expression are critical for effective AI communication, as current models often fail to capture user intent.
- Both Wen and Gay agree that Automatic Speech Recognition (ASR) models frequently miss important keywords, leading to flawed context capture and diminished trust in AI platforms.
Why it matters
The development of voice AI is crucial for enhancing user interaction in various applications, from customer service to meeting transcription. However, the current limitations in reasoning speed and contextual understanding hinder its potential. As companies invest heavily in this technology, addressing these shortcomings is essential for building user trust and ensuring effective communication. The insights from industry leaders indicate that without significant improvements, voice AI may struggle to fulfill its promise as a next-generation interface.
📖 Reader Mode
~3 min readThe theory of voice being the next big interface has picked up strong momentum, with investors pouring billions of dollars into voice AI startups working on areas ranging from model makers to enterprise customer service providers, and from meeting note-takers to AI-powered dictation.
Every week there is a new model or a tool release that claims to sound human and converse like one. However, in reality, that might not be the case. Enterprise voice AI platform PolyAI’s CTO Shawn Wen thinks that despite the release of full-duplex models — which can speak while listening to you — voice AI doesn’t have its “ChatGPT moment” yet.
“We have reached the milestone of developing full-duplex models. The next challenge is to make reasoning very fast, so that the models can fetch answers quickly and the conversation feels natural,” he told me on stage at the HumanX conference last month.
He also said that AI agents in customer service should not sound robotic and should give callers enough confidence that they can solve problems.
“I think the next stage will be slightly different because once the voice is good enough, like, and the customer is willing to engage with them for the first two or three turns, they start to build confidence, and over time, they will feel like I probably don’t have to talk to a human if the agent can solve my problem,” he said.
Alex Gay, CMO for meeting notetaker Otter, opined that speaker identification, intent capture, and typing that up with organizational knowledge is a key step for enabling automation. The company is also working on digital twins that might represent people in meetings. For that technology, he said it’s paramount that the output voice gives the same emotive expressions of talking to a human in a meeting.
“If you think about the meetings that you’re in right now, the best conversations that you have are where you can have debate, and strategic discussions, and when you feel like there’s a relationship that underpins it. If you aren’t able to have that with an avatar, then it’s just a q and a chatbot,” Gay noted.
Voice AI’s understanding and transparency
While voice AI models have improved, AI assistants often don’t understand users, or your meeting notetaker shows the wrong transcript or a summary.
Wen thinks that ASR (Automatic Speech Recognition) models often miss important keywords, and that creates an issue in capturing the whole context.
Otter’s Gay agreed with this, adding that the company keeps working on improving transcription. He also said that language is one area where voice models need to improve.
“For Otter, you know, transcription was never the end point. It was just the layer that we could start to drive some of the productivity gains on the back of. But if your original transcription didn’t have the accuracy that you needed, all the follow-up actions that you have become flawed. And the minute that starts to take action, that is wrong. You lose trust in the platform. It is critical for us to continue to improve that ASR model because all of the downstream impacts are significant,” he said.
With the new voice tools, there is also a question of transparency. Tools should declare to customers that they are being recorded or talking to AI. Otter said that it wants to instill trust in people who are in a meeting, so even for a meeting where the bot is not present, it wants to try methods like notifying everyone in the chat that the meeting is being recorded. PolyAI’s Wen also said that it’s important to establish that people are talking to an AI in enterprise calls.
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.
Ivan covers global consumer tech developments at TechCrunch. He is based out of India and has previously worked at publications including Huffington Post and The Next Web.
You can contact or verify outreach from Ivan by emailing [email protected] or via encrypted message at ivan.42 on Signal.
— Originally published at techcrunch.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from TechCrunch
See more →
AI chip startup Etched defies skeptics, hits $10.3B valuation from big-name investors
AI chip startup Etched has achieved a $10.3 billion valuation after a $300 million Series C funding round, led by Sequoia and supported by notable investors like Andreessen Horowitz. The company claims to have developed innovative low-voltage chips for AI inference, significantly enhancing performance and reducing costs, with $1 billion in orders already booked.

