Tech • AI • Robotics • Game

VIDEO
ENFR

The next generation of voice AI with Google DeepMind and Sierra AI

9.1/10
GoogleGoogle for DevelopersSeptember 24, 2026 at 11:02 PM5:21
Audio player
0:00 / 0:00

TL;DR

Two new native audio AI models are being introduced to give developers a choice between faster conversational performance and higher-precision enterprise voice agents, amid a broader push to improve latency, multilingual speech handling and real-world benchmarking.

KEY POINTS

Two-model launch for voice agents

The release centers on two native audio models. One is a lightweight, faster option built for snappy back-and-forth conversations and lighter tool use. The other is positioned as a more enterprise-focused model designed for higher precision in voice tasks, especially where multi-step function calling, stronger accuracy and higher-stakes deployment are required.

Enterprise demand is shifting toward harder tasks

Companies building customer-facing AI agents are moving beyond simple support interactions toward more complex work. That includes long-horizon reasoning, outbound agents and sales-focused workflows, reflecting how stronger large language models are expanding the range of tasks voice systems are expected to handle.

Latency is now measured in two ways

Developers are increasingly separating latency into time to first audio and time to first useful response. The first measures how quickly a model begins speaking back in a way that feels conversational. The second tracks how long it takes to deliver meaningful information after an acknowledgement such as a promise to look something up.

Faster first response is a priority

Earlier generations could post stronger benchmark results when internal reasoning modes were enabled, but that often increased delay before the model started speaking. The current push is to bring time to first audio back down to a genuinely conversational level so users feel they are in a live exchange rather than waiting on a system to think.

Native audio may improve code-switching

One of the most notable claimed advantages of native audio systems is smoother handling of multilingual speech and natural code-switching between languages. Instead of relying on a speech-to-text pipeline that can lose nuance in a "broken telephone" effect, end-to-end audio models are expected to preserve switches in language more reliably and respond more seamlessly.

Voice quality is being judged as a hierarchy

A useful framework emerging in the industry breaks voice performance into three layers: accuracy, quality and experience. Accuracy is whether the system can complete the task. Quality is whether the interaction is tolerable rather than frustrating. Experience is the highest bar, covering natural pronunciation, tone and whether the conversation is actually pleasant.

Benchmarks are evolving beyond static tests

Real-time voice evaluation is moving away from static, single-answer tests toward more dynamic setups with user simulators, noise and personas. The argument is that conversational AI should be tested in conditions closer to real deployments, where interruptions, ambiguity and shifting context matter more than matching one golden response.

Benchmark saturation is arriving quickly

Teams working on agent benchmarks say tasks initially considered difficult can become saturated within a few months as model capabilities improve. That pace is forcing benchmark designers to raise difficulty while also trying not to make public scoreboards so punishing that they discourage participation or fail to reflect practical product gains.

Dialogue is becoming a key interface

Real-time spoken interaction is increasingly seen as one of the most compelling ways to use AI because it feels seamless and reduces dependence on screens, desks or phones. That convenience, combined with rapid progress across labs, is helping push voice dialogue toward a central role in how people may use AI systems over the next year.

CONCLUSION

The latest native audio models reflect a broader industry shift from simple speech interfaces to fully deployed conversational agents that must balance speed, accuracy and user experience. The next competitive frontier is likely to be real-world voice performance, not just raw benchmark scores.

Ask a question
Full transcript

More from Google