Google's Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking models have achieved a leading score of 82.6% on the Artificial Analysis Speech to Speech Index. The models support real-time, multilingual conversations across 97 languages, with advanced capabilities in visual context processing and tool execution.
TL;DR
- Google's new models lead in speech-to-speech performance with a 82.6% score.
- The models support real-time, multilingual conversations and advanced tool execution.
- Other AI labs, including StepFun, Qwen, and TypeSafe AI, also released new audio and decision models this week.
What happened
Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, real-time dialogue models that accept streaming audio, images, video, and text. These models are designed to handle multilingual conversations, real-time visual context processing, and background tool execution while maintaining a 128K-token input context.
StepFun introduced the StepAudio 3 model family, including real-time conversation, music generation, voice audio generation, automatic speech recognition, and text-to-speech models. The real-time model showed excellent conversational dynamics and speech reasoning benchmark performance.
Qwen launched Qwen3.8-Omni-Flash, an omni-modal model for professional audio-visual content production, and Qwen3.8-LiveTranslate, a real-time simultaneous interpretation model supporting 60 languages.
Why it matters
Google's models lead the Artificial Analysis Speech to Speech Index with a 82.6% score, outperforming competitors like OpenAI's GPT-Live-1 and Grok Voice. This positions Gemini 3.8 Live as a top choice for developers and startups looking for advanced speech-to-speech capabilities.
The release of StepAudio 3 and Qwen3.8 models expands the options for developers seeking high-performance audio processing and real-time translation tools. These models are offered through APIs, making them accessible for integration into various applications.
The advancements in real-time audio and speech processing models are driving the AI industry towards more natural and efficient human-AI interactions, benefiting both developers and end-users.
Key facts
- Gemini 3.8 Live Extended Thinking leads the Artificial Analysis Speech to Speech Index with a score of 82.6%.
- Gemini 3.8 Live supports 97 languages and a 128K-token input context.
- StepAudio 3 Realtime shows excellent conversational dynamics and speech reasoning benchmark performance.
- StepAudio ASR Max recorded a 1.7 percent word-error rate in the cited evaluation.
- Qwen3.8-Omni-Flash supports text, image, audio, and video inputs with a 1-million-token context window.
- Qwen3.8-LiveTranslate supports 60 languages and features real-time speaker separation and synchronized bilingual output.
Context
The AI industry is increasingly focused on enhancing real-time interaction capabilities, driving the development of advanced speech-to-speech and audio processing models. These advancements are crucial for creating more natural and efficient human-AI interactions, which are essential for various applications, from customer service to content creation.
Google's leadership in the speech-to-speech index highlights the company's commitment to advancing AI technologies. The release of Gemini 3.8 Live and Extended Thinking models underscores the importance of real-time, multilingual capabilities in the current AI landscape.
Other AI labs, such as StepFun and Qwen, are also contributing to the growth of audio and speech processing technologies. Their models offer developers a range of options for integrating advanced audio capabilities into their applications, further driving innovation in the AI industry.
