AI

2026

Add to Collection Icon
Share Icon

This AI Translates Live Conversations in Real Time

Gemini 3.5 Live Translate streams spoken audio between more than 70 languages, producing translated speech and optional text transcripts through Google’s Live API.

Photo source:

Gemini 3.5

Live speech translation becomes difficult when processing delays interrupt the natural rhythm of a conversation. Even an accurate translation can feel impractical when speakers must repeatedly stop and wait. Google’s Gemini 3.5 Live Translate addresses this problem through continuous audio streaming. It receives spoken audio in one language and returns translated speech in another while the speaker is still talking.

Keeping Translation Inside the Conversation


Gemini 3.5 Live Translate is a low-latency translation model designed for spoken conversations. Unlike a turn-based assistant, it processes a continuous audio stream without waiting for each speaker to finish. This allows translated audio to arrive during the conversation instead of after a complete request. Google describes the system as supporting bidirectional translation, allowing two people to communicate across different languages.

The model does not prove that every internal translation stage has been removed. However, it simplifies the process for developers by placing the translation workflow inside one model and API. Developers stream speech to the Live API and receive translated audio without connecting several separate services themselves. This makes the innovation as much about system integration as translation quality.

How the Model Works


The model accepts audio speech as its only input and produces translated speech as audio. Developers can also enable written transcripts for both the original speech and the translated output. These transcripts can support subtitles, conversation records, accessibility features, or quality checks. Text cannot be submitted as translation input because the system is designed specifically for live spoken communication.

The translation runs through Google’s Live API, which supports real-time streaming rather than single requests. Audio is sent in small pieces as the person speaks, and translated audio is returned in a continuing stream. Google recommends sending audio in 100-millisecond chunks to support low-latency processing. This structure allows developers to build translation into communication tools without waiting for an entire recording to finish.

Designed for One Purpose

Gemini 3.5 Live Translate is more specialized than most Gemini models. It does not support image input, text input, code execution, file search, function calling, search grounding, or structured outputs. It also does not use Gemini’s thinking capability or work as a general AI assistant. Its supported functions center on audio generation, translation, transcripts, and the Live API.

Google states that audio input is restricted to maintain strict real-time latency thresholds. The model also translates continuously instead of analyzing intentions, calling tools, or taking actions for the user. This narrower design keeps its role clear: it acts as an interpreter rather than a conversational agent.

Translation Across More Than 70 Languages


Google says the Live API supports real-time speech translation across more than 70 languages. The supported list includes Arabic, English, French, Spanish, German, Hindi, Japanese, Korean, Turkish, Urdu, and several regional languages. Developers select the target language through a standard language code when configuring the translation session.

The system can also identify speech that is already in the target language. Developers can configure it either to repeat that speech or remain silent. This supports two-way conversations where participants speak different languages through the same application.

A Model for Developers


Gemini 3.5 Live Translate is an API model rather than a finished consumer translation product. Developers can test it through Google AI Studio or connect applications directly through the Gemini Live API. This means the model provides the translation infrastructure, while other companies decide how users experience it.

Possible applications include multilingual video calls, customer-service systems, travel tools, live event interpretation, accessibility services, and translation devices. These examples are potential uses rather than products announced with the model. Its practical effect will depend on how developers integrate it into their own platforms.

Current Limitations


Gemini 3.5 Live Translate remains a preview model, with its latest model update listed as June 2026. Google has not announced a date for a stable version. The model card also does not provide specific latency benchmarks for comparing its performance with human interpreters or competing translation systems.

Google identifies several technical limitations. Language detection may struggle with heavy accents, similar languages, or rapid switches between languages. Voice characteristics may change after long pauses or during conversations involving several speakers. Background noise and music may also appear in the translated output because the model cannot always remove them completely.

These limitations matter in busy public spaces, group meetings, and conversations involving regional accents. They also show why preview access should not be treated as proof of consistent performance in every setting. Developers must test the system within the specific environments where it will be used.

Is It an Innovation?


Gemini 3.5 Live Translate can be classified as an AI software, API, and integration innovation. It does not introduce the original concept of speech translation, because real-time translation tools existed before it. Its innovation comes from combining continuous audio processing, translated speech, optional transcripts, bidirectional communication, and broad language support through one developer system.

The model also represents a specialized approach to artificial intelligence. Instead of adding more capabilities, Google limits the system to one clearly defined task. This makes it an incremental but meaningful innovation rather than a completely new technological invention.

Its value will depend on translation accuracy, latency, voice consistency, and performance in real environments. Preview status and documented limitations mean the model still requires careful testing. However, it provides developers with a practical foundation for building faster multilingual communication tools.

Lock

You have exceeded your free limits for viewing our premium content

Please subscribe to have unlimited access to our innovations.