Gemini 3.5 Live Translate streams spoken audio between more than 70 languages, producing translated speech and optional text transcripts through Google’s Live API.
Photo source:
Gemini 3.5
Live speech translation becomes difficult when processing delays
interrupt the natural rhythm of a conversation. Even an accurate translation
can feel impractical when speakers must repeatedly stop and wait. Google’s
Gemini 3.5 Live Translate addresses this problem through continuous audio
streaming. It receives spoken audio in one language and returns translated
speech in another while the speaker is still talking.
Gemini 3.5 Live Translate is a low-latency translation model designed for
spoken conversations. Unlike a turn-based assistant, it processes a continuous
audio stream without waiting for each speaker to finish. This allows translated
audio to arrive during the conversation instead of after a complete request.
Google describes the system as supporting bidirectional translation, allowing
two people to communicate across different languages.
The model does not prove that every internal translation stage has been
removed. However, it simplifies the process for developers by placing the
translation workflow inside one model and API. Developers stream speech to the
Live API and receive translated audio without connecting several separate
services themselves. This makes the innovation as much about system integration
as translation quality.
The model accepts audio speech as its only input and produces translated
speech as audio. Developers can also enable written transcripts for both the
original speech and the translated output. These transcripts can support
subtitles, conversation records, accessibility features, or quality checks.
Text cannot be submitted as translation input because the system is designed
specifically for live spoken communication.
The translation runs through Google’s Live API, which supports real-time
streaming rather than single requests. Audio is sent in small pieces as the
person speaks, and translated audio is returned in a continuing stream. Google
recommends sending audio in 100-millisecond chunks to support low-latency
processing. This structure allows developers to build translation into
communication tools without waiting for an entire recording to finish.
Gemini 3.5 Live Translate is more specialized than most Gemini models. It
does not support image input, text input, code execution, file search, function
calling, search grounding, or structured outputs. It also does not use Gemini’s
thinking capability or work as a general AI assistant. Its supported functions
center on audio generation, translation, transcripts, and the Live API.
Google states that audio input is restricted to maintain strict real-time
latency thresholds. The model also translates continuously instead of analyzing
intentions, calling tools, or taking actions for the user. This narrower design
keeps its role clear: it acts as an interpreter rather than a conversational
agent.
Google says the Live API supports real-time speech translation across
more than 70 languages. The supported list includes Arabic, English, French,
Spanish, German, Hindi, Japanese, Korean, Turkish, Urdu, and several regional
languages. Developers select the target language through a standard language
code when configuring the translation session.
The system can also identify speech that is already in the target
language. Developers can configure it either to repeat that speech or remain
silent. This supports two-way conversations where participants speak different
languages through the same application.
Gemini 3.5 Live Translate is an API model rather than a finished consumer
translation product. Developers can test it through Google AI Studio or connect
applications directly through the Gemini Live API. This means the model
provides the translation infrastructure, while other companies decide how users
experience it.
Possible applications include multilingual video calls, customer-service
systems, travel tools, live event interpretation, accessibility services, and
translation devices. These examples are potential uses rather than products
announced with the model. Its practical effect will depend on how developers
integrate it into their own platforms.
Gemini 3.5 Live Translate remains a preview model, with its latest model
update listed as June 2026. Google has not announced a date for a stable
version. The model card also does not provide specific latency benchmarks for
comparing its performance with human interpreters or competing translation
systems.
Google identifies several technical limitations. Language detection may
struggle with heavy accents, similar languages, or rapid switches between
languages. Voice characteristics may change after long pauses or during
conversations involving several speakers. Background noise and music may also
appear in the translated output because the model cannot always remove them
completely.
These limitations matter in busy public spaces, group meetings, and
conversations involving regional accents. They also show why preview access
should not be treated as proof of consistent performance in every setting.
Developers must test the system within the specific environments where it will
be used.
Gemini 3.5 Live Translate can be classified as an AI software, API, and
integration innovation. It does not introduce the original concept of speech
translation, because real-time translation tools existed before it. Its
innovation comes from combining continuous audio processing, translated speech,
optional transcripts, bidirectional communication, and broad language support
through one developer system.
The model also represents a specialized approach to artificial
intelligence. Instead of adding more capabilities, Google limits the system to
one clearly defined task. This makes it an incremental but meaningful
innovation rather than a completely new technological invention.
Its value will depend on translation accuracy, latency, voice
consistency, and performance in real environments. Preview status and
documented limitations mean the model still requires careful testing. However,
it provides developers with a practical foundation for building faster
multilingual communication tools.
Please subscribe to have unlimited access to our innovations.