How to Translate a Live Conversation
Both sides, in real time, without either person waiting for a turn. Here's how to translate a live conversation in the browser — and which mode to use for the room you're in.
Most translation apps handle one sentence at a time: you speak, you stop, you wait, you hand the phone over. That works for asking directions and falls apart in an actual conversation. To translate a live conversation you need both directions running at once, and you need it to keep up with people who interrupt each other, trail off, and answer in three words.
Live Translate Live is a live conversation translator that runs in a browser. There are two ways to use it, and the right one depends on whether you can put a screen down between you.
The Short Version
- Open the app in any browser and pick the two languages.
- Put the device where both people can see it, or keep it in hand if you'd rather listen than read.
- Start talking. Both people speak their own language, at the same time as each other if they want.
- The translation appears on screen as it's spoken — or is played aloud in the other person's language.
Nothing to install, no app store, and no hardware. A device with a microphone and a browser is the whole requirement.
Mode 1 — Continuous, Read on a Shared Screen
This is the mode that most resembles a real conversation. The microphone stays open, both people talk in their own language, and the translation flows onto the screen as they speak. Nobody presses a button to take a turn and nobody waits for the other person's translation to finish before answering.
The screen is shared rather than personal, which is the part most tools don't do. One phone laid flat on the table between two people shows both halves of the conversation, each person reading their own language at the same time — and if there's a bigger screen in the room, the display opens in its own window with its own URL, so it can go on a second monitor, a projector, or a TV while the microphone stays on the phone near the people talking. Competing tools output to a phone in one person's hand or to earbuds in one person's ears; this one is meant to be looked at by both people at once.
Speech recognition runs server-side rather than in the browser, which is what makes it hold up against accents, background noise, and people who don't speak in tidy complete sentences. Translation carries context across turns, so a short answer or a follow-up question is translated against what came before it rather than on its own.
103 languages are supported for speech recognition, in any combination — 10,506 possible pairs, not just to and from English. See the list.
Mode 2 — Turn-Based, Heard Out Loud
Audio mode is the one to use when a shared screen isn't practical: standing up, walking, at a counter, or anywhere you'd rather hold the phone than lay it down.
You speak, your transcript appears so you can check it caught you correctly, and you tap to hear the translation spoken aloud in the other person's language. Then you hand the phone over for their turn. Spoken playback covers 74 languages today, and languages without a voice yet still work on the scrolling display.
The tap is deliberate. It means you can fix a bad transcription before anything is said out loud, and it means nothing is translated while you're standing in a noisy room not talking to anyone — which also keeps credits from draining in the background. More on audio mode.
Which Mode for Which Conversation
- Sitting across a table — continuous mode, phone flat between you. Both of you read.
- A meeting room with a screen — continuous mode on the big display, microphone on the phone.
- Standing, walking, or at a counter — audio mode, spoken aloud, phone in hand.
- Someone who can't comfortably read the screen — audio mode, so nobody has to read anything.
- Captioning one language, not translating — set both languages the same and it becomes a live transcript at a fraction of the credit cost. Same-language mode.
What's Coming: Continuous Speech-to-Speech
In development — not available yet. This section describes work in progress, not a feature you can use today. Everything above this heading is live now.
Both modes above put a step between speaking and hearing: either you read the translation on a screen, or you tap a button to play it. We're building a third mode that removes it. You speak, and your words come out as spoken audio in the other person's language continuously — trailing a few seconds behind you rather than waiting for you to finish, and running in both directions at the same time.
What that removes is the reading. In the modes that exist today, at least one person has to look at something for the conversation to work. In continuous mode neither does, which matters when you're walking somewhere together, when someone can't comfortably read the screen, or when the script isn't one they read at all.
Two honest caveats, because they'll decide whether this is the right mode for you when it lands. It will cover fewer languages than the 103 above, so the scrolling display will remain the broader option and is not going away. And it has no typed-input path — if you need to type or paste text in a language you can read but can't pronounce, that stays with the existing pipeline.
Continuous mode is an addition, not a replacement. Both modes on this page keep working exactly as they do now.
Common Questions
Can it translate both sides of a conversation at the same time?
Yes, in continuous mode. Both microphones' worth of speech are transcribed and translated in both directions simultaneously, and both halves appear on the shared display together. Neither person has to wait for the other's translation to finish before speaking. In audio mode it's deliberately turn-based instead, because that mode is built around spoken playback on a handheld phone.
How do I use it for an in-person conversation?
Open it in the browser, choose the two languages, and put the device between you with the screen facing up. The display's top half is flipped 180° by default so the person across from you reads right-side-up — that's vis-à-vis mode, and you can turn it off from the menu if you're sitting side by side. Then just talk normally.
Can it handle more than two languages in the same meeting?
Not as separate simultaneous outputs — a session is a pair of languages, and the display shows those two. Speech recognition will still transcribe whatever it hears rather than dropping a third language on the floor, but it will be translated into the session's target language, not into a third one. For a multilingual meeting, the practical setup today is one session per language pair on separate devices.
Can it translate a phone conversation?
Not a phone call itself — it doesn't tap into call audio, and no browser app can. What works is an in-person conversation, or a call on speakerphone with the app open on a second device near the speaker. If translating live calls is specifically what you need, this isn't the right tool.
What does it cost to try?
Credits start at $1 for about 15 minutes of live translation, or $3 for an hour. There's no subscription and credits don't expire. In audio mode, transcription is free — you only spend credits when you tap Translate. Full pricing.
Which services process the audio?
Stated per step, because different steps use different providers: speech is transcribed by ElevenLabs (Scribe v2 Realtime), the transcribed text is translated by Google Gemini via Google Cloud Vertex AI, spoken playback is voiced by ElevenLabs, and sign-in runs through Google. The continuous speech-to-speech mode described above will send conversation audio directly to Google when it ships. Full detail is in the privacy policy.
Try It on a Real Conversation
Credits start at $1 for about 15 minutes. No subscription.
Get Started