Skip to content
Al indsigt
  • Speech AI
  • Latency
  • Real-Time Systems

Live speech translation: where model speed fights speech quality

An ASR-to-MT-to-TTS pipeline adds latency at every stage. In the live multilingual meetings product we built, the real decision wasn't how to make latency zero, but where in the pipeline latency was allowed to live.

Udgivet

9 min. læsning

Every live speech translation pipeline passes through three stages: automatic speech recognition (ASR), machine translation (MT), and text-to-speech synthesis (TTS). Each stage adds its own latency, and the natural temptation is to assume the problem is that one of these models is too slow. Building a live, multilingual voice-and-video meetings product taught us that the real bottleneck sits somewhere else: the largest share of latency happens before the first model even starts working.

Where latency actually accumulates

Before anything can be translated, the system has to figure out when a speaker has actually finished their sentence. Detecting the "end of turn" requires a silence window — the system has to wait long enough to be confident the speaker has genuinely stopped, not just paused to breathe. That window is a physics-bound latency floor that no faster model can work around, because the constraint isn't processing speed, it's confidence that the sentence is actually over. Shorten that window and the system assumes sentences are "done" too eagerly and translates incomplete fragments; lengthen it and translation arrives more accurately but later.

Only after that window closes do translation and speech synthesis enter the picture — and there's an architectural decision here that genuinely changes perceived latency: instead of waiting for the entire translated sentence to be generated, then the entire audio clip to be synthesized, and only then starting playback, you can start playback the moment the first chunk of audio is ready, while the rest is still being generated. This doesn't change the pipeline's shape — recognition, translation and synthesis still happen in that order — but it reduces the latency the listener actually perceives, because the listener isn't waiting for "finished," they're waiting for "started."

Another point about where latency accumulates: when multiple target languages and multiple simultaneous listeners are involved, the right design is not to generate a separate translation per listener. If ten listeners are hearing the same language, you need exactly one translation and one audio clip for that language, not ten. This is partly a cost optimization, but more importantly it keeps latency from growing linearly with the number of listeners — latency grows with the number of languages present in the room, not the number of people.

Cross-talk and accents

In a real conversation, people talk over each other, pause, and abandon sentences halfway through. One approach is for the engineering team to write its own voice-activity-detection and sentence-boundary logic from scratch. The other — and the more sensible one, given how well modern speech recognition models already handle this — is to hand that responsibility to the same model that's already listening to the audio, rather than maintaining two independent sources of truth for "when did the sentence end." Two independent sources of truth means two places to be wrong, and when they disagree, whoever wrote the surrounding code has to resolve it, not the model.

Source language presents a similar problem: asking the user "what language are you speaking" is a poor user experience and doesn't even make sense in a multilingual room, since speakers may switch over the course of a session. The more sensible approach is automatic detection from among a set of plausible candidates — the languages actually present in that room — rather than guessing from every language in existence. This improves both accuracy and detection latency, because the search space is smaller.

How much lag a conversation tolerates

Here it's worth being honest: we don't have a number we're confident enough to state as "this many milliseconds is the ceiling a conversation tolerates," and a number like that shouldn't be borrowed from somewhere we're not sure of. What can be said is qualitative: latency that is consistent and predictable, even if it runs to a few seconds, reads as "an interpreter is translating" — something close to the rhythm of human simultaneous interpretation that people are already familiar with. Latency that varies — sometimes fast, sometimes slow, with no predictable pattern — is far more grating, because the listener's brain can't settle into an irregular rhythm. The practical conclusion is that the engineering goal shouldn't just be "lowest possible latency," it should also be "lowest possible variance in latency." A system that always responds with a fixed, predictable delay can produce a better experience than one that's sometimes faster but unpredictable.

Evaluating a system you can't unit-test

The biggest difference between this pipeline and an ordinary service is that "correctness" of a translation can't be checked with a simple assertion. A unit test can confirm that a function turns one input into one expected output, but it can't say whether a translation reads naturally, whether the speaker's tone survived, or whether the model fabricated something the speaker never actually said. Language models can confidently invent things that weren't said — adding a currency unit or a number out of nowhere — and this class of error is only caught by actually reviewing real samples, not by running a pre-written test suite.

What works in practice is a combination of two layers: a continuous, automated layer that monitors what's actually measurable — the time each pipeline stage takes, the error rate of calls to underlying services, whether speech synthesis starts on time — and a human, sample-based review layer that specifically looks for what no metric can show: whether the translation changed meaning, whether something was added that wasn't said. This second layer is slower and more expensive, but it's the only one that actually measures quality rather than just speed. A team that watches only automated metrics may end up with a system that always responds quickly and occasionally says things nobody said — and the metrics will never show that.

Not every kind of latency shares one budget

One last point worth making: a speech translation product doesn't necessarily have a single latency budget; it has several, depending on whether a given path is genuinely "live." The live path — the one running in the middle of an ongoing conversation — needs short, hard timeouts on every service call, because a call that's gotten slow is better off failing and falling back to something else (say, showing text instead of synthesized audio) than holding up the entire session. But if the same product also has a non-live capability — dubbing a recorded video after a session ends, for instance — that path can and should have much longer timeouts, because nobody is waiting on the other end of the line for its result. The common mistake is defining a single timeout for "calling the translation service" across the whole codebase; the correct timeout is a function of whether a human is actually waiting on the answer right now.

This distinction has a direct product consequence. A team that wants to build both a live experience and a non-live one (post-session subtitles or recording dubbing, say) on the same translation infrastructure needs to design these as two paths with different latency budgets from day one, rather than assuming "the same pipeline is fine for both." A pipeline optimized for being live usually sacrifices quality that doesn't need to be sacrificed on the non-live path; conversely, a pipeline designed for the highest possible quality usually isn't fast enough for the live experience.

Questions worth answering before building a pipeline like this

For a team just starting to build a live speech translation product, a handful of upfront questions matter more than any specific technology choice: is the silence window we've picked for detecting end-of-sentence appropriate for languages with naturally longer pauses, or was it tuned for just one? Do we start playback the moment the first audio chunk is ready, or wait for the whole clip? Do we generate one translation per target language, or one per listener? Do the live and non-live paths carry separate timeouts? And, most importantly: who reviews real output samples, and how often, to make sure the model hasn't invented something on its own? The answer to each of these can shift the user's experience from "feels like a live interpreter" to "slow and untrustworthy," or the reverse — and none of them can be guessed from a language model's documentation alone.

Har du noget, der skal bygges?

Fortæl os, hvad du arbejder på. Vi siger ærligt, om vi er det rette team til det.

Start en samtale

eller skriv til os på hello@larsima.com