RESEARCH NOTE

What Matters in Video Translation

FAMILIAR RESEARCHALL ARTICLES

ABSTRACT

Four things decide whether a dub holds up: the translation reads like it was written in the target language, the audio still sounds like the speaker, the on-screen text carries over, and the lips match. The first two gate everything else. All four turn out to be the same problem: understand the video, then generate it back.

01.The four problems

Ask what makes a dubbed video good and you get a long list. It compresses to four:

PROBLEMWHAT GOOD MEANS
Natural translationA new script in the target language that conveys the video well; not a sentence-for-sentence mapping.
Natural audioThe same person, same emotion, same emphasis, in the new language.
On-screen textAdapted for the language and the moment it appears in, without damaging the video.
Lip-syncInvisible. The viewer never notices it happened.

If the first two fail, nothing after them matters: a mistranslated sentence or an audibly wrong voice ends the dub on its own. The last two decide whether the result feels like the original or feels like a dub.

02.Natural translation: a new script, not a mapping

Every language expresses things its own way: different words, idioms, sentence structures, forms of address, and different choices about what gets said at all. Ask a model to "translate" and it tends to map sentences one-to-one from language A into language B. People do the same thing; the moment someone is framed as a translator, the task collapses into converting sentences.

A dub translated that way is understandable and wrong. What the viewer needs is a new script, written in the target language, that conveys the content of the video well.

  • Japanese tends to avoid overly direct phrasing; a natural line often has to be rebuilt indirectly, not softened word by word.
  • Korean omits subjects constantly; carrying English pronouns into Korean at English frequency reads foreign even when every word is correct.
  • Word order, register, honorifics, and what a culture leaves unsaid shift with every language; a translation is not good because the literal meaning survived.

"Well" resists a checklist, so we compress the goal to one sentence: convey the content of the video well in the target language. That is measurable:

100.3%

of expert human translation quality from our production pipeline, scored blind against expert reference edits

FRENCH, CHINESE, HINDI, INDONESIAN (FIRST PASS FOR EXPERT PROFESSIONAL TRANSLATORS)
−54.7%

fewer important translation errors from Familiar on the same clips (29 vs 64)

VS ELEVENLABS DUBBING V2 (ALPHA) · 111 PAIRED DUBS · 3 LANGUAGES

03.Natural audio: the same person, in a new language

Pronunciation is close to solved; if the speech itself sounds wrong, the dub already failed, but that bar is now routinely met. The interesting failures are past it:

  • Identity: the same person keeps sounding like the same person, in every language.
  • Emotion: what the speaker felt survives the language change.
  • Emphasis, pacing, intent: where the sentence leans, how fast it moves, what it is trying to do.

The goal compresses the same way the translation goal did: understand the content of the video and generate the audio that matches it. Not "generate language B in the voice from language A" -- the same performance, in a new language. And the scene is part of the audio: the laughter, the music, the room.

+270%

more background-sound damage from ElevenLabs (11.92 vs 3.22 dB): music, effects, ambience

DUBBING V2 (ALPHA) STUDY · SAME CLIPS
+27.8%

closer speaker resemblance from Familiar (0.461 vs 0.360), closer in all 11 languages

VS ELEVENLABS · DUBBING V1 STUDY · 418 PAIRED DUBS

04.On-screen text is part of the video

Modern video leans on on-screen text: captions, jokes, memes, a single word timed to a beat. Each one is a visual event that helps the viewer follow or enjoy the video, and each one is in the source language. Adapting it has two requirements:

  • Never damage the video. Artifacts or visual distraction make the result worse than the original.
  • Never translate it literally. You have to understand why this text appears at this exact moment, in this exact place, before you can adapt it into another language.

Both are video-understanding requirements. Today Familiar translates titles, descriptions, and subtitles alongside the dub; replacing text inside the frame itself is where the model is headed, and the same rule will govern it: if the edit can be noticed, it is not good enough.

05.Lip-sync: succeed by being unnoticed

Lip-sync carries the same hard requirement: never damage the video. The highest compliment a dub's lip-sync can earn is the question "what changed?" -- the viewer should not notice it happened.

It reads as the least important of the four problems; in the finished experience it is close to the opposite. Humans listen with their eyes: the same audio is perceived as different sounds over different mouth movements -- the McGurk effect. The brain processes mouth movement unconsciously, and a mismatched mouth is one reason dubbed video feels uncomfortable in a way viewers can't name.

The effect also works in the dub's favor. Once the mouth matches, the brain resolves ambiguity toward agreement: pronunciation reads clearer, and the timbre reads more like the same person. Comfort nobody consciously notices, doing real work.

The hard parts sit upstream of moving any pixels: knowing who is speaking (off-screen voices, overlapping speakers, people who look or sound alike) and editing the mouth without breaking the face. Systems that chain separate models to find faces and guess the speaker hit a ceiling here: put two people on screen and the state of the art lip-syncs both mouths at once.

And lips are only where this starts. A person performs a sentence with the whole face -- eyes, brow, cheeks, head -- and with the body: gestures, nods, the half-beat of listening. Editing the whole performance so it carries the new language is what world translation is built for.

06.One problem, not four

Each requirement kept compressing to the same sentence. Natural translation: write what this video says, in the target language. Natural audio: generate the performance that matches it. On-screen text: know why it is there before touching it. Lip-sync: edit the person without breaking them. Understand the video, then generate it back.

A chain of single-purpose models can't do that, because understanding does not flow through a chain: the face finder doesn't know what was said, the voice model doesn't know who is on screen, and each stage's guess hardens into the next stage's input. One model that understands the scene -- who is speaking, on and off screen, what the text is doing, what the speaker means -- and generates from that shared understanding is the other path. That is the one we build: world translation, one unified model, with the results measured on the same clips and published at /benchmark.

REFERENCES

  1. [1]Familiar vs ElevenLabs benchmark: full results, intervals, and listening examples
  2. [2]AI vs Human Translation: Tested Against Professionals
  3. [3]AI Lip Sync: What It Is, What It Misses, What Beats It
  4. [4]Why the Dub Feels Off Abroad, and How Studios Fix It (the studio's version of this argument)
  5. [5]McGurk & MacDonald, “Hearing lips and seeing voices,” Nature 264 (1976)

Dubbing is finally good. See the measurements, then talk with the team about your catalog.

FAMILIAR · THE LAUNCH FILM · 2:31