THE SHORT ANSWER
THE SHORT VERSION
- Three faults make a dub feel off: a stranger's voice, a mouth still speaking the source language, and a scene flattened under the new dialogue.
- On the same clips, ElevenLabs Dubbing v2 (Alpha) had 270% more background-sound error: the laughter, the music, the ambience (111 paired outputs, August 5, 2026).
- Familiar sounds 27.8% more like the real speaker than ElevenLabs Dubbing v1, closer in all 11 languages (418 paired outputs, August 5, 2026).
- The studio's own translators edit any translated line before it renders, and each language renders on its own approval.
01.Why does the dub feel off?
Two of the faults are older than AI: a voice actor is also a stranger's voice, and no recording session moves the actor's mouth. Audio-only AI dubbing can add a third, a flattened scene, measured below.
- A stranger's voice. Whether a voice actor or a stock voice reads the line, someone who knows the actor no longer recognizes them (AI Video Dubbing, Explained (2026)).
- A mouth speaking the wrong language. The brain combines lip movements and sound to understand speech; when they mismatch, comprehension becomes harder, or the viewer perceives a different sound entirely: the McGurk effect, part of audiovisual speech integration. A voice-only dub leaves that mismatch in every frame.
- A flattened scene. The laughter, the music, the sound effects, and the room ambience get stripped or smeared under the new speech.
- A sentence-for-sentence translation. A line mapped word for word reads foreign even when correct, and jokes land late or not at all (What Matters in Video Translation).
02.What do the measurements say?
ElevenLabs Dubbing v2 (Alpha) had 270% more background-sound error: the laughter, the music, the ambience.
11.92 VS 3.22 DB · ELEVENLABS DUBBING V2 (ALPHA) · 111 PAIRED OUTPUTS · MANDARIN, SPANISH, JAPANESE · AUGUST 5, 2026Familiar sounds 27.8% more like the real speaker than ElevenLabs Dubbing v1; all 11 languages favored Familiar.
0.4605 VS 0.3604 · ELEVENLABS DUBBING V1 · 418 PAIRED OUTPUTS · 11 LANGUAGES · AUGUST 5, 2026of the first-pass expert professional's quality, averaged across French, Chinese, Hindi, and Indonesian, scored blind against expert reference edits. A match.
PRODUCTION TRANSLATION STUDY VS FIRST-PASS EXPERT PROFESSIONAL TRANSLATORS · AUGUST 2026The two ElevenLabs cohorts and the human-translation study are never pooled. Full data, intervals, and listening examples: Familiar vs ElevenLabs and vs human translators; method in AI vs Human Translation: Tested Against Professionals (2026).
- The words, ElevenLabs Dubbing v2 (Alpha). Familiar made 54.7% fewer important spoken translation errors (29 vs 64).
- Spoken-output mistakes, ElevenLabs Dubbing v1. Familiar made 48.8% fewer review-flagged spoken-output mistakes (132 vs 258).
03.What changes when voice, lips, and scene come from one model?
One unified model generates the voice, the lips, and the scene together, with identity permanence and scene & world understanding, so the same actor's performance is preserved instead of a mismatched voice over a mismatched face.
- The voice comes from the actor's own performance. Dubbing is not text-to-speech: it takes the original audio in and preserves the speaker's identity and delivery.
- The face re-performed -- eyes, brow, cheeks, head -- not just lips. The whole face performs the new language, so the mouth viewers read matches the words they hear (AI Lip Sync: What It Is, What It Misses, What Beats It (2026)).
- The scene survives the dub. The laughter, the music and bass, the sound effects, the room ambience; spans the model detects as music pass through untouched.
- Everyone on camera gets dubbed, each in their own voice; the three-speaker movie scene on /demos is dubbed English to Spanish. Mics over the mouth, hands across the face: the model understands occlusions.
- Names, places, and catchphrases stay exactly as said through a Do Not Translate list the studio controls; jokes are re-landed in the target language.
- One render carries the translation, the actor's voice, the scene audio, and the lip-sync.
04.How does a studio run it?
- Review before render, per language. The pipeline pauses at a review step where the studio's own translators edit any translated line; each language renders on its own approval, Spanish today, the rest when their reviewers sign off (Transcript & review).
- A one-sentence context brief. Who is on screen and what the show is, set at the creator scope for a whole series or at the job scope for one episode; the narrower scope wins (Context).
- A Do Not Translate list at the same three scopes, so a character's name or catchphrase holds in episode 60 as in episode 1 (Do Not Translate).
- Webhooks. dub.ready_for_review when the translation is ready, dub.output_ready per finished language (Webhooks).
- 30 languages, any to any, in three published quality tiers, best first (Languages).
- The reference price. Full human dubbing, translation plus voice, runs $70 to $150 per finished minute; access to Familiar is by Studio contract, sized to the catalog.
| Deliverable | Format | Note |
|---|---|---|
| Finished video | mp4 with the full audio mix | Ready to post |
| Subtitles | .srt and .vtt | From the reviewed transcript; not sold alone |
| Audio on its own | The full mix, or the voice alone | For the studio's own mix |
| Title and description | Translated per language | Links, @handles, #hashtags, and chapters survive untouched |
05.What should a studio test before committing a season?
- Pick the hardest episode, one with fast dialogue, a joke, an emotional turn, proper names, and music under speech; a clean scene proves nothing about a season.
- Write the acceptance criteria before anyone listens. Native reviewers score the words, the voice, the scene, and the face against them and mark the moment immersion broke; the scorecard is in How to Run an AI Dubbing Pilot (2026).
- Include a late edit. A changed master is a new dub request; the review step is where the changed lines are checked, and the Do Not Translate list and context carry over.
QUESTIONS
Why do international viewers say the dub feels off?
Three things came apart: the voice belongs to a stranger, the mouth still speaks the original language, and the music and room under the dialogue were flattened. Viewers rarely name the cause; a mouth that disagrees with the sound makes speech harder to understand.
Is it the translation or the voice?
Both, and they compound: a sentence-for-sentence translation reads foreign even when correct, and a replaced voice removes the actor. On paired clips Familiar sounded 27.8% more like the real speaker than ElevenLabs Dubbing v1; in a separate study its translations matched first-pass professional translators at 100.3%.
Can our own translators still review the lines?
Yes. The pipeline pauses at a review step where any translated line is edited before it renders, and each language renders when its reviewer signs off: Spanish can ship while Japanese is still under review.
Do character names and catchphrases survive across episodes?
Yes. A Do Not Translate list holds names, places, and catchphrases exactly as said in the dub, the subtitles, and the translated title and description. Set at the creator scope it covers every episode of a series; a job-scope entry overrides for one upload.
What do we receive per language?
The finished video with the full audio mix, subtitles in .srt and .vtt from the reviewed transcript, the translated title and description, and the audio on its own: the full mix, or the voice alone for the studio's own mix.
REFERENCES
- [1]Familiar vs ElevenLabs benchmark: full results, intervals, and listening examples
- [2]Familiar vs human translators: the translation-quality study
- [3]Transcript & review: the review step and per-language render (API documentation)
- [4]What Matters in Video Translation (the four problems)
- [5]McGurk & MacDonald, “Hearing lips and seeing voices,” Nature 264 (1976)
Dubbing is finally good. See the measurements, then talk with the team about your catalog.
FAMILIAR · THE LAUNCH FILM · 2:31
