ABSTRACT
01."Best" is measurable
Dubbing has six dimensions, and each can be scored on the same clips:
| DIMENSION | THE QUESTION | MEASURED AS |
|---|---|---|
| Spoken meaning | Did the dub say what you said? | Important translation errors per output |
| Voice cloning | Close your eyes: is this the same person? | Resemblance to the real speaker. YouTube auto-dubbing replaces the speaker's voice with a stock stranger voice |
| The face | Mouth, jaw, and expressions re-performed to match the new words? | The search term is AI lip sync; the lips are only the floor. Audio-only tools skip this dimension |
| Scene audio | Does the noise bed (music, ambience, sound effects) survive the dub? | Spectral error between the dub's background and the source's |
| Vocal events and timing | Do laughs, sighs, and beats land where and how they landed? | Event shape correlation; transient-timing F1 |
| Workflow and coverage | Does the tool accept your clip? | A rejected clip scores zero on that clip. Length minimums, batch caps, and language count are quality dimensions too |
Two tools that tie on quality do not tie on speed, and live is the hardest setting, since nothing can be fixed in post; at Familiar, live dubbing is experimental and coming soon. Dubbing vs subtitles and voice-over: the video translation guide.
02.The method that makes comparisons honest
- Paired and black-box. The same source clip and the same target language go to each system; only the final user-facing audio is evaluated.
- One approved meaning. Each dub is transcribed with word timestamps and reviewed against it, valid paraphrases accepted.
- Acoustic measures compare the dub to the source scene.
- Equal weight per source clip; 95% intervals resample source clips.
- A rejected clip is recorded as a rejection, and separate products are separate studies.
03.The current numbers
We ran this method against ElevenLabs Dubbing twice on August 5, 2026:
- Familiar sounds 27.8% more like the real speaker; all 11 languages favored Familiar.
- ElevenLabs made 95.5% more review-flagged mistakes.
- Predicted naturalness ran 10.2% higher for Familiar (exploratory).
ElevenLabs produced 261% more background spectral error.
12.27 VS 3.40 DB · DUBBING V1Familiar scored higher on laughter and reaction shape correlation.
0.759 VS 0.221 · DUBBING V1Familiar scored higher on sound-effect and beat timing F1.
0.947 VS 0.762 · DUBBING V1The pattern held on the newest product:
ElevenLabs produced 270% more background-sound error: the laughter, the music, the ambience.
11.92 VS 3.22 DB · DUBBING V2 (ALPHA)ElevenLabs produced 221% more laughter and reaction loudness error.
3.32 VS 1.04 DB · 18 CLIPS, 54 COMPARISONS · DUBBING V2 (ALPHA)Familiar scored higher on laughter and reaction shape correlation.
0.960 VS 0.749 · DUBBING V2 (ALPHA)ElevenLabs had 56.5% more flagged dubs overall.
36 VS 23ElevenLabs source-clip coverage; it rejected clips under 11 seconds.
FAMILIAR 38/38The shipped claim: "Sounds 27.8% more like you than ElevenLabs." Clip-level breakdown: the head-to-head paper.
Translation text has its own study (August 2026): our production pipeline against expert professional translators, scored blind against expert reference edits.
- 100.3% four-language average: Familiar's production translation matches the first-pass professional's quality.
- 5× cheaper than professional human translators at $22 to $45 per spoken minute (published per-word rates, translation alone).
04.What nobody does well yet
05.Run the test yourself
Running a bake-off for an ad campaign, a channel network, or a full catalog? You do not have to trust anyone's numbers, including ours. The one-clip test is below; for a season or catalog pilot, How to Run an AI Dubbing Pilot (2026); for a multi-speaker movie scene, What Is the Best AI Dubbing for Movies?
- Pick one clip of a speaker you know with a laugh, background music, and a fast aside. Those three things break dubs.
- Dub it in two tools: same clip, same target language.
- Check the words. Have a native speaker, or a back-translation, compare the dub against what was actually said.
- Close your eyes and ask who is speaking. If the answer is "a narrator," voice cloning failed.
- Listen for the room. Is the music still there? The ambience? The laugh, and does it still sound like the speaker?
- Watch the mouth. Does the lip-sync hold, and does the whole face match the new words?
- Full human dubbing, translation plus voice, runs $70 to $150 per finished minute; one Familiar render includes the translation, the speaker's voice, the scene audio, and the lip-sync.
- Start with the demos on the landing page.
- Test it on your own footage: book a call. Access is by Studio contract, sized to your catalog, with all 30 languages, any to any (pricing).
- Live streams: experimental and coming soon; designed to return each language on its own feed over SRT or RTMP, every speaker in their own voice. Set up with the team.
REFERENCES
Dubbing is finally good. See the measurements, then talk with the team about your catalog.
FAMILIAR · THE LAUNCH FILM · 2:31
