ABSTRACT
01.Method
- Paired and black-box: the same English source clip and target language went to each system; only the final user-facing audio was evaluated.
- Meaning review: every dub transcribed with word timestamps and reviewed against the same approved target meaning.
- Acoustic measures compare the dub to the source scene: the noise bed, vocal-event envelopes, and transient timing.
- Familiar's output also carries the re-performed face; every measure below is audio against audio.
- Statistics: source clips carry equal weight, 95% intervals resample source clips, and the two cohorts are never pooled.
02.Results: ElevenLabs Dubbing v2 (Alpha) (their newest product)
ElevenLabs produced 270% more background-sound error: the laughter, the music, the ambience.
11.92 VS 3.22 dB MAEElevenLabs produced 221% more laughter and reaction loudness error.
3.32 VS 1.04 dB · 18 CLIPS / 54 COMPARISONSFamiliar's laughter and reaction shape correlation: 0.960 vs 0.749. The paired 95% interval excludes zero.
- Coverage: Familiar completed 38/38 source clips; Dubbing v2 (Alpha) rejected the 10.94-second clip under its 11-second minimum (37/38).
- Speaker resemblance: inconclusive at this sample size; a +3.5% point estimate for Familiar with a paired interval crossing zero.
03.Results: ElevenLabs Dubbing v1 (the stable product)
ElevenLabs had 95.5% more review-flagged spoken-output mistakes.
258 VS 132Familiar's laughter and reaction shape correlation: 0.759 vs 0.221, verified event regions.
Predicted naturalness for Familiar.
2.51 VS 2.28 · EXPLORATORY- Transient timing (sound effects and beats) also favored Familiar: F1 0.947 vs 0.762.
- The same output ships the re-performed face: AI lip sync and everything past it.
04.Workflow limits we hit along the way
Minimum clip length. The 10.94-second clip in our set could not be dubbed at all.
Items per batch. Four videos into six languages is already 24 outputs, more than one batch holds.
Concurrent requests. Ten videos into ten languages is 100 requests: at least four waves of batching and waiting.
- For a back catalogue, that is hours of babysitting per job.
- Familiar accepts any clip length; one upload fans out to every selected language.
05.Price
One Familiar render, per finished minute per target language: translation, the speaker's voice, music/effects/ambience preserved, and the lip-sync.
ACCESS BY STUDIO CONTRACTElevenLabs Dubbing, per dubbed minute per target language. Audio only.
CHECKOUT QUOTE · AUG 5, 2026And $3 buys audio only. To match what one Familiar render includes you still stitch on lip-sync, and the best standalone lip-sync tool, Sync.so, runs $8 a minute: $11 a minute combined. For scale, full human dubbing, translation plus voice, runs $70 to $150 per finished minute.
The full price study, with sources: How Much Does Dubbing Cost in 2026?
06.Translation text vs professional translators
A separate August 2026 study put our production translation pipeline against expert professional translators, scored blind against expert reference edits: across French, Chinese, Hindi, and Indonesian, Familiar matches their first pass.
Average of the first-pass professional's translation quality.
FRENCH, CHINESE, HINDI, INDONESIAN (FIRST PASS FOR EXPERT PROFESSIONAL TRANSLATORS)Cheaper than professional translation at $22–45 per spoken minute (published per-word rates, translation alone).
$0.15–0.30/WORD × ~150 WORDS/MIN07.Limitations
- The meaning review is model-assisted and single-pass; independent native double review is pending.
- Dubbing v2 (Alpha) speaker resemblance and naturalness were inconclusive at this sample size; the v1 identity result is not pooled with them.
- Curated real-world clips, not a probability sample of all video. Product limits and prices were observed on August 5, 2026 and may change.
REFERENCES
Dubbing is finally good. See the measurements, then talk with the team about your catalog.
FAMILIAR · THE LAUNCH FILM · 2:31
