THE SHORT ANSWER
THE SHORT VERSION
- A pilot is paired: the same source clips go through the current workflow and the candidate, in two or three target languages.
- Acceptance criteria are written before anyone listens: one pass mark per dimension, six measured dimensions plus two operational rows.
- Native reviewers score blind where the format allows it and mark the second immersion broke.
- The review step counts corrections: how many translated lines the reviewer changed before render, and of which class.
- Two operational metrics ride beside quality: revision cycles per episode, and the time from a source change to an approved replacement.
01.What is a pilot for?
A demo shows the vendor's best clip. A pilot shows the buyer's hardest footage through two workflows.
- The unit is a decision. The pilot answers whether a season, a catalog, or a creator network can move; one clean scene says nothing about episode sixty.
- The pass mark comes first. Criteria fixed before anyone listens cannot drift toward whichever dub the room happened to like.
- The design is published. Familiar's paired studies at /benchmark use it; the pilot re-runs it on the buyer's own clips.
02.Which footage goes in?
- A connected run, for a series. Two or three consecutive episodes with a recurring lead, so voice, names, and honorifics are tested for continuity across the run.
- The scenes that break dubs. Fast dialogue, a joke, an emotional turn, proper names, music under speech, and two or three people on camera talking over one another.
- One late edit. A changed line or a re-cut scene submitted after the first dub; a changed source is a new dub request, checked at the review step.
- Two or three priority languages, from the three published quality tiers at /docs/languages, best tier first.
03.What do the acceptance criteria say?
One question a native reviewer answers on every clip, one pass mark per dimension, and the metric Familiar's benchmark uses for the same thing.
| DIMENSION | THE REVIEWER'S QUESTION | FAMILIAR'S BENCHMARK MEASURES |
|---|---|---|
| Spoken meaning | Does the line say what was said? A native speaker or a back-translation checks it. | Important translation errors per output, against one approved target meaning |
| Voice cloning | Eyes closed: is this the same person? | Speaker resemblance to the real speaker |
| Scene audio | Are the music, the laughter, and the room still there under the new speech? | Background-sound error against the source scene, in dB |
| The face | Do the mouth and the whole face match the new words, on every speaker on camera? | Judged on the paired listening examples; the audio measures do not score it, and audio-only tools skip the dimension |
| Vocal events and timing | Do laughs, sighs, and reactions land where and how they landed? | Event shape correlation; transient-timing F1 |
| Coverage | Did the tool accept every clip, at every length, in every language asked for? | A rejected clip scores zero on that clip; length minimums and batch caps count against it |
| Revision cycles | How many passes before each language was approved? | Not in the benchmark; counted per episode in the pilot |
| Time to an approved replacement | How long from the late edit to an approved replacement dub? | Not in the benchmark; timed once, on the late edit |
- A pass mark is a number. Zero translation errors on names and numbers; the speaker recognized with eyes closed by three of four reviewers; no scene where the music drops out; one revision cycle per episode.
04.How is it run?
- Paired and black-box. Each source clip and target language goes through both workflows, and only the finished dub is judged, the way the published method judges only the final audio.
- Blind where the format allows it. Preference scores drift toward whatever sounds familiar; the record that holds is the second immersion broke, and what broke it.
- A target-language panel beside the in-house reviewers. Outsiders catch slang, register, and a late joke; the in-house side catches a name or a house term.
- The review step is the correction counter. Reviewers stay in the loop because the pipeline pauses for them: any translated line is edited at the review step before render, and the number and class of edited lines is a quality measure.
- Tag every changed line. The classes show what a Do Not Translate list and a one-sentence context brief would have prevented on the second run.
05.How are the results read?
A gain counts when it holds across footage types and languages; a musical scene or a fast-cut sequence may need its own threshold.
- Report footage types as separate rows. Clean single-camera scenes and fast-cut scenes averaged together hide the row that decides the catalog.
- Report per language. A gain in two of three languages is a two-language contract until the third is re-tested.
- Count the operational rows. Revision cycles per episode and time from a source change to an approved replacement decide whether a daily release schedule holds.
- The published numbers are the reference; the pilot is the proof. Familiar's results on the same dimensions:
ElevenLabs had 270% more background-sound error: the laughter, the music, the ambience.
ELEVENLABS DUBBING V2 (ALPHA) · 111 PAIRED OUTPUTS · MANDARIN, SPANISH, JAPANESE · 11.92 VS 3.22 DB · AUGUST 5, 2026Familiar sounds 27.8% more like the real speaker; all 11 languages favored Familiar.
ELEVENLABS DUBBING V1 · 418 PAIRED OUTPUTS · 11 LANGUAGES · 0.4605 VS 0.3604 · AUGUST 5, 2026Familiar made 54.7% fewer important spoken translation errors (29 vs 64).
ELEVENLABS DUBBING V2 (ALPHA) · 111 PAIRED OUTPUTS · SAME COHORT AS THE 270% · AUGUST 5, 2026Familiar's production translation matches the first-pass expert professional's quality, averaged across French, Chinese, Hindi, and Indonesian.
PRODUCTION TRANSLATION STUDY VS FIRST-PASS EXPERT PROFESSIONAL TRANSLATORS · FOUR LANGUAGES · AUGUST 202606.From pilot to contract
Access is by Studio contract, sized to the catalog, with 30 day money back on signed annual contracts; the pilot's languages become the contract's first languages.
- The corrections carry over. Names and phrases fixed in the pilot persist in the Do Not Translate list and context brief at the creator scope; the first production episode starts from them.
- Live sessions are enabled with the contract: languages, expected concurrency, and a rehearsal on the buyer's real setup over SRT, RTMP, or WHIP; broadcasts then run 5 to 9 seconds behind the original, every speaker in their own voice.
- The sibling playbooks pick up from here: studios, streaming platforms, micro-dramas, multi-country launches, and creator networks.
QUESTIONS
How long should an AI dubbing pilot take?
Long enough to dub a connected run in two or three languages, review it with native speakers, and process one late edit. Two or three consecutive episodes with a recurring lead test continuity as well as quality.
What should a dubbing pilot measure?
Six dimensions, each scored by a native reviewer against criteria written before anyone listens: spoken meaning, voice, scene audio, the face, timing, and coverage. Two operational rows ride beside them: revision cycles per episode and time from a source change to an approved replacement.
Should reviewers know which dub is which?
Blind where the format allows it. Preference scores drift toward whatever sounds familiar, so the useful record is the moment immersion broke: the timestamp, and whether a joke, a name, a voice, or two people talking at once broke it.
Can the review step be part of the pilot?
Yes. Familiar's pipeline pauses at a review step where any translated line is edited before render; the number and kind of lines the reviewer changes is a quality measure, and the corrections ship in the finished dub. Live dubs translate on the fly, with no review step.
What does Familiar publish for the same dimensions?
Paired studies, August 2026. ElevenLabs Dubbing v2 (Alpha) had 270% more background-sound error: the laughter, the music, the ambience. Familiar sounded 27.8% more like the real speaker than ElevenLabs Dubbing v1, closer in all 11 languages. Familiar's translations matched the first-pass expert professional at 100.3% across French, Chinese, Hindi, and Indonesian.
REFERENCES
- [1]Dubbing benchmarks hub: the three paired studies
- [2]The Best AI Dubbing in the World: How to Measure It (the six dimensions)
- [3]Familiar vs ElevenLabs Dubbing: Two Paired Studies (the method)
- [4]AI vs Human Translation: Tested Against Professionals (2026)
- [5]API reference: transcript and review (the review step)
Dubbing is finally good. See the measurements, then talk with the team about your catalog.
FAMILIAR · THE LAUNCH FILM · 2:31
