How is this different from ElevenLabs or HeyGen?+
We're the first to do voice + lip-sync dubbing in one model; stitching ElevenLabs and HeyGen together still sounds and looks like a stranger. And it's more than AI lip sync: the whole face performs the new language.
In our published benchmark:
- their newest product had 270% more background sound error: the laughter, the music, the ambience
- 54.7% fewer important translation errors than their newest product
- our voice sounded 27.8% more like the real speaker on their Dubbing v1, closer in all 11 languages
On the same clips, ElevenLabs:
- dropped or invented words 45 times
- broke laughs and sound effects 8 times
- lost the speaker's voice 6 times
- collided overlapping speakers 6 times
- flattened the delivery 4 times
- stripped the scene ambience 4 times
- refuses clips under 11 seconds (the API)
Stitching it yourself runs $11 a minute: $8 for Sync.so lip-sync plus $3 for ElevenLabs audio, two tools and two bills. Familiar renders voice, lip-sync, and translation in one pass.
What happens to my music and sound effects?+
They stay. The laughter, the music and bass, the sound effects, the ambience of the city, train, or party: the whole scene survives, even dense polyphony. Measured on the same clips, ElevenLabs produced 270% more background sound error. YouTube auto dubbing: 232.6% more. Hear it at
thefamiliarlab.com/benchmark.
How good are the translations themselves?+
We tested our production pipeline against expert professional translators, scored blind against expert reference edits: averaged across French, Chinese, Hindi, and Indonesian, our translations scored 100.3% of the first pass professional's quality. A match. The price is not: professional translation alone runs $22 to $45 per spoken minute, and full human dubbing, translation plus voice, $70 to $150 per finished minute.
Which languages are supported?+
30 languages, any to any, in three quality tiers (best first).
Tier 1: Cantonese, English, French, German, Indonesian, Italian, Japanese, Korean, Mandarin, Portuguese, Russian, Spanish, Thai, Vietnamese.
Tier 2: Bulgarian, Catalan, Croatian, Dutch, Greek, Lithuanian, Norwegian, Slovak, Swedish.
Tier 3: Arabic, Belarusian, Danish, Kazakh, Latvian, Serbian, Ukrainian.
What does Familiar do?+
Dubbing is finally good: world translation -- one unified model translates your voice and lip-sync together. Your videos and livestreams in every language; still you.
What was broken about dubbing?+
Until now a dub meant a stranger's voice and a mouth out of sync, because voice and lip-sync came from separate tools. One model now generates both together, so the dub stays you. Lip-sync + voice translation. Any video or livestream.
Did you actually test this against ElevenLabs?+
Yes, twice, and we published everything: the same clips went to both systems and only the finished audio was compared. 418 paired outputs across 11 languages against their Dubbing v1, then 111 against their newest Dubbing v2 (Alpha). We made 54.7% fewer important translation errors than their newest product, which also had 270% more background sound error; on Dubbing v1 our voice measured closer to the real speaker in all 11 languages. Full data and listening examples:
thefamiliarlab.com/benchmark.
How does Familiar compare with ElevenLabs' $3 a minute?+
Their checkout quoted $3 per dubbed minute, audio only. One Familiar render includes the translation, your voice, the scene audio, and the lip-sync, with no 11 second minimum, no 20 item batches, no waiting in waves: upload once and every language renders. Access is by Studio contract, sized to your catalog.
What kinds of videos work best?+
Clean, stable footage: podcasts and interviews with 2 or 3 people on camera, straight to camera videos, lectures, sermons, presentations. Everyone on camera gets dubbed. The same shape works for medical-tourism clinics, tour operators, law firms, and insurers: consults, tours, client explainers. For livestreams we dub from your clean camera feed, so quality stays practically untouched. Fast cuts, memes over faces, and shaky footage are harder for our current Alpha; the next model, built to understand the full scene, is in development.
Can you dub my livestream?+
Yes, and we set it up with you. We dub the whole stream, everyone on camera in their own voice, and push each language's feed to its own channel, seconds behind.
Book a call with us and we get your stream running.
Will it keep my catchphrases and names?+
Exactly as you say them. You keep a Do Not Translate list: catchphrases, friends' names, signature exclamations. Jokes get rewritten to land in the target language, never translated word for word.
Who's building this?+
Familiar is a research lab building world translation: one unified model with scene & world understanding, who is talking from visual and audio together, and people: emotions, how humans talk to one another, listen, react. The soul of a research lab. The heart of a product team. Our team of 7 full time (9 including advisor and part time) brings together PhDs, professors, and researchers from Stanford, Yonsei's Visual Intelligence Lab (Korea's top visual AI lab), and HKUST. We started the field of joint audio visual human models and coauthored the current state of the art in real time human animation, all while operating international channels for Veritasium, 3Blue1Brown, and Welch Labs. The published work, by area. Joint audio visual generation: JAM-Flow (ECCV 2026), Syncphony (ICLR 2026), and TTS-CtrlNet. Real time human animation: SoulX-LiveAct (ACM MM 2026, Oral). Fast video generation: FlowBlending, TCFG (CVPR 2025), Balanced Conic Rectified Flow (NeurIPS 2025), Plug-and-Play Diffusion Distillation (CVPR 2024), and HARIVO (ECCV 2024). Editing inside diffusion models: TextAway, StyleKeeper (ICCV 2025), Training-free Content Injection (WACV 2024), Visual Style Prompting (CVPR 2024 workshop, Best Paper Award), Riemannian latent geometry (NeurIPS 2023), and Asyrp (ICLR 2023). Multimodal perception: EgoNight (ICLR 2026), RealRAG (ICML 2025), Pano-R1 (ACM MM Asia 2025), UiG, and UniBind (CVPR 2024). Since 2021, over 10,000 citations and 100+ papers at ICLR, NeurIPS, CVPR, ECCV, and more. Prof. Xu Zheng, PhD from HKUST and now an assistant professor, coauthored SoulX-LiveAct, this year's leading real time human animation model. In his words: “My doctoral research develops robust and interpretable multi-modal learning algorithms spanning perception, understanding, reasoning, and generation.”
What are you building next?+
Our next version is full world translation: one unified model with scene & world understanding and permanence of identities. It reads the whole scene, even a dozen people speaking on and off screen, game windows, memes, visual effects, and knows who is talking from visual and audio together. Dubbing becomes a physical edit: the model predicts how you would have said it in the new language and renders the performance to match, not just lips but the whole face, with the full body as the goal.
Can I fix the translation?+
For videos, always: the pipeline pauses at a review screen where you edit any line before it renders; credits are only spent then. Live dubs translate on the fly, so there's no review step; your Do Not Translate list still applies.
What does it cost?+
Plans are by Studio contract, sized to your catalog:
book a call with us. For reference, full human dubbing, translation plus voice, runs $70 to $150 per finished minute.
Who owns the dub? What do you do with my video?+
The dub is yours to post anywhere. Billions of people are cut off by language; your dubs help fix that.
Will my audience accept AI dubbing?+
Audiences reject AI that fakes people and accept AI that serves them. A Familiar dub is your own voice and your own face, made with your consent; nothing fabricated, nobody replaced. Fans who want your tutorials, lessons, and streams finally get them in their language. Innovate with AI; stay loved by your audience.
Is this deepfakes?+
No. We only translate you, with your consent.
Why does Familiar exist?+
To break every barrier to understanding. Hundreds of millions can't understand their own doctor, their lawyer, their kid's teacher; language is where we start. We translate people, never fabricate them.
What files come with a dub?+
Each language arrives as a finished video with the full audio mix. You also get subtitles in .srt and .vtt from the reviewed transcript, the translated title and description, and the audio by itself: the full mix, or the voice alone for your own mix.
Can we send you a broadcast over SRT or RTMP?+
Yes. Live sessions take your feed in over SRT, RTMP, or WebRTC and send each language's dub back over any of them.
Book a call with us and we plug into your stack.
How fast is live dubbing, and which languages work live?+
Live dubbing runs 5 to 9 seconds behind the original over SRT or RTMP, every speaker in their own voice: translation takes 2 to 6 seconds depending on the quality setting, the voice about 1 second, lip-sync about 2 seconds more. Live dubs from 11 source languages today (English, Spanish, German, French, Portuguese, Italian, Dutch, Vietnamese, Arabic, Japanese, Mandarin) into all 29 others; uploaded videos dub from all 30.
Who is Familiar for?+
Movies & TV, micro-dramas, broadcasts, live shopping, corporate & education, ads, creators & podcasts, and the platforms they run on -- in every language.
Why does lip-sync matter?+
Your brain combines lip movements and sound to understand speech; when they mismatch, comprehension becomes harder, or you may perceive a different sound entirely. Scientists call it the McGurk effect, part of audiovisual speech integration. A voice-only dub leaves that mismatch in every frame; Familiar renders voice and lip-sync together, so the mouth viewers read matches the words they hear.