ABSTRACT
published papers from the team, every one linked on this page
peer-reviewed at ICLR, CVPR, ICCV, ECCV, NeurIPS, ICML, ACM Multimedia, and WACV
real-time human animation, stable for an hour: SoulX-LiveAct, coauthored by Prof. Xu Zheng
01.Joint audio-visual generation
A dub holds up only if the voice and the face are generated together; two separate models can't hold one identity. This is the field the team started, and the reason Familiar is one model rather than a stitched pipeline.

The first joint audio-visual generation model: speech and facial motion come out of one flow-matching model, so the voice and the face can never drift apart.

Audio-to-video generation that lands motion on the beat, with motion-aware training and Audio Sync Guidance, plus CycleSync, a new metric for measuring sync.

Emotion control for text-to-speech that changes over time: a frozen speech model plus a trainable ControlNet copy, with zero-shot voice cloning kept intact.
02.Real-time human animation
Live dubbing means generating a person while they speak. Prof. Xu Zheng coauthored the current state-of-the-art in real-time human animation: the most realistic real-time humans today.

Real-time human animation at 20 frames per second that stays stable for an hour, using Neighbor Forcing and ConvKV memory to stop the drift that kills long generations.
03.Fast video generation
Real-time only exists if generation is fast. Five papers on making diffusion and flow models quicker without giving up fidelity.

Stage-aware sampling that swaps between large and small models across denoising stages: up to 1.65× faster video generation at large-model fidelity.

Sharper classifier-free guidance: an SVD-based refinement of the unconditional score that steers sampling back toward the data manifold.

Straighter rectified flows trained with inversions of real data, so faster few-step sampling stays anchored to the true distribution.

Distills diffusion models down to 8 steps with an external guide: about 1% of parameters trained, no classifier-free guidance needed at inference.

Turns a frozen text-to-image model into a video generator with a motion module and new regularizers, without retraining the base model.
04.Editing inside diffusion models
Dubbing is an edit of a real performance, and this line of work goes back to the paper that found where diffusion models keep their meaning.

Found the semantic latent space hiding inside frozen diffusion models, h-space, and edited images through it with no training.

Training-free content injection through the same h-space: move the content of one image into another inside a frozen diffusion model.

Maps the geometry of the diffusion latent space with Riemannian pullback metrics, showing when and where edits actually work across timesteps.

Style without content leakage: negative visual query guidance keeps a reference image's style and leaves its objects behind.

Training-free style control: swap a reference image's self-attention into the generation and the output follows its style.

Removes on-screen text from video and restores what was behind it, end to end, with no OCR and no masks at inference.
05.Multimodal perception
A unified model has to read the scene before it can edit one: who is talking, from visual and audio together. Prof. Xu Zheng's line of work.

The first benchmark for egocentric vision at night: day-night aligned videos that expose where models fail in the dark.

Retrieval-augmented image generation: real photos pulled in as references through self-reflective contrastive learning, so unfamiliar objects render true to life.

Presents Pano-R1, a panoramic reasoning model trained with reinforcement learning on CFpano, the first benchmark for question answering across correlated 360° views.

Puts a unified model's understanding to work during generation: the model reasons about the prompt first, then generates with that reasoning infused.

One balanced embedding space that binds seven modalities, image, text, audio, video and more, aligned to LLM-augmented class centers.
06.What it adds up to
Generate voice and face together, in real time, fast enough for live video, with editing that respects the person and perception that reads the scene. That stack is Familiar Alpha today, and the measured results are published at /benchmark.
REFERENCES
- [1]JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching (ECCV 2026)
- [2]Syncphony: Synchronized Audio-to-Video Generation with Diffusion Transformers (ICLR 2026)
- [3]TTS-CtrlNet: Time Varying Emotion Aligned Text-to-Speech Generation with ControlNet (arXiv 2025)
- [4]SoulX-LiveAct: Towards Hour-Scale Real-Time Human Animation (ACM MM 2026, Oral)
- [5]FlowBlending: Stage-Aware Multi-Model Sampling for Fast and High-Fidelity Video Generation (arXiv 2026)
- [6]TCFG: Tangential Damping Classifier-free Guidance (CVPR 2025)
- [7]Balanced Conic Rectified Flow (NeurIPS 2025)
- [8]Plug-and-Play Diffusion Distillation (CVPR 2024)
- [9]HARIVO: Harnessing Text-to-Image Models for Video Generation (ECCV 2024)
- [10]Diffusion Models Already Have a Semantic Latent Space (ICLR 2023, notable top 25%)
- [11]StyleKeeper: Prevent Content Leakage using Negative Visual Query Guidance (ICCV 2025)
- [12]Visual Style Prompting with Swapping Self-Attention (CVPR 2024 workshop, Best Paper Award)
- [13]Training-free Content Injection using h-space in Diffusion Models (WACV 2024)
- [14]Understanding the Latent Space of Diffusion Models through the Lens of Riemannian Geometry (NeurIPS 2023)
- [15]TextAway: Mask-Free Video Text Removal (project page)
- [16]EgoNight: Towards Egocentric Vision Understanding at Night (ICLR 2026)
- [17]RealRAG: Retrieval-Augmented Realistic Image Generation (ICML 2025)
- [18]Omnidirectional Spatial Modeling from Correlated Panoramas (ACM MM Asia 2025)
- [19]Understanding-in-Generation (arXiv 2025)
- [20]UniBind: LLM-Augmented Unified and Balanced Representation Space to Bind Them All (CVPR 2024)
Dubbing is finally good. See the measurements, then talk with the team about your catalog.
FAMILIAR · THE LAUNCH FILM · 2:31
