THE LAB

The Research Behind Familiar: The Published Papers

FAMILIAR RESEARCHALL ARTICLES

ABSTRACT

Familiar is built by a research lab. The team's published work spans 20 papers, 16 of them at ICLR, CVPR, ICCV, ECCV, NeurIPS, ICML, ACM Multimedia, and WACV, and it covers the pieces world translation needs: joint audio-visual generation, real-time human animation, fast video generation, editing inside diffusion models, and multimodal perception. Every paper is linked in full below; beyond this selection, the team's record since 2021 runs to over 10,000 citations and 100+ papers at ICLR, NeurIPS, CVPR, ECCV, and more.
20

published papers from the team, every one linked on this page

16

peer-reviewed at ICLR, CVPR, ICCV, ECCV, NeurIPS, ICML, ACM Multimedia, and WACV

20 FPS

real-time human animation, stable for an hour: SoulX-LiveAct, coauthored by Prof. Xu Zheng

01.Joint audio-visual generation

A dub holds up only if the voice and the face are generated together; two separate models can't hold one identity. This is the field the team started, and the reason Familiar is one model rather than a stitched pipeline.

First page of JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching
JAM-Flow: Joint Audio-Motion Synthesis with Flow MatchingECCV 2026 · KWON, SHIN, JEONG, PARK, UH

The first joint audio-visual generation model: speech and facial motion come out of one flow-matching model, so the voice and the face can never drift apart.

First page of Syncphony: Synchronized Audio-to-Video Generation with Diffusion Transformers
Syncphony: Synchronized Audio-to-Video Generation with Diffusion TransformersICLR 2026 · SONG, KWON, JEONG, UH

Audio-to-video generation that lands motion on the beat, with motion-aware training and Audio Sync Guidance, plus CycleSync, a new metric for measuring sync.

First page of TTS-CtrlNet: Time Varying Emotion Aligned Text-to-Speech Generation with ControlNet
TTS-CtrlNet: Time Varying Emotion Aligned Text-to-Speech Generation with ControlNetARXIV 2025 · JEONG, LEE, KWON, UH

Emotion control for text-to-speech that changes over time: a frozen speech model plus a trainable ControlNet copy, with zero-shot voice cloning kept intact.

02.Real-time human animation

Live dubbing means generating a person while they speak. Prof. Xu Zheng coauthored the current state-of-the-art in real-time human animation: the most realistic real-time humans today.

First page of SoulX-LiveAct: Towards Hour-Scale Real-Time Human Animation with Neighbor Forcing and ConvKV Memory
SoulX-LiveAct: Towards Hour-Scale Real-Time Human Animation with Neighbor Forcing and ConvKV MemoryACM MM 2026 · ORAL · ZHEN, ZHENG, ZHANG, JIANG, YAN, TAO, YIN

Real-time human animation at 20 frames per second that stays stable for an hour, using Neighbor Forcing and ConvKV memory to stop the drift that kills long generations.

03.Fast video generation

Real-time only exists if generation is fast. Five papers on making diffusion and flow models quicker without giving up fidelity.

First page of FlowBlending: Stage-Aware Multi-Model Sampling for Fast and High-Fidelity Video Generation
FlowBlending: Stage-Aware Multi-Model Sampling for Fast and High-Fidelity Video GenerationARXIV 2026 · SONG, KWON, JEONG, UH

Stage-aware sampling that swaps between large and small models across denoising stages: up to 1.65× faster video generation at large-model fidelity.

First page of TCFG: Tangential Damping Classifier-free Guidance
TCFG: Tangential Damping Classifier-free GuidanceCVPR 2025 · KWON, KIM, UH

Sharper classifier-free guidance: an SVD-based refinement of the unconditional score that steers sampling back toward the data manifold.

First page of Balanced Conic Rectified Flow
Balanced Conic Rectified FlowNEURIPS 2025 · KIM, KWON, UH

Straighter rectified flows trained with inversions of real data, so faster few-step sampling stays anchored to the true distribution.

First page of Plug-and-Play Diffusion Distillation
Plug-and-Play Diffusion DistillationCVPR 2024 · HSIAO, KHODADADEH, DUARTE, LIN, QU, KWON, KALAROT

Distills diffusion models down to 8 steps with an external guide: about 1% of parameters trained, no classifier-free guidance needed at inference.

First page of HARIVO: Harnessing Text-to-Image Models for Video Generation
HARIVO: Harnessing Text-to-Image Models for Video GenerationECCV 2024 · KWON ET AL., WITH ADOBE RESEARCH

Turns a frozen text-to-image model into a video generator with a motion module and new regularizers, without retraining the base model.

04.Editing inside diffusion models

Dubbing is an edit of a real performance, and this line of work goes back to the paper that found where diffusion models keep their meaning.

First page of Diffusion Models Already Have a Semantic Latent Space
Diffusion Models Already Have a Semantic Latent SpaceICLR 2023, NOTABLE TOP 25% · KWON, JEONG, UH

Found the semantic latent space hiding inside frozen diffusion models, h-space, and edited images through it with no training.

First page of Training-free Content Injection using h-space in Diffusion Models
Training-free Content Injection using h-space in Diffusion ModelsWACV 2024 · JEONG, KWON, UH

Training-free content injection through the same h-space: move the content of one image into another inside a frozen diffusion model.

First page of Understanding the Latent Space of Diffusion Models through the Lens of Riemannian Geometry
Understanding the Latent Space of Diffusion Models through the Lens of Riemannian GeometryNEURIPS 2023 · PARK, KWON, CHOI, JO, UH

Maps the geometry of the diffusion latent space with Riemannian pullback metrics, showing when and where edits actually work across timesteps.

First page of StyleKeeper: Prevent Content Leakage using Negative Visual Query Guidance
StyleKeeper: Prevent Content Leakage using Negative Visual Query GuidanceICCV 2025 · JEONG, KIM, LEE, CHOI, UH

Style without content leakage: negative visual query guidance keeps a reference image's style and leaves its objects behind.

First page of Visual Style Prompting with Swapping Self-Attention
Visual Style Prompting with Swapping Self-AttentionCVPR 2024 WORKSHOP · BEST PAPER AWARD · JEONG, KIM, CHOI, LEE, UH

Training-free style control: swap a reference image's self-attention into the generation and the output follows its style.

First page of TextAway: Mask-Free Video Text Removal
TextAway: Mask-Free Video Text RemovalPROJECT PAGE · SONG, GO, GO, UH

Removes on-screen text from video and restores what was behind it, end to end, with no OCR and no masks at inference.

05.Multimodal perception

A unified model has to read the scene before it can edit one: who is talking, from visual and audio together. Prof. Xu Zheng's line of work.

First page of EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging Benchmark
EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging BenchmarkICLR 2026 · ZHANG, FU, YANG, MIAO, QIAN, ZHENG, ET AL.

The first benchmark for egocentric vision at night: day-night aligned videos that expose where models fail in the dark.

First page of RealRAG: Retrieval-Augmented Realistic Image Generation via Self-Reflective Contrastive Learning
RealRAG: Retrieval-Augmented Realistic Image Generation via Self-Reflective Contrastive LearningICML 2025 · LYU, ZHENG, JIANG, YAN, ZOU, ZHOU, ZHANG, HU

Retrieval-augmented image generation: real photos pulled in as references through self-reflective contrastive learning, so unfamiliar objects render true to life.

First page of Omnidirectional Spatial Modeling from Correlated Panoramas
Omnidirectional Spatial Modeling from Correlated PanoramasACM MM ASIA 2025 · ZHANG, FU, ZHENG

Presents Pano-R1, a panoramic reasoning model trained with reinforcement learning on CFpano, the first benchmark for question answering across correlated 360° views.

First page of Understanding-in-Generation: Reinforcing Generative Capability of Unified Model via Infusing Understanding into Generation
Understanding-in-Generation: Reinforcing Generative Capability of Unified Model via Infusing Understanding into GenerationARXIV 2025 · LYU, WONG, ZHENG, ET AL.

Puts a unified model's understanding to work during generation: the model reasons about the prompt first, then generates with that reasoning infused.

First page of UniBind: LLM-Augmented Unified and Balanced Representation Space to Bind Them All
UniBind: LLM-Augmented Unified and Balanced Representation Space to Bind Them AllCVPR 2024 · LYU, ZHENG, ZHOU, WANG

One balanced embedding space that binds seven modalities, image, text, audio, video and more, aligned to LLM-augmented class centers.

06.What it adds up to

Generate voice and face together, in real time, fast enough for live video, with editing that respects the person and perception that reads the scene. That stack is Familiar Alpha today, and the measured results are published at /benchmark.

REFERENCES

  1. [1]JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching (ECCV 2026)
  2. [2]Syncphony: Synchronized Audio-to-Video Generation with Diffusion Transformers (ICLR 2026)
  3. [3]TTS-CtrlNet: Time Varying Emotion Aligned Text-to-Speech Generation with ControlNet (arXiv 2025)
  4. [4]SoulX-LiveAct: Towards Hour-Scale Real-Time Human Animation (ACM MM 2026, Oral)
  5. [5]FlowBlending: Stage-Aware Multi-Model Sampling for Fast and High-Fidelity Video Generation (arXiv 2026)
  6. [6]TCFG: Tangential Damping Classifier-free Guidance (CVPR 2025)
  7. [7]Balanced Conic Rectified Flow (NeurIPS 2025)
  8. [8]Plug-and-Play Diffusion Distillation (CVPR 2024)
  9. [9]HARIVO: Harnessing Text-to-Image Models for Video Generation (ECCV 2024)
  10. [10]Diffusion Models Already Have a Semantic Latent Space (ICLR 2023, notable top 25%)
  11. [11]StyleKeeper: Prevent Content Leakage using Negative Visual Query Guidance (ICCV 2025)
  12. [12]Visual Style Prompting with Swapping Self-Attention (CVPR 2024 workshop, Best Paper Award)
  13. [13]Training-free Content Injection using h-space in Diffusion Models (WACV 2024)
  14. [14]Understanding the Latent Space of Diffusion Models through the Lens of Riemannian Geometry (NeurIPS 2023)
  15. [15]TextAway: Mask-Free Video Text Removal (project page)
  16. [16]EgoNight: Towards Egocentric Vision Understanding at Night (ICLR 2026)
  17. [17]RealRAG: Retrieval-Augmented Realistic Image Generation (ICML 2025)
  18. [18]Omnidirectional Spatial Modeling from Correlated Panoramas (ACM MM Asia 2025)
  19. [19]Understanding-in-Generation (arXiv 2025)
  20. [20]UniBind: LLM-Augmented Unified and Balanced Representation Space to Bind Them All (CVPR 2024)

Dubbing is finally good. See the measurements, then talk with the team about your catalog.

FAMILIAR · THE LAUNCH FILM · 2:31