TEAM

PhDs, professors and researchers

From the world's top universities in audio-visual human research.

  • Stanford
  • Yonsei
  • HKUST
  • University of Waterloo
THE TEAM

We exist to break every barrier to understanding.

An Zhu Liu
An Zhu Liu
Jibin Song
Jibin Song
Mingi Kwon
Mingi Kwon
Google Scholar ↗
Jaeseok Jeong
Jaeseok Jeong
Google Scholar ↗
Prof. Xu Zheng
Prof. Xu Zheng
Prof. Youngjung Uh
Prof. Youngjung Uh
Google Scholar ↗
RESEARCH

We're building world translation: not just voice and lips -- the scene, the sound, the human.

Joint audio-visual generation

The voice and the face generated by one model, so identity can never split: the field we started with JAM-Flow.

Real-time human animation

The most realistic real-time humans today: SoulX-LiveAct generates a person at 20 frames per second, stable for an hour.

Fast video generation

Distillation, guidance, and sampling: the research that makes diffusion models fast enough for real time.

Editing inside diffusion models

Editing inside frozen diffusion models, from finding h-space to controlling style and removing on-screen text.

Multimodal perception

Reading the scene before editing it: who is talking, from visual and audio together, in any environment.

Since 2021, over 10,000 citations and 100+ papers at ICLR, NeurIPS, CVPR, ECCV, and more.