LongCat-Video-Avatar 1.5: talking-head video from one photo
Meituan's LongCat team open-sourced a model that turns one photo and one audio track into a talking video — Whisper lip sync, 8-step distillation, MIT.
Project facts
GitHub Ecosystem- License
- MIT
- Language
- Python
- Stars
- 8,446
- Data checked
- 2026-10-01
Snapshot figures reflect the check date and may change over time.
The hard part of a talking-head video is never the script; it’s appearing on camera — lighting, shooting, editing, half a day gone. LongCat-Video-Avatar from Meituan’s LongCat team compresses all of that into two inputs: one photo, one audio track, out comes a lip-synced talking video. Version 1.5 shipped on May 21, 2026 with code and weights fully open-sourced. A Chinese short-video channel billed it as “digital humans at commodity prices,” and its official evaluation scenarios — news broadcasting, knowledge courses, commercial promotion — match exactly what those channels demonstrate.
Core features
- Photo plus audio to video: no avatar training or fine-tuning; native tasks include Audio-Text-to-Video and Audio-Text-Image-to-Video.
- Whisper-based lip sync: v1.5 swaps the Wav2Vec2 audio encoder for Whisper-Large, which the team credits for smoother lip dynamics.
- Long-video stability: official claims include full-body temporal stability and identity consistency on long generations, a classic failure mode of avatar models.
- Stylized domains: handles anime characters, animals, multi-person interactions, and object handling — not just realistic headshots.
- Single- and multi-stream audio: drives one speaker or several characters from multi-track audio, for dialogue content.
- 8-step inference: DMD2-based step distillation cuts inference to 8 NFE, pulling cost per clip down with it.
Typical use cases
- News-style updates and courses: fix one presenter image, feed the day’s script, batch-produce talking videos.
- E-commerce explainers: product imagery plus a voiceover track becomes a virtual host clip without a shoot.
- Anime and virtual-IP content: stylized characters speak too, so a dubbing session turns straight into footage.
Quick start
Code lives in meituan-longcat/LongCat-Video, weights on Hugging Face:
git clone --single-branch --branch main https://github.com/meituan-longcat/LongCat-Video
cd LongCat-Video
pip install -r requirements.txt
The README also asks for torch 2.6.0 (CUDA 12.4) and flash_attn 2.7.4.post1. To skip the environment entirely, community ComfyUI nodes (such as rookiestar28/ComfyUI-LongCat-Avatar) and RunPod images already exist.
Summary
LongCat-Video-Avatar suits teams producing talking-head content in volume without appearing on camera, and developers studying audio-driven video; if you want “upload a photo, get a clip in minutes,” a hosted service like HeyGen stays simpler. The code is MIT-licensed, about 8.4k stars as of 2026-10-01. Caveats: local inference wants serious VRAM and a CUDA setup, and rendering is offline — it cannot do live streaming. For real-time interaction, see LiveTalking.