Audio#Voice cloning#Video generation

LiveTalking: a streaming digital human that talks back live

LiveTalking by lipku is a real-time streaming digital human: synced audio-video, interruptible speech, multi-session, low-latency WebRTC, Apache-2.0.

Project facts

GitHub Ecosystem
Repositorygithub.com/lipku/LiveTalking
License
Apache-2.0
Language
Python
Stars
9,662
Data checked
2026-10-01

Snapshot figures reflect the check date and may change over time.

For always-on live streams and avatar customer service, realism matters less than latency: when a viewer says something, the avatar has to answer now. LiveTalking by lipku is a real-time interactive streaming digital-human engine with synchronized audio-video dialogue; the README states it already sees wide commercial use. It has been updated continuously since December 2023 — the latest commit, September 13, 2026, fixed a real leak: disconnected sessions previously dropped only a dict reference while render, inference, and TTS threads kept spinning, about 3 CPU cores per leaked session.

Core features

  • Swappable models: one framework drives ernerf, musetalk, wav2lip, and Ultralight-Digital-Human — pick per quality-versus-compute budget.
  • Interruptible speech: users can cut in mid-sentence; flush_talk clears the pending speech queue instead of making you wait.
  • Three output paths: WebRTC for low-latency interaction, RTMP push into a live room, or a virtual camera feeding OBS.
  • Motion choreography: custom videos play during idle stretches so the stream never freezes.
  • Multi-session and custom avatars: several sessions in parallel, with your own footage as the avatar.
  • Pluggable LLM and TTS: the LLM engine speaks to OpenAI-compatible backends including DashScope and OrcaRouter; the TTS side supports voice cloning.

Typical use cases

  • 24/7 unmanned livestreams: an LLM drafts the pitch, the avatar speaks it live, choreography fills the gaps.
  • Knowledge-base customer service: users ask by voice, the avatar answers in real time, mistakes get interrupted and re-answered.
  • E-learning and lobby displays: batch-generate lessons through the API, or drive an interactive presenter via the /human endpoint.

Quick start

git clone https://github.com/lipku/LiveTalking.git
conda create -n livetalking python=3.12
conda activate livetalking
pip install torch==2.9.1 torchvision==0.24.1 torchaudio==2.9.1 --index-url https://download.pytorch.org/whl/cu128
pip install -r requirements.txt
python app.py

Per the README, this was tested on Ubuntu 22.04 with Python 3.12, PyTorch 2.9.1, and CUDA 12.8; model weights download separately.

Summary

LiveTalking fits live, interactive scenarios; if you only need finished clips without latency constraints, an offline model like LongCat-Video-Avatar 1.5 renders better quality. Apache-2.0, written in Python, about 9.7k stars as of 2026-10-01. Three caveats: the environment is heavy and CUDA versions must line up; lip quality depends on the model you pick — wav2lip is fast but coarse, musetalk finer but hungrier; and a commercial edition exists, so the open-source path leans on issues for support.