Audio#Text to speech#Translation
VideoLingo: one-click subtitle translation and dubbing
VideoLingo turns video translation into a one-click pipeline: word-level alignment, AI segmentation, glossary translation and dubbing with voice cloning.
Project facts
GitHub Ecosystem- License
- Apache-2.0
- Language
- Python
- Stars
- 18,702
- Data checked
- 2026-10-08
Snapshot figures reflect the check date and may change over time.
Turning an English video into a subtitled release by hand means transcribing, splitting lines, machine-translating, timing and dubbing — hours of work for ten minutes of footage. VideoLingo packages that chain into an open-source, one-click pipeline (the repo had 18,702 stars as of October 8, 2026, with v3.1.1 released on September 27): yt-dlp fetches the video, Qwen3-ASR runs word-level recognition and alignment, NLP and AI handle subtitle segmentation, a glossary keeps terminology consistent, and GPT-SoVITS, CosyVoice2 and other backends produce the dub, with voice cloning from a reference clip. The project brands itself “Netflix-level subtitles”, and the README states the caveat up front: translation quality depends on the source audio, the language pair and the models you pick.
Core features
- Word-level alignment: Qwen3-ASR plus Qwen3-ForcedAligner run transcription and timing locally, so subtitle breaks follow words rather than a fixed character count.
- AI subtitle translation and dubbing in three passes: direct translation → reflection → natural rewriting, backed by a custom glossary so names and terms stay consistent.
- Dubbing with optional cloning: GPT-SoVITS, OpenAI, Edge TTS, Fish Audio, F5-TTS and CosyVoice2 backends; GPT-SoVITS and CosyVoice2 accept a reference audio clip.
- Three outputs: subtitle files only, dual-subtitle video, or fully dubbed video, across eight input languages.
- An API for agents:
uv run start.py --apistarts a local HTTP API, so scripts and agents can submit jobs, poll progress and fetch results in batch.
Typical use cases
- Knowledge re-publishers adding a subtitle pass to technical talks from YouTube.
- Teams going global producing multi-language dubs of their product videos.
- Course archives translated in batch with one consistent glossary.
Quick start
On Windows, download the source zip from the release page and double-click OneKeyStart.bat — the first run installs uv, Python 3.12 and FFmpeg for you. Everywhere else:
git clone https://github.com/Huanshere/VideoLingo.git
cd VideoLingo
uv run start.py # opens the Streamlit UI in your browser
Fill in the API URL, key and model in the sidebar, paste a YouTube link, and the first job runs end to end; Linux servers with an NVIDIA GPU can deploy the official Docker image.
Summary
For video re-publishers and teams that need batch dual subtitles, VideoLingo compresses hours into minutes; local Qwen3-ASR transcription is free, but the translation and dubbing stages still need your own LLM and TTS keys. Apache-2.0, Python 3.12. The README itself says translation quality depends on the source audio, language pair and chosen models — proofread anything important before release; Intel Mac dubbing drops the background audio by default. When the dubbed cut needs further editing, hand it to video-use and let an agent take over.