Audio#Text to speech#Translation

VideoLingo: one-click subtitle translation and dubbing

VideoLingo turns video translation into a one-click pipeline: word-level alignment, AI segmentation, glossary translation and dubbing with voice cloning.

Project facts

GitHub Ecosystem
Repositorygithub.com/Huanshere/VideoLingo
License
Apache-2.0
Language
Python
Stars
18,702
Data checked
2026-10-08

Snapshot figures reflect the check date and may change over time.

Turning an English video into a subtitled release by hand means transcribing, splitting lines, machine-translating, timing and dubbing — hours of work for ten minutes of footage. VideoLingo packages that chain into an open-source, one-click pipeline (the repo had 18,702 stars as of October 8, 2026, with v3.1.1 released on September 27): yt-dlp fetches the video, Qwen3-ASR runs word-level recognition and alignment, NLP and AI handle subtitle segmentation, a glossary keeps terminology consistent, and GPT-SoVITS, CosyVoice2 and other backends produce the dub, with voice cloning from a reference clip. The project brands itself “Netflix-level subtitles”, and the README states the caveat up front: translation quality depends on the source audio, the language pair and the models you pick.

Core features

  • Word-level alignment: Qwen3-ASR plus Qwen3-ForcedAligner run transcription and timing locally, so subtitle breaks follow words rather than a fixed character count.
  • AI subtitle translation and dubbing in three passes: direct translation → reflection → natural rewriting, backed by a custom glossary so names and terms stay consistent.
  • Dubbing with optional cloning: GPT-SoVITS, OpenAI, Edge TTS, Fish Audio, F5-TTS and CosyVoice2 backends; GPT-SoVITS and CosyVoice2 accept a reference audio clip.
  • Three outputs: subtitle files only, dual-subtitle video, or fully dubbed video, across eight input languages.
  • An API for agents: uv run start.py --api starts a local HTTP API, so scripts and agents can submit jobs, poll progress and fetch results in batch.

Typical use cases

  • Knowledge re-publishers adding a subtitle pass to technical talks from YouTube.
  • Teams going global producing multi-language dubs of their product videos.
  • Course archives translated in batch with one consistent glossary.

Quick start

On Windows, download the source zip from the release page and double-click OneKeyStart.bat — the first run installs uv, Python 3.12 and FFmpeg for you. Everywhere else:

git clone https://github.com/Huanshere/VideoLingo.git
cd VideoLingo
uv run start.py   # opens the Streamlit UI in your browser

Fill in the API URL, key and model in the sidebar, paste a YouTube link, and the first job runs end to end; Linux servers with an NVIDIA GPU can deploy the official Docker image.

Summary

For video re-publishers and teams that need batch dual subtitles, VideoLingo compresses hours into minutes; local Qwen3-ASR transcription is free, but the translation and dubbing stages still need your own LLM and TTS keys. Apache-2.0, Python 3.12. The README itself says translation quality depends on the source audio, language pair and chosen models — proofread anything important before release; Intel Mac dubbing drops the background audio by default. When the dubbed cut needs further editing, hand it to video-use and let an agent take over.