Audio#Open source#Agents#Self-hosted#Text to speech#Voice cloning#Dubbing#CLI#macOS#Windows

yovoice: an open-source voice studio that never leaves your machine

The usual way to narrate a video is pasting a script into a cloud TTS: per-character pricing, generic voices, and your unpublished text on someone else's server. yovoice runs 7 local TTS models through audio.cpp on your own hardware, clones a voice from a 1-60 second clip, and ships a full record-trim-export workflow plus a CLI and an Agent Skill. Apache-2.0, for macOS and Windows.

Project facts

GitHub Ecosystem
Repositorygithub.com/leemysw/yovoice
License
Apache-2.0
Language
TypeScript
Stars
194
Data checked
2026-09-27

Snapshot figures reflect the check date and may change over time.

Narrating a video usually means pasting your script into a cloud TTS service: per-character fees, a voice that sounds like everyone else’s, and unpublished text sitting on someone else’s servers. yovoice takes the other path — an Apache-2.0 licensed tool that runs 7 TTS models locally through the audio.cpp engine, so voiceover and audio content never leave your machine. The WeChat Channels account 对齐观察 featured it on September 27, 2026. The repo is young — created September 15, 2026 — and has collected nearly 200 stars in 12 days (194 as of 2026-09-27), at version 0.1.4.

Core features

  • Voice cloning: a 1-60 second reference clip is enough to copy a timbre; IndexTTS 2.0/2.5 add emotion control and “match the reference performance”, and 2.5 can edit pronunciation directly.
  • Design a voice with words: VoxCPM2 and Qwen3-TTS VoiceDesign do text-guided voice design — describe a delivery (low, slow) and get a new voice, no reference audio needed.
  • 7 local models: IndexTTS 2.0/2.5, VoxCPM2, OmniVoice, and Qwen3-TTS Base/CustomVoice/VoiceDesign, all inferred locally via audio.cpp, with resumable model downloads and GGUF import.
  • A complete audio workflow: import or record reference audio, trim, preview, export — all in the app. Projects, voices and settings live in ~/.yovoice and survive uninstalling.
  • CLI + Agent Skill: hand the repo’s skills/yovoice link to Claude Code or Codex, then just say “read narration.txt using voice.wav as the reference voice, calm delivery, save as narration.wav” — no desktop app required.
  • Hardware acceleration on both platforms: Metal on Apple Silicon (macOS 14+); CPU, NVIDIA CUDA, and experimental Vulkan on Windows 10/11 x64.

Typical use cases

  • Video and podcast creators: generate your own narration, regenerate per script edit, with no per-character bill and no upload of unpublished material.
  • Courses and audiobooks: keep one voice across episodes, generate long scripts in batches, export everything from History.
  • Developers: add local speech to an app, or have a coding agent batch-produce voice assets through the CLI.

Quick start

Grab the package for your platform from the repo’s Releases tab: a .dmg for macOS 14+ (Apple Silicon) or a -setup.exe for Windows 10/11 x64 (it installs WebView2 if needed). First step after launching: download your chosen model in Settings. Then type text, pick an expression mode, generate, preview, export.

To run from source:

make install
make app-run

Summary

yovoice fits creators and developers who narrate often and care about cost and script privacy; if you produce a voiceover twice a year and aren’t picky about voices, the model downloads alone will feel heavy. Three caveats: the OmniVoice model weights are CC-BY-NC, non-commercial use only; models are separate downloads and not small — pick them to match your disk; and with a project two weeks old at 0.1.4, expect to file the occasional GitHub issue. Written in TypeScript, Apache-2.0 licensed.