Media#Text to speech#Speech to text
VoiceMode: talk to Claude Code out loud
Adds two-way voice conversations to Claude Code and other agents. Runs offline with local Whisper and Kokoro, or falls back to OpenAI cloud speech services.
Project and installation docs
View projecthttps://github.com/mbailey/voicemode
Working with a coding agent mostly means typing at a screen. Sometimes your hands or eyes are busy: you’re cooking, walking to a meeting, or just done staring at a monitor. VoiceMode gives agents like Claude Code a voice channel. You speak, it listens, then answers out loud, back and forth like a phone call. The repo had about 1.4k stars as of 2026-10-06.
What it does
- Conversation: the core
conversetool handles speak-then-listen turns, stops recording when you stop talking, and keeps latency low enough to feel natural. - Local or cloud: install Whisper.cpp for speech-to-text and Kokoro for text-to-speech to run fully offline, or add an OpenAI API key as a cloud fallback. Both speak the same API, so switching needs no config change.
- Service control: the
servicetool starts and stops the local voice services. - Debugging:
VOICEMODE_SAVE_AUDIO=truesaves recordings locally so you can track down recognition problems.
Who it’s for
- Developers in long Claude Code sessions who want to keep a task moving while they step away from the keyboard.
- People with tired eyes or who find typing hard and prefer to dictate.
- Anyone who wants offline voice interaction without sending audio to the cloud.
Setup
You need a microphone and speakers. For Claude Code, the plugin is the recommended route; after installing, run /voicemode:install for dependencies and local voice services, then /voicemode:converse to start talking:
claude plugin marketplace add mbailey/voicemode
claude plugin install voicemode@voicemode
Or use the Python installer: run uvx voice-mode-install, then claude mcp add --scope user voicemode -- uvx --refresh --from voice-mode voicemode-mcp-launcher.
Our take
Most voice servers are one-way, turning text into an audio file. VoiceMode does real two-way conversation, and the local Whisper plus Kokoro pairing lets it run fully offline, which privacy-minded users will like. Caveats: it needs microphone access, Linux and WSL need system packages such as PortAudio and PulseAudio first, so the first setup can take some fiddling. Local recognition quality depends on your hardware, and the README says nothing about languages other than English. The OpenAI path costs API fees and sends audio to the cloud. Runs on Linux, macOS, Windows and NixOS. MIT-licensed.