Workplace#Text to speech#Speech to text
joinly: send an agent into Zoom, Teams and Google Meet calls
Self-hosted meeting middleware. Agents join browser video calls over MCP, read the live transcript and speak or chat, with your choice of LLM and voice.
Project and installation docs
View projecthttps://github.com/joinly-ai/joinly
Most meeting tools do one thing: hand you a transcript afterwards. joinly wants the agent in the room. Its MCP server provides tools to join, listen, speak and look at the shared screen, so an agent can answer questions during the call and use other MCP servers to open a GitHub issue or edit a Notion page as it goes. Everything ships in one Docker image you host yourself.
What it does
- Meeting control:
join_meetingtakes a link, display name and optional passcode;leave_meetingexits;mute_yourselfandunmute_yourselfhandle the mic. - Listen and look:
get_transcriptreturns the transcript, optionally for the last N minutes, and the subscribabletranscript://liveresource streams new utterances with speakers.get_participantsandget_chat_historycover attendees and chat, andget_video_snapshotgrabs the current frame, such as a screen share. - Speak:
speak_texttalks via TTS andsend_chat_messageposts in the meeting chat, with built-in handling for interruptions and multiple speakers. - Swappable parts: Whisper or Deepgram for speech-to-text, Kokoro, ElevenLabs or Deepgram for speech, and any major LLM API or local Ollama.
Who it’s for
- Developers prototyping meeting assistants that look things up or capture action items live.
- Teams that want meeting transcription without sending audio to a third-party SaaS.
Setup
Needs Docker and an LLM API key in a .env file. The image is about 2.3 GB because it bundles a browser and models. Run it as an MCP server bound to localhost:
docker pull ghcr.io/joinly-ai/joinly:latest
docker run -p 127.0.0.1:8000:8000 ghcr.io/joinly-ai/joinly:latest
Then connect with uvx joinly-client --env-file .env <MeetingUrl> or your own MCP client.
Our take
About 570 stars as of 2026-10-06, and the most complete tool design among meeting MCP servers: it reads the live transcript, talks back, and brings other MCP tools into the call. Against hosted meeting-bot services, the pitch is full self-hosting and free choice of models. The costs are real: a large image, decent hardware for local transcription, and a last push in early September 2026. On security, the README warns that the MCP server has no authentication and accepts client-supplied configuration, so bind it to localhost for a single trusted client. And tell attendees, and get their consent, before a bot records and transcribes a meeting. Licensed MIT.