Memory#Extraction#File management
MarkItDown MCP: turn PDFs and Office files into Markdown
Microsoft's official MCP package for MarkItDown. One tool takes a file path or URL and returns Markdown, so agents can read PDFs, Word, slides and spreadsheets.
Project and installation docs
View projecthttps://github.com/microsoft/markitdown
Hand an agent a 40-page PDF contract or a spreadsheet full of formulas and it often stalls before the real work starts, because it can’t read the file. MarkItDown is the file-to-Markdown converter from Microsoft’s AutoGen team, and markitdown-mcp wraps it as an MCP server. The agent passes a URI and gets back clean Markdown with headings, tables and lists intact.
What it does
- One tool: it exposes only
convert_to_markdown(uri). The URI can behttp:,https:,file:ordata:, so local files and web pages both work. - Broad format support: it reuses MarkItDown’s converters, which cover PDF, Word, PowerPoint, Excel, HTML and other common office formats.
- Three transports: STDIO by default;
--httpserves Streamable HTTP and SSE, bound to127.0.0.1unless you change it. - Container-friendly: the package ships a Dockerfile. Mount the folder you want readable at
/workdirand nothing else on the host is visible.
Who it’s for
- Knowledge workers who ask Claude Desktop or Cursor to read contracts, earnings reports or meeting attachments.
- Developers building ingestion pipelines who want every file normalized to Markdown before chunking.
Setup
Needs Python. Install and run:
pip install markitdown-mcp
markitdown-mcp
For Claude Desktop the project recommends the Docker image (build it first with docker build -t markitdown-mcp:latest . in the package folder):
{
"mcpServers": {
"markitdown": {
"command": "docker",
"args": ["run", "--rm", "-i", "markitdown-mcp:latest"]
}
}
}
Our take
The MarkItDown repo had about 189k stars as of 2026-10-06, but those belong mostly to the CLI and Python library; the MCP server is a small package under packages/markitdown-mcp. It earns a place because getting files into model-readable text is step one of nearly every knowledge workflow, and this package reduces it to a single tool call with a tiny context footprint. The catch is security. The README says plainly that there’s no authentication and that convert_to_markdown can read any file the running user can, plus anything on the network. Don’t bind HTTP mode to a public interface, and use the container with a narrow mount for sensitive folders. Licensed MIT.