Databases#RAG#Agent memory

Qdrant MCP: a vector database as semantic agent memory

Qdrant's official MCP server: two tools, store and find, with text embedded by FastEmbed. Works with a local file or a Qdrant server, for memory or snippets.

Project and installation docs

View project

https://github.com/qdrant/mcp-server-qdrant

If you want an agent to remember project conventions or reusable snippets and find them by meaning rather than keywords, you need a vector store. Qdrant MCP, Qdrant’s official server, narrows the database down to two actions: store a piece of information, and find related information by meaning. Embedding happens inside the server with FastEmbed, so there’s no separate embedding API to call. It had about 1.5k stars as of 2026-10-06.

What it does

  • Store: qdrant-store saves text with optional JSON metadata and creates the collection if it doesn’t exist; setting a default COLLECTION_NAME drops the per-call collection argument from both tools.
  • Find: qdrant-find searches by meaning; QDRANT_SEARCH_LIMIT caps results at 10 by default, and each hit comes back as its own message so the model can cite results individually.
  • Rewritable tool descriptions: TOOL_STORE_DESCRIPTION and TOOL_FIND_DESCRIPTION change how the tools are described, and the README turns it into a code-snippet library for Cursor this way, telling the model to put the actual code in metadata.code.
  • Local or remote: QDRANT_LOCAL_PATH uses a local file, QDRANT_URL connects to self-hosted Qdrant or Qdrant Cloud, over stdio, SSE or streamable HTTP; SSE listens on port 8000 by default, overridable with FASTMCP_SERVER_PORT.
  • More than one way to run it: besides uvx, there’s an official Dockerfile, a Smithery one-command install, and one-click VS Code badges. For debugging, fastmcp dev src/mcp_server_qdrant/server.py opens the MCP Inspector to watch each call’s input and output.

Who it’s for

  • Solo developers adding a cross-session, meaning-based memory to Claude or Cursor without standing up a separate retrieval service.
  • Teams already using Qdrant for RAG who want the agent to read and write the same collection, with no extra API layer in between.

Setup

Needs uv. Local file mode needs no Qdrant server:

{
  "qdrant": {
    "command": "uvx",
    "args": ["mcp-server-qdrant"],
    "env": {
      "QDRANT_LOCAL_PATH": "/path/to/qdrant/database",
      "COLLECTION_NAME": "your-collection-name",
      "EMBEDDING_MODEL": "sentence-transformers/all-MiniLM-L6-v2"
    }
  }
}

To connect to Qdrant Cloud or a self-hosted server instead, swap in QDRANT_URL and QDRANT_API_KEY; nothing else changes.

Our take

There are several vector-database servers; this one wins for official upkeep and a tiny interface, and local file mode gets you testing in minutes without standing up a Qdrant deployment or getting a separate embedding API key. That simplicity also means it isn’t a full Qdrant admin tool: no deletes, filters or collection management, and only FastEmbed models for now. The default model is English-centric, so pick a multilingual one via EMBEDDING_MODEL before storing other languages. If what you actually need is structured SQL rather than semantic memory, MCP Toolbox for Databases fits better; if you’re plugging into an existing RAG pipeline, check that its vector dimensions match this server’s embedding model first. Set QDRANT_READ_ONLY=true to drop the write tool, and leave FASTMCP_SERVER_HOST at its default 127.0.0.1 unless you mean to expose it. Licensed Apache-2.0.