Memory#RAG

knowledge-rag: offline local document search for Claude Code

Drop PDFs, Markdown and code into a folder and query them with hybrid search and reranking. Embeddings run locally: no API keys, no external services.

Project and installation docs

View project

https://github.com/lyonzin/knowledge-rag

If you have a pile of internal docs, reports and code notes and want Claude Code to answer from them, the usual choices are a cloud knowledge base or a hand-built stack of LangChain, a vector store and an embedding service. knowledge-rag packs all of that into one Python package. Put files in documents/, restart your MCP client, and the agent can call search_knowledge; embedding and reranking happen inside the local process.

What it does

  • Hybrid search: semantic vectors plus BM25, then a cross-encoder rerank, so search_knowledge returns ordered snippets in one call.
  • Index management: add_document, update_document and remove_document maintain the index, add_from_url fetches and sanitizes a web page, and reindex_documents does incremental or full rebuilds.
  • Measure retrieval: evaluate_retrieval reports MRR@5, Recall@5 and Precision@5, so you can tell whether a tweak actually helped.
  • Formats and deployment: the README lists parsers for 35 file formats. Use stdio for one client, or Streamable HTTP to share one process across several, with optional bearer auth, rate limits and Prometheus metrics.

Who it’s for

  • Developers who want a coding agent to consult local technical docs or internal standards that can’t leave the machine.
  • Security researchers; the README’s own example indexes MITRE ATT&CK and threat reports.

Setup

Needs Python 3.11 or newer. Install, then initialize to create config.yaml and a documents/ folder:

pip install knowledge-rag
knowledge-rag init

For several clients, start it with knowledge-rag --transport streamable-http and point them at http://127.0.0.1:8179/mcp.

Our take

A good project few people know: about 290 stars as of 2026-10-06. The README is refreshingly careful. It says which features touch the network (the one-time ~200 MB model download, add_from_url), doesn’t claim compliance certifications, and links a reproducible audit report. Next to RAG frameworks that want Docker, Ollama and a separate vector database, it’s far easier to stand up for a personal knowledge base. Limits: it’s designed around a single writer, so a team needs to decide who owns the index; the write tools let an agent delete documents; and retrieval quality on non-English corpora is something to check yourself with evaluate_retrieval. Licensed MIT.