Hindsight: long-term agent memory that learns from experience
Vectorize's open-source agent memory service. Retain, recall and reflect turn raw history into consolidated beliefs, with an MCP endpoint per bank.
Project and installation docs
View projecthttps://github.com/vectorize-io/hindsight
Most memory systems put chat history in a vector store and pull back whatever looks similar. The agent remembers what happened but never draws conclusions from it. Hindsight aims at that second half. It routes memories into world facts and the agent’s own experiences, then consolidates many of them into evidence-backed observations and mental models, so the agent understands your situation better over time. Every Hindsight server ships an MCP endpoint, enabled by default.
What it does
- Three operations:
retainstores new information,recallretrieves memories relevant to a question, andreflectanswers using those memories. MCP exposes all three as tools. - Separate banks: memories live in banks, and each bank has its own MCP URL,
http://localhost:8888/mcp/{bank_id}/, so you can split by user or project. - Any model:
HINDSIGHT_API_LLM_PROVIDERselects hosted providers such as OpenAI, Anthropic or Gemini, or local ones like Ollama and LM Studio. - Coding-agent package: a separate installer builds a per-repo bank from git history and past sessions for Claude Code, Codex CLI, Cursor CLI and others.
Who it’s for
- Developers building support, coaching or assistant agents that need to remember users across sessions and adapt.
- Engineering leads who want several coding agents to share one project memory.
Setup
Needs Docker and a model API key. The API runs on port 8888 and the UI on 9999:
export OPENAI_API_KEY=sk-xxx
docker run -it --pull always --name hindsight --restart unless-stopped -p 8888:8888 -p 9999:9999 \
-e HINDSIGHT_API_LLM_API_KEY=$OPENAI_API_KEY \
-v hindsight-data:/home/hindsight/.pg0 \
ghcr.io/vectorize-io/hindsight:latest
Then point your MCP client at http://localhost:8888/mcp/{bank_id}/.
Our take
About 46k stars as of 2026-10-06, and the heaviest memory option in this category. Where the official Memory server keeps a small hand-curated graph, Hindsight lets the model decide what to keep and how to generalize, which suits large, long-lived memories. The project reports LongMemEval results it says were independently reproduced; its public benchmark page has the details. The cost is operational. You run a service backed by PostgreSQL, and every retain and reflect call hits an LLM, which means token spend. If you’d rather not host it, Hindsight Cloud bills by usage. Licensed MIT.