Stop re‑teaching
your agent.
Make your agent's memory work like a brain. Hippo is long-term memory for coding agents. It's a critical layer for your AI harness that connects across your different tools (Cursor, Claude Code, Codex). It keeps your proprietary data completely local, and it actually learns over time. By strengthening memories each time they are recalled, Hippo preserves what works, lets mistakes decay, and continuously compounds your agents' intelligence.
npm install -g hippo-memory - Claude Code
- Codex
- Cursor
- OpenClaw
- OpenCode
- Pi
- any MCP client
- 98.0% R@5 on LongMemEval-S with a free local embedder (best of five settings).
- R@1 0.41 to 0.62 (the Jev reranker eval, opens in new tab) with the opt-in Jev reranker, against the free local cross-encoder. Ranking only: no answer-rate win was shown.
772 stars 39.4k downloads compared in the agent-memory-atlas
an illustrated session · the hippo lines are real hippo output
-
hippo capture-errorReal failures become lessons. Interrupts, declined permissions and empty searches are skipped.
-
hippo outcome --badMark a lesson wrong and it stops coming back.
-
hippo doctorOne command checks the install and names the fix for anything missing.
The problem
Most AI memory saves everything and searches later.
That's storage with semantic search bolted on. It's why your agent kept hitting the same deploy bug last week. And the week before.
The system saw the failure four times. It had no way to know it should remember.
How it works
Memories decay. Retrieval makes them stronger.
The layers borrow from the brain as design inspiration. On our tests, decay made no measurable difference to recall and sleep cost a little. Outcome marks and retrieval strengthening are what measured helpful; supersession is not measured yet.
New information lands here. Session-only, no decay.
Timestamped, decays by default. Retrieval strengthens it; errors stick.
Repeated episodes compress into stable patterns. The originals decay.
weak traces are forgotten; repeated ones consolidate during sleep
strength = decay over time, re-strengthened on every recall
Decay by default
Every memory fades on a one-year half-life unless it is used. We did not tune 365 days: it tied with 730 days and with decay off.
Retrieval strengthens
Use it or lose it. Each recall extends the half-life. Memories you reach for survive.
Errors stick
Tag a failure once. It decays slower and resurfaces every time you walk back into that code.
Sleep consolidates
On hippo sleep, three or more related episodes merge into one semantic pattern. The originals decay; the pattern survives. It keeps the store tidy; it has not been shown to improve recall.
Get started
Zero config. It wires itself in.
Install it, run init in your repo, and hippo detects your agent framework and patches the right config file. Next session, your agent just uses it.
Package installation alone does not enable automatic preservation on every agent. Complete the documented setup and required host trust; capture and compaction coverage depend on the integration.
Detected and patched automatically
- Claude Code
- Codex
- Cursor
- OpenClaw
- OpenCode
- Pi
In a repo, init patches the instruction file each agent already has (CLAUDE.md, AGENTS.md) and adds session hooks where the agent supports them. hippo init --scan ~ gives every git repo under your home folder a store and installs the Claude Code hooks and the OpenCode plugin, but patches no instruction files. hippo init --no-hooks --no-schedule skips the hooks and the daily run.
- 1 $
npm install -g hippo-memory - 2 $
hippo init
Then hippo sleep tidies the store when a Claude Code
session ends, and a daily run at 6:15am learns from the day's commits in every registered project.
Local-first
Your memory lives on your machine.
Proven by a globalThis.fetch spy that throws on call, across the 1000-event ingestion smoke. Not a hardcoded zero. The default recall path makes no network call either; opt-in features such as the Jev reranker, the LLM reranker and the API embedders do.
Memories live in a local .hippo/ store with markdown mirrors you can read, grep, and commit. No cloud and no account. One default to know: hippo sleep sends text to Anthropic for fact extraction when ANTHROPIC_API_KEY is set, and one config line turns that off.
Right-to-be-forgotten is a single API call. Every row carries kind, scope, owner, and provenance.
Multi-tenant keys are scrypt-hashed with an audit log on every mutation. Tenant A cannot see tenant B, proven by a negative test.
And it's not locked to one tool.
Your ChatGPT memories don't travel to Claude; your .cursorrules don't travel to Codex. Hippo is one store behind all of them.
Receipts
Numbers, including the losses.
every claim links to its source
Compare
Learn what is wrong. Stop repeating it.
How hippo compares with nine other memory tools on the features that define a memory lifecycle.
| Feature | Hippo | MemPalace | Mem0 | Basic Memory | gbrain | Zep | Letta | Cognee | Memoria | EverMind |
|---|---|---|---|---|---|---|---|---|---|---|
| Decay by default | Yes | No | No | No | No | No | No | No | No | No |
| Retrieval strengthening | Yes | No | No | No | No | No | No | Partial | No | Partial |
| Cross-tool import (ChatGPT/Claude/Cursor) | Yes | No | No | No | Partial | ? | No | Partial | No | Partial |
| Auto-hook install | Yes | No | No | No | No | No | No | No | No | No |
| Zero runtime deps | Yes | No | No | No | No | No | No | No | Yes | No |
| LongMemEval (best published) | 98.0% any / 88.5% all R@5* | 96.6% raw / 100% reranked R@5 | 94.4** | N/A | 95.53% all R@5* | 90.2% accuracy** | N/A | N/A | 88.78% accuracy** | 83.00% accuracy** |
scroll for more tools →
* LongMemEval retrieval recall at 5. "Any" counts a hit when one answer session is in the top 5, "all" only when every one is. gbrain leads on "all".
** Answer accuracy with a reader model, a different metric from recall.
Most of the others save everything and search it later, or build a knowledge graph. Hippo learns what is wrong and stops repeating it.
Is this just RAG?
No. RAG searches a fixed corpus. Hippo's store changes as your agent works: a memory marked wrong drops out of the top results, a newer fact supersedes the old one, and memories that keep getting recalled last longer while unused ones fade on a half-life. Recall itself is search: BM25, plus embeddings if you install them.
Which agents does hippo work with?
hippo init detects Claude Code, Codex, Cursor, OpenClaw, OpenCode and Pi, and wires itself into each one's instruction file, hooks or plugin. It only patches instruction files that already exist. Any MCP client can use the MCP server, and other tools can call the CLI or the HTTP API that hippo serve starts.
Where does hippo keep my data?
On your machine, in SQLite: .hippo/hippo.db in each project, plus a global store in ~/.hippo/ for lessons shared across projects, with markdown mirrors you can read and commit. Recall makes no network call by default. Text goes to an outside provider only through features that use one: an API embedder, the Jev or LLM reranker, hippo refine, and the fact extraction hippo sleep runs through Anthropic's API whenever ANTHROPIC_API_KEY is set in its environment. To turn that last one off, set {"extraction":{"enabled":false}} in .hippo/config.json.
How is hippo different from mem0?
mem0 uses a language model to extract memories, OpenAI by default in its open-source library, and memories stored through its hosted MCP server live in your Mem0 account (mem0 docs (opens in new tab), checked 2026-09-28). Hippo stores memories in SQLite on your machine, needs no account and no model, and hippo init wires it into the coding agents it finds. mem0's platform and hippo both mark an older fact superseded when a newer one replaces it. Hippo also lets you mark a recalled memory wrong with hippo outcome --bad, and it drops out of the top results.
What happens when a memory turns out to be wrong?
Mark it, and it drops out of the top results. hippo outcome --bad weakens the memories from the last recall, hippo supersede <id> "<new fact>" replaces one with a newer version, and hippo reject <id> --reason "<why>" stops that value from returning at all. On the synthetic E1 test, where every mark is correct, plain BM25 plus the outcome mark cut how often a marked-bad memory stayed in the top five from 71.9% to 0.0%. Real marks are noisier, because --bad marks the whole recall batch.
Has hippo been shown to make agents better at their work?
Not yet. The published numbers measure retrieval: whether the right memory comes back, and whether a memory marked wrong stays out of the results. A paired test that runs real agent sessions with and without hippo is under way. Every measurement, including failed runs and one retracted claim, is indexed in docs/evals (opens in new tab).
One command. Every repo gets memory.
Zero config. SQLite under the hood, zero runtime deps. The scan gives every git repo under your home folder a store and wires in Claude Code and OpenCode; run hippo init inside a repo for Codex, Cursor, OpenClaw and Pi.
npm install -g hippo-memory hippo init --scan ~