Skip to content

Stop re‑teaching
your agent.

Make your agent's memory work like a brain. Hippo is long-term memory for coding agents. It's a critical layer for your AI harness that connects across your different tools (Cursor, Claude Code, Codex). It keeps your proprietary data completely local, and it actually learns over time. By strengthening memories each time they are recalled, Hippo preserves what works, lets mistakes decay, and continuously compounds your agents' intelligence.

$ npm install -g hippo-memory
  • Claude Code
  • Codex
  • Cursor
  • OpenClaw
  • OpenCode
  • Pi
  • any MCP client

772 stars 39.4k downloads compared in the agent-memory-atlas

claude · with hippo hooks
Monday · billing-service
> add the refunds endpoint
● Bash(npm install stripe)
⎿ lockfile is pnpm-lock.yaml; npm install would rewrite it
hippo · stored error memory (observed)
Tuesday · new session
> add a webhook for failed payments
hippo · 2 memories in context
· billing uses pnpm; never run npm install here
● Bash(pnpm add stripe) ✓
> /compact
Hippo saved your task snapshot before compacting.

an illustrated session · the hippo lines are real hippo output

  • hippo capture-error

    Real failures become lessons. Interrupts, declined permissions and empty searches are skipped.

  • hippo outcome --bad

    Mark a lesson wrong and it stops coming back.

  • hippo doctor

    One command checks the install and names the fix for anything missing.

The problem

Most AI memory saves everything and searches later.

That's storage with semantic search bolted on. It's why your agent kept hitting the same deploy bug last week. And the week before.

The system saw the failure four times. It had no way to know it should remember.

How it works

Memories decay. Retrieval makes them stronger.

The layers borrow from the brain as design inspiration. On our tests, decay made no measurable difference to recall and sleep cost a little. Outcome marks and retrieval strengthening are what measured helpful; supersession is not measured yet.

Buffer

New information lands here. Session-only, no decay.

Episodic

Timestamped, decays by default. Retrieval strengthens it; errors stick.

Semantic

Repeated episodes compress into stable patterns. The originals decay.

weak traces are forgotten; repeated ones consolidate during sleep

strength = decay over time, re-strengthened on every recall

365d half-life

Decay by default

Every memory fades on a one-year half-life unless it is used. We did not tune 365 days: it tied with 730 days and with decay off.

+2d / recall

Retrieval strengthens

Use it or lose it. Each recall extends the half-life. Memories you reach for survive.

2x half-life

Errors stick

Tag a failure once. It decays slower and resurfaces every time you walk back into that code.

3+ → 1

Sleep consolidates

On hippo sleep, three or more related episodes merge into one semantic pattern. The originals decay; the pattern survives. It keeps the store tidy; it has not been shown to improve recall.

Get started

Zero config. It wires itself in.

Install it, run init in your repo, and hippo detects your agent framework and patches the right config file. Next session, your agent just uses it.

Package installation alone does not enable automatic preservation on every agent. Complete the documented setup and required host trust; capture and compaction coverage depend on the integration.

Detected and patched automatically

  • Claude Code
  • Codex
  • Cursor
  • OpenClaw
  • OpenCode
  • Pi

In a repo, init patches the instruction file each agent already has (CLAUDE.md, AGENTS.md) and adds session hooks where the agent supports them. hippo init --scan ~ gives every git repo under your home folder a store and installs the Claude Code hooks and the OpenCode plugin, but patches no instruction files. hippo init --no-hooks --no-schedule skips the hooks and the daily run.

  1. 1
    $ npm install -g hippo-memory
  2. 2
    $ hippo init

Then hippo sleep tidies the store when a Claude Code session ends, and a daily run at 6:15am learns from the day's commits in every registered project.

Local-first

Your memory lives on your machine.

0 outbound HTTP

Proven by a globalThis.fetch spy that throws on call, across the 1000-event ingestion smoke. Not a hardcoded zero. The default recall path makes no network call either; opt-in features such as the Jev reranker, the LLM reranker and the API embedders do.

SQLite on disk

Memories live in a local .hippo/ store with markdown mirrors you can read, grep, and commit. No cloud and no account. One default to know: hippo sleep sends text to Anthropic for fact extraction when ANTHROPIC_API_KEY is set, and one config line turns that off.

1 call to forget

Right-to-be-forgotten is a single API call. Every row carries kind, scope, owner, and provenance.

tenant-safe by default

Multi-tenant keys are scrypt-hashed with an audit log on every mutation. Tenant A cannot see tenant B, proven by a negative test.

And it's not locked to one tool.

Your ChatGPT memories don't travel to Claude; your .cursorrules don't travel to Codex. Hippo is one store behind all of them.

imports from ChatGPTCLAUDE.md.cursorrulesSlackmarkdown

Compare

Learn what is wrong. Stop repeating it.

How hippo compares with nine other memory tools on the features that define a memory lifecycle.

Feature Hippo MemPalace Mem0 Basic Memory gbrain Zep Letta Cognee Memoria EverMind
Decay by default YesNoNoNoNoNoNoNoNoNo
Retrieval strengthening YesNoNoNoNoNoNoPartialNoPartial
Cross-tool import (ChatGPT/Claude/Cursor) YesNoNoNoPartial?NoPartialNoPartial
Auto-hook install YesNoNoNoNoNoNoNoNoNo
Zero runtime deps YesNoNoNoNoNoNoNoYesNo
LongMemEval (best published) 98.0% any / 88.5% all R@5*96.6% raw / 100% reranked R@594.4**N/A95.53% all R@5*90.2% accuracy**N/AN/A88.78% accuracy**83.00% accuracy**

scroll for more tools →

* LongMemEval retrieval recall at 5. "Any" counts a hit when one answer session is in the top 5, "all" only when every one is. gbrain leads on "all".

** Answer accuracy with a reader model, a different metric from recall.

Most of the others save everything and search it later, or build a knowledge graph. Hippo learns what is wrong and stops repeating it.

Every feature, with sources and check dates, on GitHub →

FAQ

Questions, answered.

All 14 questions →

Is this just RAG?

No. RAG searches a fixed corpus. Hippo's store changes as your agent works: a memory marked wrong drops out of the top results, a newer fact supersedes the old one, and memories that keep getting recalled last longer while unused ones fade on a half-life. Recall itself is search: BM25, plus embeddings if you install them.

Which agents does hippo work with?

hippo init detects Claude Code, Codex, Cursor, OpenClaw, OpenCode and Pi, and wires itself into each one's instruction file, hooks or plugin. It only patches instruction files that already exist. Any MCP client can use the MCP server, and other tools can call the CLI or the HTTP API that hippo serve starts.

Where does hippo keep my data?

On your machine, in SQLite: .hippo/hippo.db in each project, plus a global store in ~/.hippo/ for lessons shared across projects, with markdown mirrors you can read and commit. Recall makes no network call by default. Text goes to an outside provider only through features that use one: an API embedder, the Jev or LLM reranker, hippo refine, and the fact extraction hippo sleep runs through Anthropic's API whenever ANTHROPIC_API_KEY is set in its environment. To turn that last one off, set {"extraction":{"enabled":false}} in .hippo/config.json.

How is hippo different from mem0?

mem0 uses a language model to extract memories, OpenAI by default in its open-source library, and memories stored through its hosted MCP server live in your Mem0 account (mem0 docs (opens in new tab), checked 2026-09-28). Hippo stores memories in SQLite on your machine, needs no account and no model, and hippo init wires it into the coding agents it finds. mem0's platform and hippo both mark an older fact superseded when a newer one replaces it. Hippo also lets you mark a recalled memory wrong with hippo outcome --bad, and it drops out of the top results.

What happens when a memory turns out to be wrong?

Mark it, and it drops out of the top results. hippo outcome --bad weakens the memories from the last recall, hippo supersede <id> "<new fact>" replaces one with a newer version, and hippo reject <id> --reason "<why>" stops that value from returning at all. On the synthetic E1 test, where every mark is correct, plain BM25 plus the outcome mark cut how often a marked-bad memory stayed in the top five from 71.9% to 0.0%. Real marks are noisier, because --bad marks the whole recall batch.

Has hippo been shown to make agents better at their work?

Not yet. The published numbers measure retrieval: whether the right memory comes back, and whether a memory marked wrong stays out of the results. A paired test that runs real agent sessions with and without hippo is under way. Every measurement, including failed runs and one retracted claim, is indexed in docs/evals (opens in new tab).

One command. Every repo gets memory.

Zero config. SQLite under the hood, zero runtime deps. The scan gives every git repo under your home folder a store and wires in Claude Code and OpenCode; run hippo init inside a repo for Codex, Cursor, OpenClaw and Pi.

$ npm install -g hippo-memory
$ hippo init --scan ~