How it works
How hippo decides what your agent forgets
Every memory has a strength. It fades on a half-life, grows each time the memory is recalled, and drops when you mark the memory wrong. A newer fact can supersede an older one, and memories that contradict each other are flagged.
The design borrows from the hippocampus, but that is inspiration, not evidence. In our mechanism audit, outcome marks and retrieval strengthening measured helping; supersession is not measured yet, and decay and sleep have not been shown to make recall better.
Memories decay. Retrieval makes them stronger.
The layers borrow from the brain as design inspiration. On our tests, decay made no measurable difference to recall and sleep cost a little. Outcome marks and retrieval strengthening are what measured helpful; supersession is not measured yet.
New information lands here. Session-only, no decay.
Timestamped, decays by default. Retrieval strengthens it; errors stick.
Repeated episodes compress into stable patterns. The originals decay.
weak traces are forgotten; repeated ones consolidate during sleep
strength = decay over time, re-strengthened on every recall
Decay by default
Every memory fades on a one-year half-life unless it is used. We did not tune 365 days: it tied with 730 days and with decay off.
Retrieval strengthens
Use it or lose it. Each recall extends the half-life. Memories you reach for survive.
Errors stick
Tag a failure once. It decays slower and resurfaces every time you walk back into that code.
Sleep consolidates
On hippo sleep, three or more related episodes merge into one semantic pattern. The originals decay; the pattern survives. It keeps the store tidy; it has not been shown to improve recall.
Each mechanism, with what was measured
Decay: a 365-day half-life
Every memory has a half-life, 365 days by default, and its strength halves each time one passes without use. Until 1.46.0 the default was 7 days. On a pre-registered test, the current version of a fact was in the top five 29% of the time at 7 days and 75% at 365 days, and 730 days and decay switched off both tied with 365 (result (opens in new tab)). So 365 was not tuned, and on that test decay made no measurable difference to recall. Set defaultHalfLifeDays in .hippo/config.json to choose your own, or pin a memory so it never decays.
Recall strengthens
Each recall adds 2 days to a memory's half-life. Recall ranks memories by relevance, strength and recency, and the strongest matches fill the token budget first.
Errors stick
A memory stored with --error, or captured from a failed tool call, gets twice the half-life.
Outcome marks
hippo outcome --good or --bad marks the memories from the last recall. After 5 good marks and no bad ones, a memory's effective half-life is about 42% longer; after 3 bad marks and no good ones, it decays nearly twice as fast. On the synthetic E1 test, where every mark is correct, plain BM25 plus the outcome mark cut how often a marked-bad memory stayed in the top five from 71.9% to 0.0% (mechanism audit, round 2 (opens in new tab)). Real marks are noisier, because --bad marks the whole recall batch.
Supersession and invalidation
hippo supersede <id> "<new fact>" stores the new version and links the old one to it; the old one leaves recall but stays on record. hippo decide "<decision>" --supersedes <id> replaces an earlier architectural decision, halving its half-life and marking it stale. hippo learn --git weakens memories about a tool that a migration commit replaced, and hippo reject <id> --reason "<why>" stops a value from coming back at all.
Conflicts
When two memories overlap in content and contradict each other, hippo records an open conflict; shared tags alone do not count. hippo sleep refreshes the list, hippo conflicts shows it, and each conflict is mirrored under .hippo/conflicts/. Resolving one keeps one memory and weakens or deletes the other.
Sleep and dormant memories
hippo sleep merges three or more related episodes into one pattern, and the originals decay. It keeps the store tidy but has not been shown to improve recall: in round 2 of the mechanism audit, a slept LongMemEval store scored 3.6 points lower at hit@5 than the same store never slept. A memory that fades below the threshold goes dormant instead of being deleted. hippo dormant restore <id> brings it back, and a dormant memory nobody restores is deleted after 180 days by default.
Confidence
Each memory is verified, observed, inferred or stale, and hippo context shows the level beside each memory. A memory unretrieved for 30 days or more is marked stale at the next sleep, and a recall wakes it back to observed.
What happens when a memory turns out to be wrong?
Mark it, and it drops out of the top results. hippo outcome --bad weakens the memories from the last recall, hippo supersede <id> "<new fact>" replaces one with a newer version, and hippo reject <id> --reason "<why>" stops that value from returning at all. On the synthetic E1 test, where every mark is correct, plain BM25 plus the outcome mark cut how often a marked-bad memory stayed in the top five from 71.9% to 0.0%. Real marks are noisier, because --bad marks the whole recall batch.
Is this just RAG?
No. RAG searches a fixed corpus. Hippo's store changes as your agent works: a memory marked wrong drops out of the top results, a newer fact supersedes the old one, and memories that keep getting recalled last longer while unused ones fade on a half-life. Recall itself is search: BM25, plus embeddings if you install them.
Does it need embeddings?
No. Recall runs on BM25 out of the box, with no model and no network call, and a default install has no embedder. Embeddings are an optional install for hybrid search. On LongMemEval-S, where each question gets its own haystack, the benchmark scripts (not hippo recall) fuse BM25 with the free local MiniLM embedder and reach 98.0% recall@5, counting a hit when any answer session is in the top five. On LongMemEval's oracle split with one pooled store, BM25 alone scored 74.0% recall@5 in v0.11. The two runs use different setups, so they are not a before and after.