OMNI
Your AI agent pays to read the same output over and over. OMNI stops that.
One small program between your terminal and your agent. Local, no API key, no proxy. Install it and you never type its name again.
brew install fajarhide/tap/omni && omni init
Inside Claude Code, two lines and the agent does the rest:
/plugin marketplace add fajarhide/omni
/plugin install omni@omni
What that buys, measured
| a file your agent reads twice | 97.2% off the second read |
git log -15 | 94% smaller, every commit kept |
cargo test, 490 passed and 10 failed | 92.9% smaller, the failures kept |
| build and test output across the corpus | 78.0% |
| the tool definitions in every request | 4,940 bytes lighter |
Every one of those replays on your own history. That is the point of the rest of this page.
The problem, in one screen
Your agent runs a test suite. Four hundred lines come back, one of them matters.
$ cargo test
Compiling omni v0.7.5
Running unittests src/lib.rs
running 412 tests
test pipeline::scorer::tests::scores_errors_critical ... ok
... 409 more lines of "ok" ...
test result: FAILED. 411 passed; 1 failed
The failure survives. The 406 lines of ok do not. A handle on the last line brings
every one of them back, byte for byte, if anything ever needs them.
The part nobody else does
Filtering noise is the easy half, and several tools do it. Here is the harder half, and it is where most of OMNI’s saving comes from.
Your agent reads a file. Three turns later it reads the same file again, because nothing remembered the first read. You pay full price both times.
OMNI remembers. The second read comes back as one line:
[OMNI: 178 lines already shown, omni retrieve 0000000000000000]
A 7.6 KB file read twice costs 7.6 KB and then 214 bytes. Nothing was deleted: those lines are already in your agent’s context from the first read, so sending them again buys nothing. The handle is there in case they scroll out of reach.
This is the ledger, and on real command histories it does more work than every filter combined.
Prove it on your own machine
Most tools in this space ask you to trust a number from someone else’s laptop. Run these instead:
omni stats # what OMNI did on your history, in counted bytes
omni retrieve <handle> # any handle from any marker, printed back byte for byte
Every figure on this site comes from a corpus you can rebuild. Benchmarks has the method and the exact command for each row, and every head-to-head we have run, including the runs that did not favour us and the arms that failed to run at all.
What you get
| Longer sessions | Less context spent on ceremony means more turns before you hit the wall, and fewer compactions that lose your thread. |
| Lower bills | 14.9% fewer bytes across 6,656 real commands. On file reads, 25.0%. On git, 22.1%. On build and test output, 78.0%. |
| Nothing lost | Everything removed is archived locally. omni retrieve <handle> prints it back. |
| Nothing invented | If OMNI cannot understand output, it hands it back untouched rather than guessing. |
| Memory between sessions | Close your editor, come back tomorrow, switch from Claude Code to Codex: the project context is still there. |
| Nothing to change | No proxy, no API key, no command to prefix. Install it and use your terminal normally. |
Where it actually helps
Where OMNI helps walks through the situations with the real numbers attached, including where it stands aside and why that is the right call.
Start here
Just want it working. Install takes about five minutes. Then read Reading the markers, which is the one page worth your time, because the markers are how OMNI tells you what it did.
Want to understand it first. What OMNI is, then How it decides what to cut, then The ledger.
Three things it will not do
It will not send anything anywhere. Every stage runs on your machine and the archive is a SQLite file in your home directory.
It will not sit between you and your model. There is no proxy and no API key handed to a local process. That was decided against on purpose, and the reasoning is written down.
It will not quietly guess. A stage that failed to understand its input hands the input back unchanged. Structured data like JSON and YAML is never touched at all. Anything removed leaves a marker saying so. Those three rules outrank compression, in that order, every time they conflict.
What the numbers actually say
OMNI is selective, and that is where its leverage comes from. It goes after the class that dominates an agent’s context, the same file read again and again. A file your agent reads twice comes back 97.2% smaller the second time, and that one is a property of the mechanism: it reproduces on any machine, on demand.
Across a whole corpus the figure is a property of the corpus instead. On the 9,478
command corpus frozen as 0b63218ef78a1edb the ledger takes 1.5% off file reads,
because file reads there average 2.1 KB. An earlier week whose file reads averaged
12.4 KB gave up twenty times more, on the same code, which is what makes the byte figure
a property of the week rather than of OMNI.
The share of available repetition the ledger actually took is 10.7% in aggregate on this corpus. The comparable figure for that earlier week was measured before #760, when the benchmark could not see the ledger’s own guards, so the two are not comparable and this page no longer pretends they are.
A file that changed between the two reads still folds around the change. Each fold keeps the line count of what it replaced, so the lines you did not see moved stay on the numbers your editor gives them.
The corpus is frozen on disk and its hash ships in docs/benchmarks/, so unlike every
figure we published before it, this one can be checked against the same bytes next
release. Benchmarks publishes each run with its corpus, and what
any of them is worth to you depends on how much your own week repeats itself.
Where there is nothing safe to take it takes nothing. A two-line git status has no
ceremony to drop and no repeats to fold, and a JSON payload a later step parses is never
touched at all, so OMNI hands those straight back rather than inventing a saving to
report.
The 14.9% in the table above is a different corpus on purpose: the same harness over a week of ordinary work, with every one of those hands-back counted in alongside the wins. It is an average over that mix, not a promise for yours. The per-class rows are what predict your own workload, and on the frozen corpus their capture rate runs from 6.8% on infra to 12.4% on the mixed bucket, so find the classes you actually run. Both corpora, the method, and every unflattering figure we have are on Benchmarks.
All four arms run now, and on this corpus OMNI is not the top arm. The tool shipping the same cross-turn dedup takes 5.8% of these bytes where our ledger takes 3.0%, over the same filters and the same blocks, so that gap is the dedup engine on its own. Our own figure read 4.9% until #760, when the benchmark stopped measuring a ledger it had never told which command produced the payload. Our filter tier is the weakest of the four at 1.4%, which is why our own ledger scores higher when it is stacked on a competitor’s filters than on ours. The table is generated by the same run as every other figure here, and it is on Benchmarks.
If you want a number that describes your machine rather than someone else’s, run
omni stats after a few days.
Where to ask
Discord for questions, and especially for the case this project cares about most: OMNI stating a result its input does not support. The issue tracker works too. A report with the raw and distilled output side by side gets fixed either way.
What OMNI is
A small program on your machine that edits what your AI agent reads, before the agent reads it.
That is the whole idea. Everything else on this page is about the rules it follows while doing it, and the rules are more interesting than the editing.
The problem it exists for
An agent working in a terminal spends most of its context on output nobody chose to send it:
- a test run is 400 lines of
okand one line that matters - a build is a compile log wrapped around a one-word verdict
- a file gets read, then read again three turns later, because nothing remembered the first read
None of that is free. It fills the context window, which ends your session sooner, and you pay for it again every time the conversation is compacted.
The obvious fixes are all worse than the problem:
| The fix | Why it fails |
|---|---|
| Truncate long output | It cuts the end, and the end is where the verdict lives |
| Ask a model to summarise | An inference call per command, and a summariser that can be wrong |
| Tell the agent to be careful | Works until the agent is busy, which is always |
OMNI is the fourth option: a program that knows what cargo test output looks like,
remembers what your agent has already been shown, and never guesses when it is unsure.
Where it sits
Every serious agent host can run a program when a tool finishes and use what that
program returns. Claude Code calls it a PostToolUse hook, Cursor and the others have
their own name for the same idea. OMNI installs itself there, and in the matching slot
before the tool runs, which it uses only to hand a matched command to itself. The
command still runs unchanged; the shell never knows.
Two consequences follow from that position, and they are the reason this shape was chosen over a proxy.
It sees output, not requests. Your API key never passes through it, no request is delayed waiting on it, and if it dies the host carries on with the raw bytes.
It cannot help where the host will not let it. A host that does not apply a hook’s rewrite to its built-in shell tool will show the agent the same bytes no matter how good the filters get. That is not a bug to fix in OMNI, it is a property of the host, and Supported agents says which host is on which tier.
What it does to a command
Four things, in order, and any of them may decide to do nothing:
- Refuse. JSON, YAML, base64, terraform plans, anything a later step is going to parse: handed back untouched. See What it refuses to touch.
- Filter. A distiller that understands this tool keeps the verdict and the failures and drops the ceremony. There are 12 of them, covering build, test, git and other version control, search, cloud, database, JavaScript and TypeScript tooling, file reads, security scanners and system operations, plus a generic fallback.
- Collapse. Long runs of near-identical lines become one line saying how many there were.
- Fold. Lines the agent has already been shown become a handle instead of a repeat. This is the ledger, and on real corpora it does more work than the filters do.
Then the raw input goes into the archive, and the agent gets the result plus a marker saying what happened.
If you would rather see this as situations than as stages, Where OMNI helps has six of them with the measured saving on each.
What it does to itself
Everything above is about output. There is a second thing OMNI edits, and for a long time it did not edit it at all: its own weight.
OMNI registers as an MCP server, and tool definitions sit in the prefix of every request of every session it is attached to. A prefix byte is not paid once. It is carried from the first request and re-read on every one after it, where a byte removed from tool output was inserted somewhere in the middle and is read fewer times.
Measured across 229 sessions, sixteen of the twenty-five tools OMNI advertised had never been called once, and those sixteen were 4,940 bytes. The distillers remove a median of 4,942 bytes from tool output in a session that pushes real volume through the hook. Two bytes apart, and the prefix side is the one carried from the start.
So OMNI now tells a host about the tools its tier actually uses. A tool that spends as much context describing itself as it saves is not a token-efficiency tool, and noticing that required pointing its own measurement at itself.
What it is not
Not a compressor. It is not trying to make output small. It is trying to make output that an agent can act on, next to a number a human can check. Those pull in different directions more often than you would expect, and when they conflict the number loses.
Not a summariser. No model runs inside the pipeline. The budget for a hook is single-digit milliseconds and nothing with an inference call fits in it.
Not a memory product, though it has one. omni remember, omni goal and the
session handoff exist because the same agent that reads too much also forgets
everything between sessions. Memory across sessions covers that
half.
The rule it is most serious about
A stage that recognised nothing hands back what it was given.
The failure this project keeps having to fix is not lost bytes. It is a confident
summary of input that was never parsed: a find that reported 99% saved by throwing
away the file paths that were the answer, a cargo test that said 1 passed about a
run cargo itself called 490 passed, a dev server reported as a passing test suite.
Every one of those compressed beautifully. All of them were wrong. So the trait
that every distiller implements returns Option<String>, and a distiller that
failed to parse returns None and the caller hands back the raw bytes. It is
enforced by the type rather than by the author remembering.
Where OMNI helps
Eleven situations, with the measured number attached to each. Two of them are cases where OMNI stands aside, and those are in here on purpose: a tool that claims to help everywhere is a tool nobody can predict, and knowing where it declines is what makes the rest worth trusting.
Every figure comes from the same replay of 6,656 real commands described in Benchmarks, so they are averages over a real mix rather than a good day picked out of a log.
1. The agent keeps re-reading the same files
The situation. You ask for a refactor. The agent reads auth.rs, wanders off to
check a caller, comes back and reads auth.rs again. Six turns later it reads it a
third time. Every read is charged at full price, and none of the repeats told it
anything the first one did not.
What OMNI does. The second read comes back as a marker with a handle. The lines are already in the agent’s context; sending them again is paying twice for one fact.
The number: 25.0% off file reads across the corpus, and up to 97.2% off a single repeated read of one file.
This is the biggest single win in the whole product and it is invisible while it works, which is why the marker exists.
2. A test suite fails and you cannot see why
The situation. 412 tests, one failure, and the failure is on line 388 of the output. Your agent reads all 412 lines to find it, and if the run is long enough the host truncates the tail, which is exactly where the verdict lives.
What OMNI does. The test distiller keeps the tally and every failure with its assertion and file position, and drops the passing lines.
The number: 78.0% off build and test output.
This is the case where filtering, not the ledger, does the work. Test output is enormously repetitive within one run, so there is real ceremony to remove before anything has been seen twice.
3. git log and git diff fill the screen
The situation. One commit’s Author, Date and wrapped body is five lines. Fifteen
commits is a screen and a half, and your agent wanted the subjects.
What OMNI does. Every commit is kept, as one hash subject line. Nothing is
summarised away and no commit disappears; the envelope around each one goes.
The number: 22.1% across git and gh on the corpus, and 94% on a verbose
git log -15 specifically.
4. Your session dies at the context limit, repeatedly
The situation. Long debugging session, and about two hours in the conversation compacts. The agent loses the thread, re-reads files it had already understood, and you re-explain the task.
What OMNI does. Two things. Less context spent per command means the wall arrives
later. And memory across sessions survives the compaction: project
knowledge, recurring error patterns, and the goal you pinned with omni goal are in
SQLite, not in the context window.
Where the limit is, and why it is deliberate. OMNI cannot stop a compaction. When one happens it drops what it had shown you on purpose, because a handle is only honest while the agent is still holding those lines, and compaction is exactly when that stops being true. That rule is what keeps every marker true on the other side of one.
5. You switch agents, or machines, mid-project
The situation. You start in Claude Code, move to Codex CLI for a change, and both of them start from nothing.
What OMNI does. The store is one SQLite file keyed by project path, not by agent.
A second agent working in the same directory reads the same project knowledge, and the
ledger’s project scope will hand it a handle for output an earlier session already
produced. That marker says not shown here rather than already shown,
because this agent has genuinely never seen those bytes and the wording has to be true.
The number: 3.7% of post-filter bytes repeat across sessions, against 19.1% within one. So this is a real bonus on top of the in-session saving rather than the main event, and it arrives without anyone configuring anything.
What is not keyed on the agent yet. Two agents in one repository share that history by side effect
rather than by design. The marker used to say from an earlier session, which reads as
your earlier session when it was someone else’s, and worse, as a claim the content had
already arrived; it now says not shown here. The ledger is
straight about what is and is not keyed on the agent today.
6. kubectl get pods -o json | jq
The situation. You pipe structured output into something that parses it.
What OMNI does: nothing. JSON, YAML, NDJSON, CSV and TSV pass through byte for byte. A compressor that reformats a payload the next command is about to parse has not saved you anything, it has broken your pipeline.
The number: 0%, by design. See What it refuses to touch.
7. You read one big file in several passes
The situation. A file is longer than one read, so the agent takes it at an offset, then another, then another. Each window repeats the head of the file, because that is what a window at an offset contains.
What OMNI does. It folds the repeated head and moves the line numbering to match, so the lines you can still see are numbered where the file really has them. That second half matters: a fold that renumbers what is under it is worse than no fold, and it is why this case was refused for a release until the numbering could be kept true.
The number: 0.0% before, 4.7% after, measured on four overlapping windows of one markdown file. Source files are unaffected, since the readfile distiller reaches those first at 46.6% either way.
8. You dispatch a subagent
The situation. Your agent spawns a helper to do a scoped job. The helper starts with an empty context and reads a file the parent already read.
What OMNI does. It gives the helper its own view. Claude Code hands a subagent the parent’s session id, so a ledger keyed on the session alone would answer the helper with the parent’s history and tell it 200 lines were already shown, about bytes that context had never received. The helper now sees either the content or a marker that says plainly nothing was shown here.
The number: no ratio, and that is the point. This is a correctness case. The saving was never the problem; the claim was.
9. You follow a marker to get the content back
The situation. A marker says omni retrieve <handle>. You run it, or your agent
does, and reads the result.
What OMNI does. It hands those bytes over whole. Before, they went back through the pipeline, hashed the same, and were folded into the very marker that sent you there, so following the instruction returned the instruction.
The number: one delivery, not an exemption. The next repeat folds again, which matters because 15.05% of the archive on a real installation has been pulled at least once, and exempting all of it would trade a false claim for a lost saving.
10. Your context gets compacted mid-session
The situation. The session runs long, the host compacts the conversation, and half of what your agent was holding is gone.
What OMNI does. It forgets. The ledger’s whole licence is that the agent still holds the bytes a handle replaces, and compaction is where that stops being true, so the shown-set goes with it. Nothing after a compaction claims you have already seen something you no longer have.
The number: no ratio. It costs savings on purpose, and it is the trade that keeps the markers true.
11. Every request carries a tool list you never call
The situation. OMNI registers as an MCP server, and tool definitions sit in the prefix of every request of every session. Unlike output, a prefix byte is not paid once: it is re-read on every request after the first.
What OMNI does. It advertises the tools your host’s tier actually uses, nine instead
of twenty-five, with OMNI_MCP_TOOLS=all to restore the rest and omni doctor naming
which set is in force.
The number: 4,940 bytes off every request. Measured across 229 sessions: sixteen of the twenty-five had never been called once.
And one more where nothing happens
kubectl get pods with 35 pods returns a table where every row is a fact. There is no
ceremony to drop and nothing has been seen before, so OMNI hands back all 35 rows and
reports a 0% saving.
Most calls in this corpus look like this, and that is the shape of the tool. OMNI is not a thing that shrinks everything a little. It stands aside until there is something worth taking, then takes a great deal: on this corpus, 78.0% off build and test output and 25.0% off file reads. The 14.9% aggregate counts every stand-aside in alongside those wins, which is why the per-class row is the one to read for your own workload.
What this adds up to
| Class of command | Calls in the corpus | Saved |
|---|---|---|
| build and test | 69 | 78.0% |
| file reads | 699 | 25.0% |
git, gh | 661 | 22.1% |
search (grep, rg, find) | 828 | 13.3% |
infra (kubectl, az, docker) | 254 | 8.2% |
| everything else | 4,145 | 6.9% |
| all of it | 6,656 | 14.9% |
Run omni stats after a few days and you get this table for your own history, which is
the only version of it that describes your work.
How it decides what to cut
The pipeline is fixed and every payload walks the same stages:
Read → Guard → Score → Distill → [Collapse] → Ledger → Route → Persist
None of them is allowed to invent anything, and each one is allowed to decline. Collapse is bracketed because it is a fallback rather than a step: it runs only when the distilled form failed to beat the guardrail. The pipeline, stage by stage has the diagram and the reasoning.
Guard
The gate. It answers one question: is this payload something a later step is going to parse? If yes, nothing downstream runs and the bytes come back exactly as they arrived. What it refuses to touch is the whole of this stage and it is worth its own page, because “OMNI did nothing” is usually this working correctly rather than a failure.
Score
Every line gets a relevance tier. The scorer is a pure function of the text, the command that produced it, and whatever session history exists.
| tier | weight | what lands here |
|---|---|---|
| Critical | 1.0 | errors, failures, the verdict line, anything naming a file and a line number |
| Important | 0.7 | warnings, counts, state that changed |
| Noise | 0.1 | progress, timing, decoration, repeated ceremony |
The tiering happens before any distiller sees the block, which matters when you are debugging why a distiller behaved oddly: the tier may already have decided the outcome, so probe the segment tiers before rewriting the distiller.
Distill
Now a tool-specific filter runs, chosen by matching the command. The cargo test
distiller keeps the counts and every failure with its assertion. The git distiller
keeps the changed paths. The search distiller keeps the match lines with their
filenames.
Each one implements the same trait, and the signature is the design:
fn distill(&self, segments: &[OutputSegment], input: &str,
session: Option<&SessionState>) -> Option<String>;
Option, not String. A distiller that did not understand its input returns None
and the caller hands back the raw bytes. That is the difference between “I read this
and here is what matters” and “I recognised nothing and here is a confident summary
of it”, and it is enforced by the compiler for all 12 rather than by each author
remembering to check.
Collapse
Runs of near-identical lines become one line stating the count. Twenty
Downloading foo v1.2.3 lines become one.
Two things about this stage surprise people. It runs after the distiller and only
when the distiller did not earn its keep: both hooks distill the raw bytes, ask
beats_guardrail, and reach for the collapsed form only if that fails. So a distiller
always reads the original output, never [N similar lines collapsed] markers. And
which collapse mode fires is chosen by specificity, so a kubectl command piped into
grep may take the infrastructure path rather than the log path.
Ledger
Everything above judges this payload on its own. The ledger is the one stage that judges it against what the agent has already been shown, replacing a run of repeated lines with a marker and a handle. It is the largest single source of savings and it has its own page: The ledger.
Persist
The raw input is archived, keyed by SHA-256, and the marker the agent sees carries a handle into that archive. Covered in Nothing is deleted.
Archiving happens even when the projection saved nothing. A block is worth remembering because it may be seen again, not because it compressed today.
What decides the order
Correctness beats compression at every stage, and the order they win in is written down:
- Never fabricate. A stage that recognised nothing hands back what it was given. A failed command passes through verbatim. Structured payloads are never touched.
- Never lose the answer quietly. Anything dropped leaves a marker, and where the content allows, a handle that retrieves it.
- Then compress, as hard as the first two allow and no harder.
The reason that ordering is explicit is that the project has broken it before. A
kubectl table once came out as k8s: 2 pods because a pod table is an enumeration
where every row is a datum. It reported a large saving. There was no noise in the
input to remove, so the saving was the answer.
Nothing is deleted
Every byte OMNI removes is written to a local SQLite archive first, keyed by its SHA-256. The agent gets a marker carrying a 16 character handle, and the handle brings the original back byte for byte.
[OMNI: 406 lines omitted, omni retrieve 0000000000000000 for full output]
omni retrieve <handle>
That works from any shell, in any session, on any host, and it does not re-run your
command. Where MCP is wired, the agent can do it itself with the omni_retrieve tool
without asking you.
Why this is the load-bearing rule
Filtering output is a bet that the removed part did not matter. The archive is what makes the bet safe to lose. It changes the worst case from “the answer is gone” to “the answer costs one retrieval”, and that difference is what lets the rest of the pipeline be aggressive at all.
It also changes what a bug means here. A distiller that cuts too much is a bad trade. A handle that does not resolve is a broken promise, and it is the one defect this mechanism cannot have.
The one rule the archive enforces on everything else
A run is archived before its marker is written, and a failed archive means the run stays verbatim.
The order matters. Writing the marker first and archiving second would produce, on any
write failure, a marker pointing at content that was never stored: output that looks
like it can be recovered and cannot. That happened once, store_rewind returned a key
even when the write had failed, and the fix was to make the marker conditional on the
archive rather than the other way round.
So when you see a handle, the content behind it exists. That is not a hope, it is the order of two statements.
What it costs
Disk, and a write on every distillation that removed something.
The archive is capped rather than unbounded: archiving every lossy distillation measured 83.1 MB over 30 days, and capping the archived block at 64 KB brought it to 13.3 MB while still covering 3,604 of 3,657 rows. The cap was chosen from that measurement rather than picked.
Traces used for benchmarking are pruned separately, at seven days by default. That prune is why no published figure here can be re-derived after a week, and why every number in Benchmarks names the window it was measured in.
Where it lives
~/.omni/omni.db, a single SQLite file. It never leaves the machine.
omni stats # what it has been doing
omni diff # the last command, raw against distilled
omni retrieve <handle>
omni diff is the quickest way to develop trust in this: run a noisy command, then
look at exactly what the agent was handed instead.
The ledger
Every distiller answers the same question one command at a time: given this output, what can be dropped.
The ledger answers a different one: given everything already shown in this session, what is this output repeating.
The two are orthogonal, and on real corpora the second one is worth more. Replayed over 6,656 traces, 22.9% of raw bytes were lines the agent had already been shown, and 22.4% still were after every distiller had run. Filtering barely dents repetition, because repetition is not noise. Each line is perfectly good signal. It is just signal that was already delivered.
What it does
A run of consecutive lines that were all emitted earlier becomes one marker naming the count and a handle. Everything else passes through byte for byte.
[OMNI: 40 lines already shown, omni retrieve 0000000000000000]
In a Read, the marker is followed by one ⋮ per line it replaced:
[OMNI: 40 lines already shown, omni retrieve 0000000000000000]
⋮
⋮ (38 more)
⋮
That looks like padding and it is doing real work. The editor numbers whatever it is
handed, counting from the line the read started at, so a view with fewer lines than the
file puts every surviving line on a number it is not at. Keeping the count means the
survivors keep their own numbers, which is what lets a fold sit in the middle of a file at
all. Without it a fold had to stop at the first line that survived, and on this repo’s
CHANGELOG.md that was 4.4% against the 76.5% its runs were worth.
It reaches the class nothing else can. File reads are the largest class in the corpus, and the filters save 0.0% of them, correctly: you cannot strip lines from a file the agent asked to see without guessing which parts it meant. The ledger takes 25.0% of that same class without guessing anything, because those lines were already delivered once.
Two scopes, two different claims
They are not the same statement and the marker says which one it is making.
| origin | marker | what it means |
|---|---|---|
| session | N lines already shown | the agent is still holding these bytes, so the handle is free unless it chooses to re-read |
| project | N lines not shown here | these went to a different session of this project and this agent has never seen them |
The distinction is the whole reason the project scope exists. An earlier design cancelled it on the grounds that a handle for another session’s content is a lie, which was right about the wording and wrong about the remedy: the fix is to stop saying “already shown”, not to stop remembering.
Because the project claim is not free, it carries a higher bar. A session-origin run must save 150 bytes over its marker; a project-origin run must save three times that, since the agent has no choice about paying a retrieval if it needs the content.
And the project scope may only ever fold part of a reply. A session fold can take
the whole of one, because the reader really is holding those bytes. A project fold
cannot: the reader has never seen them, so replacing everything leaves it with one
marker, no content, and no way to check the only claim it was given. A subagent’s
first command came back as a single line saying 40 lines were identical to an earlier
session, and a reviewer dispatched onto a pull request had to pipe files through
base64 to read the code it was sent to review.
The condition is what the fold would leave behind, not who is reading. “Has this reader seen anything yet” sounds like the same test and is not: every whole-output project fold recorded on this machine that names a session happened in a session already holding between 261 and 1,369 lines of its own. A reader having seen something says nothing about whether it has seen these lines, which is exactly what the project scope answers for. Refusing the class costs 10 folds and 32,104 bytes, 1.51% of every byte this store has folded, against the 678,585 bytes the partial arm keeps earning.
The two floors that decide nothing folds at all
Both bars above ask whether a run outgrows the marker replacing it. Two floors are checked before either of them, and between them they explain most of the cases where output comes back untouched and looks like the ledger is off.
Output under 264 bytes never reaches the ledger. Below that there is no run long enough to be worth a handle, so the whole stage is skipped.
A fold that covers the entire output needs 1024 bytes. The bars assume the agent still holds the rest of the output beside the marker and can decide whether the handle is worth spending. Cover everything and there is nothing beside it, so needing any part of the payload costs a retrieval the agent had no say in. Every whole-output fold this machine recorded was under 1 KB, and four of the four were retrieved within nine seconds, against a 0.85% retrieve rate across all 5,178 distillations in the same store. They saved 2,680 bytes, then spent 319 bytes of marker plus four extra tool calls handing back the same 2,999. The floor is the top of that measured range rather than a knee, because nothing above it was observed either way. n=4, one machine.
The premise everything else follows from
The agent is still holding these bytes.
That single statement is what licenses replacing forty lines with a handle. Every rule below is either a consequence of it or a defence of the moment it stops being true. When you find yourself asking why the ledger does something, ask what it would take for the premise to be false, and the answer is usually there.
It is also why this is a cache invalidation problem and not a memory system. The ledger does not store knowledge. It stores receipts.
The three readers the premise fails for
Every rule worth knowing here is a defence of the moment the premise stops being true. There are exactly three readers it fails for, and the ledger answers each differently.
A subagent. Claude Code hands a helper the parent’s session id, so a ledger keyed on the session alone would answer it with the parent’s history and claim 200 lines were already shown to a context that had received none of them. The scope is the reader, not the session, so a helper accumulates its own and falls through to the project scope for anything else, where the wording says plainly that nothing was shown here.
A context that was compacted. The host says so before it happens, and the ledger forgets that session’s shown-set at that moment. It costs savings on purpose. Nothing after a compaction claims you already have something you no longer hold.
A reader following a handle. Asking for bytes back is proof the reader does not have them, so the delivery answering a pull is handed over whole. Before, it went through the pipeline, hashed the same, and came back as the very marker that sent the reader there. One delivery, not an exemption: the next repeat folds again.
The pattern is worth more than the three cases. When the ledger surprises you, ask which reader is holding the bytes, and whether anything told OMNI that reader had changed.
The flow, one command at a time
Structured payloads never get this far: the same format sniff that gates collapse gates this stage too.
Two details are easy to read past and are the whole correctness story.
The archive happens before the marker, so a handle never names content that was
not stored. And what gets recorded is what was delivered, not what arrived: a run
that became a marker never reached the agent, so recording it would let the next
occurrence claim already shown about bytes nobody received. That was a real defect
(#465) and it cut both ways, because
session origin charges a third of what project origin does, so the false claim also
made the ledger three times more willing to fold.
How it remembers
Three verbs, and each one is a different table or a different trigger.
Store
Two tables, on purpose.
| holds | size | |
|---|---|---|
ledger_lines | (scope, line_hash, ts, agent_id) | 16 bytes of hash per line |
rewind_store | the actual bytes of a folded run, keyed by their SHA-256 | the content, once per distinct block |
Recording every emitted line is cheap because the line itself is never stored, only its hash. The content only goes to the archive when a handle is actually issued.
The hash is taken on the trimmed line, so the same line reached through sed -n
and through cat is one line rather than two.
Recording is unconditional; folding is not. A block is worth remembering because it may show up again, not because it compressed today. So a command whose output is entirely new still writes its lines, and pays for itself the next time.
Retrieve
omni retrieve <handle>
An exact lookup on a content address. There is no candidate set, no ranking, no merging of results, and no search: one handle names one block of bytes. The handle is derived from the content, so identical output is one row however many commands produced it.
Nothing is ever pulled back automatically. The marker is a pointer, and the agent decides whether the content is worth a retrieval. That is the trade the whole design rests on: the worst case is not “the answer is gone”, it is “the answer costs one round trip”.
Where MCP is wired the agent calls omni_retrieve itself. Otherwise it runs the shell
command the marker printed.
Forget
Time, plus one event.
At compaction, the session scope is dropped entirely. Compaction is the moment inside a session where the agent stops holding what it was shown, so every claim the session scope could make becomes false at once. Forgetting costs a missed reduction. Not forgetting means telling an agent it has content its context no longer contains, which is the defect, not the cost.
At 30 days, both scopes prune on the same window. A session scope cannot outlive its session, so the ordinary retention window already bounds it. The project scope is the one that could grow without limit, and the honest bound on it is the same window: content nobody has produced in a month is content this project has stopped emitting, and a handle for it buys a retrieval of something the agent will not recognise either.
A repeat refreshes the timestamp rather than being ignored, so output that is still being produced does not age out on the strength of when it was first seen.
There is no eviction by size, and that is deliberate. Evicting by size drops the oldest rows of the busiest project first, which is exactly where the repeats are.
What two agents in one repo share
The session scope is one agent’s, because a host session id belongs to one host. The project scope is keyed on the working directory and nothing else, so two agents running in the same repository write into one history and read from it.
That is sharing by side effect rather than by design. Nothing in the ledger knows which agent it is talking to, so a project-origin marker can hand agent B a handle for lines only agent A was ever shown. The higher bar means the trade is priced as a retrieval either way.
The wording used to make that worse. from an earlier session states where the lines
came from, and a reader took it as your earlier session, which it need not be, and
then as a claim they had already seen the content. A run marker now says
not shown here and states the only thing the reader has to act on, which is that
these bytes never arrived (#567).
As of #509 the agent is recorded on
every line, and nothing keys on it yet. The measurement decides that: keying the scope
on (project, agent) would end the cross-agent case together with whatever reuse in
it is genuinely free, and the corpus says the effect is currently latent rather than
live. The column is what makes it possible to ask.
The rules it inherits
Append-only. It only ever shortens the output of the command in flight and never rewrites anything already delivered. That is what keeps the upstream prompt cache intact: a cache works on a prefix, so shortening the suffix costs nothing while retroactive compaction would destroy it.
Deterministic. The same ledger state renders byte-identical output. The handle is
a content address and carries no timestamp. An earlier design used
{timestamp}_{hash} and made 4 of 73 repeated inputs emit different bytes.
Nothing is lost. Stated above and enforced by the order of two writes. The general rule and what it costs are in Nothing is deleted.
Failures are never folded. A line stating a failure is exempt however often it has been shown. “You have seen this already” is sound for informational lines and wrong for the error channel, where the repetition is the signal: the same TypeError on a re-run means the bug is still there. Eliding it delivers source context and no statement of what went wrong, which an agent reasonably reads as the failure being fixed. Marking the line unseen rather than filtering it afterwards also splits the run around it, so the frames either side still fold.
Unknown means untouched. Structured payloads never reach the ledger at all.
What it is worth
From the same replay, the ledger is 12.2 points on top of OMNI’s own filters and 11.4 points on top of a competitor’s, which is the clearest statement that it is orthogonal to whose patterns run:
| bytes | saved | |
|---|---|---|
| omni, filters only | 6,469,047 to 6,292,856 | 2.7% |
rtk pipe | 6,469,047 to 6,067,012 | 6.2% |
lean-ctx compress | 6,469,047 to 6,073,757 | 6.1% |
| omni, with the ledger | 6,469,047 to 5,506,627 | 14.9% |
rtk pipe + omni’s ledger | 6,469,047 to 5,333,483 | 17.6% |
The last row is deliberate. A reader who wants the largest possible number would run their filters with our ledger, and saying so is cheaper than being caught not saying it.
What it refuses to touch
Before anything else runs, the payload is classified. If it looks like something a later step is going to parse, the whole pipeline stands down and the bytes come back exactly as they arrived.
Four kinds are recognised: JSON, YAML, CSV and TSV. Recognising any of them ends the matter.
This is the stage people mistake for a failure. kubectl get pods -o json coming back
at full length is not OMNI missing an opportunity, it is OMNI declining one.
Why declining is the right answer
A distilled JSON document is not a smaller JSON document. It is a broken one. The
jq two steps later fails, the agent reads the failure, and the cost of that round
trip is larger than anything the compression could have saved.
So the gate is deliberately biased. Bracketed but unparseable input, truncated JSON, JSON carrying comments: all treated as structured. Compression cannot repair a malformed payload but it can certainly make it worse.
How it decides, and where it has been wrong
JSON: a whole document that parses. Above a size threshold a full serde_json
parse would blow the latency budget, so bracket shape alone decides. Free text almost
never carries "key":, which is the cheap signal for the ambiguous cases.
YAML: key-shaped lines, plus one rule that exists because of a real failure.
A block scalar (config.hcl: |) hands the rest of the block to whatever the value
happens to be: Vault HCL, a shell script, a PEM certificate. Those lines carry no
key: and are not YAML-shaped, so a naive sniff calls them prose. One embedded
ConfigMap sank a whole 608-line kubectl kustomize manifest that way: the sniff said
“not YAML”, the gate stood down, and the manifest went down the lossy path. Lines
introduced by a block indicator are now skipped rather than judged.
CSV and TSV: a consistent delimiter count across a minimum number of rows. One row proves nothing.
Turning it off, and when to
OMNI_PASSTHROUGH=1 <your command>
Skips the pipeline entirely. Use it when you are debugging OMNI itself and need to see what a command really printed, or when reading a file whose exact bytes matter.
The prefix works on every path, including inside an agent, but not for the reason it
looks like. A hook is a separate process the host spawned, so it inherits the host’s
environment and never sees a variable you assign in front of a command. What it does
see is the command string, so OMNI reads the assignment there. Two consequences worth
knowing: only a leading assignment counts, the same position a shell would apply
it in, and echo OMNI_PASSTHROUGH=1 mentions the name without setting anything and is
still distilled. Exporting it for the whole session works the ordinary way.
This is the single most useful environment variable here, and it is the first thing to reach for when you suspect OMNI has changed something it should not have. If the output is identical with and without it, OMNI was not involved.
Things that look like this gate and are not
Negative savings on small output. A short payload can come back a few percent larger, because the marker costs more than the compression saves. Expected, not a defect.
A command whose output arrives intact anyway. Most calls are handed back untouched, because taking anything would be unsafe or would not pay for its own marker. That is the pipeline working.
kubectl binary streams. SPDY corrupts those with or without OMNI in the picture.
Shell quoting. Word splitting is your shell, not this program.
What it costs
Not zero. Here is the whole bill.
Latency
Median of 12 runs each, release binary, measured end to end through the post-hook:
| fresh database | 205 MB database | |
|---|---|---|
git status (496 B) | 21.1 ms | 60.7 ms |
cargo test (16.5 KB) | 24.5 ms | 64.5 ms |
Payload size barely matters. Database size does, and that is the number to watch as your archive grows.
The distillation itself is single-digit milliseconds. Almost all of the rest is the archive write. Earlier releases measured 82 ms and 276 ms on the same machine, and the difference was three fixes rather than faster hardware: a tokenizer loaded per command for a reporting column, 249 line-filter regexes compiled whether or not their filter matched, and a connection pool opening four SQLite handles in a process that exits after one payload.
Measure latency by removal, not by a unit-test timer. A microbenchmark in the suite reported 66 ms for work that an A/B on the release binary put at 34.3 ms. Only the second kind of number is quotable.
Memory
Flat. The pipeline works on streams, so a 20,000 line log does not cost more resident memory than a short one.
Disk
One SQLite file at ~/.omni/omni.db.
Archived content is capped at 64 KB per block. That cap came from a measurement: archiving every lossy distillation cost 83.1 MB over 30 days, and the cap brought it to 13.3 MB while still covering 3,604 of 3,657 rows.
Benchmark traces are pruned at seven days (OMNI_TRACE_RETENTION_DAYS). That prune
is why no published figure can be re-derived a week after it was measured.
Tokens
The thing you came for, and it has two halves.
What it saves. Over 6,656 real commands on 0.7.3: 14.9% fewer bytes across the whole mix. By class, the spread is enormous:
| class | filters | with the ledger |
|---|---|---|
| build and test | 76.9% | 78.0% |
| file reads | 0.0% | 25.0% |
git, gh | 4.4% | 22.1% |
| search | 4.8% | 13.3% |
| infra | 4.4% | 8.2% |
| everything else | 0.6% | 6.9% |
What it costs. Every marker is bytes the agent pays for, and on short output a marker can cost more than the fold saves outright. That is why a fold has to clear a floor before it is allowed, and why most calls are handed back untouched instead. The pipeline’s latency is paid on every command whether or not anything is taken.
There is also a cost no byte count can express: a retrieval. When the agent needs content behind a handle, it pays a round trip it would not have paid if the bytes had simply arrived. Project-scope folds carry three times the profitability bar for exactly that reason.
The cost that is not OMNI’s to pay
On a flat-rate plan, compression does not reduce a bill at all. What it buys is session lifetime and fewer re-runs. Prompt-cache reads bill about a tenth of fresh input, so bytes saved once are not dollars saved per turn.
This is why the project’s own primary measure is context-window pressure for the same job, and reduction percentage is a diagnostic rather than a headline. See Where OMNI is going.
If it panics
It fails open. The raw output passes through and your agent never sees an error. Every
hook runs inside catch_unwind, and a database that will not open costs session
context rather than the whole pipeline.
Install
Get the binary
macOS and Linux, via Homebrew:
brew install fajarhide/tap/omni
macOS, Linux, WSL:
curl -fsSL omni.weekndlabs.com/install | bash
Windows, PowerShell:
irm omni.weekndlabs.com/install.ps1 | iex
From source, which needs the toolchain pinned in rust-toolchain.toml:
git clone https://github.com/fajarhide/omni
cd omni
cargo build --release
From inside Claude Code, if you would rather have the agent do the rest:
/plugin marketplace add fajarhide/omni
/plugin install omni@omni
On any agent that reads skills, the same skill installs through the skills directory CLI, and is listed at skills.sh/fajarhide/skills/omni:
npx skills add fajarhide/skills --skill omni
Either way that installs a skill, not the binary. The skill carries the install commands below, the verification step, and how to read the markers, so the agent stops guessing at any of the three. Everything on this page still applies; the plugin only means someone else types it.
Wire it into your agent
omni init # the host you are running in, or a menu if you have a terminal
omni init --claude # or --cursor, --codex, --gemini, and 11 more
omni init --all # every host, and a .vscode/mcp.json in the current directory
omni init writes hooks and registers the MCP server. It is idempotent, so running it
again after an upgrade is the right move rather than a risk.
On Claude Code the hooks are the part that shortens output, and the MCP server is a convenience with a cache cost. Keeping it is fine; knowing what it costs is in Supported agents.
With no terminal to prompt on, which is how an agent runs it, omni init configures
the host it is running inside instead of failing on the absent menu. It says which
host it picked. If it cannot name the host, a plain shell for instance, it stops and
asks for a flag rather than installing into somewhere nobody asked for.
Every supported flag is in init. Which hosts get what is in Supported agents, and that page matters more than it sounds: a host that cannot rewrite its own shell tool’s output will not show the agent distilled bytes however well the pipeline works.
Verify
omni doctor
This is not optional ceremony. It checks the binary is on PATH, the database opens,
the hooks are actually installed where the host reads them, and the MCP server is
registered. omni doctor --fix repairs what it can.
Codex CLI needs one extra step. It runs only hooks it has been told to trust and
skips the rest silently. After omni init --codex, start codex once and approve
them under “Hooks need review”. omni doctor will keep failing until you do.
Confirm it is really running
omni doctor says the wiring is correct. This says the wiring is being used:
cat some-long-file.txt # through your agent, not this shell
omni diff # raw against distilled, for the last command
omni stats
If omni stats shows rows and omni diff shows a difference, the hook is live.
A trap worth knowing now rather than later: the numbers in omni stats are split by
agent_id, and a row recorded under terminal is TTY output no model ever read.
When you are judging whether OMNI is earning its place, look at the rows for your
actual host.
Upgrade
omni update # Homebrew installs
brew upgrade omni
Re-run omni init afterwards if a release changes the hook contract. The changelog
says when that happens.
Remove it
omni init --uninstall # hooks and MCP registration for one host
omni reset --all # every integration, and offers to wipe omni.db
omni reset without flags gives an interactive menu. Neither command touches your
shell configuration, because OMNI never wrote any.
Your first hour
Assumes omni init and omni doctor are done. Nothing here changes configuration.
What the hour buys is the ability to check OMNI instead of trusting it. By the end you will be able to see any cut side by side with the original, pull back anything it removed, and tell one of its markers apart from a line that merely looks like one. That last skill is the one that makes the other two worth having.
See a distillation happen
Ask your agent to run something noisy. A test suite or a build is ideal.
Then, in your own shell:
omni diff
Raw on one side, distilled on the other, for the last command. This is the fastest way to develop either trust or suspicion, and both are useful.
Try one by hand
omni exec cargo test
omni exec runs a command through the whole pipeline and prints the result with a
footer. It is the harness every bug report in this project is asked to use, because it
takes the host out of the picture.
The argument form is exact: omni exec cargo test, not omni exec -- cargo test
and not a quoted string. Both of those fail with “No such file or directory”.
Look at the numbers
omni stats
It leads with session lifetime, how many commands a session carries before the host closes it, because that is what the context window actually costs you. The distillation percentage below it is a diagnostic for one host’s pipeline.
Every absolute figure it prints is in bytes, which are counted. It used to report tokens, and those were the same byte counts divided by a constant calibrated against another vendor’s tokenizer, so the unit could not be defended even though the arithmetic was fine. Percentages were never affected: the divisor cancels in a ratio.
omni stats --view detail # per command, per route, per session, per agent
omni stats --rerun # which distillers cost a re-run
omni dashboard # the same numbers in a browser, on 127.0.0.1 only
--rerun is the interesting one. Reduction percentage cannot tell you whether a
distiller removed something the agent then had to go and fetch again; this can.
Pull back something it removed
Every marker names a handle. Run it:
omni retrieve 0000000000000000
That exact handle is the documentation example and is refused by name, which is the point of this section. Copy a real one out of a marker in your own output and you get the bytes back verbatim, and the exit code tells you which happened: 0 when the handle resolved, 1 when it did not.
That pair is the fastest trust check there is. A tool that removes things and cannot give them back is a tool you have to take on faith.
Tell a real marker from one that is just text
This page is full of markers, so is OMNI’s own source, and so is any bug report that quotes one. Searching your transcript for the marker shape will find all of them.
The handle is what separates them. Worked examples everywhere in this manual use the
reserved 0000000000000000, which no real fold can ever be assigned, so:
omni retrieve <handle-from-your-output> # exit 0, and the content
omni retrieve 0000000000000000 # exit 1, "the documentation example"
If you are measuring whether OMNI did anything at all on a run, that exit code is the
answer and grepping for [OMNI is not.
Pin what you are working on
omni goal set 'Migrate the billing service off the legacy queue'
The scorer favours output related to that goal, and the agent is reminded of it rather
than drifting. omni goal show to check, omni goal clear to drop it.
Turn it off for one command
OMNI_PASSTHROUGH=1 kubectl get pods -o yaml
The first thing to reach for when you suspect OMNI changed something it should not have. If the output is identical with and without it, OMNI was not involved.
Things worth knowing before they bite
Reading a file through your shell may arrive distilled. Since the hook really does
rewrite Bash output, a cat or sed of a source file can come back folded. Use your
agent’s file-reading tool, or OMNI_PASSTHROUGH=1, when you need exact bytes.
A matched command may be rewritten before it runs. The pre-hook turns some
commands into omni exec, redirection included, so the log file you later read is the
distilled one. Break the prefix (env cargo test, or true && cargo test) when you
need the raw log on disk.
Do not judge OMNI by output you read through OMNI. A cargo test read through the
hook once reported “1 failed” for a 398-pass green suite. Redirect to a file with
passthrough on before making any claim about a result.
When to ask for help
If output ever looks shorter than it should, if a row is missing, or if OMNI reports a
success for something that failed, that is worth reporting. Reproduce it with
omni exec first, and read the whole distilled output rather than a grep of it:
grepping hides the headers that often make the output lossless after all.
Reading the markers
A marker is OMNI telling you what it did. There are only a few shapes, and knowing them is the difference between trusting the tool and suspecting it.
The shapes
[OMNI: 406 lines omitted, omni retrieve 0000000000000000 for full output]
Content was cut and archived. The 16 characters are a handle:
omni retrieve <handle> prints the original back, byte for byte, from any
shell in any session.
[OMNI: 40 lines already shown, omni retrieve 0000000000000000]
The ledger. These lines were emitted earlier in this session, so the claim is that the agent is still holding them and the handle costs nothing unless it wants to re-read.
[OMNI: 40 lines identical to an earlier run, omni retrieve 0000000000000000]
The ledger again, for the case where the identity is the answer. You ran the same command a second time and it printed the same lines, so the marker says that rather than “already shown”: a poll re-run to find out whether a value moved is answered by the value not having moved, and a re-read is answered by the file not having changed.
[OMNI: 40 lines not shown here, omni retrieve 0000000000000000]
Also the ledger, different claim. These lines went to a different session of this project, and this agent has never seen them. The wording is deliberately not “already shown”, because that would be false. Folding them is a bet that the agent will not need them, and it carries three times the profitability bar for that reason.
That other session may also have been a different agent. The project history is keyed on the directory, so anything running in this repository contributes to it. See what two agents share.
[OMNI: identical to the 40 lines already shown, omni retrieve 0000000000000000]
[OMNI: identical to 40 lines from an earlier session, none shown here, omni retrieve 0000000000000000]
The same two claims, for a reply that is repeated in full. When the fold covers
every line, the marker is the whole output rather than a gap inside it, so it says
identical to and you get one line where a re-run would have printed the same
hundreds. Anything less than the whole reply keeps the wording above.
[OMNI: 40 lines already shown, omni retrieve 0000000000000000]
⋮
⋮
In a file read, and only there, the marker is followed by one ⋮ per line it replaced.
Your editor numbers the lines it is handed, counting from the line the read began at, so a
shorter view would put every surviving line on a number it is not at. Keeping the count is
what lets a fold sit between two pieces of content you still need. The filler is never
retrievable content: omni retrieve on the handle above prints the real lines.
[OMNI: 40 lines already shown from charlie.tf, omni retrieve 0000000000000000]
Any of the four can carry from <source>, naming the command whose output first
showed those lines. It appears only when that command is not the one you just
ran, which is the case you cannot resolve from the marker alone: reading one file
and having a block elided because a different file showed it earlier. Without the
clause, comparing two files to check a shared block matches is answered by deleting
the evidence.
Running the same command again carries no clause, because the marker above already says so in words. Marker length decides what is worth folding at all, so a source on every marker would cost savings on the common case to label the rare one.
[N similar lines collapsed]
Collapse. A run of near-identical lines, replaced by a count.
[OMNI Active] ⏺ 93.7% reduction (2.3 KB → 147 B) 3ms
The footer, on omni exec and pipe mode. Input size, output size, and how long the
pipeline took.
[Partial signal]
The pipeline recognised some of the output but not all of it.
[OMNI: output truncated, 1200 of 4000 lines kept, 2800 dropped from the middle, omni retrieve 0000000000000000]
[OMNI: output truncated, 50000 of 180000 bytes kept]
The last cut before a reply leaves, at 50 KB. The line form keeps the start and the end, because a build or a test run puts its verdict last, and says how much of the middle went. The byte form is what you get when there is no line structure inside the budget, and it keeps the start only, so a single enormous line loses its tail. The handle is there when the dropped part was archived; without one, the marker still says what was removed.
[OMNI: 2 sensitive value(s) redacted]
Credential redaction on command output. A line that assigns a value to a sensitive name
had that value replaced with [REDACTED], and this counts the lines. Quotes and the
punctuation around the value stay, so a redacted line still parses.
The name decides, matched per underscore-separated word so PASSED and AUTHORS are not
caught: SECRET, TOKEN, PASSWORD, PASSWD, PASS, AUTH, CREDS, CREDENTIAL,
CREDENTIALS, DATABASE_URL, REDIS_URL, MONGO_URL, CLIENT_SECRET, ACCESS_KEY,
PRIVATE_KEY, and anything starting API_, AWS_, GITHUB_, ANTHROPIC_, OPENAI_
or GEMINI_. Some values are left alone because they hold no credential
whatever the name: an empty value, a shell expansion like $DB_PASSWORD, an unquoted call
or null.
KEY on its own is the weak pattern, since key= is ordinary code. Under it the value has
a say too, and a short lowercase identifier or a { expression passes. Under every other
pattern the value never argues, so hunter2 and decrypt("ghp_x") are cut. When in doubt
the value is cut: hiding a harmless value costs a re-read, and printing a real one cannot
be undone.
Redaction protects what the agent receives, not your disk. OMNI’s local trace store keeps the raw output, secret included, for seven days.
[OMNI: Re-injecting critical files due to Warning pressure]
When the context is under pressure (Warning or Critical), OMNI periodically puts up to
three of your pinned_files back in front of the agent, each cut to about 400 characters.
Nothing was removed; what follows the line is added.
Reading a percentage correctly
The worst bugs in this project’s history reported the highest reductions. A distiller that deletes the answer compresses beautifully.
So a large number is not on its own good news. omni diff is the check:
omni diff # the last command, raw against distilled
If a 99% saving turns out to have removed the file paths that were the answer, that is a bug worth reporting, and it is the exact class this project cares about most.
When there is no marker at all
Most of the time, and that is the pipeline working rather than failing. OMNI hands the output straight back whenever taking anything would be unsafe or would not pay:
- The payload is JSON, YAML, CSV or TSV. Never touched, on purpose.
- The command failed. A non-zero exit passes through verbatim.
- There was no noise to remove. A
kubectl get podstable is an enumeration where every row is a datum. - The output was too short to be worth a marker.
Getting content back
omni retrieve <handle>
Works on every host, with or without MCP. Agents with the MCP server wired can call
omni_retrieve themselves without asking you.
One boundary a handle cannot promise: the archive is a rolling 30 day window, so
omni retrieve on content older than that will not resolve. Verbatim traces are
shorter still at seven days.
Telling a real marker from a printed one
Markers appear in prose too. This page is full of them, so is OMNI’s source, and so is any bug report that quotes one. That matters if you are measuring whether OMNI was active on a run, because searching a transcript for the marker shape will find the examples as readily as the folds.
The handle is what separates them. Every worked example in this manual and in OMNI’s
own source uses one reserved value, 0000000000000000, which no real fold can ever
be assigned:
omni retrieve 0000000000000000 # exit 1, "the documentation example"
omni retrieve <handle-you-found> # exit 0 if OMNI really folded it
So the exit code answers the question, and a marker copied out of documentation cannot be mistaken for evidence that anything was shortened.
Seeing what it saved
omni stats
Everything on this page reads the same aggregation, so a figure in the share card cannot drift from the one in the report.
The report
omni stats # last 30 days, the default
omni stats --today # or --hour, --week, --month
omni stats --view detail # commands, routes, sessions, agents
omni stats --limit 0 # every command, not just the top ones
omni stats --view projects # broken down per project path
omni stats --json # machine readable
It leads with the bytes that never reached your model, then one line per engine:
5.1 MB never reached your model
folded 1.7 MB 41% of what it folded 906 folds
distilled 3.4 MB 48% of what it distilled 1,056 calls
left alone 0 by design 15,718 calls
The two percentages come from different populations and may never be added or
averaged. The distiller took 48% off the calls it distilled; the ledger took 41%
off the payloads it folded. Byte totals may be summed, which is what the headline
does. Until 0.7.7 the report read distillations alone and never read the ledger at
all, so the engine removing the most bytes was missing and the one percentage printed
was the distiller divided by 15,718 calls it had deliberately declined.
left alone reads 0, not the 31 MB that passed through. Those bytes are neither
a win nor a loss, and putting them in the savings column makes a reader think something
went missing.
The fold percentage covers the folds that record the payload they came out of.
payload_bytes arrived in a migration defaulting to zero, so older rows carry a saving
with no base. The report says how many of the folds it divided, and prints no percentage
at all when none of them do.
Session lifetime, the per-period table, top commands and the agent split all moved to
--view detail, which also names why each call was declined, out of passthrough_events,
which is what turns a 94% passthrough share from an accusation into an explanation.
What the numbers are counted in
Bytes, and they are counted rather than derived. Every absolute figure the report
prints is a byte total out of distillations, and every percentage is a ratio of two of
them.
They used to be tokens, which were those same byte counts divided by 3.6, a constant
calibrated against cl100k_base. That is GPT’s encoding, so the unit could not be
defended even though the arithmetic was sound. Percentages were never affected: the
divisor cancels in a ratio, which is why the reduction figures did not move when the
absolute ones did.
One block is still an estimate and says so. The context breakdown accumulates file sizes
from metadata, so Context Breakdown is exact for what it counts and is not a token
count in disguise.
If you parse --json, the commands[].tokens_saved field is now bytes_saved. It
held bytes under the old name for one release, which is a machine-readable surface
asserting the wrong unit, so it was renamed rather than left lying. Consumers have to
follow.
Reading it without fooling yourself
Split by agent_id before quoting anything. Rows recorded under terminal are
TTY bytes no model ever read. On one installation those were 73% of every byte OMNI
claimed to have saved. omni stats excludes them now, but the same trap waits for
anyone querying the database directly.
A high percentage is not automatically good. The worst defects in this project’s
history reported the highest reductions, because deleting the answer compresses very
well. Pair any number with omni diff on a real command.
A low aggregate is usually correct, and it is not the number to judge OMNI by. Most calls are handed back untouched because taking anything would be unsafe or would not pay: structured payloads, failed commands and enumerations all pass through by design. The per-command rows are where the work shows, so sort by what a class actually saved rather than reading the average.
The check a percentage cannot make
omni stats --rerun
Which distillers cost a re-run. If a distiller removes something the agent then has to go and fetch again, the reduction was not a saving, it was a deferral. Nothing in a byte count can see that.
Sharing it
omni stats --share # copy-pasteable summary of your own measured savings
omni stats --card # the same summary written as an image
Both come from your own database, which is the point. A ratio claim in someone else’s README cannot be verified before installing.
In a browser
omni dashboard # http://127.0.0.1:7717
omni dashboard --port 8080
Read-only, same database, binds loopback and nothing else.
Digging further
omni stats --view detail # per-command and per-route breakdown
omni query errors in last 5 commands
omni query warnings from cargo
omni query timeline today
omni patterns # errors that keep coming back
omni patterns --tool cargo
omni_history gives the same per-call rows to an MCP client, on the tiers that advertise
it. On a Full-tier host it costs more in the prefix than it earns, so the route there is
omni stats --view detail, which folds repeated commands into one row with a count. There
is no omni history subcommand; this page listed one until 0.7.4.
omni query speaks a small fixed query language rather than free text. The supported
forms are listed in its own help.
Querying the database directly
~/.omni/omni.db is plain SQLite and there is nothing stopping you.
Never read
sqlite3output through the Bash hook while investigating OMNI. The pipeline can fold the rows you are trying to count, and aLIKEfilter that catches the wrong rows has already put a wrong figure into a published issue. List the rows before quoting any aggregate over them, and setOMNI_PASSTHROUGH=1.
Memory across sessions
The same agent that reads too much also forgets everything the moment you restart it. OMNI carries three kinds of memory, and they are kept for different lengths of time on purpose.
What is kept, and for how long
| tier | what | kept |
|---|---|---|
| Permanent | project knowledge, recurring error patterns, engrams, goal memory | until you delete it, except goal memory which honours its own ttl_days |
| Working, 30 days | sessions, distillation rows, hot files, the archive, the event index, the ledger | rolling window |
| Verbatim, 7 days | execution traces and the session transcript | shorter on purpose, two orders of magnitude heavier per row |
The short answer to “will OMNI still know my project after a month away” is yes for
the conclusions and no for the raw bytes. The boundary that matters in practice:
omni retrieve on content archived more than 30 days ago will not resolve.
The ledger has one more way of forgetting that is not on a clock. At compaction its session half is dropped entirely, because compaction is where the agent stops holding what it was shown, and every “already shown” claim becomes false at the same moment. If folding seems to stop after a long session compacts, that is this, working. The project half survives, and The ledger explains the split.
Pinning a goal
omni goal set 'Migrate the billing service off the legacy queue'
omni goal show
omni goal clear
The scorer favours output related to the goal, and the agent is reminded of it on every prompt rather than drifting off task over a long session.
Facts worth keeping
omni remember 'The staging database ignores migrations run outside the deploy job'
Agents with MCP wired call omni_remember themselves, and pull facts back with
omni_recall, which is a semantic search across engrams, stored knowledge and
distillation history. Both are advertised on the MCP-only and Handoff-first tiers.
A Full-tier host such as Claude Code writes through omni remember in the shell OMNI
already hooks, and reaches omni_recall, which has no CLI equivalent, with
OMNI_MCP_TOOLS=all.
Store what is not derivable from the code: a decision and its reason, a gotcha, a constraint that no file states. Do not store what the repository already records.
Carrying a session across a restart
Session context is injected at session start, so a new agent knows which files were hot and what the last active error was. If the host closes or you switch tools, the project context is still there.
omni session --status
omni session --history
omni session --resume # resume an interrupted session
omni session --transcript
omni session --health
For moving to a machine or a host that shares no database, omni_handoff exports the
current session state as portable markdown you can paste into a new session. It is an
MCP tool only; the CLI subcommand was removed. It is outside the set advertised by
default, so set OMNI_MCP_TOOLS=all before reaching for it.
Engrams
Digests of finished subtasks, written as work completes rather than reconstructed later.
omni engram
omni engram --json
Knowledge that outlives a session
omni query errors in last 5 commands
omni patterns # errors that keep coming back across sessions
omni_insight ranks the same recurring issues project-wide, and is an MCP tool with no
CLI equivalent, and it is outside the default advertised set, so it needs
OMNI_MCP_TOOLS=all. It was listed in the block above as though you could run it.
What it cannot do
It is per machine. There is no sync, no server, and no shared store between people.
~/.omni/omni.db is the whole of it, and a remote archive was
explicitly not built rather than merely not built
yet.
When something looks wrong
Work down this page in order. The first three sections rule out the look-alikes, which is where most suspicions end.
First, is OMNI even involved
OMNI_PASSTHROUGH=1 <the command>
Identical output with and without it means OMNI did nothing. That is the end of the investigation, and it settles more cases than anything else here.
Then check which path ran, because they are not the same:
omni --version && ls -la "$(which omni)" # the installed binary, not your checkout
omni doctor
A closed issue still bites if the fix is unreleased.
Things that look like a bug and are not
Structured payload untouched. JSON, YAML, CSV, TSV, base64, terraform plans and
anything destined for jq pass through by design. Not a missed opportunity.
Negative savings on small output, roughly -1% to -4%. The marker costs more
than the compression saves on a short payload.
Most calls saving nothing. Expected. Taking anything would have been unsafe or would not have paid for its own marker.
File reads showing zero token savings in a session that read many files. OMNI’s surface on most hosts is shell output. Your agent’s own file-reading tool, skill files and the system prompt are outside it.
kubectl binary streams corrupting. SPDY does that with or without OMNI.
Shell word splitting and quoting. That is your shell.
The traps that produce false conclusions
Do not judge OMNI by output you read through OMNI. A cargo test read through the
hook once reported “1 failed” for a suite cargo itself called 398 passed. Redirect to
a file with OMNI_PASSTHROUGH=1 before making any claim about a result.
Do not grep the distilled output. Grepping hides the group headers that often make output lossless after all. A 116 line search result looked like it had dropped every filename until the full payload showed a filename header per group with matches indented under it. Read the whole thing.
Output is not deterministic against a warm database. Session history feeds the scorer, so the same command can distill differently on two runs. Isolate it:
OMNI_DB_PATH=/tmp/probe.db omni exec <command>
A failed reproduction is not a verdict. If a bug does not reproduce, read the
dispatch path in the source before concluding anything. A pipe that appeared to be
discarded turned out to be the pre-hook wrapping the entire command string, so
distillation landed upstream of the caller’s tail. Three hand-built reproductions
had come back clean.
Common problems
The hook is installed but nothing is distilled.
omni doctor checks the wiring. Then check the host’s tier: a Handoff-first or
MCP-only host cannot rewrite its built-in shell tool’s output at all. See
Supported agents.
Codex CLI does nothing after omni init --codex.
It runs only hooks it has been told to trust and skips the rest silently. Start
codex once and approve them under “Hooks need review”.
Warnings in the terminal that the agent never mentions. Hook rejections are recorded by the host as attachments that never enter the model’s context. The agent can genuinely believe the hook is fine while your screen fills with warnings. On Claude Code:
grep -c hook_error_during_execution ~/.claude/projects/<project>/<session>.jsonl
The attachment carries the host’s verbatim reason.
Commands feel slow. Expected, and it grows with database size rather than payload size: about 21 ms against a fresh database and 61 ms against a 205 MB one.
omni exec appears to hang.
A warm shared database serialises writes. Give it its own with OMNI_DB_PATH.
Reporting it
Worth reporting, in this order of importance:
- A false claim. OMNI asserting a result its input does not support: a success reported for a failure, a count that does not match the runner’s own.
- Lost signal. Something needed was dropped without a marker saying so.
- Noise. Verbose but harmless.
A good report carries the raw output and the distilled output side by side, including
the [OMNI Active] footer, the exact omni exec command, and omni --version. The
footer is often the point: the worst bugs here report the highest reductions.
Reproduce on a synthetic command where you can, so there is nothing to redact. Real terminal output carries hostnames, account ids and internal addresses more often than people expect.
Tracker: https://github.com/fajarhide/omni/issues
Discord: https://discord.gg/zHTuvZhF2M, if you would rather ask before filing.
Commands
Every subcommand, grouped the way omni --help groups them: by what you are trying
to do, not alphabetically.
omni <COMMAND> [FLAGS]
cmd | omni # distill any command's output through a pipe
Set up
| command | what it does |
|---|---|
init | Install OMNI into your agent, hooks and MCP |
doctor | Check the install is healthy, and fix what is not |
update | Upgrade to the latest release |
reset | Uninstall cleanly, keeping a backup of your config |
See what it saved
| command | what it does |
|---|---|
stats | How many tokens were cut, and from which commands |
retrieve | Print the content a marker archived, by its handle |
dashboard | The same numbers in a browser, on 127.0.0.1 |
diff | The last command’s output, before against after |
session | What this session has spent, and on what |
Tune it
| command | what it does |
|---|---|
exec | Run one command through OMNI, to see what it would do |
query | Search past distillations |
patterns | Errors that keep coming back |
Memory
| command | what it does |
|---|---|
remember | Save a fact for future sessions |
engram | Digests of finished subtasks |
goal | Pin a north-star goal so scoring favours it |
version | Version and environment details |
Hook entry points
Not for typing. These are what an agent host invokes, and they are documented in Hooks.
omni --pre-hook omni --post-hook omni --hook
omni --session-start omni --session-end omni --pre-compact
omni --mcp
A note on how flags are parsed
A match on the first argument routes the subcommand and hands the module the raw
env::args(), so every module parses its own flags and declares its own accepted
set. cli::check_flags rejects anything outside that set, which is what stops
omni stats --detial printing the default overview and exiting 0.
Per-command help is real and worth reading: omni <command> --help. Where this
reference and the help disagree, this records what the source accepts.
omni init
Installs OMNI into an agent: writes the hook configuration where that host reads it, and registers the MCP server.
omni init # interactive menu, or the current host when there is no terminal
omni init --claude
omni init --all
Idempotent. Running it again after an upgrade is the right move, not a risk.
With no flags
On a terminal, a menu. Without one, which is how an agent runs it, the menu cannot
be drawn, so omni init configures the host it is running inside and prints which
one that is. A host it cannot name from the environment, a plain shell included,
gets an error listing the flags rather than a guess: installing into a host nobody
asked for is the worse of the two failures.
Hosts
One flag per host. Each writes that host’s own configuration format in that host’s own location.
| flag | host |
|---|---|
--claude | Claude Code (Anthropic) |
--cursor | Cursor |
--zed | Zed |
--cline | Cline |
--roo, --roo-code | Roo Code |
--copilot | GitHub Copilot CLI |
--gemini | Gemini CLI |
--opencode | OpenCode |
--codex | Codex CLI |
--openclaw | OpenClaw |
--antigravity | Antigravity IDE, and generic webhook |
--hermes | Hermes Agent |
--vscode | VS Code (MCP) |
--pi | Pi Agent |
What each host actually lets OMNI do differs a great deal. See Supported agents before assuming a flag buys shell distillation.
Modes
| flag | effect |
|---|---|
--all | Every host above. Also writes .vscode/mcp.json in the current directory. |
--hook | Hooks only, no MCP registration |
--mcp | MCP registration only, no hooks |
--status | Report what is currently installed, change nothing |
--uninstall | Remove OMNI’s hooks and MCP server |
--help, -h | Help |
After running it
omni doctor
Always. init reports what it wrote; doctor reports whether the host is reading it.
Codex CLI needs one more step. It runs only hooks it has been told to trust and
skips the rest without a word. Start codex once and approve them under “Hooks need
review”. omni doctor fails until you do.
Notes
--all is the only flag that writes into the current directory. Everything else
touches your home configuration only.
An unrecognised host flag does not always fail loudly: a misspelled one has been known to run the interactive default and exit 0 while installing nothing that was asked for. Read what it printed.
omni doctor
Checks that the installation is healthy, and repairs what it can.
omni doctor
omni doctor --fix
It covers the binary’s version and accessibility, the configuration directory and database, hook installation per host, MCP server registration, and signal loading.
Flags
| flag | effect |
|---|---|
--fix | Repair configuration and integration issues automatically |
--detail | Print every integration row, not only the ones needing attention |
--json | Machine readable |
--help, -h | Help |
Reading the output
Host tiers. doctor prints the tier for every installed host, and the tier is
the honest ceiling on what OMNI can do there. A Handoff-first or MCP-only host cannot
rewrite its built-in shell tool’s output, so no amount of pipeline work will move its
distillation numbers. See Supported agents.
[N UNRELEASED]. A build compiled from a tree whose CHANGELOG.md has entries
under ## [Unreleased] says so, and tells you to cut a tag. On a release build there
is no such line. This exists so a binary that was tagged without moving the changelog
entries accuses itself rather than shipping quietly.
Live retention counts. How much is in each memory tier right now.
What it does not check
That the host is actually applying the rewrite. doctor verifies the configuration is
where the host reads it, which is not the same as the host honouring it. The proof for
that is a distillation row in the database under your host’s agent_id, or the host’s
own session transcript.
On Claude Code, a hook payload the host rejected is recorded as an attachment that never reaches the model, so the agent can believe everything is fine while your terminal fills with warnings:
grep -c hook_error_during_execution ~/.claude/projects/<project>/<session>.jsonl
omni stats
Token savings analytics, read from your own database.
omni stats
Leads with the bytes that never reached your model, then one line per engine, each percentage against its own base, then the command classes those bytes came from. The aggregate below them mixes in every call OMNI deliberately declined, so it is a diagnostic for one host’s pipeline rather than a product claim.
Every view draws the same frame, OMNI · <view> · <window> between two rules, and
--view context carries no window because it reads the live session rather than a
period.
Flags
| flag | effect |
|---|---|
--since <window> | hour, today, week, month (default), all |
--view <name> | summary (default), detail, projects, context, rerun, folds, share |
--limit <n> | Rows in a table view, default 10, 0 for all. Read by detail, projects and rerun; a table it cuts says how many rows it hid |
--json | Machine readable, scoped by --since |
--card | Write the summary as an image, sized for social posts |
--help, -h | Help |
Every earlier spelling still resolves: --detail, --today, --day, -d, --week,
-w, --month, -m, --hour, -H, --all-commands, --project, --context,
--rerun, --share and --view commands. They are not listed because there is one way
to say each thing now, and they print no deprecation notice: the rename was ours, not
yours. --view commands is in that list rather than the table above because it renders
the detail view and always did.
--json and --card are output formats rather than views. --card outranks everything,
since naming it can only mean writing the file; --json outranks --view, since there is
one machine-readable report and it is not per view. Both used to be read as views, which is
how --view detail --card came to write no image.
--view folds is the ledger’s own calibration
A fold is a claim: the reader does not need these bytes. A retrieve on that marker’s handle is the reader disagreeing. This view puts the two side by side, per marker shape, so the guard with the highest rate is the one to look at.
omni stats --view folds # the last 30 days
omni stats --view calibration --since week
calibration is the same view under the word most readers reach for first.
There is no threshold and no verdict colour in it. No bar has been measured yet, and a number that looks judged when nothing judged it is the defect this project exists to fight. A store written before markers were recorded says so rather than printing a confident zero.
--rerun is the one to know
Reduction percentage cannot tell you whether a distiller removed something the agent then had to fetch again. If it did, the reduction was a deferral, not a saving. This flag is the check that percentage cannot make.
Traps
Terminal rows are not tokens. Output written to a TTY is read by a human, not a
model. On one installation those rows were 73% of every byte OMNI claimed to have
saved. stats excludes them now, and so does the benchmark harness, but anyone
querying ~/.omni/omni.db directly has to filter by agent_id themselves.
A high number deserves suspicion. The worst defects in this project reported the
highest reductions, because deleting the answer compresses very well. Pair any figure
with omni diff on a real command.
A low aggregate is usually right, and it is not the number to judge OMNI by. Most calls are handed back untouched by design, so read the per-class rows to see where the work actually happened.
--share and --card cannot drift from the report. Both read the same
aggregation as omni stats itself, which was a deliberate choice after an earlier
version computed them separately.
omni exec
Runs one command through the full pipeline and prints the result, with a footer showing what it cost.
omni exec cargo test
cargo test: 411 passed, 1 failed
FAILED ledger::tests::renders_identical_bytes_for_identical_state
[OMNI Active] ⏺ 93.7% reduction (2.3 KB → 147 B) 3ms
This is the harness every bug report in this project is asked to use, because it takes
the host out of the picture. If a corruption survives omni exec, it is OMNI.
The argument form is exact
omni exec cargo test # correct
omni exec -- cargo test # fails: No such file or directory
omni exec 'cargo test' # works, single-string form
omni exec sh -c 'a; b' # works, split-argv form
The -- form is the one people reach for and the one that does not work.
Flags
| flag | effect |
|---|---|
--session <id> | Forward a host session id, which is what scopes the ledger |
--agent <id> | Record the run under a given agent_id |
--help, -h | Help |
Both are what the pre-hook uses when it rewrites a command into omni exec.
--session is worth knowing when you are investigating ledger behaviour: it is the
only way to drive two distinct sessions by hand and see the difference between an
already shown fold and a not shown here fold.
Isolate the database while probing
Output is not deterministic against a warm database, because session history feeds the scorer. Give each probe its own:
OMNI_DB_PATH=/tmp/probe.db omni exec <command>
A warm shared database also serialises writes, which is the usual reason omni exec
appears to hang.
Related
omni diff shows the same before and after for the last command the hook
processed, which is what you want when the interesting command already ran.
omni retrieve
Prints the content a marker archived.
omni retrieve <handle>
The handle is the 16 characters inside a marker:
[OMNI: 406 lines omitted, omni retrieve 0000000000000000 for full output]
[OMNI: 40 lines already shown, omni retrieve 0000000000000000]
It returns the original bytes. Not a summary, not a re-run of your command, and not an approximation.
Works on every host, in any session, whether or not MCP is wired. Agents with the MCP
server registered call omni_retrieve instead and never have to ask you.
What can go wrong
The handle does not resolve. The archive is a rolling 30 day window, so content older than that is gone. Verbatim execution traces are pruned sooner still, at seven days.
A handle that fails to resolve inside the window is a serious bug rather than an inconvenience, because a marker promising retrievable content is the one thing this mechanism cannot get wrong. Report it.
You typed the marker text, not the handle. Only the hex, no brackets, no prefix.
Why it can promise this
A run is archived before its marker is written, and a failed archive leaves the run verbatim rather than producing a marker. So a handle you can see is a handle whose content exists. That ordering was a fix, not the original design: an earlier version returned a key even when the write had failed.
omni session
Session state: what this session has spent, on what, and how to carry it across a restart.
omni session --status
Flags
| flag | effect |
|---|---|
--status | Current session status |
--history | Recent session history |
--health | Visual session health dashboard |
--transcript | Transcript of the recent session |
--clear | Reset the current session |
--continue | Continue a stale session |
--resume | Resume an interrupted session |
--inject | Emit session context for an agent to consume |
--json | Machine readable |
--help, -h | Help |
omni sessions is accepted as an alias.
What a session is here
The scope key is the host’s session id, not an internal timestamp. That distinction was a real defect: an internal wall-clock id once covered 16 projects in one value, which would let the ledger tell one session it had been shown output that went to another.
That is also why omni exec takes --session: without a forwarded host id there is
no ledger scope, and for a while the exec path therefore ran no ledger at all.
Continuity across a restart
Session context is injected at session start, so a new agent knows which files were hot and what the last active error was. Restarting your editor or switching hosts does not lose the project context.
--inject is the manual form of that, for a host wired to consume it.
For crossing to a machine that shares no database, use the omni_handoff MCP tool,
which exports the state as portable markdown. The CLI subcommand of that name was
removed; the MCP tool is unchanged. It is outside the default advertised set, so it
needs OMNI_MCP_TOOLS=all.
Retention
Sessions are in the 30 day working tier. The verbatim transcript is in the 7 day tier, because it is two orders of magnitude heavier per row.
Everything else
The commands that need a paragraph rather than a page.
update
omni update
Fetches the latest release from GitHub and upgrades. Homebrew installations only at present; other install methods upgrade through their own channel.
Re-run omni init afterwards if the release notes say the hook contract changed.
reset
omni reset # interactive menu
omni reset --all # every integration, and offers to wipe omni.db
omni reset --claude # one host
Per-host flags mirror init: --claude, --cursor, --zed, --cline,
--roo / --roo-code, --copilot, --gemini, --opencode, --codex,
--antigravity, --hermes, --pi.
--all is the only one that offers to delete your database, and it asks first. It
keeps a backup of the configuration it removes.
dashboard
omni dashboard
omni dashboard --port 8080 # default 7717
The same numbers omni stats prints, in a browser. Read-only, reads the same
database, and binds 127.0.0.1 and nothing else. Ctrl-C stops it.
diff
omni diff
The last command’s output, raw against distilled. The fastest way to build trust in what OMNI is doing, and the first thing to run when a result looks wrong.
query
omni query errors in last 5 commands
omni query warnings from cargo
omni query context for src/main.rs
omni query timeline today
omni query timeline today --json
A small fixed query language over distillation history, not free text. Four forms are
supported and they are the four above. --json for machine-readable output.
patterns
omni patterns
omni patterns --tool cargo
Errors that keep coming back across sessions. --tool <name> scopes to one tool.
Useful for the question “have I hit this before”, which is the one a fresh session cannot answer on its own.
remember
omni remember 'The staging database ignores migrations run outside the deploy job'
Stores a fact in persistent memory, retrievable later through omni_recall or the
session context injection.
Worth storing: a decision and its reason, a gotcha, a constraint no file states. Not worth storing: anything the repository already records.
engram
omni engram
omni engram --json
Digests of finished subtasks, written as work completes.
goal
omni goal set 'Migrate the billing service off the legacy queue'
omni goal show
omni goal clear
Pins a north-star goal. The scorer favours output related to it, and the agent is
reminded of it rather than drifting over a long session. Goal memory honours its own
ttl_days rather than the standard retention tiers.
set is the default subcommand, so omni goal 'some text' also works.
version
omni version
omni version --json
Version and environment details: build date, git hash, and the paths OMNI resolved for its configuration and database. Worth including in any bug report.
Environment variables
Every OMNI_* variable the binary reads. Grouped by why you would reach for one.
The one you will actually use
| variable | effect |
|---|---|
OMNI_PASSTHROUGH=1 | Skip the pipeline entirely. Raw output, every time. |
This is the first thing to reach for when you suspect OMNI changed something it should not have, and the thing to set when you need exact bytes from a file read through your shell. Identical output with and without it means OMNI was not involved.
Where things live
| variable | effect |
|---|---|
OMNI_HOME | Puts the whole tree, config and data, in one directory |
OMNI_CONFIG_HOME | Config directory, when you want it split from data |
OMNI_DATA_HOME | Data directory, likewise |
OMNI_DB_PATH | Path to the SQLite database |
OMNI_TRANSCRIPT_DIR | Where session transcripts are written |
OMNI_DB_PATH earns its own note. Point it at a scratch file whenever you are probing
OMNI’s behaviour by hand:
OMNI_DB_PATH=/tmp/probe.db omni exec <command>
Output is not deterministic against a warm database, because session history feeds the
scorer, and a shared warm database serialises writes, which is the usual reason
omni exec looks like it has hung. It is also required when running the test suite
against a live installation.
Commands run through the MCP server
| variable | effect |
|---|---|
OMNI_RUN_TIMEOUT_SECS | How long omni_run waits for a command. Default 60. |
The default sits below every host MCP timeout we know of, so a stalled command comes back as a sentence naming itself rather than the host’s idle-timeout error. Raise it when a build legitimately takes longer, and remember the host has a deadline of its own: Cursor’s is 120 seconds, and nothing OMNI does can extend it.
Retention
| variable | effect |
|---|---|
OMNI_TRACE_RETENTION_DAYS | Days of verbatim execution traces. Default 7. |
OMNI_SESSION_TTL | Session time to live, in minutes |
Hold the trace window open while a measurement is in flight:
OMNI_TRACE_RETENTION_DAYS=90 ...
Seven days is why no published benchmark figure can be re-derived a week after it was measured, including by the people who published it. Raise it before you start, not after.
Context pressure
| variable | effect |
|---|---|
OMNI_CONTEXT_WINDOW | Context window size hint, in tokens |
OMNI_PRESSURE_WARN | Warning threshold, as a share of the window |
OMNI_PRESSURE_CRITICAL | Critical threshold |
OMNI estimates how full the session’s context is and injects a warning past these thresholds. Set the window to match the model you are actually running.
Session behaviour
| variable | effect |
|---|---|
OMNI_FRESH | Force a fresh session rather than continuing one |
OMNI_CONTINUE | Set internally by the dispatcher to mark a continued session |
OMNI_SUBAGENT=1 | Sub-agent mode |
OMNI_AGENT_ID | Agent identity, recorded on every row |
OMNI_AGENT_ID is the one to understand before quoting any number. Every distillation
row carries it, and rows recorded under terminal are TTY bytes no model ever read.
Blending those with hook rows once made 73% of a published saving fictional. When
several agents run side by side, give each its own id.
Loops
| variable | effect |
|---|---|
OMNI_LOOP_ID | Loop identifier. Alphanumeric and dash, 64 characters. |
OMNI_LOOP_GOAL | Goal string, 500 characters, no shell metacharacters |
OMNI_LOOP_BUDGET | Token budget per iteration, up to 10M |
OMNI_LOOP_ITERATION | Current iteration number. Default 0. |
See Loop engineering.
Output
| variable | effect |
|---|---|
OMNI_QUIET=1 | Suppress the stderr stats line in pipe mode |
OMNI_OUTPUT_JSON | JSON output from the pipe path |
OMNI_EXPORT_CSV | Export session data as CSV at session end |
Build and internal
Not for setting by hand. Listed so that seeing one in a stack trace or a generated config is not a mystery.
| variable | set by |
|---|---|
OMNI_BIN | Written into the generated Hermes plugin, naming the binary path |
OMNI_CMD | The command being processed, falling back to CMD |
OMNI_GIT_HASH, OMNI_BUILD_DATE | Stamped at build time, reported by omni version |
OMNI_UNRELEASED_ENTRIES | Computed by build.rs from CHANGELOG.md, so a binary built from an untagged tree says so in omni doctor |
OMNI_PI_PACKAGE_SOURCE | Package source for the Pi agent integration |
OMNI_DATA_HOME_UNSET_FOR_TEST | Test fixture only |
Benchmarking
| variable | effect |
|---|---|
OMNI_BENCH_DB | Database to replay from |
OMNI_BENCH_ALL=1 | Replay the wider population including terminal output |
OMNI_BENCH_RTK | Path to an rtk binary, adding the head-to-head arm |
OMNI_BENCH_ALL exists so the harness can say which population it measured rather than
leaving it to be inferred. Including terminal output printed 79.1% where the
model-facing population printed 43.3%, on the same data.
MCP tools
omni init registers OMNI as an MCP server, which gives the agent tools it can call
itself without going through you. This page describes all 25. Your host is told about
a subset.
What your host is told about
Tool definitions sit in the prefix of every request, so a tool nobody calls is re-read on every request of every session rather than paid for once. OMNI advertises the set your host’s tier can use:
| tier | advertised |
|---|---|
| Full | omni_retrieve, omni_explain_savings |
| Handoff-first, and any host OMNI does not recognise | those two, plus omni_remember, omni_recall, omni_run, omni_find_noise, omni_context_breakdown, omni_history |
| MCP-only | omni_remember, omni_recall, omni_retrieve, omni_knowledge, plus omni_run: no hook means the host’s own tool output is never rewritten, so omni_run is the only path by which the model reads less |
Full-tier hosts get the shortest list because they are the ones whose shell OMNI already hooks. Measured across 256 recorded sessions, those two tools carry 138 of the 149 calls in 467 bytes, while the other six cost 1,631 bytes for 11 calls in the life of the corpus. A Handoff-first host never has its built-in tool output rewritten, so MCP is the only door OMNI has there and nothing is priced away.
On Claude Code the shortest list is still not free, and the bytes are the smaller half of why: the host discards the whole prompt cache when an MCP server connects or disconnects with its tools loaded. Supported agents has what that means and how to run the hooks without it.
An unadvertised tool is not callable on that tier either, so the way back is the CLI or the override:
| tool | on a Full-tier host |
|---|---|
omni_run | omni exec <command> |
omni_remember | omni remember '<fact>' |
omni_context_breakdown | omni stats --view context |
omni_history | omni stats --view detail, which folds repeated commands into one row with a count instead of listing every call |
omni_recall, omni_find_noise | no CLI equivalent; OMNI_MCP_TOOLS=all |
OMNI_MCP_TOOLS=all advertises all 25, and omni doctor says which set is in force and
which host it resolved:
MCP tools: 2 of 25 advertised to claude_code (OMNI_MCP_TOOLS=all restores the rest)
Confirm the list against your own binary rather than this page:
{ echo '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"p","version":"1"}}}'
echo '{"jsonrpc":"2.0","method":"notifications/initialized"}'
echo '{"jsonrpc":"2.0","id":2,"method":"tools/list","params":{}}'; } \
| omni --mcp | tail -1 | jq -r '.result.tools[].name'
Getting content back
| tool | what it does |
|---|---|
omni_retrieve | Retrieve full content a marker omitted, by its handle |
omni_run | Run a shell command and return distilled output |
omni_signal_extract | Extract signal from raw text, without the hook pipeline |
omni_run matters most on hosts that cannot rewrite their built-in shell tool. There,
it is the only path to distilled output, which is why omni init --cursor installs a
rule telling the agent to prefer it.
Understanding what OMNI did
| tool | what it does |
|---|---|
omni_explain_savings | Route, filter, input and output bytes, savings % per recent command |
omni_history | Recent distillations with per-call savings and ratios |
omni_context_breakdown | Token breakdown by source for the current turn |
omni_density | How much signal against noise in a piece of text |
omni_budget | Token budget usage and compression efficiency for this session |
These four are the right answer to “is OMNI helping here”. Pull the numbers rather than forming an impression.
Memory
| tool | what it does |
|---|---|
omni_remember | Store a decision, gotcha or constraint |
omni_recall | Semantic search across engrams, knowledge and distillation history |
omni_knowledge | Query or store cross-session project knowledge |
omni_insight | Top recurring issues and error patterns across the project |
omni_adaptive_insights | Retrieval patterns, as a judgement on distillation effectiveness |
omni_handoff | Export session state as portable markdown, no network needed |
omni_handoff is MCP only. The CLI subcommand of that name was removed.
Session and search
| tool | what it does |
|---|---|
omni_session | Session state: status, context, clear |
omni_search | Search this session’s history |
omni_query | Query distillation history with the fixed query forms |
omni_agents | Other agents currently active on this project |
Tuning
| tool | what it does |
|---|---|
omni_find_noise | Analyse recent raw traces for repetitive noise |
Advisory only, and the learner treats “repeated” as “noise”. It has suggested stripping
^metadata:,^spec:, code fences and^\[stderr\], which are structure and the error channel. Never paste its output anywhere without reading it line by line.
Loops
| tool | what it does |
|---|---|
omni_loop_status | One-call status check for an orchestrator before each iteration |
omni_loop_memory | Read and write loop memory that survives session restarts |
omni_set_loop_context | Update loop context dynamically |
omni_budget_status | Budget status for this iteration. Call before expensive work. |
omni_verify | As a checker sub-agent, evaluate the maker agent’s recent work |
See Loop engineering.
A tool that is not one
omni_auto_noise appears as a string in the server source and is not a tool. It is
a filter name passed to the TOML generator. Calling it returns -32602 tool not found.
It has been miscounted before: a source grep for "omni_*" returns 27, and 27 is
therefore wrong wherever it appears. Run the tools/list call above for the count.
What left the surface
omni_context was advertised until 0.7.7 and had never been called once across
253 recorded sessions, so it cost 189 bytes in the prefix of every request for a
capability nobody reached for. It is omni context <file> now, which costs nothing
per request:
omni context src/ledger/mod.rs
An agent can still reach it through omni_run.
Hooks
The entry points an agent host invokes. You never type these; omni init writes them
into the host’s configuration.
| entry point | when the host calls it |
|---|---|
omni --pre-hook | Before a tool runs |
omni --post-hook | After a tool produces output |
omni --hook | Universal dispatcher, for hosts with one hook slot |
omni --session-start | Session begins |
omni --session-end | Session ends |
omni --pre-compact | Before the host compacts the conversation |
omni --mcp | Run as an MCP server over stdio |
cmd | omni | Pipe mode, no host involved |
One call, both hooks
The shell runs whatever the pre-hook handed it and never knows OMNI exists. Only the reply is rewritten, which is why nothing here can change what your command did.
What each one does
Pre-tool decides whether a command should be routed through OMNI at all, and can
rewrite it into omni exec. That rewrite wraps the entire command string,
redirection included, which is why a matched command’s log file on disk can turn out
to be the distilled version. Break the prefix (env cargo test) when you need the raw
log.
Post-tool is the main event: the raw output arrives, the pipeline runs, and the distilled result is handed back for the host to substitute.
Post-tool-failure exists because a failed command must pass through verbatim, and
hosts disagree wildly about how they say a command failed. Claude Code sends a plain
string, Error: Exit code N. Others carry structured error flags. Reading only one
shape is a bug this project has had.
Session start injects project context: hot files, the last active error, stored knowledge, the pinned goal.
Session end writes the summary and can export CSV.
Pre-compact is the host’s warning that the conversation is about to be shortened.
Two doors into one pipeline
post_tool and pipe are separate entry points that run the same stages, and keeping
them in step has been a recurring source of bugs. Three separate fixes each corrected
one copy and left the other. The ledger stage existed in post_tool for a release
before pipe had it at all, so a command the pre-hook rewrote into omni exec got
the filters and nothing else.
If you are changing pipeline behaviour, change both, or check why not.
Why it never crashes your agent
Every hook runs inside catch_unwind, at the highest entry point. A panic in one
stage costs that distillation, not the session. A database that will not open costs
session context, not the pipeline.
That is the fail open rule, and it has one sharp edge worth stating: failing open
means handing back the raw bytes. It does not mean emitting a cheerful summary. A
distiller that parsed nothing returning 0 tests passed is failing closed, and
confidently.
What a host has to do for any of this to matter
Register the hook, and then honour what it returns.
The second half is not guaranteed. OMNI once emitted its distilled output under a key Claude Code ignores, so nothing was applied on that path for two releases while OMNI recorded a saving and printed a footer for each one. The fix corrected the key and left the value shape wrong, and the symptom survived unnoticed.
Two things that taught, both non-obvious:
- The rewrite is validated against the host tool’s own output schema, one shape per tool. There is no universal shape.
- The fields are independent. A rejected rewrite still lets the context message through, so the savings footer prints for a distillation that was reverted.
So the proof that a hook is working is not the footer and it is not omni stats. It
is the host’s own session transcript:
grep -c hook_error_during_execution ~/.claude/projects/<project>/<session>.jsonl
A warning you can see is not a warning the agent can see. Those attachments never enter the model’s context, so an agent can tell you the hook is fine while your terminal fills with rejections.
Testing a hook by hand
Feed it a payload directly rather than guessing which path ran:
echo '<host payload json>' | omni --post-hook
omni exec and the post-hook route differently, so a result from one is not evidence
about the other.
The payload shape, which differs per tool
Getting this wrong fails silently and identically: the hook exits 0, prints nothing, and a
probe reads that as 0.0% saved. There is no error to notice, so a distiller that is in
fact cutting 96% can be written off as not firing.
Bash puts the output at the top of tool_response:
{ "session_id": "s1", "tool_name": "Bash",
"tool_input": { "command": "cat server.log" },
"tool_response": { "content": "line one\nline two\n" } }
Read wraps it in file, and the extra keys are not decoration. startLine is what the
host counts cat -n numbering from, so a fold that removes lines above the survivors has
to move it:
{ "session_id": "s1", "tool_name": "Read",
"tool_input": { "path": "notes.txt" },
"tool_response": { "file": { "filePath": "notes.txt", "content": "...",
"startLine": 1, "numLines": 40, "totalLines": 400 } } }
The reply goes under hookSpecificOutput.updatedToolOutput, and Claude Code validates it
against the host tool’s own output schema. A wrapped Read gets a file reply back,
which is the shape the host accepts.
A bare tool_response.content does not get a bare reply. Verified rather than assumed:
it comes back as {status, result}, which is OMNI’s own shape and is what #187 was about.
So the bare form is fine for asking what the ledger did, and its reply is not what a real
Read would accept.
Both Read shapes are real and they reach different stages. A Read payload written
with a bare tool_response.content is accepted and reaches the ledger, while
tool_response.file.content reaches the readfile distiller and the startLine
adjustment. Neither is wrong; they answer different questions. A probe aimed at one and
built on the other returns a clean nothing and looks like a verdict.
Supported agents
Which host you run decides what OMNI can do, and the ceiling is the host’s, not the pipeline’s. This page is worth reading before judging whether OMNI is earning its place.
The tiers
| tier | hosts | what you get |
|---|---|---|
| Full | Claude Code, Codex CLI, Gemini CLI, OpenClaw, Hermes, Pi, Aider (pipe) | The host applies OMNI’s rewrite, so the model reads distilled output from its own built-in tools. |
| Handoff-first | Cursor, Windsurf | The host cannot rewrite built-in tool output. omni_run distils anything routed through it, and omni init --cursor installs the rule that makes the agent reach for it. |
| MCP-only | Cline, Roo, OpenCode, VS Code, Zed, Copilot, Antigravity | Memory, recall and session state, plus omni_run. The host’s own tool output is never rewritten, so omni_run is the only path by which the model reads less. |
omni doctor # prints the tier for every installed host
Savings are only ever counted where the model actually received less. A host that cannot apply the rewrite will not move the distillation numbers however good the filters get, and claiming otherwise would be the same defect as a distiller reporting a saving it did not make.
Installing for each
omni init --claude omni init --cursor omni init --zed
omni init --cline omni init --roo omni init --roo-code
omni init --copilot omni init --gemini omni init --opencode
omni init --codex omni init --openclaw omni init --antigravity
omni init --hermes omni init --vscode omni init --pi
omni init --all
Host-specific notes
Codex CLI runs only hooks it has been told to trust, and skips the rest without a
word. After omni init --codex, start codex once and approve them under “Hooks need
review”. omni doctor fails until you do. This has bitten before: Codex ran zero
hooks for a whole release while everything looked correctly installed.
Cursor cannot rewrite its built-in shell tool’s output. Intercepting the shell by denying execution and returning output as a hook message is technically possible and was rejected: it tells the agent its command was blocked, loses the exit code, moves execution semantics into OMNI, and bypasses the host’s approval flow.
Claude Code matches more than Bash. The post-tool matcher is
Bash|Read|Grep|WebFetch, which is what finally let the file-read, search and fetch
distillers run at all. Three of them had been fully written and tested and had never
executed on a real session.
It is also the host where omni init installs two things and only one of them does the
work. The hooks are what shortens output. The MCP server is a convenience that puts
omni_retrieve and omni_explain_savings where the agent can call them, and on this
host it is a trade against the prompt cache: those two definitions are 471 bytes of JSON
in the prefix of every request, and the host discards the whole cache when an MCP server
connects or disconnects with its tools loaded, which a server process can do by exiting
and reconnecting mid-session without you touching anything. omni init --claude
registers it. The plugin route adds no tool definitions at all. To keep the hooks and drop
the trade, install that half on its own:
omni init --hook # hooks, no MCP registration
omni init --mcp # the MCP server on its own, if you change your mind
omni doctor then reports the MCP server as not registered rather than as a fault, and
neither it nor --fix puts it back (#757). Naming the host, omni init --claude, still
installs both.
OpenClaw is Full on a later turn, not the current one. Its tool_result_persist
hook rewrites the tool result OpenClaw persists, so the model reads the distilled bytes
every time the transcript is re-read, while the turn that ran the command still sees the
raw output. That is where a tool result’s cost sits anyway, since it is re-read many
times, but it is a narrower Full than Claude Code’s.
Hermes hands OMNI every tool result, not only the terminal, which is the widest reach of any host here: it is what runs the file-read, search and fetch distillers on a host that has them. It also has its own integration page: Hermes Agent.
Windows is supported. Paths, line endings and the .exe suffix are handled, and
the CI matrix includes windows-latest.
Several agents at once
Give each its own identity so the numbers stay separable:
OMNI_AGENT_ID=claude ...
OMNI_AGENT_ID=cursor ...
omni_agents reports which agents are currently active on the project. Every
distillation row carries the id, and any figure that blends them is describing a mix
rather than a product.
Adding a host
The agent modules live in src/agents/, one file per host, and each writes that
host’s own configuration format in that host’s own location. The pattern is small and
mostly mechanical.
The part that is not mechanical is verification. “Provider unreachable” is not a reason to leave a hook path unverified: serve the API and fake only the model. Hooks that had never run in production have been found on three hosts by doing exactly that.
Architecture
Local, deterministic, and the same input always produces the same output. Nothing leaves the machine at any stage.
src/
├── main.rs CLI dispatch, and the single command list
├── lib.rs library re-exports, so the crate is testable
├── paths.rs path resolution
├── agents/ one file per host: claude, cursor, codex, hermes, pi, …
├── cli/ one file per subcommand
├── distillers/ 11 content filters
├── graph/ code graph indexing
├── guard/ safety, limits, trust bounds, env hygiene
├── hooks/ the entry points, and the dispatcher that routes them
├── ledger/ cross-turn line dedup
├── mcp/ the MCP server and its 25 tools
├── pipeline/ scorer, collapse, registry, format gate
├── session/ tracking, learning, adaptive thresholds
├── store/ SQLite and transcripts
└── util/ command families, token estimation
About 46,000 lines of Rust.
Design rules that the code actually enforces
Library first. main.rs is a thin entry point. Logic lives in lib.rs and its
submodules, so OMNI can be tested as a crate.
Single source of truth. Command-to-behaviour mapping is centralised in
pipeline/registry.rs. Duplicated matches!(cmd, ...) blocks in distillers and
scorers are the thing that rule exists to stop. Magic numbers live as named constants
in pipeline/mod.rs or guard/limits.rs.
IO separated from logic. Scoring and filtering are pure functions over &str.
They do no filesystem or network work.
Panic safety. Every hook runs inside catch_unwind at the highest entry point, so
one failing hook cannot take down the host agent.
Graceful degradation. If the database will not open, hooks still work, without session context.
Deterministic. No randomness anywhere. The ledger’s handle is a content address
and carries no timestamp, because an earlier {timestamp}_{hash} form made 4 of 73
repeated inputs emit different bytes.
The database
One SQLite file, ~/.omni/omni.db.
| table | holds |
|---|---|
sessions | session state, task and domain hints |
distillations | every distillation: filter, bytes in and out, route, score, latency, agent |
file_access | hot file tracking per session |
rewind_store | compressed content by SHA-256, with a retrieval counter |
session_events | FTS5 full-text index |
ledger_lines | which lines a scope has been shown |
ledger_folds | one row per marker issued: which scope, and which agent’s bytes it drew on |
passthrough_events | telemetry for commands that bypassed the pipeline |
unhandled_tools | tools OMNI does not support natively yet |
execution_traces | raw input and distilled output per command |
session_summaries | per-session metrics |
project_knowledge | cross-session semantic memory |
agent_sessions | shared state across multiple agents |
passthrough_events and unhandled_tools are worth knowing about: they are how a
coverage gap becomes visible instead of staying a guess.
Cross-platform
The CI matrix includes windows-latest, and four rules keep it green:
- No hardcoded separators.
PathBufandpush, never/or\\. - No exact
\nmatching in assertions. Use.lines(), or normalise\r\nfirst. The ledger splits withsplit_inclusive('\n')rather thanlines()for exactly this reason:lines()drops the terminator, so rebuilding with\nwould silently rewrite every CRLF payload on Windows. - No assuming the binary is
./omni. Usestd::env::consts::EXE_SUFFIX. - Environment variables are case-insensitive on Windows. Use
eq_ignore_ascii_casewhen readingstd::env::vars().
Build
cargo build --release
cargo test --all
make ci
The toolchain is pinned in rust-toolchain.toml, currently 1.97.0, and the pin is
load-bearing. A release once produced no binaries at all because release.yml asked
for stable per cross target while the pin said otherwise, and every cross-compile
died with can't find crate for core before compiling a line. ci.yml stayed green
throughout, because it only builds host-native.
The pipeline, stage by stage
Read → Guard → Score → Distill → [Collapse] → Ledger → Route → Persist
The order is fixed. This page is about what each stage may and may not do, which is where the bugs live.
The brackets around Collapse are the part people get wrong, including this page until recently. It is a fallback, not a step.
Guard
pipeline::format::sniff classifies the payload. Some(Structured) ends the pipeline
and the bytes pass through.
Four kinds: JSON, YAML, CSV, TSV. The bias is deliberate: bracketed but unparseable, truncated, or comment-bearing JSON all count as structured, because compression cannot repair a malformed payload and can certainly make it worse.
Above a size threshold, bracket shape alone decides JSON, since a full serde_json
parse would blow the latency budget.
The YAML sniffer skips lines introduced by a block scalar indicator (key: |). One
embedded ConfigMap once sank a 608-line kubectl kustomize manifest: the block’s
contents carried no key:, so the sniff said “not YAML” and the manifest went down
the lossy path.
Score
scorer::score_with_command(input, cmd, session) returns Vec<OutputSegment> with
tiers: Critical 1.0, Important 0.7, Noise 0.1.
semantic::is_critical tiers the block before any distiller runs. When a distiller
behaves oddly, probe the segment tiers first; the tier may already have decided the
outcome, and a guard added to the distiller will not move it.
Pure function. No IO.
Collapse
Runs of near-identical lines become [N similar lines collapsed].
It runs after Distill, and only when Distill did not earn its keep. Both hooks
score and distill the raw content, then ask beats_guardrail; only if that fails
does the collapsed form get used instead. A distiller therefore sees the original
text, never collapse markers.
This page said the opposite until 0.7.4, which was true before #116 and wrong for two
releases after it. The behaviour is pinned by
kubectl_table_distills_from_raw_not_collapse_markers: the bug it guards against is a
column parser reading [30 similar lines collapsed] as a pod row.
The mode is picked by specificity. A kubectl … | grep payload exercises the
Infra path rather than the Log path, so a fixture chosen to test a collapse guard can
pass with the guard removed. Check which mode your fixture actually reaches.
Distill
registry::resolve_profile(command) picks the distiller, then:
fn distill(&self, segments: &[OutputSegment], input: &str,
session: Option<&SessionState>) -> Option<String>;
Option, and that is the whole design. A distiller that parsed nothing returns None
and the caller hands back the raw bytes. The invariant lives in the trait rather than
in each author remembering to call a helper, so it holds for all 12 by construction.
The TOML layer that used to short-circuit this stage was retired in 0.7.4, so the Rust code is now the only thing that can claim a command.
The ledger
After distillation, ledger::Ledger replaces runs of lines the scope has already been
shown with a handle. Gated on the same format sniff as collapse.
It is append-only, which is what keeps the upstream prompt cache intact: a cache works on a prefix, so shortening the suffix costs nothing while retroactive compaction would destroy it.
See The ledger for the two scopes and their different claims.
Persist
The raw input is archived by SHA-256 and the marker carries the handle.
Order matters and is not negotiable: archive, then write the marker. A failed archive leaves the run verbatim. Doing it the other way round produces, on any write failure, a marker pointing at content that was never stored.
Recording is unconditional even when the projection saved nothing, because a block is worth remembering in case it is seen again.
Two doors, one pipeline
hooks/post_tool.rs and hooks/pipe.rs both run these stages. Keeping them in step
is a live maintenance problem: three separate fixes each corrected one copy and left
the other, and the ledger stage existed in one for a release before the other had it
at all.
Change both, or write down why not.
Adding a stage
Do not, unless the measurement says so. The pipeline earns its shape from a replay harness, and the useful pattern is to price a proposal before building it:
- “Route a pipeline by its last stage” sounds obviously right, and would have handed
871 of 1,035 recorded pipelines to
head,tailorsed, all verbatim passthroughs, stopping distillation on them entirely. - Quote-aware chain splitting keeps 205 of 2,928 routed commands, which is what justified 25 lines of scanner over a 5-line naive split.
- An import-graph signal for the scorer sized at 196 traces, until the graph itself turned out to be wrong. Corrected, it sized at 26, against a 542 ms build on a 10 ms budget.
A measurement that kills a design is the measurement working.
Adding a distiller
The most common change here. Five steps, and the fourth is the one that matters.
1. The module
src/distillers/my_type.rs:
use crate::pipeline::{OutputSegment, SessionState};
use super::Distiller;
pub struct MyDistiller;
impl Distiller for MyDistiller {
fn distill(
&self,
segments: &[OutputSegment],
input: &str,
session: Option<&SessionState>,
) -> Option<String> {
// Return None the moment you are not sure you parsed this.
todo!()
}
}
Return None whenever parsing failed. That hands back the raw bytes, which is the
correct answer and the one the whole design rests on.
Never return a success string from a zero state. vitest: ✓ 0/0 passed for output
that was actually a dev server is failing closed, confidently, and it is the exact
defect this project keeps fixing.
2. Register it
src/distillers/mod.rs:
pub mod my_type;
// in get_distiller():
ContentType::MyType => Box::new(my_type::MyDistiller),
Routing belongs in pipeline/registry.rs. Do not add a matches!(cmd, ...) block
inside the distiller; that duplication is what the registry exists to prevent.
3. A realistic fixture
tests/fixtures/my_type_example.txt. Real output from the real tool, not something
hand-written to be easy to parse.
4. A snapshot test, and prove it can fail
snapshot_test!(test_my_type_distillation, "my_type_example.txt", ContentType::MyType);
cargo test
cargo insta review
Then break the rule deliberately and watch the test go red, before restoring it. A check that cannot fail proves nothing, and this repo has shipped two regression tests that could not fail.
Two specific ways a test here passes for the wrong reason:
- Your fixture reaches a different collapse mode than you think. A
kubectl … | grepfixture exercises Infra, not Log, so a guard you are testing may never be consulted. - “No rewrite from the hook” is not proof the distiller punted. It can mean the format gate fired, or the guardrail rejected the result.
A distiller can also return a near-copy rather than the exact input, so detect “this
did not help” with beats_guardrail rather than comparing against the input.
5. Gates
cargo fmt
cargo clippy -- -D warnings
OMNI_DB_PATH=/tmp/t.db cargo test
OMNI_DB_PATH is not optional. Parallel tests competing for ~/.omni/omni.db cause
SQLite locks: 79 seconds green against an isolated database, 433 seconds and then a
hang against the live one.
Before you write any of it
Measure the workload. ~/.omni/omni.db prices a proposal in one query, and the answer
is often the opposite of the request.
“Improve the python3 distiller” turned into two facts in two queries: python3 was
already reporting 97.2%, and the savings were the collapse fallback deleting data
rows. The obvious feature, a traceback distiller, died on 9 of 7,506 traces
containing a traceback.
-- distillations.filter_name is the command's first token
-- execution_traces holds raw_input and distilled_output in full
Read the rows before quoting an aggregate over them. A LIKE filter that caught the
wrong rows has already put a wrong figure into a published issue.
And never read sqlite3 output through the Bash hook while doing this. The pipeline
can fold the rows you are counting.
The bar the result has to clear
Not “did it compress”. These:
- Would the agent still have the answer?
- Does anything dropped leave a marker?
- Does the reported number describe what actually happened?
A patch that raises reduction percentage while removing signal is the project’s own recurring defect, shipped again with your name on it.
Testing
OMNI_DB_PATH=/tmp/omni-test.db cargo test
Start with that line. It is not a suggestion.
The two guardrails that waste the most time
Isolate the database. Parallel integration tests competing for ~/.omni/omni.db
cause SQLite locks and hangs. Measured: 79 seconds green against an isolated database,
433 seconds and then a hang against the live one. tests/hook_e2e.rs has an
omni_cmd() helper that spawns the binary with a unique OMNI_DB_PATH from a
NamedTempFile. Use it.
Lock early, release fast. Rust mutexes are not reentrant, so nested or redundant
lock() calls on session_arc deadlock. Open a scope, take what you need, let the
guard drop before doing anything that might lock again.
If cargo test runs over a minute on macOS or Linux, suspect one of those two. Check
pipe mode and the E2E tests first; they are the heaviest.
Suites
cargo test # everything
cargo test --test hook_e2e # binary spawn, end to end
cargo test --test savings_assertions # per-filter savings thresholds
cargo test --test security_tests
cargo test distillers::tests # snapshots
cargo insta review # approve snapshot changes
tests/smoke_test.sh ./target/debug/omni
tests/fixtures/ holds 45 realistic tool outputs. Add real output from the real tool,
not something shaped to be easy to parse.
Naming
Inside #[cfg(test)], drop the test_ prefix. The attribute already says it is a
test.
fn returns_default_when_config_missing()
fn excludes_sensitive_data_from_summary()
fn preserves_errors_during_collapse()
fn renders_identical_bytes_for_identical_state()
Not test_config_ok, not handles_it, not valid_json. Start with a verb, say what
the behaviour is, English only.
Design
One behavioural assertion per test. Arrange, act, assert, with the sections visible.
Test observable behaviour, not internals: assert_eq!(result.status, Status::Ready)
rather than assert!(internal_cache.len() > 0).
Every non-trivial feature carries a happy path, an edge case, a malformed input case,
a regression case if it is a fix, and an explicit no-panic case. Malformed input must
return Err, never panic.
Prove the test can fail
Break the rule deliberately, watch it go red, restore it.
This repo has shipped two regression tests that could not fail. Both looked correct. Both passed with the fix reverted.
Three specific ways a green test here means nothing:
Your fixture reaches a different code path than you think. Collapse mode is picked
by specificity: a kubectl … | grep fixture exercises Infra, not Log, so a
collapse-guard test passes with the guard removed.
“No rewrite from the hook” is not proof the distiller punted. It can equally mean the format gate fired or the guardrail rejected the output.
A distiller can return a near-copy rather than the exact input. Detect “this did
not help” with beats_guardrail, not output == input.
Proving a refactor changed nothing
For behaviour-preserving work, diff distiller output over the whole recorded corpus, about 5,100 commands times 11 probes, then break one arm deliberately to show the harness has teeth. A differential harness that cannot detect a planted difference is not evidence.
Gates
make ci # fmt + clippy + test + security + binary-check
Or individually:
cargo fmt --all -- --check
cargo clippy --all-targets --all-features -- -D warnings
cargo test --all
Zero clippy warnings. Not “few”.
source "$HOME/.cargo/env" first, so you get the pinned 1.97.0 rather than a Homebrew
cargo that ignores rust-toolchain.toml and will keep drifting from CI.
Never weaken a check to make it pass
Not an assertion, not a security check, not a threshold. If a test is in the way, it is either wrong, in which case fix it and say why, or right, in which case the code is wrong.
Benchmarks
One developer’s real command history, replayed. Every figure in a section comes from the same run as the rest of that section, including the ones that do not flatter us.
Two runs are published here and they disagree by a factor of fifteen. That is the most useful thing on the page: the number OMNI reports is a property of the week it replays, not a constant. Read the corpus line before the figure, every time.
Correction, 2026-09-03: every ledger figure below is overstated
The replay built its ledger without ever telling it which command produced the payload
(#760), so every rule in the ledger that reads the command was inert here while running
normally in production: the from clause, the line budget that leaves a tail -5 alone,
and the wording a re-run gets. The harness measured a ledger with its guards switched off.
Re-measured on the same frozen corpus, same commit, with only that fixed:
| arm | as published | with the guards visible |
|---|---|---|
| omni, with the ledger | 4.9% | 3.0% |
| rtk + our ledger | 5.7% | 3.7% |
| caveman + our ledger | 5.6% | 3.7% |
| headroom dedup | 5.8% | 5.8% |
| lean-ctx compress | 4.8% | 4.8% |
| omni, filters only | 1.4% | 1.4% |
Per class, with the ledger: file read 4.3% to 1.5%, other 4.6% to 2.8%, git 8.6% to 6.4%, search 4.2% to 4.0%, infra 3.8% to 3.6%, build and test 11.1% to 10.7%. The capture rate goes 23.3% to 10.7%.
Only the arms that use our ledger move, which is the check that the change is what it claims to be.
Where the 1.9 points went, measured by switching each guard off on the same corpus and the same commit. Every one of them exists to stop a fold that should not happen, so this is the price of the ledger behaving, not a regression:
| ledger arm | aggregate |
|---|---|
| guards invisible to the harness, as published in 0.7.8 | 4.9% |
naming the source switched off (from clause, re-run wording) | 4.3% |
line budget switched off (tail -5, sed -n 60,200p fold again) | 3.4% |
| shipped, every guard on | 3.0% |
Naming the source costs 1.3 points and the line budget 0.4. The rest is the two interacting: a marker that carries a source is longer, and marker length decides which runs are worth folding at all.
competitor engines are untouched.
The tables below are left as they were measured, because deleting a published number is
worse than labelling it. Artifacts carry harness.ledger_knows_the_command from #760 on,
and one written without that field was measured this way.
The current run, 2026-08-24, and the first one that will still exist next month
Corpus: 9,478 traces, 8,458,937 bytes, 70 sessions, all agent_id='claude_code',
0 terminal rows, 0 errored. Frozen on disk and hashed as 0b63218ef78a1edb, replayed
in 7.1 s on OMNI 0.7.8.
Every run above this line on this page was measured on execution_traces, which prunes
at seven days, so none of them can be re-derived. This one is a file. That is the whole
of #704: a release-over-release delta was previously a code change and a corpus change
added together, with no way to separate them.
0.8% from the filters. 2.4% with the ledger, on 0.7.10. 98.8% of calls saved
nothing, 1.2% shrank, and no call came back larger. Tokens, cl100k_base as a
proxy for a vocabulary Anthropic does not publish: 2,404,625 to 2,384,061, 0.9%.
Both numbers fell from 0.7.9, where they read 1.4% and 3.0%, and that is the release rather than a regression: 0.7.10 closed thirteen classes of false claim, and every point given up was being earned by a fold that should not have been made. The ledger arm was itself re-measured under #760, and this section first published 5.1%.
search reads 0.0% captured, and that is a refusal rather than a failure. The
class was 10.2% on 0.7.9. #814 reported what those folds were: a grep reply with
9 of its 11 matches behind a handle, which reads as a file with two matches. #815
made a grep reply fold whole or not at all, and a whole fold needs the identical
reply to have been printed before. This corpus contains no repeat of a grep
command inside one session, across 1,069 grep, rg and ag traces, so nothing
in it qualifies any more. The refusal cost 4.6 KB over 810 calls, 0.054% of
the corpus. Checked against the 0.7.10 binary, an identical re-run still folds to
one marker and a partially-seen reply now folds nothing.
| Class | Calls | Input | Filters | + ledger | Available | Captured |
|---|---|---|---|---|---|---|
| other | 6,457 | 4.81 MB | 0.8% | 2.8% | 15.9% | 12.6% |
| file read | 1,056 | 1.89 MB | 0.0% | 1.5% | 17.7% | 8.5% |
| git | 899 | 0.86 MB | 0.0% | 1.3% | 19.8% | 6.7% |
| search | 810 | 0.77 MB | 3.4% | 3.4% | 6.5% | 0.0% |
| infra | 215 | 0.14 MB | 2.4% | 2.7% | 5.5% | 5.2% |
| build and test | 41 | 0.02 MB | 9.0% | 10.7% | 21.7% | 8.6% |
| aggregate | 9,478 | 8.49 MB | 0.8% | 2.4% | 15.7% | 10.3% |
| Arm | bytes | saved |
|---|---|---|
| headroom dedup, omni’s filters | 8,486,830 to 8,061,205 | 5.0% |
lean-ctx compress | 8,486,830 to 8,076,957 | 4.8% |
| caveman + omni’s ledger | 8,486,830 to 8,174,311 | 3.7% |
| rtk + omni’s ledger | 8,486,830 to 8,174,685 | 3.7% |
| omni, with the ledger | 8,486,830 to 8,281,114 | 2.4% |
caveman compress | 8,486,830 to 8,311,999 | 2.1% |
rtk pipe | 8,486,830 to 8,308,491 | 2.1% |
| omni, filters only | 8,486,830 to 8,417,541 | 0.8% |
Measured by make bench over 9,478 traces (8.42 MB, 70 sessions), corpus 0b63218ef78a1edb, OMNI 0.7.10.
available and captured are new, and captured is the figure that survives a
change of workload. The ledger substitutes lines it has already delivered, so it can
never fold what was not repeated; available is that ceiling and captured is the share
of it taken. Between this corpus and the 0.7.5 run below, the file-read saving moves by a
factor of twenty. The capture rate barely moves. Only one of those two is a statement
about OMNI.
Why the savings column is so much lower than the run below: this corpus is mostly
one-line shell plumbing. 4,815 of the 9,478 calls begin with cd, carrying 3.98 MB at
1.2%, and the median repeated run is 10 bytes. It was a week of driving OMNI’s own
development from a terminal, so the payloads are git, gh, sed and sqlite3 output
that is either tiny, structured, or seen once. Repetition is 16.3% of raw bytes here
against 80.6% in the window below.
What it does not measure. The harness replays execution_traces, which holds shell
commands only: zero of these 9,478 rows is a Read payload. The ledger’s file-read path,
including the count-preserving fold added in #664, is not exercised by any figure on this
page. That gap is the honest reason the fold’s own numbers live in
changelog.d/664.changed.md with their own corpus instead of here.
All four arms answer now, and the table above is the result. #711 fixed the
headroom arm. The caveman arm was calling caveman tools compress, and nothing on the
measuring machine answers to a bare caveman: ~/.caveman/bin holds caveman-engine,
caveman-shrink and four others. Pointed at caveman-shrink, tools compress is not
a subcommand it knows, so it falls through to shrink mode and hands back the input with
"ratio":0. An arm reporting zero on every trace reads as an arm that lost, which is
#711’s failure wearing a different binary. caveman-engine compress is the entry point
that works (#712).
On this corpus OMNI is last of the four rows that carry a ledger, and both halves are
behind. headroom’s dedup takes 5.0% where ours takes 2.4% over identical filters and
identical blocks, so that 2.6 points is the dedup engine alone. It read 0.9 points until
#760, which is the correction at the top of this page and not a change in the code. Our filter tier is the
weakest of the four at 0.8%, against 2.1% for rtk and caveman and 4.8% for lean-ctx,
and that shortfall is what carries rtk + omni's ledger and caveman + omni's ledger
above our own stack: the ledger is identical in all three rows and only the filters
underneath it differ.
Both of our rows fell from 0.7.9, where they read 3.0% and 1.4%. That is #815 and #832 refusing folds and rewrites that were producing a wrong answer, rather than the engine getting worse. headroom’s row fell with them, 5.8% to 5.0%, because that arm runs its dedup over our filters, so the gap between the two ledgers narrowed slightly, 2.8 points to 2.6.
lean-ctx beating our filters by 4.0 points is the largest single gap here and it is not argued away. It is a deep compressor rather than a per-command filter, and the same shape showed up on the 0.7.5 corpus below at a much larger magnitude.
Versions: rtk 0.45.0, lean-ctx 3.9.18, caveman bin-v1.0.0, headroom 0.34.0. No
lean-ctx + our ledger row: its preview reports compressed_bytes and never emits the
text, so the row could only be estimated, and an estimate beside four measurements is
the blend this harness exists to avoid.
The aggregate moved from 5.1% to 4.9% on the same corpus, and it was paid for on purpose. Replaying the frozen corpus at the two commits either side of #728 puts the whole move on that one change: 5.1% and 24.1% captured before it, 4.9% and 23.3% after. #728 stopped the project scope folding a whole reply where it should fold part of one, and folding less is the point of that fix. The corpus hash did not change, so the delta is code alone.
The 0.7.5 run, 2026-08-14, which can no longer be re-derived
execution_traces prunes at seven days, so the corpus behind every figure in this section
is gone. It is kept because deleting a published number is worse than labelling it, and
because the gap between the two runs is the point.
Corpus: 5,984 traces, 23,086,649 bytes, 2026-08-11 11:03:00 to 2026-08-14
18:11:10 UTC, all agent_id='claude_code', 123 terminal rows excluded from 6,107,
0 errored. Replayed in 1,238 s.
That run’s headline
32.6% fewer bytes from the filters. 69.6% with the ledger. 23,086,649 to 15,557,823 to 7,026,021.
| tokens, filters only | 7,682,124 to 4,874,124, 36.6% |
| bytes per token | 3.005 raw, 3.192 distilled (the shipped estimate is 3.6) |
| calls that saved nothing | 96.1%, 5,748 of 5,984 |
| calls that shrank | 3.9%, 236 |
| calls that grew | 0 |
| ledger folds | 882 calls, 3,231 session markers, 86 project markers |
| raw bytes already shown once | 68.4% before filters, 64.7% after |
| that repetition, by scope | 67.3% same session, 1.1% earlier session, same project |
Filtering and repetition are orthogonal. That is the argument for the ledger, and on this corpus the ledger is worth more than twice what the filters are.
Read the corpus before the number. This window is unusual and it inflates everything below. 148 of the 5,984 calls carry 64.7% of all bytes, 286 groups of byte-identical payloads account for 80.6% of the total, and the single largest contributor is five traces of exactly 820,000 bytes whose content is one sentence repeated to fill. It is the week this machine did nothing but develop and benchmark OMNI. A corpus of ordinary work reads far lower: the same harness on 6,656 traces in August 2026 read 2.7% and 14.9%.
Which commands benefit
| class | calls | input | filters | + ledger |
|---|---|---|---|---|
| other | 3,703 | 11.05 MB | 29.1% | 56.2% |
file read (cat, sed, head, tail) | 884 | 10.93 MB | 39.2% | 89.6% |
search (grep, rg, find) | 600 | 540 KB | 2.3% | 4.3% |
git, gh | 696 | 475 KB | 2.5% | 7.0% |
infra (kubectl, az, docker) | 65 | 70 KB | 0.0% | 6.8% |
| build and test | 36 | 24 KB | 10.8% | 10.8% |
| aggregate | 5,984 | 23.09 MB | 32.6% | 69.6% |
infra reads 0.0% from the filters on purpose. It was 1.7% one release ago, bought
by summarising kubectl get pods tables, which deleted the pod names that were the
answer. That saving is gone and the rows are back (#562). What remains for infra is
the ledger, which folds a listing the agent has already seen and needs the rows
intact to do it.
By shell shape:
| form | calls | input | saved |
|---|---|---|---|
| bare program | 782 | 10,683,924 | 40.2% |
| chain | 2,024 | 9,843,901 | 32.5% |
cd prefix | 1,655 | 1,476,727 | 0.4% |
VAR= assignment | 952 | 567,135 | 0.4% |
| pipe only | 571 | 514,962 | 4.6% |
Top commands by input bytes, filters only:
| command | calls | input | output | saved |
|---|---|---|---|---|
tail | 441 | 9,558,272 | 5,599,666 | 41.4% |
zsh | 282 | 8,391,102 | 5,202,443 | 38.0% |
cd | 1,727 | 1,503,873 | 1,497,294 | 0.4% |
cat | 119 | 770,972 | 442,285 | 42.6% |
export | 535 | 429,265 | 428,731 | 0.1% |
grep | 447 | 410,575 | 402,804 | 1.9% |
sed | 217 | 381,331 | 381,331 | 0.0% |
git | 401 | 262,113 | 256,786 | 2.0% |
gh | 238 | 145,950 | 140,095 | 4.0% |
kubectl | 68 | 71,129 | 71,129 | 0.0% |
Byte-sink and token-sink rankings disagree at the tail: bash enters the token top
15 where kubectl sits in the byte one.
Head to head, one corpus
Identical bytes into every arm. Versions: rtk 0.45.0, lean-ctx 3.9.18, caveman 1.1.0
(binaries bin-v1.0.0), headroom at cross_turn_dedup.py.
| bytes | saved | claimed | |
|---|---|---|---|
rtk pipe | 23,086,649 to 21,655,277 | 6.2% | |
caveman tools compress | 23,086,649 to 21,516,757 | 6.8% | |
| omni, filters only | 23,086,649 to 15,557,823 | 32.6% | |
lean-ctx compress | 23,086,649 to 11,678,975 | 49.4% | 425 of 5,984 |
| headroom dedup, our filters | 23,086,649 to 7,905,764 | 65.8% | |
| omni, with the ledger | 23,086,649 to 7,026,021 | 69.6% | |
| rtk + our ledger | 23,086,649 to 8,906,376 | 61.4% | |
| caveman + our ledger | 23,086,649 to 8,844,105 | 61.7% |
headroom is 3.8 points behind our ledger and that is the only close race here. Both arms run the same filters over the same blocks, so the gap is the dedup engine and nothing else.
lean-ctx beats our filters by 16.8 points, 49.4% against 32.6%, over 425 calls to our 236. That is not argued away: this corpus is a few enormous repetitive payloads, which is exactly the shape a deep-and-narrow compressor is built for.
No lean-ctx + our ledger row: its preview reports compressed_bytes and never emits
the text, so that row could only be estimated.
Single fixtures
From tests/fixtures/, same build, reproducible by hand. “Delivered” includes the
marker.
| command | input | delivered | saved |
|---|---|---|---|
docker build (heavy noise) | 9,207 B | 102 B | 98.9% |
cargo build (large, successful) | 3,220 B | 62 B | 98.1% |
cargo test (490 passed, 10 failed) | 16,515 B | 1,153 B | 93.0% |
git status (dirty) | 496 B | 165 B | 66.7% |
git diff (multi-file) | 397 B | 247 B | 37.8% |
kubectl get pods (mixed) | 840 B | 840 B | 0.0% |
kubectl get pods reading 0.0% is the design, not a gap. For one release it read
73.5%, because a summariser that had been shadowed since #110 became live when #510
retired the TOML layer, and a 10 row table arrived as three lines with seven pod names
deleted. A count of pods cannot be turned back into a pod name (#562).
docker build is the opposite case and worth the contrast: 251 lines of per-layer
DEBUG and INFO become docker build: ✓ complete (50 layers, 50 cached), and the build
did succeed. Noise, not an enumeration.
Method
OMNI_BENCH_DB=~/.omni/omni.db \
cargo test --release --test bench_replay -- --ignored --nocapture
Snapshot the database first if the figure has to be quotable: hooks write to it while the
replay reads, and execution_traces prunes at seven days, so a run against the live file
measures a corpus that is already different from the one anybody else would get.
sqlite3 ~/.omni/omni.db ".backup /tmp/bench-corpus.db"
OMNI_BENCH_DB=/tmp/bench-corpus.db \
cargo test --release --test bench_replay -- --ignored --nocapture
| corpus | execution_traces.raw_input, real usage, replayed. Not synthetic. Shell commands only: no Read payload is in this table, so no figure here covers the ledger’s file-read path |
| population | calls whose result reached a model. OMNI_BENCH_ALL=1 widens it |
| state | session: None, store: None, HOME at an empty temp dir |
| path | run_inner, the same pipeline the hook and omni exec run, markers included |
| binary | release build |
| arms | OMNI_BENCH_RTK, _LEANCTX, _CAVEMAN, _HEADROOM, each off unless it names a binary, so CI never needs a competitor installed |
Terminal output is excluded, and it is worth two different headlines. On an
installation carrying it, it was 68% of raw bytes: 79.1% including it against 43.3%
model-facing. The harness and omni stats both counted it until that was fixed, and
both now print which population they used.
Every figure comes from one run. This file once published 15.7% and 16.1% from two replays a day apart without saying so.
Every window closes. execution_traces prunes at 7 days, so this corpus is gone a
week after it was measured. Hold one open with OMNI_TRACE_RETENTION_DAYS.
Old figures are deleted, not kept for comparison. Releases keep changing the rule that decides whether the ledger folds a run, so an older number describes a pipeline that no longer exists, and printing both invites a reader to read two programs as a trend.
Latency was not re-measured on this build, so no table is printed rather than an older one relabelled. The method that produced the last one: median of 12 runs per payload, release binary, end to end through the post-hook, against a fresh database and a large one. Payload size barely mattered; database size did. Measure by removal, never with a microbenchmark: a unit-test timer once said 66 ms for work an A/B on the release binary put at 34.3 ms.
What no figure here can tell you
Whether the removed lines were signal.
Measure your own
omni stats
omni stats --share
Both read the same aggregation, so the share card cannot drift from the report. Terminal output is excluded from both.
Where OMNI is going
Direction only. The queue lives on the
Now / Next / Later board and the
shipped history lives in CHANGELOG.md. Copying either one here is how an earlier
version of this page spent six weeks announcing v0.6.0 as in progress while 0.6.8
shipped.
The goal
OMNI removes noise from what an agent reads, without removing the answer and without overstating what it removed.
Compression is the easy half. A distiller that deletes a whole kubectl table and
reports 99% saved compressed perfectly and did the job wrongly. So the target is not a
reduction percentage. It is output an agent can act on, next to a number a human can
reproduce.
Three properties, in the order they win when they conflict:
- Never fabricate. A stage that recognised nothing hands back what it was given. A failed command passes through verbatim. Structured payloads are never touched.
- Never lose the answer quietly. Anything dropped leaves a marker and, where the content allows, a handle.
- Then compress, as hard as the first two allow and no harder.
The number that decides progress
Primary: context-window pressure for the same job. Conversation growth, turns before compaction, and the cost of recovering task state in a new chat. That is the meter a user watches and the one OMNI is bought to move.
Secondary: distill %. Always scoped by agent_id, always model-facing only. A
diagnostic for one host’s pipeline, not a product claim.
Why the swap away from blended reduction. On the reporting corpus, 81% of calls are
passthrough and correctly do nothing, so a blended percentage describes the command
mix more than the product. terminal rows are TTY bytes no model reads. Prompt-cache
reads bill about a tenth of fresh input, so bytes saved once are not dollars saved per
turn. And on a flat-rate plan compression does not reduce a bill at all; what it buys
is session lifetime and fewer re-runs.
The gate on any public headline number. It cites the agent_id it covers, the
corpus it was measured on, and a command a reader can run to reproduce it. A figure
that blends terminal with hook agents, or counts a rewrite the host never applied,
does not ship.
Non-goals
Recorded with dates, because the useful part of a rejected option is the reason.
| not building | why | decided |
|---|---|---|
| An HTTP proxy in front of the model | It puts OMNI on the request path and routes the user’s API key through a local process. The hook is the product, and the absence of that friction is most of the advantage. | 2026-07-23 |
| A model or ML compressor inside the pipeline | Hooks have a sub-10 ms budget. Nothing with an inference call meets it. | 2026-07-23 |
| Chasing a higher reduction % with more aggressive distillers | The failure mode this project keeps shipping is a confident summary that deleted the answer. More aggression buys the number and costs the product, and on a host that cannot rewrite built-in tool output it buys nothing at all. | 2026-08-07 |
| Claiming shell distillation on a Handoff-first or MCP-only host | The host does not apply the rewrite, so the model reads the same bytes it always did. Saying otherwise is the same defect as a distiller reporting a saving it did not make. | 2026-08-07 |
| Intercepting a host’s shell by denying it and returning output as a hook message | Technically possible on Cursor. It tells the agent its command was blocked, loses the exit code, moves execution semantics into OMNI, and bypasses the host’s approval flow. | 2026-08-07 |
| Filter marketplace, team mode, remote archive, IDE extension | Ecosystem features for a tool whose core claims are not all true yet. Worth reopening once the axes below are done. | 2026-07-29 |
| A user or project filter tier on disk | It let a checkout decide what an agent is shown, behind a trust gate that hashed one file and admitted another. Deleted rather than repaired, and the whole layer was worth 804 bytes over 6,656 commands. | 2026-08-11 |
The three axes
A change that moves none of these can still be worth making, but it is maintenance, not direction.
1. Correctness: nothing asserted that was not parsed
Closed. The invariant moved off the authors and into the trait: distill returns
Option<String>, so a distiller that parsed nothing returns None and the caller
hands back raw bytes. It holds for all 12 by construction.
What is not closed is the class this axis exists for. Returning None proves a
distiller knew it had failed. It proves nothing about one that parsed something and
summarised it wrongly.
Check: no open bug describes OMNI asserting a result it did not parse, and that stays true across a full release cycle. This class has been filed against nine separate releases, so a quiet month is not evidence.
2. Coverage: the hook reaches the tools agents use
Closed for Claude Code. The post-tool matcher is Bash|Read|Grep|WebFetch, so the
three distillers that had never run now do. Still open for hosts whose matcher
vocabulary is narrower.
Check: the installed hook configuration names more than one matcher, and the
database holds distillation rows for a tool other than Bash.
3. Proof: every published number can be reproduced
The open one. The blending is fixed: terminal runs are excluded from the
model-facing figure and duplicate rows are gone. What remains is that numbers cannot
outlive their corpus. execution_traces prunes at seven days, so any published figure
stops being re-derivable a week after it is measured, which is the opposite of what
this axis asks for.
Check: every published figure states its agent_id, its corpus, and the command
that reproduces it.
Off the axes
Dependency and CI hygiene, README translation sync, dead-code removal, packaging and release mechanics. Real work, regularly done, and deliberately not direction.
Contributing
The most useful contributions are a distiller for a tool not covered, a signal for a tool whose noise is line-shaped, and a reproduction of any case where OMNI’s output claims more than its input supports.
The third is worth more than it sounds. See CONTRIBUTING.md in the repository.
Releasing
make ci # fmt + clippy + test + security + binary-check
make bump VERSION=x.y.z
make release VERSION=x.y.z
make release-sha VERSION=x.y.z # after the tag has actually built
The order, and why it is not negotiable
Cut the changelog first, then bump.
bump_version.sh does not touch CHANGELOG.md, and build.rs counts what is still
uncut in the tree it compiles: the bullets under ## [Unreleased] plus the fragments in
changelog.d/. Tag without folding them and the released binary tells every user
[N UNRELEASED] … cut a tag. It accuses itself.
So:
make changelog-cut VERSION=x.y.z # folds changelog.d/ into ## [x.y.z] - <date>
git commit -am "docs(changelog): cut x.y.z"
make bump VERSION=x.y.z
A correctly cut build prints omni vx.y.z [AHEAD/RC] with no UNRELEASED line. Verify
that before pushing the tag.
The half of that line to trust afterwards is the missing UNRELEASED, not the
label. guard::update::get_status caches the newest known release in
~/.omni/update_cache.json for 14400 seconds, so for four hours after a tag a machine
that ran omni doctor beforehand still holds the previous version and reports Ahead
whatever it is running. Observed on 0.7.5, where the freshly installed release printed
[AHEAD/RC]. build.rs computes the unreleased count from the tree with no cache, so
that half is always current; delete the cache file if you want the label to mean
something.
Day to day, the entry goes in changelog.d/<issue>.<section>.md as the work merges, not
into CHANGELOG.md and not at tag time. One file per entry means two branches never
write the same path, which is what stopped every parallel branch conflicting on
## [Unreleased]. The format is Keep a Changelog and SemVer, and the entries here are
unusually detailed on purpose: each states the measured evidence, the wrong number that
was published, and the mechanism. A one-line entry is a regression in that file’s
quality.
The first cut needed one manual tidy and it is done. 0.7.5 folded three fragments beside
seven bullets written into ## [Unreleased] before the convention existed, and arrived
with two ### Changed and two ### Fixed under one version heading. Merging those four
into two was the only hand edit. Check a cut by word count rather than by eye: 2,252
words across the old section plus the fragments, 2,252 in the folded section. A
reordering that drops a bullet body looks correct in a heading-level diff.
CI green does not mean the release will build
The 0.6.2 tag produced no binaries at all. release.yml asked for stable per
cross target while rust-toolchain.toml pinned a version, so every cross-compile died
with can't find crate for core before compiling a line. ci.yml stayed green
throughout, because it only builds host-native.
The fix is that cross targets belong in rust-toolchain.toml, and its targets list
has to stay in sync with the release matrix.
After tagging, watch the release workflow actually produce artifacts before running
make release-sha or announcing anything.
Things that look like failures and are not
omni-release.sh ends in an interactive read -p, so an automated run has to pipe
echo y | into it.
It pushes main and the tag together. main is branch-protected, and a maintainer
token bypasses it: the push prints “Changes must be made through a pull request” and
succeeds anyway with rc=0. That line is not an error.
The Homebrew step
update_homebrew_sha.sh pushes to two repositories: the tap, and omni.rb back
to main. Check the tap clone is clean and synced with its remote first, or the run
aborts partway and leaves the formula half updated.
Afterwards, verify the formula’s SHAs against the release’s published SHA256SUMS
rather than trusting the script’s own success line, then confirm:
brew info fajarhide/tap/omni # expect: x.y.z → stable <new>
Before merging anything into a release
CI green is not review-clean. Read the review comments, automated and human, validate each one against the code rather than assuming the reviewer is right or wrong, and fix or reply. A green pipeline says the tests passed. It says nothing about a correctness bug a reviewer flagged.
Branch shape
One branch per batch, not per issue. N parallel branches cost N full CI runs of about
eleven minutes each, serialised. Batch a lane into one branch, one commit per issue, one
pull request with several Closes #N lines.
Split only when a reviewer would genuinely need them apart, or when one is risky enough to be reverted alone.
That conflict used to be CHANGELOG.md, every time. It is gone: entries are files in
changelog.d/ now, and two branches never write the same path. What remains is the CI
cost, which is why batching still pays.
Closes #N must be in the pull request body before the merge. GitHub evaluates
the keyword at merge time only; adding it afterwards does nothing, silently.
Hermes Agent
OMNI plugs into Hermes twice: a plugin on the hook path, and the MCP server.
| layer | mechanism | what changes |
|---|---|---|
| hooks | ~/.hermes/plugins/omni-signal-engine/__init__.py calling omni --pre-hook, --post-hook, --session-start | terminal tool output is distilled before it enters Hermes’ context |
| MCP | mcp_servers.omni running omni --mcp | OMNI’s MCP tools become first-class Hermes tools |
Prerequisites
brew install fajarhide/tap/omni
omni --version
omni doctor
export HERMES_VENV="${HERMES_HOME:-$HOME/.hermes}/hermes-agent/venv"
export HERMES_PY="$HERMES_VENV/bin/python"
"$HERMES_PY" --version # 3.11 or newer
The venv Python is needed because hermes plugins enable runs inside it.
Install
omni init --hermes
hermes plugins enable omni-signal-engine
hermes gateway restart
"$HERMES_PY" -m pip install hermes-omni-plugin
omni init --hermes is idempotent. It installs the plugin scaffold, registers the MCP
server in ~/.hermes/config.yaml if it is not already there, enables Hermes
compression when that is safe, and writes Hermes-oriented defaults to
~/.omni/config.toml without overwriting an existing OMNI config.
Use either
hermes-omni-pluginor theomni init --hermesscaffold, not both at once, or you get duplicate plugin registrations.
Config
# ~/.hermes/config.yaml
plugins:
enabled:
- omni-signal-engine
mcp_servers:
omni:
command: "/opt/homebrew/bin/omni"
args: ["--mcp"]
env:
OMNI_AGENT_ID: "hermes"
compression:
enabled: true
threshold: 0.50 # compress at 50% context usage
target_ratio: 0.20 # keep 20%
Three things have to be true: plugins.enabled contains omni-signal-engine,
mcp_servers.omni points at the real binary, and compression.enabled is on so
Hermes’ own compaction and OMNI’s pressure warnings line up rather than fighting.
OMNI_AGENT_ID: "hermes" matters more than it looks. Without it, Hermes’ rows blend
with every other host’s and no figure about either is meaningful.
Verify
omni doctor
hermes plugins list | grep omni # expect: omni-signal-engine enabled
hermes tools list | grep mcp_omni_ # the advertised set, after a restart
Then a functional check on a real fixture:
cat tests/fixtures/cargo_test_500.txt | omni --post-hook 2>&1 | head -20
# passing test lines stripped, failures preserved
For a live test, run something noisy through Hermes’ terminal tool
(terminal("npm install", timeout=120)) and compare the tool result size against raw
npm output. Confirm with omni stats.
Read the list rather than a number written down. This guide has said 27, which came from grepping the server source and counted a filter name as a tool, and then 25, which is the whole surface and not what a host is told about. Hermes is a Full-tier host, so it is advertised
omni_retrieveandomni_explain_savings, the two that pay for their place in the prefix of every request.OMNI_MCP_TOOLS=allserves the whole surface.
Where OMNI helps and where it does not
| output | OMNI’s effect |
|---|---|
npm install, cargo build, docker build | large, 70% and up. Progress, cache hits and layer hashes are pure ceremony. |
| test runs | large. The verdict and the failures survive, the ok lines do not. |
| file reads | nothing from the filters, a great deal from the ledger on re-reads |
kubectl -o json, terraform plans | nothing, deliberately. Structured payloads pass through. |
| short commands | nothing, or slightly negative. The marker costs more than the saving. |
Use the MCP tools as Hermes’ controls over all of it: omni_explain_savings to see
what a recent command actually cost, omni_retrieve to get folded content back, and
omni_budget to see where the session’s tokens went. That tool is outside the set
advertised by default, so it needs OMNI_MCP_TOOLS=all.
After a Hermes upgrade
hermes plugins list | grep omni
hermes tools list | grep mcp_omni_
omni doctor
An upgrade can reset plugins.enabled or move the venv. Both fail quietly: the plugin
simply stops being called, and nothing announces it.
Loop engineering
Running an agent in a loop, where each iteration adds to a context window that does not grow. OMNI’s part is tracking what the loop has spent and carrying memory across iterations that would otherwise reset.
Setting a loop up
export OMNI_LOOP_ID=$(uuidgen)
export OMNI_LOOP_GOAL="Migrate the billing service off the legacy queue"
export OMNI_LOOP_BUDGET=100000
export OMNI_LOOP_ITERATION=0
| variable | constraint |
|---|---|
OMNI_LOOP_ID | alphanumeric and dash, 64 characters |
OMNI_LOOP_GOAL | 500 characters, no shell metacharacters |
OMNI_LOOP_BUDGET | token budget per iteration, up to 10M |
OMNI_LOOP_ITERATION | current iteration, default 0 |
OMNI_SUBAGENT=1 | sub-agent mode |
OMNI_AGENT_ID | identity, so traces stay separable |
Budget
The budget is estimated context window usage per iteration, not a spend limit.
| loop shape | budget | what OMNI does |
|---|---|---|
| quick fix, 1 to 5 iterations | 200,000 | passive tracking |
| feature work, 5 to 20 | 100,000 | active distillation, engrams |
| large refactor, 20 to 100 | 80,000 | aggressive distillation, predictive warnings |
| marathon, 100+ | 60,000 | maximum compression, loop memory persistence |
Warnings fire at 65% and critical at 82%, adjustable with OMNI_PRESSURE_WARN and
OMNI_PRESSURE_CRITICAL.
Do not set a budget above 1M: warnings will never fire before real exhaustion. Do not set one below 30K: the agent will compact constantly and lose short-term memory.
The goal string also shifts distillation aggressiveness. A goal containing “test” preserves test detail, “debug” keeps error context, “refactor” compresses harder.
Tools an orchestrator calls
None of these are advertised by default. OMNI tells a host about the tools its tier
actually uses, and the loop tools are outside that set, so an orchestrator that calls them
needs OMNI_MCP_TOOLS=all in its environment. omni doctor prints which set is in force.
The MCP tools reference has the per-tier lists.
| tool | when |
|---|---|
omni_loop_status | once before each iteration, the cheapest full picture |
omni_budget_status | before anything expensive |
omni_set_loop_context | when the goal or scope shifts mid-loop |
omni_loop_memory | read and write memory that survives a session restart |
omni_verify | as a checker, to evaluate the maker’s recent work |
Maker and checker
Two agents, one shared context layer.
LOOP_ID=$(uuidgen)
# the loop tools are outside the default advertised set
export OMNI_MCP_TOOLS=all
# maker
export OMNI_AGENT_ID=maker OMNI_LOOP_ID=$LOOP_ID
claude "Implement: $GOAL"
# checker
export OMNI_AGENT_ID=checker OMNI_SUBAGENT=1
RESULT=$(claude "Verify the implementation of: $GOAL. Use the omni_verify tool.")
case "$RESULT" in
*PASS*) echo "verification passed" ;;
*) echo "checker found issues" ;;
esac
Distinct OMNI_AGENT_ID values are what keep the two from contaminating each other.
Traces are tagged by agent, so omni_verify can read across sessions while writes stay
isolated.
Four things that make it work: give the checker specific measurable criteria, keep
last_n_calls between 5 and 20, escalate to a human after three consecutive checker
failures, and remember that every interaction is logged so the audit trail is real.
Monitoring
omni stats # real-time
omni stats --view detail
omni stats --json # for an orchestrator to read
omni doctor # health
omni handoff is not a CLI subcommand. It was removed. The omni_handoff MCP tool
is unchanged, so session export is reachable from an MCP client rather than a shell.
A caution about the numbers
Every figure a loop reports is scoped by agent_id. If the orchestrator and the agents
share one id, the maker’s savings and the checker’s are one number and neither is
meaningful. Set the id per role before the first iteration, not after you notice.