
Can a decision model prune your coding agent's context? What fast-jev-compaction keeps, costs and drops
Tools in this post
A decision model can prune a coding agent's context, but the evidence so far says it is not yet doing the choosing. fast-jev-compaction, a Claude Code plugin with 7,341 GitHub stars on 3 October 2026, asks Jev two yes-or-no questions per tool call and deletes what scores below 0.5. Independent tests timed a compaction at 0.65 to 1.6 seconds, against a 30-second median for Claude Code's own summary. At the default threshold, though, every independent replay I found kept almost no tool results, and on one dataset a fake scorer that answered 0 to everything removed about as much: 88.5% against 87.7%.
The compaction explainer covered how the plugin works and the cache objection. This post covers what the extra pass costs, what a scorer should never drop, and what the replays found about calibration.
What does fast-jev-compaction replace in a Claude Code session?
The summary step of compaction, both /compact and the automatic one. It is a Claude Code mod that hooks the compaction event and returns the original messages minus the tool calls and results Jev scored low. It also forces a compaction at 60% of the context window. If Jev fails, or the cut is under 25%, Claude Code's own summary runs instead.
Reading the code at its last commit, on 17 September: the plugin posts the conversation to TypeSafe's hosted API, with the last three user prompts as the goal and each tool result replaced by a note such as ok, 4213 chars (omitted). Jev never sees what a tool returned. For each call it answers whether the call still matters and whether its full output must stay. A result above 0.5 stays whole. A call above 0.5 keeps its first 300 characters and a truncation note. Anything else is deleted with its call, and nothing marks the gap.
Three things to check before installing. The README describes the early-access path, while Claude Code 2.1.287 made mods generally available on 1 October and ignores the flag it tells you to set. The public mods reference lists only a skip result for the compaction event, so I could not confirm that replacing messages still works on current versions. And the npm package called fast-jev-compaction is a fork by a different publisher.
How is scoring each tool call different from summarising the history?
A summary rewrites the history into new prose, so whatever survives is the model's paraphrase. Scoring deletes chosen tool calls and results and rewrites nothing, so whatever survives is exact. The cost moves from one slow LLM call to one or more fast classifier calls. The risk moves from distortion to deletion: a summary blurs a fact, a scorer removes it outright.
Claude Code already does some deletion before it summarises. Its docs say it "clears older tool outputs first, then summarizes the conversation if needed". The Claude API offers the same rule-based deletion as context editing, which removes the oldest tool results by age. fast-jev-compaction replaces the age rule with a probability.
What does the extra decision pass cost in latency and money per compaction?
Under a cent and about a second. Independent replays measured 0.65 to 1.6 seconds per compaction and, by my arithmetic, $0.0007 to $0.003 of Jev at TypeSafe's $0.042 per million input tokens. Claude Code's summary took a median 30 seconds in one test. The larger costs are indirect: more frequent compactions, a broken prompt cache, and a full transcript on resume.
| Measurement | Result | Measured by |
|---|---|---|
| Per compaction, plugin against built-in | 743 ms median against 30,375 ms, n=6 | Independent, issue #89 |
| Per compaction, 60 public sessions | 0.65 s median, about three cents of Jev in all | Independent, staysup.io |
| One large session | 1.24 s for 333,671 to 47,797 tokens | Independent, issue #65 |
| Test session | 440 to 615 ms | The author |
The README's "one fast request" is several. The plugin resends the whole state, up to about 25,000 tokens, with every batch of questions and runs the batches at once. One user saw 9 requests for a single compaction. At the default caps that is about $0.0013 a request, so around one cent for a large session.
The indirect costs:
- Earlier compactions. The 60% trigger fires long before Claude Code's default, which waits until about 967K tokens on a 1M model. Each compaction breaks the conversation's prompt cache.
- Resume. In issue #89, plugin-compacted sessions reloaded 166,000 to 168,000 tokens after
--resume, against about 28,000 for natively compacted ones, n=12 against n=10. - Re-reads. Deleted content comes back only by running the tool again. In the LaMR paper, an over-aggressive pruner raised total tokens by 6.6% on SWE-Bench Verified with Opus 4.6.
Local runtimes do not cut these costs yet. The plugin never passes a custom endpoint, so Ollaya or laya-mlx work only through the underlying library. Laya's 512 to 1,024-token context is far below the plugin's 25,000-token state.
Which parts of a session should never be dropped by a scorer?
Anything the agent cannot get back by re-running a tool: records of edits and other side effects, errors, subagent results, answers the user gave, and output from calls whose result has since changed. Also the task and the latest turns, which every pruning paper I read keeps verbatim. fast-jev-compaction protects the first message, the newest six messages and all text, but not edits, errors or subagent results.
Its state tells Jev that "the assistant can always re-run a tool or re-read a file". That is false for a deploy, a sent message or an API response that has moved on. In the issues, users propose rules ahead of the scorer:
- Keep edits. "Edit results are tiny (195 calls, 30k characters across all 89 logs), so keeping Edit / Write calls by rule costs almost nothing", one replay found.
- Keep receipts. People porting the plugin to Hermes Agent found "Jev scores receipt-bearing results at 0.25 (lowest of all fixtures)", for outputs carrying a message ID or commit.
- Mark what was cut. Without a marker, one session's agent reported finished work for 22 minutes "without calling a single tool. None of the reported artifacts existed." That is a single report and I have not reproduced it.
The papers agree on the shape. CliffCompaction keeps "the system prompt and the first user message (the task description)" and the most recent turns, keeps short results, and reduces long calls to signatures the agent can reissue. LaMR keeps code structure: "imports, definitions, scope headers, and control-flow siblings". The pattern is rules for what must stay, and a scorer only for the rest. The API's context editing has that rule built in, as an exclude_tools list.
Where does this sit next to native compaction and context editing?
In place of Claude Code's summary step, after its rule-based clearing, on your machine. Context editing deletes on Anthropic's servers by age, with an exclude list and a placeholder. Server-side compaction summarises. fast-jev-compaction is the only one of them that chooses by a model's probability, and the only one that leaves no marker.
| Approach | Chooses by | Rewrites | Always keeps | Extra cost |
|---|---|---|---|---|
| Claude Code native | Age, then an LLM summary | The whole history | CLAUDE.md, plan, up to 5 recent files, reloaded | One LLM request |
| API context editing | Age: oldest results first | Nothing; leaves a placeholder | Text, tool calls, last 3 results, excluded tools | No model call |
| fast-jev-compaction | Jev's probability per call | Nothing | First message, newest 6 messages, all text | One or more Jev requests |
| CliffCompaction, a paper | Length: results over 500 characters | Nothing | Task, recent turns, call signatures | No model call |
| pi-context-prune | An LLM summary per batch | Tool results | Originals, retrievable by ID | One LLM call per batch |
The sibling tools keep a way back. pi-context-prune, an extension for the Pi coding agent, keeps every original output behind a query tool. Another Jev-based plugin, fast-compact, saves every cut to a file and leaves a pointer. Its author reports keeping "75% of the outputs that turned out to be needed, vs 69% for a plain cut of the same size".
What can go wrong when the scorer is miscalibrated on your data?
It deletes almost everything, and the toast does not say so. At the shipped 0.5 threshold, independent replays found almost no tool results kept, with Jev's result scores mostly between 0.06 and 0.24. In one comparison its ranking did no better than keeping the newest outputs. Lowering the threshold swings results sharply: in one test, moving it by 0.05 changed the outcome by 81 points.
| Replay | Finding | Who |
|---|---|---|
| 256 scored calls | 0 results at 0.5 or above. Jev removed 87.7%, a fake scorer answering 0 removed 88.5%, and both dropped about 6.6% of results that were needed again | Issue #26 |
| 246 labelled pairs | AUC 0.544 for the shipped questions, 0.847 for "keep the largest outputs" | Comment on issue #26 |
| 60 public sessions, 1,729 candidates | "At matched budgets we couldn't distinguish it from recency" | staysup.io |
| 183-message session | 141 calls deleted, 4 kept, all of them pinned | Issue #56 |
The question wording is the likeliest cause. The result question packs two judgments into one: the output is still needed, and "re-running the tool would not do". Dropping the second clause lifted a load-bearing file read from 0.22 to 0.81 in one test, while noise stayed at 0.03. TypeSafe's own model notes list "Hiding several judgments inside one question" as a thing to avoid, and warn that "Accuracy falls as the state grows with content unrelated to the decision". The plugin sends up to 25,000 tokens of history, and Jev is judging outputs it cannot see.
The samples are small, the labels are proxies for "needed later", and several sessions were not in English, which TypeSafe says Jev handles less well. The author has not replied in any issue, and nothing has been merged since 17 September. If you run it, read the decisions: log line, which prints both probabilities for every call, before trusting the toast.
On Stackness, as of 3 October 2026, no real profile lists fast-jev-compaction and one lists Jev, mine, against 8 for Claude Code (data sources). The AI tools developers list beside their agents do not include a pruning plugin yet.
Key numbers
- 7,341 GitHub stars and no commits since 17 September 2026 for fast-jev-compaction (GitHub API, 3 October 2026).
- 743 ms median per compaction against 30,375 ms for Claude Code's summary, n=6, measured independently (issue #89, 22 September 2026).
- 0 of 256 tool results scored 0.5 or above in one replay, and a fake scorer answering 0 removed 88.5% against Jev's 87.7% (issue #26, 18 September 2026).
- 0.042 * *permillioninputtokensforJev([TypeSafe](https : //docs.typesafe.ai/models)), about * *0.0007 to $0.003 per compaction by my arithmetic.
- 166,000 to 168,000 tokens reloaded on
--resumeafter a plugin compaction, against about 28,000 after a native one (issue #89). - 0 real Stackness profiles list fast-jev-compaction and 1 lists Jev, as of 3 October 2026 (data sources).
Quick answers
Can you prune coding agent context with a decision model? Yes. fast-jev-compaction asks Jev whether each tool call and result still matters and deletes the rest in about a second. At its default threshold, though, independent replays found it kept almost nothing, so the speed comes from deleting, not from choosing well.
Is pruning better than summarising? It keeps what survives exact, where a summary paraphrases. It also deletes outright, so it needs rules that protect edits, errors, receipts and recent turns, and ideally a marker or a file the agent can recover cuts from.
How much does fast-jev-compaction cost? Under a cent of Jev per compaction in independent replays. The larger costs are earlier compactions, each breaking the prompt cache, and a full transcript reload on --resume.
Does fast-jev-compaction work with a local model? Not as shipped. The plugin always calls TypeSafe's hosted API. Its library accepts another endpoint, but Laya's context of 512 to 1,024 tokens is far below the plugin's 25,000-token state.
What threshold should I use? There is no tested answer. 0.5 kept almost nothing in every replay I found, and small changes swing the outcome widely. Read the decisions log on your own sessions before you settle on one.
Tools in this post
Use any of these tools?
Put them on a Stackness profile, say how you use each one and see who pairs them the same way. It takes a couple of minutes.


