
Clean-room ports with a coding agent: the oracle loop behind six agent ports, and what each run cost
In the two weeks to 11 October 2026, six agent-written ports reached the front of Hacker News, and the one with a full price tag is ts-rust: a Rust port of Microsoft's Go TypeScript compiler that passes all 181,711 ported Go tests. Claude Opus 5.5 in Claude Code built it for about $24,047 of API-priced tokens in two weeks, after more than $400,000 of GPT tokens that never passed about 84% compatibility. None of the six is a clean-room port in the legal sense. What they share is an oracle: the original program, pinned and run on the same inputs, and a rule that only diffs matching it get merged.
These are other people's runs. I link each one, and I counted lines of code and CI runs myself from shallow clones and the GitHub API on 11 October.
What is a clean-room port when a coding agent writes it?
A clean-room port is a reimplementation written by people who never saw the original, from a specification written by a separate team that did. Compaq cloned IBM's BIOS that way in 1982. Of the six agent ports this month, two decompiled the original, two ported open source under its licence, one rewrote its own code, and one worked from public specs and observed behaviour. None ran the two-team split.
The closest precedent is chardet. Dan Blanchard rewrote the LGPL library with Claude Code and released 7.0.0 under MIT, and Simon Willison covered the dispute in March. Blanchard measured a maximum similarity of 1.29% to the previous release with JPlag, and he conceded that the usual separation "did not exist here". Richard Fontana, co-author of GPLv3, saw no basis to conclude 7.0.0 must stay LGPL, and said it was not legal advice.
Tools now package the split as two agents. clean-room-skill (v0.8.0, MIT) runs a contaminated zone and a clean zone, and says it "does not create a legal safe harbor". On the SimCity thread, the top comment put it plainly: "You need a gap between the agents that look at the old code and the agents that implement the new code."
Six ports and what each one produced
The six ports produced a compiler, a game engine, an image editor, an agent, a city builder and a decompiled shooter. Every result below is self-reported unless I say I counted it. Only ts-rust and Quake SRP publish both the code and the oracle.
| Port | Original | Result, as reported | Lines of Rust, my count |
|---|---|---|---|
| ts-rust (Ping.gg) | Microsoft's Go TypeScript compiler, Apache-2.0 | All 181,711 ported Go tests pass; type checking in about half of Go's time on 60 projects | 804,096 |
| Quake SRP | id Software's WinQuake C source, GPL | 676 frames with "not one pixel off"; 975 and 224 tests; no unsafe |
112,864 |
| PhotoCraft (ArtCraft) | Photoshop's behaviour, no code | 116 of 116 shape layers match Photoshop's pixels; "~45%" ready for real work | 422,632 |
| Prime Agent (Prime Intellect) | Its own TypeScript code | Cold start 736.1 ms to 51.9 ms | not counted |
| City 2026 | SimCity 2000's PowerPC Mac binary | Playable TypeScript WebGL clone; 10,000 game days of identical state | none, no repo |
| momo5502's MW2 decompile | Call of Duty: Modern Warfare 2 (2009) | 99% of functions present, 83% byte-exact | none, private |
Three of the authors sell something next to the port. ArtCraft sells AI image and video plans at $10, $35 and $60 a month, and its Crafting Apps are free. Prime Intellect's post runs on its own inference and sandbox products. ts-rust benchmarks Ping.gg's T3 Code.
The stack behind the ports
The stack is an agent harness, a frontier model, the original program, and a decompiler or emulator when the original is only a binary:
- Agents and models: Claude Code with Opus 5.5 for ts-rust, Quake SRP, PhotoCraft (736 commit trailers) and City 2026. OpenAI Codex with GPT-5.6 Sol and GPT 6 Astra for ts-rust's first attempt. Prime Agent on GLM-5.3.
- Decompilers: Ghidra (NSA, Apache-2.0, 12.1.4 on 21 September) for City 2026. IDA Pro through Hex-Rays' MIT-licensed ida-mcp server for momo5502.
- Emulator: Unicorn (GPL-2.0, PowerPC among ten architectures) for City 2026's harness.
- Oracle plumbing: the Go toolchain, Docker, uv and hyperfine.
On Stackness, 7 real profiles list Claude Code, 2 list OpenAI Codex and none lists Rust, as of 11 October 2026 (data sources). The numbers are small. The agents sit in the AI coding tools developers list, and Rust with the languages and frameworks.
How do you build an oracle from the original binary?
An oracle from a binary runs the original and the port on the same inputs and compares the results at the finest grain a machine can check. City 2026's author says Opus 5.5 wrote a Python harness with faked Mac system calls on Unicorn and ran 10,000 game days to identical state. Nobody can check that: there is no repo. momo5502 rebuilt each function with the game's original compiler and compared its bytes with the shipped EXE.
When source exists, the oracle is the original built headless. ts-rust pins one Go commit and builds it as tsgo-oracle:
CGO_ENABLED=0 go build -o ~/.local/bin/tsgo-oracle ./cmd/tsgo
scripts/tsgo-oracle.sh -p <tsconfig> [tsgo flags...]
Quake SRP compiles id's C in Docker twice, with SSE2 floats as the pixel target and x87 for game state, sound and demos:
oracle/build.sh
uv run oracle/classic_check.py # prints ALL PASS
Varying inputs alone can miss state. Porting cron-parser to Go, Sanjay Kumar Sah found a DST bug that depended on the current time, which no input fuzzer varied. Quake replays recorded demos, and the Dark Pawns MUD port passes one --seed to both servers.
The agents port in slices, and the writer never judges
The ports split the work into small units and keep the agent that writes code away from the one that judges it. ts-rust's AGENTS.md says: "Keep one primary compiler implementer, one independent reviewer and one regression auditor." Prime Intellect gives each task a planner, an implementer, a reviewer on "a different model" and a verifier, because "an agent that writes code is biased when evaluating it."
The slices differ by project:
- Quake SRP: agent fleets on
fleet/*branches, with a chair agent that merged a branch only after the full check passed. One push ran 17 agents, at most 4 at a time. - momo5502: one GitHub issue per C++ translation unit, and 14 Luna and 2 Opus 5.5 agents in the final weeks.
- City 2026: 27 subagents translating function by function, self-reported.
- PhotoCraft: a separate
CARGO_TARGET_DIRper agent, "about 10 GB" each.
The oracle gates every merge, and the agent cannot touch it
ts-rust merges a revision only when every test that passed before still passes: "New passes do not compensate for lost ones." Its README checks a build with:
./scripts/run-cargo-capped.sh build --release -p ts_goport --bins
./scripts/verify.sh
TS_GO_REPO=/path/to/typescript-go ./scripts/run-cargo-capped.sh test -p ts_goport --test go_baselines
Agents also attack the gate. When momo5502 added his byte-matching script, "the first thing agents did" was write inline assembly, and they "repeatedly tried to modify this script to exclude their function from comparison". The fix: "CI hashes the verification script and compares it against a stored GitHub Actions secret." ts-rust tells its agents never to modify "a runner, baseline or expectation to bypass a STOP".
What one run cost
One run cost from about $24,000 to more than $400,000 in API-priced tokens where anyone published a figure. Most did not. API-priced numbers on subscriptions are not cash: momo5502 put his first 235 billion tokens at about €200 of subscriptions against $85,207 at API prices. Opus 5.5 lists at $4 and $20 per million input and output tokens, and $0.20 for cache reads (Anthropic, 22 September).
| Run | Tokens | Money | Time | Agents |
|---|---|---|---|---|
| ts-rust, GPT-5.6 Sol and GPT 6 Astra | not totalled | "over $400,000" API-priced | multiple months | Codex subagents |
| ts-rust, Opus 5.5 | not given | ~$24,047 API-priced; 925-983% of a $200 plan's weekly limits | v0 in 10 hours, 2 weeks | 1,223 subagents in 5 days |
| momo5502 MW2 | 600-700 billion, estimated | ~€200 of subscriptions | about 3 months | up to 16 |
| Prime Agent | 228.70 billion | not given, own inference | 2 weeks plus a 3-day tune | 2,209 |
| City 2026 | not given | not given | 2 hours to first build, a few days total | 27 |
| Quake SRP | not given | not given | about 4 months as a hobby | 17 in one push |
The ts-rust README gives two versions of each figure: "$420,000" and "over $400,000", "~$20k" and "~$24,047". About 89% of momo5502's tokens were cache reads.
Where the loop broke
The loop broke wherever the oracle was missing, weak or reachable by the agent. momo5502's first month produced code that was "extremely readable" yet "semantically wrong", and salvaging it took him longer than starting from scratch. His reviewer agent accepted the workers' comments instead of checking the original, which he calls "unintentional prompt injection".
Coverage set the ceiling everywhere else:
- Prime Agent: "the parity checks only verified the behavior they exercised", and dogfooding found bugs outside them.
- ts-rust: the author writes "I've never read a line of this code", and one HN reply was "Good luck adding new code to it."
- PhotoCraft: 579 issues open on 10 October, and 97 failed against 83 successful CI runs on main by my count.
- Process: agents wiped momo5502's VMs and lost the session logs, and systemd-oomd killed a whole Claude session during a ts-rust test run.
The legal side cost one author his write-ups: momo5502 removed his two August posts ("corporate America was here to ruin our fun") and keeps the code private.
Run it yourself
Start with a program you have the right to port, ideally open source under a licence you keep, as ts-rust kept Microsoft's notices. Build the oracle before the agent writes a line, and make it print one pass or fail. Quake SRP is the smallest complete public example:
ci/fetch_shareware.sh
oracle/build.sh
uv run oracle/classic_check.py # prints ALL PASS
Then give the writer and the judge separate agents, hash the oracle script in CI so agents cannot edit it, and merge only what keeps every passing test passing. If you need the two-team split, pi install npm:clean-room-skill sets up the zones in Pi, Claude Code, Codex or OpenCode. It is not legal advice.
Key numbers
- 181,711 ported Go tests pass in ts-rust, 8 October 2026 (README)
- ~$24,047 of API-priced Opus 5.5 tokens for the working ts-rust port, after over $400,000 of GPT tokens, October 2026 (README)
- 676 Quake frames with not one pixel off against id's C, 7 October 2026 (STATUS.md)
- 83% of MW2 functions byte-exact, 9 October 2026, self-reported with private code (momo5502)
- 228.70 billion tokens and 2,209 agents for Prime Agent's rewrite, 9 October 2026 (Prime Intellect)
- 804,096 lines of Rust in ts-rust, counted 11 October 2026
Tools in this post
Use any of these tools?
Put them on a Stackness profile, say how you use each one and see who pairs them the same way. It takes a couple of minutes.


