Stackness
How do you tell whether a hosted model got nerfed? The frozen eval panel, and what livenerf measures

How do you tell whether a hosted model got nerfed? The frozen eval panel, and what livenerf measures

To tell whether a hosted model was nerfed, rerun a frozen set of prompts on a schedule, grade them by exact match and compare each prompt with its own baseline. As of 2 October 2026 nobody has shown a drop that way for Claude Opus 5.5. livenerf, a public attempt with a pre-registered decision rule, is on day 8 of a 10-day baseline and has published no drift result. Its rules allow no call before about 24 October. Its 78-question panel can detect a 7.5-point change per 10-day window.

The first "Opus 5.5 already nerfed?" post on r/ClaudeCode is timestamped 31 minutes before Anthropic's launch post on Reddit on 22 September, flaired as humor. A search of r/ClaudeAI and r/ClaudeCode on 2 October turned up 48 posts with a nerf claim in the title, and none comes with anything rerunnable.

How can you tell whether a hosted model was nerfed after launch?

Rerun the same prompts and compare. A frozen prompt panel, rerun on a schedule and graded without a judge model, is the only method that works on any hosted model and that someone else can repeat. Log probability tracking is far cheaper and more sensitive, but only where the API returns logprobs, and Claude's API reference has no such field. Vote counts and one-off canary prompts cannot separate a change from noise.

Method Example Detects Misses
Frozen panel, rerun and graded by exact match livenerf, MarginLab A pass-rate or token shift larger than the panel's noise The cause, small shifts, other surfaces
Log probability tracking Chauvin et al. Changes "as small as one step of fine-tuning" Any API without logprobs
Private composite score BridgeBench Nerf Bench Unknown: tasks and weights are private Anything a third party could check
Votes, sentiment, canary prompts nerf.watch, forum threads That people are complaining, or a swap to a much weaker model Whether anything subtle changed

The logprob paper watched 189 endpoints from 10 providers hourly for over four months and flagged 37 suspected changes. None was on the 19 OpenAI endpoints, and 34 were on open-weight models served by third parties. It puts the cost at $0.14 per endpoint per year.

A panel that changes between runs produces false alarms. In April 2026 BridgeBench announced that Claude Opus 4.6 had fallen from 83.3% to 68.3% on its hallucination benchmark, in a post that passed a million views. According to VentureBeat's report, the first run had six tasks and the second thirty, and on the six shared tasks the score moved from 87.6% to 85.4%. That part is second-hand, because I could not reach the original rebuttal.

What does livenerf measure?

Opus 5.5 at effort high, served through headless Claude Code on a Claude Max subscription, on 78 questions graded by exact match, run once a day for 30 days. It does not measure the raw API, coding, tool use or any other effort level. It is built on Inspect, and every number below is the author's own, from the repository at its 1 October commit.

  • Panel. 12 GPQA Diamond questions, 59 from MMLU-Pro, 4 competition math and 3 AIME, kept from 2,336 screened. The model got 97% of the candidates always right or always wrong, and those cannot show a drop.
  • Grading. Exact match, with no LLM judge, "since the judge would drift too".
  • Pinning. Claude Code is pinned at 2.1.280, and effort is passed explicitly.
  • Control. Opus 5 answers the 12 GPQA questions through the same harness. If both models move together, the harness or platform changed.
  • Rejected samples. Anything a fallback model served is thrown out. In a pilot, part of 3 of 5 samples was served by Opus 4.8.
  • Decision rule. A change of at least 3 points, outside the 99% interval, in two consecutive 10-day windows, with no matching move in the control.
  • Cost. About 115,000 output tokens a run, roughly $2.50 a day at API list price by my arithmetic.

Read the validation run before trusting a result in either direction:

Change against Opus 5.5 at high Accuracy Output tokens
Effort medium -4.2 ± 3.9 points -26%
Effort low -8.3 ± 4.5 points -62%
Opus 5 swapped in -3.8 ± 6.3 points -23%

None of the accuracy differences is significant, including a swap to the previous model. The token drops for lower effort are. The README says: "This instrument can't detect a same-family model swap of that size in a validation's worth of samples." A model that starts thinking less shows up in tokens before it shows up in accuracy, and livenerf logs tokens but builds no decision rule on them.

As of 1 October it had 8 of 30 days and no results row. Reading its published chart, the daily scores run from 52.6% to 64.1%, each with an interval of about 11 points either way. All eight days are baseline, so a dip lowers the reference and does not count as drift. The raw logs are not public. Issue 9 reports that the published analysis code cannot fire the rule within the 30 days. It was closed on 1 October, and the fix was not in the repository at that commit. The author's Reddit updates also say more than the repository does: they name SWE-bench, which the panel does not contain, and reported "no nerf is detected" on day 4, before any comparison was possible.

Why do user reports of degradation and vendor benchmarks disagree?

They measure different things. A vendor benchmark runs a pinned model at max effort through a fixed harness. A user meets the model through a product whose default effort, system prompt, router and safety fallbacks can all change. Every degradation a vendor has confirmed was in that layer. I found no vendor confirming a weight change or quantization behind a pinned model ID.

When What changed Size Source
August 2025 A routing bug sent Sonnet 4 requests to the wrong servers 16% of requests at the worst hour, about 30% of Claude Code users touched Anthropic postmortem
March 2026 Claude Code's default effort went from high to medium 34 days before the revert Anthropic postmortem
April 2026 One system prompt line capping text between tool calls "A 3% drop" on one eval Same
July 2026 Reasoning effort experiments in Codex Reverted OpenAI staff on X

Anthropic's April note says "The API was not impacted". Its versioning docs draw the same line: "Anthropic does not update the weights or configuration of an existing model ID", but the "request router, safety classifiers, and sampling logic" around it can change.

Opus 5.5 adds two gaps by design. Its launch benchmarks ran at max effort, while Claude Code and the apps default to medium. And its safeguards reroute flagged requests to Opus 4.8 or Opus 5, so the model that answers is not always the one you named. The harness checklist for Opus 5.5 covers what else moved.

Run-to-run noise accounts for some of the rest. In livenerf's own check, two passes of the same model at the same settings on the same night differed by 13.7 points. I found no controlled study of people perceiving a decline in an unchanged model, only hypotheses.

What belongs in a frozen prompt panel, and how many prompts is enough?

Prompts the model fails some of the time, graded by exact match, taken from the work you depend on. One run of 78 pass or fail prompts has a 95% interval of about 11 points either way, so a single rerun proves nothing. Pairing each prompt with its own baseline and repeating it is what gets livenerf to 7.5 points. A 3-point change needs about 969 questions.

What the sources agree on:

  • Middling prompts only. A prompt the model always passes cannot show a drop. livenerf's first panel scored 21 of 21 and was dropped.
  • Your own workload. livenerf has no coding questions, and MarginLab has only coding.
  • Pin the harness and set effort explicitly. A harness update looks the same as a model change.
  • Log output tokens and the served model for every sample, and run a control model through the same harness.
  • Prove it can see something. Run it against lower effort or an older model before trusting a quiet result.

The arithmetic for one run of pass or fail prompts at a 60% pass rate:

Prompts per run 95% interval of one run
50 ± 13.6 points
78 ± 10.9
200 ± 6.8
1,000 ± 3.0

The 969 figure is the worked example in Evan Miller's Adding error bars to evals: a 3-point difference, 80% power, 5% false positives. MarginLab publishes the same kind of numbers for its tracker: with 50 cases a day it needs a 12.8% change to reach p below 0.05, and 2.3% with 1,200 cases. Inspect handles the repeats: its epochs setting reruns each sample, and stderr(cluster=...) and ci_wilson() report the error bars.

How often should a drift check run, and what do you do when it trips?

Daily, at a fixed hour, with the decision taken on windows of a week or more. That is what livenerf and MarginLab both do. When it trips, investigate before you conclude anything: rerun it, rule out your own harness, check a control model, read the tokens and the served model, and only then pin, fall back or report.

  1. Wait for a second window. livenerf requires two in a row.
  2. Check what you changed. A CLI update, an SDK update or a new default. Anthropic's April 2026 incident was three Claude Code changes.
  3. Check the control. A second provider is a control too: the 2025 routing bug peaked at 16% of first-party Sonnet 4 requests and 0.18% on Amazon Bedrock.
  4. Read tokens and served model. Fewer output tokens at the same pass rate is the earlier signal.
  5. Read the status page. A wave of Opus 5.5 complaints followed a one-hour outage on 29 September.
  6. Pin what you can. A model ID pins weights. It does not pin the serving stack, and inside a subscription product it pins nothing.
  7. Report with examples. Anthropic credits "specific, reproducible examples" for its April fixes.

In a stack, this is a scheduled job next to your other evals: a frozen panel, a pinned harness and a log. As of 2 October 2026, 5 real profiles on Stackness list Claude Opus, and all 5 also list Claude Code (data sources). The numbers are small, but they say the surface to test is the harness, which is the one livenerf chose. Inspect and livenerf joined the AI tools developers list on Stackness with this post, so neither has users yet.

Key numbers

  • 78 questions in livenerf's panel, kept from 2,336 screened (livenerf, 24 September 2026, self-reported).
  • 7.5 points is the smallest change it can detect per 10-day window, at 80% power and a 99% test.
  • -3.8 ± 6.3 points: Opus 5 swapped for Opus 5.5 in livenerf's validation, not distinguishable.
  • 8 of 30 days collected and 0 drift results as of 1 October 2026. The first possible call is about 24 October 2026.
  • 16% of Sonnet 4 requests misrouted at the worst hour of the 2025 bug (Anthropic, 17 September 2025).
  • 969 questions to detect a 3-point difference (Miller, 1 November 2024).

Quick answers

How to tell if a model was nerfed? Rerun a frozen set of prompts on a schedule, grade them by exact match and compare each prompt with its own baseline over at least two windows. One rerun cannot tell you, because a 78-prompt run swings about 11 points either way by chance.

Has Opus 5.5 been nerfed? Nobody has shown it as of 2 October 2026. livenerf is still collecting its baseline, and BridgeBench's board showed Opus 5.5 at 94.2% of its launch score, inside its own 10% band.

Can vendors change a model behind a pinned ID? Anthropic says it does not change weights or configuration for an existing model ID. The router, safety classifiers, sampling logic and every product default around it can change.

How many prompts does a drift panel need? About 969 single-shot prompts for a 3-point change, by Miller's worked example. Fewer if you pair each prompt with its baseline and repeat it, which is how livenerf reaches 7.5 points with 78.

Should I use logprobs to detect model changes? Where the API returns them, yes. It costs cents a year and catches small changes. Claude's API reference has no logprobs field.

Tools in this post

Use any of these tools?

Put them on a Stackness profile, say how you use each one and see who pairs them the same way. It takes a couple of minutes.

Show my stack