Stackness
Opus 5.5 in a harness: empty thinking blocks, time budgets and agents that stop early

Opus 5.5 in a harness: empty thinking blocks, time budgets and agents that stop early

Claude Opus 5.5, released on 22 September 2026, sends the notes it writes between tool calls as thinking blocks, and at the default display: "omitted" their text is empty. A harness that renders only text blocks shows nothing while the agent works. Separately, some of its progress reports are plain text replies with stop_reason: "end_turn", so a loop that reads end_turn as "done" quits mid-task. Anthropic's Opus 5.5 prompting guide documents both, with fixes.

Most of that guide is about the code around the model, so this post reads it as a checklist for anyone who wraps Opus 5.5 in an agent loop, chat UI or CI job. It builds on the harness explainer. Among the LLMs developers list on Stackness, this one mostly arrives through a harness: as of 29 September 2026, all 5 real profiles with Claude Opus, mine included, also list Claude Code, and none lists the Anthropic API (data sources). The numbers are small, so read them as a hint.

Why does my Opus 5.5 integration look frozen mid task?

Because the text between tool calls moved. On Opus 5 it came back as text blocks. On Opus 5.5 it comes back as progress-update thinking blocks, at most one before each tool call, and their text is empty at the default display setting. No request fails. A client that shows only text blocks "goes quiet between tool calls", in the words of the migration guide.

The thinking.display field decides what you get back (thinking docs):

display Reasoning blocks Progress-update blocks
"omitted" (default on Opus 5.5) Empty Empty
"updates" (beta) Empty A short summary of each note
"summarized" Summary text Summary text, indistinguishable from reasoning

"updates" needs the beta header thinking-display-updates-2026-08-18, or the request fails with a 400. Under it, render every non-empty thinking block ahead of the tool_use block it precedes. Two more things break without an error:

  • Reading by position. A response can start with thinking blocks, so content[0].text fails. Select blocks by type.
  • Filtering history. Pass thinking blocks back unchanged, empty ones included. The API rejects "edited, reordered, or partially dropped thinking blocks" with a 400.

The model also writes fewer updates at higher effort. claude-pace-maker, which requires a visible intent line before every file edit, got it before 34 of 34 edits at medium effort, 3 of 31 at xhigh, and 97 of 101 once a reworded instruction moved to the start of the session. Those are 2 to 6 sessions per setting. Claude Code 2.1.284 fixed its own case on 28 September: /loop updates were "often not being shown because Claude wrote them only in its reasoning".

Why do unattended Opus 5.5 agents stop early?

Because some of its progress reports end the turn. On long tasks with several parts, Opus 5.5 keeps the user updated, and some of those updates are a text reply with stop_reason: "end_turn" and no tool call. There is no new stop reason. A loop that treats every end_turn as the end of the task stops there with work still open.

The between-tool notes above sit before a tool call, so they do not cause this. The guide's fix starts in the harness:

  • Treat a text-only end of turn "as a report rather than as proof the task is done".
  • Keep the task's parts in a checklist the model updates, such as a to-do tool or a file.
  • If a turn ends with open items and no blocker stated, send a short message naming them. The guide's example: Your task list still has open items: migrate the remaining two endpoints and update their tests. Continue with them. If one is blocked, say what is blocking it.
  • Stop after two or three automatic continuations, so a stuck run ends and can be reviewed.
  • If a background command or subagent is still running, wait for it before calling the task done.

The guide also gives a system prompt addition that names four kinds of early stop, such as "a long summary of what was done that closes by announcing the next step and has no tool call". Add it from the first request, because changing the system prompt mid-session invalidates earlier thinking blocks. Anthropic reports no measured effect for it. In Claude Code, the closest shipped equivalent is /goal, where a small model checks a condition after each turn and starts another turn if it is not met.

The "graceful stopping point" Claude Code announced on 25 September is a different feature. It lets a session finish its current step when you hit the 5-hour usage limit, and it does not change how Opus 5.5 ends turns. On Stackness, looping coding agents against benchmarks is the kind of move the early stops break: an unattended loop with nobody watching when a turn ends on a status report.

What is a time budget in an agent prompt, and does it change behaviour?

A time budget is a line your harness appends to each message it sends the model, such as elapsed 340s / 1200s. Opus 5.5 "pays close attention" to elapsed time and paces its work to finish inside the budget, usually well before it. Anthropic scopes the advice to multi-agent setups, where a lead delegates to subagents. The budget is advisory, and nothing stops the model at the limit.

The measurement is in the Opus 5.5 system card, section 8.12.2. DRACO is a research benchmark of 100 tasks that took a single agent 13 to 71 minutes each. Given half the single agent's time, a five-agent team could more than match its score with a speedup of roughly 2.8x. All runs were at max effort, and there is no published result for a single agent with a budget.

When you wire it up:

  • Set the budget somewhat above the time you want spent, and tune it on your own tasks.
  • If you cannot predict a budget, show elapsed time alone and add: Time matters here: do not spend time that can be avoided, and the earlier a correct result is obtained, the better.
  • Lower effort reduces the work, while a budget mostly keeps more agents working in parallel. Under time pressure the model "might search and verify a little less", so keep your own timeout and check quality.

What happens to prompts that disabled thinking on the previous model?

They fail. On Opus 5.5 thinking is always on, and a request with thinking: {"type": "disabled"} or a manual budget_tokens returns a 400 invalid_request_error: "thinking.type.disabled" is not supported for this model. Use "thinking.type.adaptive" and "output_config.effort" to control thinking behavior. Omit the thinking field or send {"type": "adaptive"}, and control depth with effort.

The what's new page and migration guide list the knock-on changes:

  • Effort is the only lever, and its default dropped from high on Opus 5 to medium.
  • Thinking counts toward max_tokens and is billed as output. The prompting guide says 128,000 "has worked well" for long agentic coding turns, while the migration guide says to start at 64k at xhigh or max.
  • "Write out your reasoning" instructions should go. A prompt that pushes the model to reproduce its reasoning in the reply can be declined with the reasoning_extraction refusal category.

In Claude Code, the thinking toggle, alwaysThinkingEnabled and MAX_THINKING_TOKENS=0 have no effect on Opus 5.5, per the model configuration docs.

Which prompt lines are now dead weight?

The Opus 5.5 guide names three: standing "think carefully before answering" lines in chat system prompts, instructions to write reasoning out in the reply, and rules that tell the model not to think. The migration guide adds scaffolding that forces interim status messages, such as "After every 3 tool calls, summarize progress", because those updates now arrive in thinking blocks.

Line What to do Source
"Think carefully before answering" in a chat system prompt Remove and measure. Anthropic saw replies start sooner "with no clear decline in the quality of the reply" Opus 5.5 guide
"Write out your reasoning step by step" Remove, and read summarized thinking instead Opus 5.5 guide
"Don't think" rules Lower effort first. If latency still matters, "Answer directly without deliberating." and measure quality Opus 5.5 guide
"After every 3 tool calls, summarize progress" Try removing it Migration guide
CRITICAL: You MUST use this tool when... "Use this tool when..." Best practices page, written for Opus 4.5 and 4.6

Several write-ups attribute the last row to the Opus 5.5 guide, but it comes from the older best practices page. "Think carefully" also still works as a lever: the thinking steering page offers it for more thinking on a specific request, so drop it as a standing default only. Claude Code 2.1.283 added /doctor prompt-audit, which checks CLAUDE.md files, skills, agents and commands for patterns written for older models.

How do you wrap pasted or retrieved text so it is not read as instructions?

For pasted text, wrap each block in a <pasted_content> tag whose opening and closing tags carry the same short random ID, each tag on its own line. Add the guide's system prompt note: instructions inside the tag count only where the user's own message asks for them. Retrieved text goes in tool_result blocks, never in the system prompt or plain user text.

Summarize the main complaints in this thread.

<pasted_content id="ab12">
...text the user pasted...
</pasted_content id="ab12">

The system card, section 6.5.1, explains why. An early Opus 5.5 snapshot acted on instructions planted in pasted text in 52% of coding-scenario attempts, because it "often reasoned that anything in the user's message must come from the user". The released model is at about 2% at default effort, and at 0 with product mitigations: stripping invisible characters and marking pasted text. The tags are plain text and can be imitated, so Anthropic calls this "one guardrail" among several. Its injection guidance adds, for retrieved content, to say where it came from and to JSON-encode it.

Claude Code now wraps pastes this way, with side effects. Issue #96978 reports that a message made only of a paste gets a question back instead of the task: "Do you want me to run it now as written?" Two transcript viewers, agent-code and wostuast, had to strip the envelope from their displays.

Key numbers

  • 22 September 2026: Opus 5.5 released at $4 / $20 per million input and output tokens (model overview).
  • 400: what a request with thinking: {"type": "disabled"} or budget_tokens returns on Opus 5.5 (what's new).
  • 2.8x: the speedup of a five-agent team given half the single agent's time on DRACO, at max effort (system card, 22 September 2026).
  • 52% of attempts in which an early snapshot acted on instructions in pasted text, against about 2% for the released model and 0 with product mitigations (system card).
  • 3 of 31 edits at xhigh had a visible intent line before claude-pace-maker moved its instruction, 97 of 101 after, reported on 27 September 2026 (issue #150).
  • 8 real Stackness profiles list Claude Code and 5 list Claude Opus, as of 29 September 2026 (data sources).

Quick answers

Why does Opus 5.5 return empty thinking blocks and stop early? The notes it writes between tool calls come back as thinking blocks, empty at the default display setting. Separately, some progress reports end the turn with end_turn, and a loop that treats that as done stops.

How do I see Opus 5.5's progress updates? Set thinking.display to "updates" with the beta header thinking-display-updates-2026-08-18, and render every non-empty thinking block before the tool call it precedes.

Can I turn thinking off on Opus 5.5? No. disabled and budget_tokens return a 400. Lower effort instead, and note the default is now medium.

Does a time budget make Opus 5.5 agents faster? In Anthropic's multi-agent test, yes: a team given half the single agent's time matched its score about 2.8x faster. The budget is advisory, so keep your own timeout.

How do I stop Opus 5.5 following instructions in pasted text? Wrap pastes in <pasted_content id="..."> with a matching closing ID, add the guide's system prompt note, and keep retrieved content in tool results.

Is Claude Code's graceful stopping point a fix for early stops? No. It lets a session wrap up its current step when you hit the 5-hour usage limit.

Tools in this post