Playbooks#Coding agent#Prompts

Prompting Opus 5.5: the official guide says subtract

The official Opus 5.5 prompting guide is mostly about deleting instructions: effort is the main control now. Which level fits which work, and what to cut.

A black knob on a silver 0–8 scale dial, pointer resting near the low end — a metaphor for setting the right effort level (photo by Retired electrician, CC0, Wikimedia Commons)

Paste your Opus 5 prompts into 5.5 unchanged and two things tend to happen: turns run longer and cost more, and a few instructions that always worked now get ignored — or rejected outright. The reason is spelled out in Anthropic’s official prompting guide, Prompting Claude Opus 5.5, published shortly after the model’s launch: thinking is always on in this generation and cannot be turned off, and the effort setting — not your prompt — controls how much the model reasons. Behavior you used to coax out with prompt language is now either the effort dial’s job or something to delete.

Fact-checked as of 2026-10-02. Every figure and quote below comes from the official documentation linked in the text.

Key points

  • Opus 5.5 defaults to medium effort; in Anthropic’s testing it matched or beat Opus 5 at high on multistep coding in real repositories, in fewer steps and with fewer tokens.
  • The guide’s bias is deletion: removing “think carefully before answering” made replies start sooner with no clear quality decline, and asking the model to write out its reasoning can be declined outright by the reasoning_extraction classifier.
  • In Claude Code, change effort with /effort or --effort; the ultrathink keyword deepens reasoning for a single turn without touching your saved level.

Why old prompts misbehave

Adaptive thinking is always on in 5.5, and effort is the primary control trading off intelligence, latency, and cost. The catch: at a given level, 5.5 thinks more per turn than Opus 5 did — especially at xhigh. Carry over your old effort setting and you get longer turns and more output tokens without asking for them. Anthropic’s own diagnosis starts there: if turns got longer and pricier after upgrading, suspect the level before the prompt.

The offsetting good news, per the same document: 5.5 generates output tokens more than 30 percent faster than Opus 5 and tends to finish the same task with fewer tokens. On quality, medium matched or beat Opus 5 at high on multistep coding tasks in real repositories, and on several coding evaluations low comes close to Opus 5’s high at much lower cost. Most tasks don’t need a longer prompt — they need the dial set right.

Anthropic states it flatly: to get less thinking, lower the effort level first. Lowering effort reduces thinking more reliably than prompt instructions do.

Choosing an effort level

The 5.5 series supports five levels: low, medium, high, xhigh, max. Opus 5.5 defaults to medium (Opus 5 defaulted to high); Sonnet 5.5 defaults to high on the API. Claude Code’s documentation maps each level to a use case:

Level Fits Examples
low Quick exchanges you review each result Brainstorming, small changes
medium 5.5’s default; day-to-day work with clear scope Routine coding, knowledge work
high Verification matters, edge cases likely Bug fixes
xhigh Deeper reasoning at higher token spend Tough problems
max Hard autonomous problems Security-vulnerability hunting

The supporting rules, all from the official docs: keep xhigh and max for work where you’ve measured a quality gain — Anthropic specifically warns that max may overthink. Set max_tokens generously, because thinking tokens count toward the output limit; for agentic coding the recommendation is 128,000, the model’s maximum. Changing the top-level effort between requests invalidates the prompt cache — to vary effort per turn on the API, use the per-message effort change (beta), which preserves it. Sonnet 5.5 additionally supports thinking: {"type": "between_tools"} to skip up-front thinking at low through high (a 400 error at xhigh and max).

As for method, the official advice is one sentence: don’t carry settings over from an earlier model — run an effort sweep across all five levels on your own evals.

Delete these instructions first

The guide spends two full sections on removal, which is unusual for a prompting guide.

“Think carefully before answering” — cut it. The model decides for itself how much to think, and effort is the switch. In testing in a chat product, removing that line made replies start sooner with no clear decline in quality.

“Write out your reasoning” — cut it, and expect possible refusals. 5.5 runs safety classifiers, including reasoning_extraction, which declines requests that push the model to reproduce its internal reasoning in the response text. To follow the model’s thinking, set display: "summarized" and read the summarized thinking blocks; asking for a short explanation of an answer or a summary of actions taken is still fine.

Anything built around thinking: disabled — 5.5 returns a 400 error at every level. Re-test the mitigations you added for thinking-disabled Opus 5 (permission to speak before tool calls, no internal tags); most of those instructions can go too.

Multi-turn re-thinking — in chat, 5.5 sometimes goes back over an earlier answer while thinking about a new message, even a short follow-up. Anthropic offers two sentences for the end of your system prompt, quoted verbatim: “Once you have answered something, treat that answer as done. On later turns, focus your thinking on what the user is asking now, and don’t go back over an earlier answer unless the user asks about it or points out a problem with it.” Tested effect: faster reply starts, no quality hit. The trade-off — the model becomes less likely to flag its own earlier mistakes — means leaving it out for long analyses and agentic tasks where a later step can expose an earlier error.

Four lines worth adding

Deletion first, but the official guides do recommend adding a few things back.

Stop the premature wrap-up. On long tasks the model may post a summary, announce the next step, and end its turn — a place an unattended agent loop dies. The fix: state the completion condition up front, keep a checklist, and when the turn ends with items open, send a short continuation message (the doc’s example: “Your task list still has open items: migrate the remaining two endpoints and update their tests. Continue with them.”). Cap automatic continuations at two or three, then review by hand. There’s also a longer system-prompt paragraph that names four early-stop patterns to avoid; it must be added from the first request of the session — adding it mid-session edits the conversation prefix and invalidates earlier thinking blocks. Expect somewhat more tool calls and tokens, and keep your own confirmation step for risky actions.

Explore before acting in multi-app workflows. For agents working across email, docs, and spreadsheets, the information a task depends on often sits somewhere the request didn’t point to. One sentence in the system prompt — Anthropic’s “Before taking any action, explore broadly with tool calls…” — measurably increased correct completions at both medium and max, at the cost of slightly more tool calls. Since it tells the model to act on what it finds, keep untrusted content out of the records it searches.

Give multiagent teams a time budget. 5.5 pays close attention to elapsed time. Have your harness append a line like “elapsed 340s / 1200s” to each message; in Anthropic’s evaluations, small agent teams finished considerably sooner with answer quality comparable to a single agent. The budget is advisory — keep your own timeout for a hard stop — and know that under time pressure the model may search and verify a little less.

Tag pasted text. Text a user pastes from elsewhere can carry instructions. Wrap each pasted block in opening and closing tags carrying the same random ID, plus a system note; that sharpens resistance to indirect prompt injection. It can make the model slightly more cautious at times, so measure it on your own tasks.

One more tested line, from the companion Sonnet 5.5 guide: at xhigh and max, the model starts its own review rounds after finishing and pads the diff with tests, docs, and small files. Anthropic’s recommended system-prompt addition reads: “When the work the user asked for is done and its checks pass, stop and report. Don’t start extra rounds of review or hardening on your own, and don’t launch reviewer sub-agents unless the user asked for a review. If you think a deeper review is worth doing, say so at the end.” On coding tasks at max effort, that stopped the model from launching reviewer subagents and cut session cost by about a third with no change in quality. Summaries circulating on Chinese social media attribute this figure to the Opus 5.5 guide; it’s actually from the Sonnet 5.5 one.

Inside Claude Code

In Claude Code you don’t tune effort through prompts at all. /effort sets it directly (/effort auto clears a saved level and returns to the model default); claude --effort high covers one session; CLAUDE_CODE_EFFORT_LEVEL and per-model modelSettings persist it (settings won’t accept max). The arrow keys in the /model picker adjust it too, and 5.5 defaults to medium here as well. For a one-off hard problem, the ultrathink keyword deepens reasoning for that turn only.

As for the prompt itself, Claude Code’s official best practices point the same direction as the 5.5 guide: be specific — name the file, the scenario, and how to verify (“write a test for foo.py covering the edge case where the user is logged out, avoid mocks” beats “add tests”). Give the model a check it can run — tests, a build, a screenshot comparison; in Anthropic’s words, it’s “the difference between a session you watch and one you walk away from.” And after two corrections on the same problem, /clear and restart with a better prompt: “A clean session with a better prompt almost always outperforms a long session with accumulated corrections.” Apply the same subtraction to CLAUDE.md — for each line ask whether removing it would cause mistakes; a bloated memory file makes the model ignore your real instructions.

Charts and screenshots

Reading dense charts is where 5.5 breaks most from its predecessor: at its lowest effort setting it read values off dense charts more accurately than Opus 5 did at its highest, using a small fraction of the output tokens. Don’t raise effort for chart work. For harder inputs, feed technical drawings at higher resolution, and for the densest material run the model as an agent with a PIL/OpenCV container so it can crop, zoom, measure, and verify — more accurate and cheaper than maxing out.

Frontend style requests deserve the same specificity: “avoid the generic AI look” is, in the guide’s own framing, swapping one default aesthetic for another. Name the concrete things you don’t want — blue-purple gradients, the overused fonts — and extend the list as you review output.

One number to double-check

A widely shared Chinese summary credits the stop-condition line to the Opus 5.5 guide and its one-third cost saving to a max-effort test there. The figure is real — but it comes from the Sonnet 5.5 guide, and the exact experimental claim is narrower: on coding tasks at max effort, the instruction stopped self-started reviewer subagents and cut session cost by about a third with no change in quality. We couldn’t find it in the Opus 5.5 guide, the effort documentation, or the migration guide. The neighboring claims that do have Opus 5.5 test results behind them: removing think-carefully lines speeds up replies with no clear quality decline, and a periodic “say what you’re doing” reminder roughly halved the share of agentic tasks with a long silent stretch at no measurable cost change.

Three steps to start

  1. Run /effort auto in Claude Code to confirm 5.5 is back on its medium default rather than an old saved level.
  2. Delete “think carefully” and “write out your reasoning” from your project prompts, rerun an old task, and compare steps and tokens.
  3. Give the next long task a completion condition and a check that can actually run, and let the model continue until the checklist is clear.

Get the dial right and delete the stand-in instructions first — prompt technique means something different after that. For examples of how others prompt 5.5, see our prompt gallery breakdown and the video workflow playbook.