Claude Opus 5.5 API
Claude Opus 5.5 launched on 2026-09-22 at a 20% lower official price than Opus 5, with cache reads down from $0.50 to $0.20. This page covers the official benchmarks, how it compares with Opus 5 and Sonnet 5.5, which effort level to use, third-party and community reviews, the 400 errors you will hit migrating from Opus 5, and the Claude Code setup.
Price check: official, OpenRouter and Wokey
USD per million tokens. Official and OpenRouter rates were checked by hand on 2026-09-29, excluding tax; the Wokey rate is read from the live price list.
| Channel | Input | Cache read | Output |
|---|---|---|---|
| Official API | $4 | $0.2 | $20 |
| OpenRouter | $4 | $0.2 | $20 |
| Wokey | $0.6 | $0.03 | $3 |
· Official pricing page · OpenRouter endpoint list
Use it in Claude Code
This model supports /v1/messages on Wokey. Set these variables, then start claude.
export ANTHROPIC_BASE_URL="https://api.wokey.ai"
export ANTHROPIC_AUTH_TOKEN="$WOKEY_API_KEY"
export ANTHROPIC_MODEL="claude-opus-5-5"Claude Opus 5.5 at a glance
- Release date
- 2026-09-22, the first Claude 5.5 model; Sonnet 5.5 followed on 09-28
- Positioning
- Built for long-running agentic coding and knowledge work. Anthropic says it performs at the level of Fable 5.1 on most work, costs 40% less to run than Opus 5 and outputs more than 30% faster
- Context and output
- 1M context with no beta header; 128K max output, or 300K on the Batch API with the output-300k-2026-03-24 beta header
- Input and cutoff
- Text and image input; knowledge cutoff June 2026
- Thinking and effort
- Adaptive thinking is always on and cannot be disabled; five effort levels (low / medium / high / xhigh / max), with medium as the API default (Opus 5 defaults to high)
- Official price
- $4 input / $20 output per million tokens (Opus 5: $5 / $25); cache read $0.20, 5% of input (Opus 5: $0.50); 5-minute cache write $5, 1-hour $8; Batch 50% off
- Safeguards
- Most cyber tasks are handed to Opus 4.8, and biology research use needs the Life Sciences Verification Program; refusals return stop_reason: "refusal"
- Model ID
- claude-opus-5-5 (anthropic.claude-opus-5-5 on Bedrock; the same ID on Google Cloud and Microsoft Foundry)
· Sources: Anthropic · Overview · Migration guide
Official benchmarks: Opus 5.5 vs Opus 5 vs Fable 5.1 vs Sonnet 5.5 vs GPT-6 Astra
From Anthropic’s launch page; these are vendor-reported results. Opus 5.5 ran at max effort (xhigh on Terminal-Bench) and GPT-6 Astra scores are OpenAI-reported. The Sonnet 5.5 column comes from the Sonnet 5.5 launch page; “—” means no score was published.
| Benchmark | Opus 5.5 | Opus 5 | Fable 5.1 | Sonnet 5.5 | GPT-6 Astra |
|---|---|---|---|---|---|
| Terminal-Bench 4.0 | 66.4% | 52.3% | 55.8% | 70.6% | 57.9% |
| FrontierCode v1.1 (Main) | 54.4% | 48.0% | 50.3% | — | 53.3% |
| CursorBench 4.0 | 57.8% | 46.6% | 51.8% | 55.5% | — |
| GDPval-AA v2.1 | 1846 | 1708 | 1735 | 1844 | 1542 |
| AutomationBench | 40.0% | 26.9% | 31.4% | — | 41.4% |
| Humanity's Last Exam (tools) | 67.7% | 63.6% | 65.6% | 64.5% | 57.2% |
| Terminal-Bench-Science 0.1 | 58.7% | 29.0% | 52.6% | — | 64.6% |
| OSWorld 2.1 | 81.8% | 74.0% | 80.7% | 80.1% | — |
| Chartography (tools) | 89.0% | 83.4% | 88.4% | — | — |
Anthropic notes that production safeguards were on during testing and cyber tasks fell back to Opus 4.8 when they triggered, which likely lowered Opus 5.5’s scores. In Artificial Analysis’ independent run, Opus 5.5 scored 59.6% on Terminal-Bench, level with GPT-6 Astra and 11 points above Opus 5, and 61.4% on Humanity’s Last Exam.
Which to use: Opus 5.5, Opus 5 or Sonnet 5.5
vs Opus 5
Cheaper and ahead on every published benchmark: Terminal-Bench 4.0 rose from 52.3% to 66.4% and GDPval-AA from 1708 to 1846. Anthropic says that at default effort it already beats Opus 5 at max on Terminal-Bench for about a fifth of the cost. Check the migration list below first: disabling thinking, manual budget_tokens and forced tool use all return 400, and the default effort dropped from high to medium.
vs Sonnet 5.5
Use Opus 5.5 for complex, open-ended work that needs sustained judgment and Sonnet 5.5 for well-scoped everyday coding and documents. Opus 5.5 lists at twice Sonnet 5.5’s price, but both read cache at $0.20. At max effort Artificial Analysis measured Sonnet 5.5 using about 60% more output tokens per task than Opus 5.5, so a task cost $7.60 on Sonnet 5.5 versus $5.98 on Opus 5.5. A common pattern is Opus 5.5 as the orchestrator with Sonnet 5.5 sub-agents.
Choosing effort
Anthropic recommends a fresh effort sweep on your own evals instead of carrying over Opus 5 settings, and a large max_tokens at higher levels, starting at 64K for xhigh and max. On Artificial Analysis’ Intelligence Index, medium scored 51 at $1.34 per task, high 54 at $1.82 and max 58 at $5.98; medium through max all sit on the cost/intelligence Pareto frontier. Community reports show max can exhaust the 128K output limit while still thinking, so it is rarely worth making the default.
Third-party and community reviews
- Artificial Analysis: Intelligence Index 58 at max, #1 by several points; it leads six of the ten evaluations, and GDPval-AA is 111 Elo above Fable 5.1. At max it uses 1.6x Opus 5’s output tokens per task, but the price cut keeps a task at $5.98, level with Opus 5 (max) at $5.86. Every Wokey meter is 15% of the official rate, so the same task comes to about $0.90 at max and $0.20 at medium.
- Early-access customers (Anthropic launch page): Deloitte: caught 72% of known bugs in code review versus 56% for Opus 5 at high effort. Box: about a third of Opus 5’s tokens. Factory: the first model they would default to at medium effort. GitHub: among the fewest tokens and steps they measured. These are customer statements quoted by the vendor.
- Hacker News (1,804-point thread): Most of the thread debates the “pacing the frontier” wording, which many read as at odds with the release cadence. On the model itself, commenters describe it as a cost, efficiency and writing-style improvement rather than a capability jump, and welcome the price cut; one user saw max effort run out of the 128K token budget while still thinking, twice, with no answer.
Migrating from Opus 5: requests that return 400
These changes come from Anthropic’s migration guide. Check your requests and response parsing before switching the model to claude-opus-5-5.
- thinking: {"type": "disabled"} and {"type": "enabled", "budget_tokens": N} both return 400. Remove the thinking field or send {"type": "adaptive"}, and control depth with effort. Through Wokey the gateway strips these thinking fields before forwarding, so the request does not fail, but thinking still runs and bills as output tokens.
- tool_choice of any or tool returns 400, including on count_tokens. Use auto with strict: true or structured outputs; strict needs additionalProperties: false on every object in the schema.
- The default effort changed from high to medium, so a request without effort runs one level lower than on Opus 5. Set it explicitly.
- max_tokens covers thinking plus text. Re-size requests that used to disable thinking; start at 64K for xhigh or max.
- Responses can start with thinking blocks, so read text blocks by type instead of content[0].text. thinking.display defaults to "omitted" with empty text; set display: "summarized" to get readable summaries.
- Text between tool calls now comes back inside thinking blocks; to show progress, set display: "updates" (beta) or "summarized". Pass thinking blocks back unchanged in tool-use loops.
- Thinking blocks are bound to the model and the conversation: for accounts created on or after 2026-08-31, replaying them after editing history returns 400, and other models ignore them after a route change.
- The Claude API and Google Cloud reject computer_20251124; use computer_toolset_20260801 (Bedrock is unaffected). The minimum cacheable prompt drops to 512 tokens; non-default temperature, top_p or top_k and a trailing assistant prefill still return 400.
Request body after migration (/v1/messages)
{
"model": "claude-opus-5-5",
"max_tokens": 16000,
"thinking": { "type": "adaptive", "display": "summarized" },
"output_config": { "effort": "high" },
"messages": [{ "role": "user", "content": "Review this diff for concurrency bugs" }]
}- Model ID
claude-opus-5-5- Vendors
- anthropic
- Supply status
- Currently callable
Catalog pricing snapshot
Prices are per 1M tokens; request-time pricing controls settlement.
| Meter | Wokey | Official reference | Savings |
|---|---|---|---|
| Input | $0.6 | $4 | 85% |
| Output | $3 | $20 | 85% |
| Cache read | $0.03 | $0.2 | 85% |
| Cache write | $1.2 | $8 | 85% |
Official price source: https://www.anthropic.com/claude-opus-5-5
Model capabilities
- Context
- 1,000,000 tokens
- Max output
- 128,000 tokens
- Streaming
- Supported
- Tools
- Supported
- Vision
- Supported
- Supported API forms
/v1/messages
Send a request
/v1/messages
curl https://api.wokey.ai/v1/messages \
-H "x-api-key: $WOKEY_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{"model":"claude-opus-5-5","max_tokens":1024,"messages":[{"role":"user","content":"Hello"}]}'Use Claude Opus 5.5 in your tools
The gateway is https://api.wokey.ai (OpenAI-compatible clients usually take https://api.wokey.ai/v1), and the model ID is claude-opus-5-5. Open the guide for the client you use.
- Use Claude Opus 5.5 in Claude Code: Anthropic Messages · base URL and API key setup
- Use Claude Opus 5.5 in Claude Desktop: Anthropic Messages · base URL and API key setup
- Claude API cost calculator: what Claude Opus 5.5 costs per month: Compare monthly cost on the Anthropic API, OpenRouter and Wokey for your usage and cache hit rate.
- OpenRouter alternative: Wokey vs OpenRouter: Compare token prices, payment fees, API forms, and upstream sources.
- Verifiable AI API: check a Claude Opus 5.5 response came from the official upstream: Which responses carry a tee.proof, and what a proof does and does not show.
- Verify a Claude Opus 5.5 response online: AI API relay verifier: Paste a response with its tee.proof and check it locally in your browser; nothing is uploaded.
Frequently asked questions
When was Claude Opus 5.5 released, and how is it different from Opus 5?
Claude Opus 5.5 was released on 2026-09-22 with a June 2026 knowledge cutoff. The official price dropped from Opus 5’s $5 input / $25 output per million tokens to $4 / $20, and cache reads from $0.50 to $0.20. In the official benchmarks Terminal-Bench 4.0 rose from 52.3% to 66.4% and GDPval-AA from 1708 to 1846. Adaptive thinking is always on and cannot be disabled, and the API default effort changed from high to medium.
Should I use Claude Opus 5.5 or Sonnet 5.5?
Use Opus 5.5 for complex, open-ended tasks that need sustained judgment, and Sonnet 5.5 for well-scoped everyday coding, bug fixes and document work at half the official price. Both read cache at $0.20. The official benchmark gaps are all under 5 points and Sonnet 5.5 is ahead on Terminal-Bench (70.6% vs 66.4%). At max effort Artificial Analysis measured $7.60 per task on Sonnet 5.5 versus $5.98 on Opus 5.5, so Opus 5.5 is not always the more expensive choice at the top setting.
Why do disabling thinking and forced tool use return 400 on Claude Opus 5.5?
Adaptive thinking is always on in Opus 5.5, so thinking: {"type": "disabled"} and manual budget_tokens both return 400; remove the thinking field and control depth with effort. tool_choice any and tool also return 400; use auto with strict: true or structured outputs. Through Wokey the gateway removes unsupported thinking fields automatically, but forced tool use still returns 400.
Which effort should I use for Claude Opus 5.5, and how do I use it in Claude Code?
The API default is medium. Artificial Analysis measured $1.34 per task and an Intelligence Index of 51 at medium, $1.82 and 54 at high, and $5.98 and 58 at max. Start at medium or high, use xhigh or max only where your evals show a gain, and set max_tokens to 64K or more there. To use it in Claude Code through Wokey, set ANTHROPIC_BASE_URL to https://api.wokey.ai, ANTHROPIC_AUTH_TOKEN to your Wokey API key and ANTHROPIC_MODEL to claude-opus-5-5.
Is there a cheaper Claude Opus 5.5 API than the official price?
Wokey serves claude-opus-5-5 over the native Anthropic /v1/messages API at $0.60 input / $3.00 output per million tokens (official $4 / $20). With "Official verification" turned on for the API key, passthrough Messages responses carry a tee.proof signature you can check locally at /tools/verify to confirm the response is the official upstream's original output.
How much does the Claude Opus 5.5 API cost, and how much does it save versus the official API?
Wokey lists input at $0.6 per 1M tokens and output at $3 per 1M tokens; the official references are $4 and $20. That is 85% lower for input and 85% lower for output. Request-time pricing controls settlement, and the vendor pricing source is linked on this page.
What is the Claude Opus 5.5 API model ID?
The model ID for Claude Opus 5.5 on Wokey is claude-opus-5-5. Use it exactly as written in the request body's model field or in your client's model setting (for example ANTHROPIC_MODEL in Claude Code, or model in the Codex CLI config.toml); GET https://api.wokey.ai/v1/models also lists it.
Which API forms does Claude Opus 5.5 support, and how do I call it through Wokey?
Claude Opus 5.5 supports /v1/messages (Messages requests). Set the API base URL to https://api.wokey.ai, authenticate with a Wokey API key, and use the canonical model ID claude-opus-5-5; you can start with /v1/messages. The Send a request section includes a curl example for every supported API form.
How do I use Claude Opus 5.5 in Claude Code?
Claude Code calls Claude Opus 5.5 over /v1/messages. Set ANTHROPIC_BASE_URL="https://api.wokey.ai", ANTHROPIC_AUTH_TOKEN to your Wokey API key, and ANTHROPIC_MODEL="claude-opus-5-5", then start claude. Claude Desktop also uses /v1/messages; the setup guides cover both step by step.
What are the context window and maximum output for Claude Opus 5.5?
The catalog context window is 1,000,000 tokens and the maximum output is 128,000 tokens. These are separate limits, and a request must also satisfy the upstream constraints in effect when it is processed.
How are cache reads and cache writes priced for Claude Opus 5.5?
Cache read: Wokey lists $0.03 per 1M tokens and the official reference is $0.2 per 1M tokens. Cache write: Wokey lists $1.2 per 1M tokens and the official reference is $8 per 1M tokens.
How does calling Claude Opus 5.5 through Wokey differ from OpenRouter?
Wokey lists Claude Opus 5.5 at $0.6 input / $3 output per 1M tokens, against an official reference of $4 / $20. OpenRouter states it adds no markup on inference and charges its fee when you buy credits (5.5% by card). Wokey carries a curated set of models and has no per-provider routing parameters, so OpenRouter fits better if you need a wider catalog. The OpenRouter alternative page has the full comparison.
How do Claude Opus 5.5 and Claude Opus 5 differ in price and context?
In the published price comparison, Claude Opus 5.5 lists input/output at $0.6 / $3 with a 1,000,000-token context window; Claude Opus 5 lists $0.75 / $3.75 with a 1,000,000-token context window. This is a factual price-and-capacity comparison, not a quality ranking; also compare their supported API forms and capabilities.
Related models and documentation
Get API key · claude-opus-5-5