Claude Sonnet 5.5 API
Claude Sonnet 5.5 launched on 2026-09-28 at the same official price as Sonnet 5. This page covers the official benchmarks, how it compares with Sonnet 5 and Opus 5.5, which effort level to use, third-party and community reviews, the 400 errors you will hit migrating from Sonnet 5, and the Claude Code setup.
Price check: official, OpenRouter and Wokey
USD per million tokens. Official and OpenRouter rates were checked by hand on 2026-09-29, excluding tax; the Wokey rate is read from the live price list.
| Channel | Input | Cache read | Output |
|---|---|---|---|
| Official API | $2 | $0.2 | $10 |
| OpenRouter | $2 | $0.2 | $10 |
| Wokey | $0.3 | $0.03 | $1.5 |
· Official pricing page · OpenRouter endpoint list
Use it in Claude Code
This model supports /v1/messages on Wokey. Set these variables, then start claude.
export ANTHROPIC_BASE_URL="https://api.wokey.ai"
export ANTHROPIC_AUTH_TOKEN="$WOKEY_API_KEY"
export ANTHROPIC_MODEL="claude-sonnet-5-5"Claude Sonnet 5.5 at a glance
- Release date
- 2026-09-28, the second Claude 5.5 model (Opus 5.5 shipped 09-22; Anthropic says Haiku 5.5 follows in the coming weeks)
- Positioning
- A direct upgrade over Sonnet 5: 30%+ faster and up to 30% cheaper for most work because it uses fewer tokens. Anthropic calls it the fastest Sonnet to date
- Context and output
- 1M context; 128K max output, or 300K on the Batch API with the output-300k-2026-03-24 beta header
- Input and cutoff
- Text and image input; knowledge cutoff June 2026
- Default effort
- high on the API, medium in Claude Code and the Claude apps; five recalibrated levels: low / medium / high / xhigh / max
- Official price
- $2 input / $10 output per million tokens, same as Sonnet 5; cache read $0.20, 5-minute cache write $2.50, 1-hour $4; Batch $1 / $5
- Model ID
- claude-sonnet-5-5 (anthropic.claude-sonnet-5-5 on Bedrock)
· Sources: Anthropic · What’s new
Official benchmarks: Sonnet 5.5 vs Sonnet 5 vs Opus 5.5 vs GPT-6 Sol
From Anthropic’s launch page; these are vendor-reported results. “—” means no score was published. Anthropic did not publish SWE-bench.
| Benchmark | Sonnet 5.5 | Sonnet 5 | Opus 5.5 | GPT-6 Sol |
|---|---|---|---|---|
| Terminal-Bench 4.0 | 70.6% | 10.3% | 66.4% | — |
| CursorBench 4.0 | 55.5% | 34.1% | 57.8% | — |
| OSWorld 2.1 | 80.1% | 57.0% | 81.8% | — |
| Humanity's Last Exam (tools) | 64.5% | 54.9% | 67.7% | — |
| Chartography | 61.6% | 15.6% | 64.4% | 53.6% |
| GDPval-AA v2.1 | 1844 | 1449 | 1846 | 1487 |
| AA-Briefcase v1.1 | 1811 | 1359 | 1822 | 1483 |
In Artificial Analysis’ independent run, Terminal-Bench came out at 64% for Sonnet 5.5 and 60% for Opus 5.5, lower than the official figures but in the same order. Anthropic also states that Opus 5.5 remains clearly stronger at complex, open-ended work that needs sustained judgment.
Which to use: Sonnet 5.5, Sonnet 5 or Opus 5.5
vs Sonnet 5
Same price, faster, and ahead on every published benchmark, so it is a drop-in upgrade. Check the migration list below first: disabling thinking, forced tool use and non-default sampling parameters all return 400 on Sonnet 5.5.
vs Opus 5.5
Use Sonnet 5.5 for well-scoped everyday coding, bug fixes, and documents, spreadsheets and slides; use Opus 5.5 for complex, open-ended work that needs sustained judgment. Opus 5.5 lists at twice the price ($4 / $20), but both models read cache at $0.20. A common pattern is Opus 5.5 as the orchestrator with Sonnet 5.5 sub-agents.
Choosing effort
Anthropic recommends starting agentic coding at medium and moving hard tasks to high; use medium or low for chat. Artificial Analysis found high to be its most competitive setting; at max it averaged about 193K output tokens per task, the highest they have measured. A common community view: if you need xhigh or max, try Opus 5.5 at low or medium instead.
Third-party and community reviews
- Artificial Analysis: Intelligence Index 56, second only to Opus 5.5 (max) and 18 points above Sonnet 5; hallucination rate 47% versus 59% for Opus 5.5. An index task averaged $7.60 at official prices, about 50% more than Sonnet 5, which leaves it off the cost/intelligence Pareto frontier. Every Wokey meter is 15% of the official rate, so the same task comes to about $1.14.
- Hacker News (545-point thread): Positive: noticeably fast, with agentic coding close to Opus. Critical: at high or xhigh effort you may as well run Opus 5.5 at low or medium, so the two overlap; cache reads cost the same as Opus; some false refusals on security questions. Simon Willison saw max effort burn 128K thinking tokens in 15 minutes.
Migrating from Sonnet 5: requests that return 400
These changes come from Anthropic’s What’s new page. Check your requests for them before switching the model to claude-sonnet-5-5.
- thinking: {"type": "disabled"} returns 400. To skip up-front thinking, use {"type": "between_tools"}; it only works at low, medium or high effort and returns 400 at xhigh or max.
- tool_choice of any or tool returns 400. Use auto and add strict: true to the tool definition to enforce the argument shape.
- Non-default temperature, top_p or top_k returns 400; remove those fields.
- Thinking blocks are bound to the model, conversation prefix and account: replaying them after editing history returns 400, and blocks sent from another account are silently dropped.
- The computer_20251124 tool version is no longer supported.
- Text between tool calls now comes back inside thinking blocks; if your client hides thinking, progress updates between tool calls disappear.
- The minimum cacheable prompt drops from 1,024 to 512 tokens; refusals return stop_reason: "refusal".
Request body that skips up-front thinking (/v1/messages)
{
"model": "claude-sonnet-5-5",
"max_tokens": 4096,
"thinking": { "type": "between_tools" },
"output_config": { "effort": "medium" },
"messages": [{ "role": "user", "content": "Fix the failing test in src/app.ts" }]
}- Model ID
claude-sonnet-5-5- Vendors
- anthropic
- Supply status
- Currently callable
Catalog pricing snapshot
Prices are per 1M tokens; request-time pricing controls settlement.
| Meter | Wokey | Official reference | Savings |
|---|---|---|---|
| Input | $0.3 | $2 | 85% |
| Output | $1.5 | $10 | 85% |
| Cache read | $0.03 | $0.2 | 85% |
| Cache write | $0.6 | $4 | 85% |
Official price source: https://www.anthropic.com/claude-sonnet-5-5
Model capabilities
- Context
- 1,000,000 tokens
- Max output
- 128,000 tokens
- Streaming
- Supported
- Tools
- Supported
- Vision
- Supported
- Supported API forms
/v1/messages
Send a request
/v1/messages
curl https://api.wokey.ai/v1/messages \
-H "x-api-key: $WOKEY_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{"model":"claude-sonnet-5-5","max_tokens":1024,"messages":[{"role":"user","content":"Hello"}]}'Use Claude Sonnet 5.5 in your tools
The gateway is https://api.wokey.ai (OpenAI-compatible clients usually take https://api.wokey.ai/v1), and the model ID is claude-sonnet-5-5. Open the guide for the client you use.
- Use Claude Sonnet 5.5 in Claude Code: Anthropic Messages · base URL and API key setup
- Use Claude Sonnet 5.5 in Claude Desktop: Anthropic Messages · base URL and API key setup
- Claude API cost calculator: what Claude Sonnet 5.5 costs per month: Compare monthly cost on the Anthropic API, OpenRouter and Wokey for your usage and cache hit rate.
- OpenRouter alternative: Wokey vs OpenRouter: Compare token prices, payment fees, API forms, and upstream sources.
- Verifiable AI API: check a Claude Sonnet 5.5 response came from the official upstream: Which responses carry a tee.proof, and what a proof does and does not show.
- Verify a Claude Sonnet 5.5 response online: AI API relay verifier: Paste a response with its tee.proof and check it locally in your browser; nothing is uploaded.
Frequently asked questions
When was Claude Sonnet 5.5 released, and how is it different from Sonnet 5?
Claude Sonnet 5.5 was released on 2026-09-28 with a June 2026 knowledge cutoff. It keeps Sonnet 5’s official price ($2 input / $10 output per million tokens), runs more than 30% faster and uses fewer tokens on most tasks. In the official benchmarks Terminal-Bench 4.0 rose from 10.3% to 70.6% and OSWorld 2.1 from 57.0% to 80.1%.
Should I use Claude Sonnet 5.5 or Opus 5.5?
Use Sonnet 5.5 for well-scoped everyday coding, bug fixes and document work at half the official price of Opus 5.5; use Opus 5.5 for complex, open-ended tasks that need sustained judgment. The official benchmark gaps are all under 5 points, Sonnet 5.5 is ahead on Terminal-Bench, and GDPval-AA is nearly tied (1844 vs 1846). If Sonnet 5.5 only finishes a task at xhigh or max, try Opus 5.5 at low or medium.
Why does disabling thinking return 400 on Claude Sonnet 5.5?
Sonnet 5.5 no longer accepts thinking: {"type": "disabled"}. Use thinking: {"type": "between_tools"} to skip up-front thinking and only think between tool calls; it works at low, medium and high effort and returns 400 at xhigh and max. tool_choice any / tool and non-default temperature, top_p or top_k also return 400. Through Wokey the gateway converts these for you: disabled becomes between_tools at low to high effort, while disabled at xhigh / max and manual budget_tokens fall back to default adaptive thinking; temperature, top_p and top_k are removed. tool_choice any / tool has no equivalent and still returns 400, so switch to auto.
How do I use Sonnet 5.5 in Claude Code, and which effort should I pick?
Claude Code runs Sonnet 5.5 at medium effort by default. To use it through Wokey, set ANTHROPIC_BASE_URL to https://api.wokey.ai, ANTHROPIC_AUTH_TOKEN to your Wokey API key and ANTHROPIC_MODEL to claude-sonnet-5-5. Anthropic recommends starting agentic coding at medium and moving hard tasks to high.
Is there a cheaper Claude Sonnet 5.5 API than the official price?
Wokey serves claude-sonnet-5-5 over the native Anthropic /v1/messages API at $0.30 input / $1.50 output per million tokens (official $2 / $10). With "Official verification" turned on for the API key, passthrough Messages responses carry a tee.proof signature you can check locally at /tools/verify to confirm the response is the official upstream's original output.
How much does the Claude Sonnet 5.5 API cost, and how much does it save versus the official API?
Wokey lists input at $0.3 per 1M tokens and output at $1.5 per 1M tokens; the official references are $2 and $10. That is 85% lower for input and 85% lower for output. Request-time pricing controls settlement, and the vendor pricing source is linked on this page.
What is the Claude Sonnet 5.5 API model ID?
The model ID for Claude Sonnet 5.5 on Wokey is claude-sonnet-5-5. Use it exactly as written in the request body's model field or in your client's model setting (for example ANTHROPIC_MODEL in Claude Code, or model in the Codex CLI config.toml); GET https://api.wokey.ai/v1/models also lists it.
Which API forms does Claude Sonnet 5.5 support, and how do I call it through Wokey?
Claude Sonnet 5.5 supports /v1/messages (Messages requests). Set the API base URL to https://api.wokey.ai, authenticate with a Wokey API key, and use the canonical model ID claude-sonnet-5-5; you can start with /v1/messages. The Send a request section includes a curl example for every supported API form.
How do I use Claude Sonnet 5.5 in Claude Code?
Claude Code calls Claude Sonnet 5.5 over /v1/messages. Set ANTHROPIC_BASE_URL="https://api.wokey.ai", ANTHROPIC_AUTH_TOKEN to your Wokey API key, and ANTHROPIC_MODEL="claude-sonnet-5-5", then start claude. Claude Desktop also uses /v1/messages; the setup guides cover both step by step.
What are the context window and maximum output for Claude Sonnet 5.5?
The catalog context window is 1,000,000 tokens and the maximum output is 128,000 tokens. These are separate limits, and a request must also satisfy the upstream constraints in effect when it is processed.
How are cache reads and cache writes priced for Claude Sonnet 5.5?
Cache read: Wokey lists $0.03 per 1M tokens and the official reference is $0.2 per 1M tokens. Cache write: Wokey lists $0.6 per 1M tokens and the official reference is $4 per 1M tokens.
How does calling Claude Sonnet 5.5 through Wokey differ from OpenRouter?
Wokey lists Claude Sonnet 5.5 at $0.3 input / $1.5 output per 1M tokens, against an official reference of $2 / $10. OpenRouter states it adds no markup on inference and charges its fee when you buy credits (5.5% by card). Wokey carries a curated set of models and has no per-provider routing parameters, so OpenRouter fits better if you need a wider catalog. The OpenRouter alternative page has the full comparison.
How do Claude Sonnet 5.5 and Claude Sonnet 5 differ in price and context?
In the published price comparison, Claude Sonnet 5.5 lists input/output at $0.3 / $1.5 with a 1,000,000-token context window; Claude Sonnet 5 lists $0.36 / $1.8 with a 1,000,000-token context window. This is a factual price-and-capacity comparison, not a quality ranking; also compare their supported API forms and capabilities.
Related models and documentation
Get API key · claude-sonnet-5-5