Claude Sonnet 5.5 API

Claude Sonnet 5.5 launched on 2026-09-28 at the same official price as Sonnet 5. This page covers the official benchmarks, how it compares with Sonnet 5 and Opus 5.5, which effort level to use, third-party and community reviews, the 400 errors you will hit migrating from Sonnet 5, and the Claude Code setup.

Price check: official, OpenRouter and Wokey

USD per million tokens. Official and OpenRouter rates were checked by hand on 2026-09-29, excluding tax; the Wokey rate is read from the live price list.

ChannelInputCache readOutput
Official API$2$0.2$10
OpenRouter$2$0.2$10
Wokey$0.3$0.03$1.5

· Official pricing page · OpenRouter endpoint list

Use it in Claude Code

This model supports /v1/messages on Wokey. Set these variables, then start claude.

export ANTHROPIC_BASE_URL="https://api.wokey.ai"
export ANTHROPIC_AUTH_TOKEN="$WOKEY_API_KEY"
export ANTHROPIC_MODEL="claude-sonnet-5-5"

Claude Sonnet 5.5 at a glance

Release date
2026-09-28, the second Claude 5.5 model (Opus 5.5 shipped 09-22; Anthropic says Haiku 5.5 follows in the coming weeks)
Positioning
A direct upgrade over Sonnet 5: 30%+ faster and up to 30% cheaper for most work because it uses fewer tokens. Anthropic calls it the fastest Sonnet to date
Context and output
1M context; 128K max output, or 300K on the Batch API with the output-300k-2026-03-24 beta header
Input and cutoff
Text and image input; knowledge cutoff June 2026
Default effort
high on the API, medium in Claude Code and the Claude apps; five recalibrated levels: low / medium / high / xhigh / max
Official price
$2 input / $10 output per million tokens, same as Sonnet 5; cache read $0.20, 5-minute cache write $2.50, 1-hour $4; Batch $1 / $5
Model ID
claude-sonnet-5-5 (anthropic.claude-sonnet-5-5 on Bedrock)

· Sources: Anthropic · What’s new

Official benchmarks: Sonnet 5.5 vs Sonnet 5 vs Opus 5.5 vs GPT-6 Sol

From Anthropic’s launch page; these are vendor-reported results. “—” means no score was published. Anthropic did not publish SWE-bench.

BenchmarkSonnet 5.5Sonnet 5Opus 5.5GPT-6 Sol
Terminal-Bench 4.070.6%10.3%66.4%—
CursorBench 4.055.5%34.1%57.8%—
OSWorld 2.180.1%57.0%81.8%—
Humanity's Last Exam (tools)64.5%54.9%67.7%—
Chartography61.6%15.6%64.4%53.6%
GDPval-AA v2.11844144918461487
AA-Briefcase v1.11811135918221483

In Artificial Analysis’ independent run, Terminal-Bench came out at 64% for Sonnet 5.5 and 60% for Opus 5.5, lower than the official figures but in the same order. Anthropic also states that Opus 5.5 remains clearly stronger at complex, open-ended work that needs sustained judgment.

Which to use: Sonnet 5.5, Sonnet 5 or Opus 5.5

vs Sonnet 5

Same price, faster, and ahead on every published benchmark, so it is a drop-in upgrade. Check the migration list below first: disabling thinking, forced tool use and non-default sampling parameters all return 400 on Sonnet 5.5.

vs Opus 5.5

Use Sonnet 5.5 for well-scoped everyday coding, bug fixes, and documents, spreadsheets and slides; use Opus 5.5 for complex, open-ended work that needs sustained judgment. Opus 5.5 lists at twice the price ($4 / $20), but both models read cache at $0.20. A common pattern is Opus 5.5 as the orchestrator with Sonnet 5.5 sub-agents.

Choosing effort

Anthropic recommends starting agentic coding at medium and moving hard tasks to high; use medium or low for chat. Artificial Analysis found high to be its most competitive setting; at max it averaged about 193K output tokens per task, the highest they have measured. A common community view: if you need xhigh or max, try Opus 5.5 at low or medium instead.

Effort documentation

Third-party and community reviews

  • Artificial Analysis: Intelligence Index 56, second only to Opus 5.5 (max) and 18 points above Sonnet 5; hallucination rate 47% versus 59% for Opus 5.5. An index task averaged $7.60 at official prices, about 50% more than Sonnet 5, which leaves it off the cost/intelligence Pareto frontier. Every Wokey meter is 15% of the official rate, so the same task comes to about $1.14.
  • Hacker News (545-point thread): Positive: noticeably fast, with agentic coding close to Opus. Critical: at high or xhigh effort you may as well run Opus 5.5 at low or medium, so the two overlap; cache reads cost the same as Opus; some false refusals on security questions. Simon Willison saw max effort burn 128K thinking tokens in 15 minutes.

Migrating from Sonnet 5: requests that return 400

These changes come from Anthropic’s What’s new page. Check your requests for them before switching the model to claude-sonnet-5-5.

  • thinking: {"type": "disabled"} returns 400. To skip up-front thinking, use {"type": "between_tools"}; it only works at low, medium or high effort and returns 400 at xhigh or max.
  • tool_choice of any or tool returns 400. Use auto and add strict: true to the tool definition to enforce the argument shape.
  • Non-default temperature, top_p or top_k returns 400; remove those fields.
  • Thinking blocks are bound to the model, conversation prefix and account: replaying them after editing history returns 400, and blocks sent from another account are silently dropped.
  • The computer_20251124 tool version is no longer supported.
  • Text between tool calls now comes back inside thinking blocks; if your client hides thinking, progress updates between tool calls disappear.
  • The minimum cacheable prompt drops from 1,024 to 512 tokens; refusals return stop_reason: "refusal".

Request body that skips up-front thinking (/v1/messages)

{
  "model": "claude-sonnet-5-5",
  "max_tokens": 4096,
  "thinking": { "type": "between_tools" },
  "output_config": { "effort": "medium" },
  "messages": [{ "role": "user", "content": "Fix the failing test in src/app.ts" }]
}
Model ID
claude-sonnet-5-5
Vendors
anthropic
Supply status
Currently callable

Catalog pricing snapshot

Prices are per 1M tokens; request-time pricing controls settlement.

MeterWokeyOfficial referenceSavings
Input$0.3$285%
Output$1.5$1085%
Cache read$0.03$0.285%
Cache write$0.6$485%

Official price source: https://www.anthropic.com/claude-sonnet-5-5

Model capabilities

Context
1,000,000 tokens
Max output
128,000 tokens
Streaming
Supported
Tools
Supported
Vision
Supported
Supported API forms
/v1/messages

Send a request

/v1/messages

curl https://api.wokey.ai/v1/messages \
  -H "x-api-key: $WOKEY_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{"model":"claude-sonnet-5-5","max_tokens":1024,"messages":[{"role":"user","content":"Hello"}]}'

Use Claude Sonnet 5.5 in your tools

The gateway is https://api.wokey.ai (OpenAI-compatible clients usually take https://api.wokey.ai/v1), and the model ID is claude-sonnet-5-5. Open the guide for the client you use.

Frequently asked questions

When was Claude Sonnet 5.5 released, and how is it different from Sonnet 5?

Claude Sonnet 5.5 was released on 2026-09-28 with a June 2026 knowledge cutoff. It keeps Sonnet 5’s official price ($2 input / $10 output per million tokens), runs more than 30% faster and uses fewer tokens on most tasks. In the official benchmarks Terminal-Bench 4.0 rose from 10.3% to 70.6% and OSWorld 2.1 from 57.0% to 80.1%.

Should I use Claude Sonnet 5.5 or Opus 5.5?

Use Sonnet 5.5 for well-scoped everyday coding, bug fixes and document work at half the official price of Opus 5.5; use Opus 5.5 for complex, open-ended tasks that need sustained judgment. The official benchmark gaps are all under 5 points, Sonnet 5.5 is ahead on Terminal-Bench, and GDPval-AA is nearly tied (1844 vs 1846). If Sonnet 5.5 only finishes a task at xhigh or max, try Opus 5.5 at low or medium.

Why does disabling thinking return 400 on Claude Sonnet 5.5?

Sonnet 5.5 no longer accepts thinking: {"type": "disabled"}. Use thinking: {"type": "between_tools"} to skip up-front thinking and only think between tool calls; it works at low, medium and high effort and returns 400 at xhigh and max. tool_choice any / tool and non-default temperature, top_p or top_k also return 400. Through Wokey the gateway converts these for you: disabled becomes between_tools at low to high effort, while disabled at xhigh / max and manual budget_tokens fall back to default adaptive thinking; temperature, top_p and top_k are removed. tool_choice any / tool has no equivalent and still returns 400, so switch to auto.

How do I use Sonnet 5.5 in Claude Code, and which effort should I pick?

Claude Code runs Sonnet 5.5 at medium effort by default. To use it through Wokey, set ANTHROPIC_BASE_URL to https://api.wokey.ai, ANTHROPIC_AUTH_TOKEN to your Wokey API key and ANTHROPIC_MODEL to claude-sonnet-5-5. Anthropic recommends starting agentic coding at medium and moving hard tasks to high.

Is there a cheaper Claude Sonnet 5.5 API than the official price?

Wokey serves claude-sonnet-5-5 over the native Anthropic /v1/messages API at $0.30 input / $1.50 output per million tokens (official $2 / $10). With "Official verification" turned on for the API key, passthrough Messages responses carry a tee.proof signature you can check locally at /tools/verify to confirm the response is the official upstream's original output.

How much does the Claude Sonnet 5.5 API cost, and how much does it save versus the official API?

Wokey lists input at $0.3 per 1M tokens and output at $1.5 per 1M tokens; the official references are $2 and $10. That is 85% lower for input and 85% lower for output. Request-time pricing controls settlement, and the vendor pricing source is linked on this page.

What is the Claude Sonnet 5.5 API model ID?

The model ID for Claude Sonnet 5.5 on Wokey is claude-sonnet-5-5. Use it exactly as written in the request body's model field or in your client's model setting (for example ANTHROPIC_MODEL in Claude Code, or model in the Codex CLI config.toml); GET https://api.wokey.ai/v1/models also lists it.

Which API forms does Claude Sonnet 5.5 support, and how do I call it through Wokey?

Claude Sonnet 5.5 supports /v1/messages (Messages requests). Set the API base URL to https://api.wokey.ai, authenticate with a Wokey API key, and use the canonical model ID claude-sonnet-5-5; you can start with /v1/messages. The Send a request section includes a curl example for every supported API form.

How do I use Claude Sonnet 5.5 in Claude Code?

Claude Code calls Claude Sonnet 5.5 over /v1/messages. Set ANTHROPIC_BASE_URL="https://api.wokey.ai", ANTHROPIC_AUTH_TOKEN to your Wokey API key, and ANTHROPIC_MODEL="claude-sonnet-5-5", then start claude. Claude Desktop also uses /v1/messages; the setup guides cover both step by step.

What are the context window and maximum output for Claude Sonnet 5.5?

The catalog context window is 1,000,000 tokens and the maximum output is 128,000 tokens. These are separate limits, and a request must also satisfy the upstream constraints in effect when it is processed.

How are cache reads and cache writes priced for Claude Sonnet 5.5?

Cache read: Wokey lists $0.03 per 1M tokens and the official reference is $0.2 per 1M tokens. Cache write: Wokey lists $0.6 per 1M tokens and the official reference is $4 per 1M tokens.

How does calling Claude Sonnet 5.5 through Wokey differ from OpenRouter?

Wokey lists Claude Sonnet 5.5 at $0.3 input / $1.5 output per 1M tokens, against an official reference of $2 / $10. OpenRouter states it adds no markup on inference and charges its fee when you buy credits (5.5% by card). Wokey carries a curated set of models and has no per-provider routing parameters, so OpenRouter fits better if you need a wider catalog. The OpenRouter alternative page has the full comparison.

How do Claude Sonnet 5.5 and Claude Sonnet 5 differ in price and context?

In the published price comparison, Claude Sonnet 5.5 lists input/output at $0.3 / $1.5 with a 1,000,000-token context window; Claude Sonnet 5 lists $0.36 / $1.8 with a 1,000,000-token context window. This is a factual price-and-capacity comparison, not a quality ranking; also compare their supported API forms and capabilities.

Related models and documentation

Get API key · claude-sonnet-5-5