Claude Sonnet 5.5 on Run BiOS
claude-sonnet-5-5
Top up $20 and we add 50%. That gives you about 5.3M input tokens extra on Claude Sonnet 5.5, on us.
Create an accountPricing
USD per 1M tokens| Input | |
| Output | |
| Cached input | |
| Cache write (5m) | |
| Cache write (1h) |
5% off: Launch offer on Claude Opus 5.5. Promotional price on input, output, cache reads & cache writes. Standard prices shown struck through.
Billed per token with no minimum and no monthly fee. Price your workload
Capabilities and limits
- Context window
- 1M tokens
- Max output
- 125K tokens
- Modality
- multimodal
- Tool calling
- Supported
- Vision
- Supported
- Reasoning
- Optional
- Reasoning effort
- low, medium, high, xhigh, max (default: high)
- Prompt caching
- Supported (on_request)
- Cache minimum
- 512 tokens
Measured performance
7d- Time to first token (p50)
- 1.5 s
- Time to first token (p95)
- 6.0 s
- Throughput (avg)
- 150.7 tok/s
- Throughput (p50)
- 500.0 tok/s
Measured on real Run BiOS traffic over the trailing 7d. Your figures will vary with prompt shape and region.
curl https://api.runbios.ai/v1/chat/completions \
-H "Authorization: Bearer $BIOS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-5-5",
"messages": [{ "role": "user", "content": "Hello" }]
}'Existing OpenAI SDK code works by changing the base URL and the key. API overview
Model card
Curated by Run BiOS · benchmark figures, where present, are the card’s own, not our measurements
Claude Sonnet 5.5 is the second model in Anthropic's Claude 5.5 family — a faster, lower-cost complement to Claude Opus 5.5, with a 1M-token context window and up to 128K output tokens. Anthropic describes it as its best combination of speed and intelligence: strongest at well-scoped everyday work, fixing bugs, and producing polished documents, slides, and spreadsheets, with a sharp eye for design.
It is a clear upgrade over Claude Sonnet 5. Anthropic reports that it generates output more than 30% faster — its fastest Sonnet model to date — and, at the same per-token prices, typically needs far fewer tokens for the same work, costing up to 30% less per task. On several evaluations, Sonnet 5.5 at its highest effort performs comparably to Opus 5.5, although Anthropic notes that Opus 5.5 remains clearly stronger at complex, open-ended work that requires sustained judgment.
Capabilities
- Adaptive thinking — on by default; effort controls how deeply the model reasons. Anthropic recalibrated the effort levels for this model, so a setting does not produce the same amount of thinking as it did on Sonnet 5. On Run BiOS, Sonnet 5.5 exposes the
low,medium,high,xhigh, andmaxreasoning-effort levels, and defaults tohigh— the same default Anthropic uses. - Lowest thinking setting — turning reasoning off (
reasoning_effort: "none", orthinking: {"type": "disabled"}on the Messages API) runs Sonnet 5.5 at its lowest setting, which keeps up-front thinking off while still allowing short progress notes between tool calls. - Long context — a 1M-token context window and up to 128K output tokens per request.
- Multimodal input — text and image input, with text output. It is the first Sonnet model to beat Pokémon Red working only from screenshots.
- Tool use and agents — in early testing it batched tool calls together more than Sonnet 5, finishing tasks in fewer steps and at lower cost.
- Prompt caching — caching engages from a 512-token cacheable prefix, and cache reads cost a tenth of the input price.
Benchmark results
Results reported in Anthropic's launch materials and the Claude Sonnet 5.5 System Card, alongside the two reference models from Anthropic's own comparison.
| Benchmark | Claude Sonnet 5.5 | Claude Sonnet 5 | Claude Opus 5.5 |
|---|---|---|---|
| Terminal-Bench 4.0 (agentic coding) | 70.6% | 10.3% | 66.4% |
| FrontierCode 1.1 Main (agentic coding) | 52.1% (xhigh) · 46.2% (max) | 42.4% | 54.4% |
| CursorBench 4.0 (agentic coding) | 55.5% | 34.1% | 57.8% |
| GDPval-AA v2.1 (knowledge work) | 1844 Elo | 1449 Elo | 1846 Elo |
| AA-Briefcase v1.1 (long-horizon knowledge work) | 1811 Elo | 1359 Elo | 1822 Elo |
| Humanity's Last Exam (with tools) | 64.5% | 54.9% | 67.7% |
| OSWorld 2.1 (computer use, partial credit) | 80.1% | 57.0% | 81.8% |
| Chartography (chart recognition, no tools) | 61.6% | 15.6% | 64.4% |
Notes on the numbers, as reported:
- On GDPval-AA — real-world tasks across 44 occupations and nine industries — Sonnet 5.5 scores within two points of Opus 5.5 and about 400 points above Sonnet 5.
- On several benchmarks, Sonnet 5.5 at
lowormediumeffort beats Sonnet 5's best score for about a tenth of the cost per task. - On FrontierCode, Sonnet 5.5 scores lower at
maxthan atxhigh: atmaxit more often ran a multi-agent code-review step, which in some cases led to timeouts or out-of-scope edits that the benchmark penalises. - Terminal-Bench 4.0 is reported for Opus 5.5 at
xhigheffort, its highest score.
Working with the model
Early testers highlighted speed, token efficiency, and judgment:
- Slack reported that, without prompt changes, Sonnet 5.5 beat Sonnet 5 on almost all of its offline Slackbot evaluations in fewer steps and with about 14% fewer output tokens.
- Across 118 real app builds, Base44 found Sonnet 5.5 produced apps level with Opus 5 in 3.6 iterations per build on average, against 7.7 for Opus 5, with the fewest failed tool calls of any model it compared.
- Given a public company's earnings materials and a slide template, Sonnet 5.5 produced a 10-slide operating review that two experts judged ready to send as a first draft.
Behaviours to plan for when moving from Claude Sonnet 5:
- Re-tune effort. Anthropic recommends starting at
high, usingmediumfor well-specified agentic coding and multi-step tool use, andmediumorlowfor chat and other latency-sensitive work. - Forced tool use is not supported.
tool_choiceofautoandnonework;required,any, or a named tool return an error. To make the model call a tool, describe in the prompt when the tool applies. - Text between tool calls may arrive as thinking. Longer progress notes written between tool calls come back as
thinkingblocks, which are empty at the default display setting unless reasoning is turned off. - Thinking blocks are tied to the model and the conversation. Keep conversation history append-only when replaying Sonnet 5.5 thinking blocks.
Safety and evaluations
Anthropic's pre-deployment evaluation is documented in the Claude Sonnet 5.5 System Card, covering Responsible Scaling Policy evaluations, cyber, safeguards and harmlessness, agentic safety, alignment, model welfare, and capabilities. On Anthropic's automated behavioural audit — roughly 1,850 scenarios — Sonnet 5.5 improves on or matches Sonnet 5 on most measures of alignment, resistance to misuse, and honesty, and it is the least likely of any Anthropic model to probe the limits of its containers.
Because its cybersecurity capabilities are comparable to Claude Opus 5's, Sonnet 5.5 is the first Sonnet model to launch with cyber safeguards like those on Anthropic's most capable models; its biology safeguards are the same as Sonnet 5's. Both target a narrow set of high-risk requests — routine software development and most life-sciences work are unaffected — and it also ships with classifiers that prevent reasoning extraction.