1M contextVisionTool callingReasoningPrompt caching

Pricing

USD per 1M tokens
Input$4.00 $3.80
Output$20.00 $19.00
Cached input$0.20 $0.1995% below fresh input
Cache write (5m)$5.00 $4.75
Cache write (1h)$8.00 $7.60

5% off: Launch offer on Claude Opus 5.5. Promotional price on input, output, cache reads & cache writes. Standard prices shown struck through.

Billed per token with no minimum and no monthly fee. Price your workload

Capabilities and limits

Context window
1M tokens
Max output
125K tokens
Modality
text
Tool calling
Supported
Vision
Supported
Reasoning
Always on
Reasoning effort
low, medium, high, xhigh, max (default: medium)
Prompt caching
Supported (on_request)

Measured performance

7d
Time to first token (p50)
6.0 s
Time to first token (p95)
12 s
Throughput (avg)
83.3 tok/s
Throughput (p50)
150.0 tok/s

Measured on real Run BiOS traffic over the trailing 7d. Your figures will vary with prompt shape and region.

Call it now: OpenAI-compatible
curl https://api.runbios.ai/v1/chat/completions \
  -H "Authorization: Bearer $BIOS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic/claude-opus-5-5",
    "messages": [{ "role": "user", "content": "Hello" }]
  }'

Existing OpenAI SDK code works by changing the base URL and the key. API overview

Model card

Curated by Run BiOS · benchmark figures, where present, are the card’s own, not our measurements

Claude Opus 5.5 is the first model in Anthropic's Claude 5.5 family and its most capable Opus-tier model — built for long-running agentic coding, computer use, and professional knowledge work, with a 1M-token context window and up to 128K output tokens. Anthropic reports that it performs at the level of Claude Fable 5.1 on most work while costing about 40% less to run than Claude Opus 5 on typical workloads, and that it generates output more than 30% faster than Opus 5.

It is an upgrade to Claude Opus 5, and Anthropic's system card reports that it scored higher on every evaluation in its capability summary, with the largest gains in agentic coding, visual reasoning, computer use, and long-horizon professional work. Much of that improvement is available below maximum reasoning effort.

Capabilities

  • Adaptive thinking, always on — the model decides how much to reason on each task; effort is the control. Thinking cannot be switched off on this model. On Run BiOS, Opus 5.5 exposes the low, medium, high, xhigh, and max reasoning-effort levels, and defaults to medium — the same default Anthropic uses.
  • Long context — a 1M-token context window and up to 128K output tokens per request.
  • Multimodal input — text and image input, with text output.
  • Tool use and agents — built for sustained, multi-step agentic work: planning, calling tools, verifying its own work, and staying on task over many hours.
  • Clearer communication — Anthropic trained Opus 5.5 to write more plainly than Opus 5: it puts the most important information first, avoids jargon, follows the writing rules you give it, and reports what it did, what it found, and what it needs as it works.
  • Prompt caching — cache reads are priced far below fresh input, which matters most in long, context-heavy agentic sessions where cached context makes up the majority of tokens.

Benchmark results

Results reported in Anthropic's launch materials and the Claude Opus 5.5 System Card. Unless noted, Opus 5.5 results use adaptive thinking at max effort.

Benchmark Claude Opus 5.5 Claude Fable 5.1 Claude Opus 5
Terminal-Bench 4.0 (agentic coding) 66.4% (xhigh) 55.8% 52.3%
FrontierCode v1.1 Main (agentic coding) 54.4% 50.3% 48.0%
CursorBench 4.0 (agentic coding) 57.8% 51.8% 46.6%
GDPval-AA v2.1 (knowledge work) 1846 Elo 1735 Elo 1708 Elo
AutomationBench (business workflows) 40.0% 31.4% 26.9%
Humanity's Last Exam (with tools) 67.7% 65.6% 63.6%
Terminal-Bench-Science 0.1 (agentic research) 58.7% 52.6% 29.0%
OSWorld 2.1 (computer use, partial credit) 81.8% 80.7% 74.0%
Chartography (chart recognition, with tools) 89.0% 88.4% 83.4%

Notes on the numbers, as reported:

  • Opus 5.5 was evaluated with its production safeguards enabled; when a safeguard intervened on a cybersecurity, biology, or frontier-AI-development task, an earlier model completed it, which Anthropic says likely understates Opus 5.5's scores.
  • At its default medium effort, Opus 5.5 scores 54.6% on FrontierCode and 52.5% on CursorBench — above every other model's best score in Anthropic's comparison — at a fraction of the cost per task.
  • On Terminal-Bench 4.0, Opus 5.5 at default effort beats Opus 5 at max effort for about a fifth of the cost.
  • AutomationBench was run by Zapier without fallback models, so safeguard interventions counted as failures.
  • Anthropic cautions that at this level of capability, benchmark margins are a less reliable guide to real-world differences than they used to be.

Working with the model

Anthropic and its early testers highlight long, sprawling jobs as the model's strength:

  • An early tester audited and fixed a 200,000-line codebase in under three hours, where Opus 5 took over 20 hours and used 2.5× as many tokens; another completed a 680,000-line code migration in less than a day.
  • Asked to translate HAProxy from C into Rust, Opus 5.5 produced a rewrite that passed nearly all of HAProxy's own regression tests, finishing faster and at about half the cost of Fable 5.1.
  • In a research test where every figure and quote was checked against sources, 16 of 18 Opus 5.5 reports cleared Anthropic's quality bar; neither Fable 5.1 nor Opus 5 cleared it in any attempt.

Behaviours to plan for when moving from Claude Opus 5:

  • Thinking cannot be disabled. A request that asks to turn reasoning off (reasoning_effort: "none", or thinking: {"type": "disabled"} on the Messages API) is served with reasoning on at the default effort. Use low effort for the fastest, cheapest responses.
  • Forced tool use is not supported. tool_choice of auto and none work; required, any, or a named tool return an error. To make the model call a tool, describe in the prompt when the tool applies.
  • Text between tool calls may arrive as thinking. Longer progress notes written between tool calls come back as thinking blocks, which are empty at the default display setting.
  • Thinking blocks are tied to the model and the conversation. Keep conversation history append-only when replaying Opus 5.5 thinking blocks.

Safety and evaluations

Anthropic's pre-deployment evaluation is documented in the Claude Opus 5.5 System Card, and the model was tested before release by external evaluators including METR and Frontier Design. On Anthropic's automated behavioural audit — nearly 2,000 simulated scenarios — Opus 5.5 scored better than any recent Claude model on nearly every measure of misaligned behaviour, and it is Anthropic's strongest model on most measures of honesty. It is much less likely than recent models to take hard-to-reverse actions or act outside the boundaries it has been given, and it matches or beats Opus 5 on prompt-injection resistance in every setting Anthropic tested.

Because its capabilities in biology and cybersecurity are comparable to Claude Mythos 5.1, Anthropic deploys Opus 5.5 with safety classifiers similar to those on Claude Fable 5.1. These target a narrow set of high-risk requests; routine software development and most research work are unaffected, but some requests in those areas may be declined.