266K contextVisionTool callingReasoningPrompt caching

Pricing

USD per 1M tokens
Input$2.00
Output$12.00
Cached input$0.2090% below fresh input
Cache write (5m)$2.50

Billed per token with no minimum and no monthly fee. Price your workload

Capabilities and limits

Context window
266K tokens
Max output
125K tokens
Modality
multimodal
Tool calling
Supported
Vision
Supported
Reasoning
Optional
Reasoning effort
low, medium, high, xhigh, max (default: medium)
Prompt caching
Supported (automatic)
Cache minimum
1K tokens

Measured performance

7d
Time to first token (p50)
800 ms
Time to first token (p95)
800 ms
Throughput (avg)
72.8 tok/s
Throughput (p50)
50.0 tok/s

Measured on real Run BiOS traffic over the trailing 7d. Your figures will vary with prompt shape and region.

Call it now: OpenAI-compatible
curl https://api.runbios.ai/v1/chat/completions \
  -H "Authorization: Bearer $BIOS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.6-terra",
    "messages": [{ "role": "user", "content": "Hello" }]
  }'

Existing OpenAI SDK code works by changing the base URL and the key. API overview

Model card

Curated by Run BiOS · benchmark figures, where present, are the card’s own, not our measurements

GPT-5.6 Terra is designed for workloads that balance intelligence and cost. It roughly corresponds to the mini model tier used in earlier GPT-5 families.

Model details

Property Value
Developer OpenAI
Model ID gpt-5.6-terra
OpenAI default snapshot gpt-5.6-terra
Input Text, image
Output Text
Context window 272,000 tokens on Run BiOS (OpenAI's full window is 1,050,000 tokens)
Max output tokens 128,000
Knowledge cutoff Feb 16, 2026
Reasoning Yes

Details above are from OpenAI's official model page; context window and input types are as served on Run BiOS.

On Run BiOS

  • Configurable reasoning. Reasoning levels low, medium, high, xhigh, max; the default is medium. Turn reasoning off with reasoning_effort: "none" (OpenAI API) or thinking: {"type": "disabled"} (Anthropic Messages API).
  • Reasoning summaries. When reasoning is on, OpenAI may return a summary of the model's reasoning: in reasoning_content on the OpenAI API, and in thinking blocks on the Anthropic Messages API when thinking.display is summarized. OpenAI does not return a summary for every request.
  • Tool calling. Function tools with tool_choice auto, required, a named function or none, including parallel tool calls, and together with reasoning.
  • Structured outputs. response_format with json_schema or json_object.
  • Image input. Text and image input, text output.
  • Sampling parameters. With reasoning off, temperature, top_p and seed are supported. With reasoning on, sampling parameters are not applied; the request is still served. stop is not supported by this model.
  • Prompt caching. Automatic for repeated prompt prefixes of at least 1,024 tokens. Cache reads bill at 0.1× the input rate and cache writes at 1.25×, as OpenAI charges.
  • Two APIs. Works with the OpenAI Chat Completions API (/v1/chat/completions) and the Anthropic Messages API (/v1/messages).