266K contextVisionTool callingReasoningPrompt caching

Pricing

USD per 1M tokens
Input$2.50
Output$15.00
Cached input$0.2590% below fresh input

Billed per token with no minimum and no monthly fee. Price your workload

Capabilities and limits

Context window
266K tokens
Max output
125K tokens
Modality
multimodal
Tool calling
Supported
Vision
Supported
Reasoning
Optional
Reasoning effort
low, medium, high, xhigh (default: medium)
Prompt caching
Supported (automatic)

Measured performance

7d
Time to first token (p50)
800 ms
Time to first token (p95)
1.5 s
Throughput (avg)
179.9 tok/s
Throughput (p50)
100.0 tok/s

Measured on real Run BiOS traffic over the trailing 7d. Your figures will vary with prompt shape and region.

Call it now: OpenAI-compatible
curl https://api.runbios.ai/v1/chat/completions \
  -H "Authorization: Bearer $BIOS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.4",
    "messages": [{ "role": "user", "content": "Hello" }]
  }'

Existing OpenAI SDK code works by changing the base URL and the key. API overview

Model card

Curated by Run BiOS · benchmark figures, where present, are the card’s own, not our measurements

GPT-5.4 is a flagship model for complex professional work. OpenAI describes it as a more affordable model for coding and professional work.

Model details

Property Value
Developer OpenAI
Model ID gpt-5.4
OpenAI default snapshot gpt-5.4-2026-03-05
Input Text, image
Output Text
Context window 272,000 tokens on Run BiOS (OpenAI's full window is 1,050,000 tokens)
Max output tokens 128,000
Knowledge cutoff Aug 31, 2025
Reasoning Yes

Details above are from OpenAI's official model page; context window and input types are as served on Run BiOS.

On Run BiOS

  • Configurable reasoning. Reasoning levels low, medium, high, xhigh; the default is medium. Turn reasoning off with reasoning_effort: "none" (OpenAI API) or thinking: {"type": "disabled"} (Anthropic Messages API).
  • Reasoning summaries. When reasoning is on, OpenAI may return a summary of the model's reasoning: in reasoning_content on the OpenAI API, and in thinking blocks on the Anthropic Messages API when thinking.display is summarized. OpenAI does not return a summary for every request.
  • Tool calling. Function tools with tool_choice auto, required, a named function or none, including parallel tool calls, and together with reasoning.
  • Structured outputs. response_format with json_schema or json_object.
  • Image input. Text and image input, text output.
  • Sampling parameters. With reasoning off, temperature, top_p, frequency_penalty, presence_penalty and seed are supported. With reasoning on, sampling parameters are not applied; the request is still served. stop is not supported by this model.
  • Prompt caching. Automatic for repeated prompt prefixes; the minimum cacheable length varies with the request. Cache reads bill at the cached-input rate, and there is no separate charge for cache writes. Send a stable prompt_cache_key to improve cache hits.
  • Two APIs. Works with the OpenAI Chat Completions API (/v1/chat/completions) and the Anthropic Messages API (/v1/messages).