195K contextTool callingReasoningPrompt caching

Pricing

USD per 1M tokens
Input$1.10
Output$4.40
Cached input$0.5550% below fresh input

Billed per token with no minimum and no monthly fee. Price your workload

Capabilities and limits

Context window
195K tokens
Max output
98K tokens
Modality
text
Tool calling
Supported
Vision
Not supported
Reasoning
Always on
Reasoning effort
low, medium, high (default: medium)
Prompt caching
Supported (automatic)

Measured performance

7d
Time to first token (p50)
1.5 s
Time to first token (p95)
3.0 s
Throughput (avg)
859.4 tok/s
Throughput (p50)
300.0 tok/s

Measured on real Run BiOS traffic over the trailing 7d. Your figures will vary with prompt shape and region.

Call it now: OpenAI-compatible
curl https://api.runbios.ai/v1/chat/completions \
  -H "Authorization: Bearer $BIOS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "o3-mini",
    "messages": [{ "role": "user", "content": "Hello" }]
  }'

Existing OpenAI SDK code works by changing the base URL and the key. API overview

Model card

Curated by Run BiOS · benchmark figures, where present, are the card’s own, not our measurements

o3-mini is a small reasoning model, providing high intelligence at the same cost and latency targets as o1-mini. It supports Structured Outputs and function calling.

Model details

Property Value
Developer OpenAI
Model ID o3-mini
OpenAI default snapshot o3-mini-2025-01-31
Input Text
Output Text
Context window 200,000 tokens
Max output tokens 100,000
Knowledge cutoff Oct 01, 2023
Reasoning Yes

Details above are from OpenAI's official model page; context window and input types are as served on Run BiOS.

On Run BiOS

  • Always reasons. Reasoning levels low, medium, high; the default is medium. This model cannot turn reasoning off: a request for no reasoning runs at the lowest available level.
  • Reasoning summaries. When reasoning is on, OpenAI may return a summary of the model's reasoning: in reasoning_content on the OpenAI API, and in thinking blocks on the Anthropic Messages API when thinking.display is summarized. OpenAI does not return a summary for every request.
  • Tool calling. Function tools with tool_choice auto, required, a named function or none, including parallel tool calls, and together with reasoning.
  • Structured outputs. response_format with json_schema or json_object.
  • Text input only on Run BiOS.
  • Sampling parameters. temperature, top_p, penalties, stop and seed are not applied to this model on Run BiOS, because it always reasons; the request is still served.
  • Prompt caching. Automatic for repeated prompt prefixes; the minimum cacheable length varies with the request. Cache reads bill at the cached-input rate, and there is no separate charge for cache writes. Send a stable prompt_cache_key to improve cache hits.
  • Two APIs. Works with the OpenAI Chat Completions API (/v1/chat/completions) and the Anthropic Messages API (/v1/messages).