o3-mini on Run BiOS
o3-mini
+50%First top-up bonus
Top up $20 and we add 50%. That gives you about 9.1M input tokens extra on o3-mini, on us.
Create an account195K contextTool callingReasoningPrompt caching
Pricing
USD per 1M tokens| Input | $1.10 |
| Output | $4.40 |
| Cached input | $0.5550% below fresh input |
Billed per token with no minimum and no monthly fee. Price your workload
Capabilities and limits
- Context window
- 195K tokens
- Max output
- 98K tokens
- Modality
- text
- Tool calling
- Supported
- Vision
- Not supported
- Reasoning
- Always on
- Reasoning effort
- low, medium, high (default: medium)
- Prompt caching
- Supported (automatic)
Measured performance
7d- Time to first token (p50)
- 1.5 s
- Time to first token (p95)
- 3.0 s
- Throughput (avg)
- 859.4 tok/s
- Throughput (p50)
- 300.0 tok/s
Measured on real Run BiOS traffic over the trailing 7d. Your figures will vary with prompt shape and region.
Call it now: OpenAI-compatible
curl https://api.runbios.ai/v1/chat/completions \
-H "Authorization: Bearer $BIOS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "o3-mini",
"messages": [{ "role": "user", "content": "Hello" }]
}'Existing OpenAI SDK code works by changing the base URL and the key. API overview
Model card
Curated by Run BiOS · benchmark figures, where present, are the card’s own, not our measurements
o3-mini is a small reasoning model, providing high intelligence at the same cost and latency targets as o1-mini. It supports Structured Outputs and function calling.
Model details
| Property | Value |
|---|---|
| Developer | OpenAI |
| Model ID | o3-mini |
| OpenAI default snapshot | o3-mini-2025-01-31 |
| Input | Text |
| Output | Text |
| Context window | 200,000 tokens |
| Max output tokens | 100,000 |
| Knowledge cutoff | Oct 01, 2023 |
| Reasoning | Yes |
Details above are from OpenAI's official model page; context window and input types are as served on Run BiOS.
On Run BiOS
- Always reasons. Reasoning levels
low,medium,high; the default ismedium. This model cannot turn reasoning off: a request for no reasoning runs at the lowest available level. - Reasoning summaries. When reasoning is on, OpenAI may return a summary of the model's reasoning: in
reasoning_contenton the OpenAI API, and in thinking blocks on the Anthropic Messages API whenthinking.displayissummarized. OpenAI does not return a summary for every request. - Tool calling. Function tools with
tool_choiceauto, required, a named function or none, including parallel tool calls, and together with reasoning. - Structured outputs.
response_formatwithjson_schemaorjson_object. - Text input only on Run BiOS.
- Sampling parameters.
temperature,top_p, penalties,stopandseedare not applied to this model on Run BiOS, because it always reasons; the request is still served. - Prompt caching. Automatic for repeated prompt prefixes; the minimum cacheable length varies with the request. Cache reads bill at the cached-input rate, and there is no separate charge for cache writes. Send a stable
prompt_cache_keyto improve cache hits. - Two APIs. Works with the OpenAI Chat Completions API (
/v1/chat/completions) and the Anthropic Messages API (/v1/messages).