Run BiOS Adaptive on Run BiOS
runbios/bios-adaptive
Top up $20 and we add 50%. That gives you at least 8.0M input tokens extra on Run BiOS Adaptive, on us.
Create an accountOne id that picks the model for you
runbios/bios-adaptive is ours, and it is not a single model: you call it like any other id and each request is answered at the depth the work needs. That is why the rates below are a band rather than one figure — what you pay follows the model the request lands on. How Run BiOS Adaptive works.
Pricing
USD per 1M tokens| Input | $0.44 – $1.25 |
| Output | $1.32 – $4.00 |
| Cached input | $0.1468% below fresh input |
Billed per token with no minimum and no monthly fee. Price your workload
Capabilities and limits
- Context window
- 1M tokens
- Max output
- 63K tokens
- Modality
- multimodal
- Tool calling
- Supported
- Vision
- Supported
- Video
- Supported
- Reasoning
- Optional
- Reasoning effort
- low, medium, high
- Prompt caching
- Supported (automatic)
Measured performance
7d- Time to first token (p50)
- 1.5 s
- Time to first token (p95)
- 3.0 s
- Throughput (avg)
- 101.3 tok/s
- Throughput (p50)
- 200.0 tok/s
Measured on real Run BiOS traffic over the trailing 7d. Your figures will vary with prompt shape and region.
curl https://api.runbios.ai/v1/chat/completions \
-H "Authorization: Bearer $BIOS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "runbios/bios-adaptive",
"messages": [{ "role": "user", "content": "Hello" }]
}'Existing OpenAI SDK code works by changing the base URL and the key. API overview
Model card
Curated by Run BiOS · benchmark figures, where present, are the card’s own, not our measurements
Run BiOS Adaptive is the default front door of the Run BiOS platform: one model id that always answers at the strength the request actually needs. Instead of asking you to choose a model per workload, it reads each query and applies exactly the right amount of intelligence to it — no more, no less.
What It Is
Every query deserves exactly the right amount of intelligence. Run BiOS Adaptive makes sure you always get the best accuracy, matching the right resources to the right query so you achieve maximum quality at maximum saving. Your quality never drops — only your bill does.
Simple prompts are answered simply; hard problems get deeper reasoning automatically. You integrate once against a single model id and the platform keeps the answer quality at the frontier as the underlying model landscape moves.
Capabilities
- Adaptive intelligence — Per-query depth of reasoning, from instant answers for simple requests to extended step-by-step thinking for hard ones. You can also pin the depth yourself with the standard reasoning effort ladder (
low,medium,high). - Multimodal input — Accepts text, image, and video inputs and generates text output.
- Function calling — Native structured tool use for agentic workflows.
- Long context — A 1,048,576-token context window, so full codebases, long documents, and extended conversations fit in a single request.
- Prompt caching — Repeated prefixes are served from cache at a reduced rate.
When to Use It
Use Run BiOS Adaptive as the default for production traffic: it is the model id that needs no revisiting. Pick a specific named model instead when you have a hard requirement on one — a compliance review pinned to a single model, a benchmark you reproduce, or a price/performance point you want to hold constant.
Behavior Notes
- Reasoning effort is optional and configurable; when a request names no effort, the platform applies a balanced default.
- Identity — Run BiOS Adaptive presents itself as Run BiOS Worker in conversation.