One id that picks the model for you

runbios/bios-adaptive is ours, and it is not a single model: you call it like any other id and each request is answered at the depth the work needs. That is why the rates below are a band rather than one figure — what you pay follows the model the request lands on. How Run BiOS Adaptive works.

1M contextVisionVideoTool callingReasoningPrompt caching

Pricing

USD per 1M tokens
Input$0.44 – $1.25
Output$1.32 – $4.00
Cached input$0.1468% below fresh input

Billed per token with no minimum and no monthly fee. Price your workload

Capabilities and limits

Context window
1M tokens
Max output
63K tokens
Modality
multimodal
Tool calling
Supported
Vision
Supported
Video
Supported
Reasoning
Optional
Reasoning effort
low, medium, high
Prompt caching
Supported (automatic)

Measured performance

7d
Time to first token (p50)
1.5 s
Time to first token (p95)
3.0 s
Throughput (avg)
101.3 tok/s
Throughput (p50)
200.0 tok/s

Measured on real Run BiOS traffic over the trailing 7d. Your figures will vary with prompt shape and region.

Call it now: OpenAI-compatible
curl https://api.runbios.ai/v1/chat/completions \
  -H "Authorization: Bearer $BIOS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "runbios/bios-adaptive",
    "messages": [{ "role": "user", "content": "Hello" }]
  }'

Existing OpenAI SDK code works by changing the base URL and the key. API overview

Model card

Curated by Run BiOS · benchmark figures, where present, are the card’s own, not our measurements

Run BiOS Adaptive is the default front door of the Run BiOS platform: one model id that always answers at the strength the request actually needs. Instead of asking you to choose a model per workload, it reads each query and applies exactly the right amount of intelligence to it — no more, no less.

What It Is

Every query deserves exactly the right amount of intelligence. Run BiOS Adaptive makes sure you always get the best accuracy, matching the right resources to the right query so you achieve maximum quality at maximum saving. Your quality never drops — only your bill does.

Simple prompts are answered simply; hard problems get deeper reasoning automatically. You integrate once against a single model id and the platform keeps the answer quality at the frontier as the underlying model landscape moves.

Capabilities

  • Adaptive intelligence — Per-query depth of reasoning, from instant answers for simple requests to extended step-by-step thinking for hard ones. You can also pin the depth yourself with the standard reasoning effort ladder (low, medium, high).
  • Multimodal input — Accepts text, image, and video inputs and generates text output.
  • Function calling — Native structured tool use for agentic workflows.
  • Long context — A 1,048,576-token context window, so full codebases, long documents, and extended conversations fit in a single request.
  • Prompt caching — Repeated prefixes are served from cache at a reduced rate.

When to Use It

Use Run BiOS Adaptive as the default for production traffic: it is the model id that needs no revisiting. Pick a specific named model instead when you have a hard requirement on one — a compliance review pinned to a single model, a benchmark you reproduce, or a price/performance point you want to hold constant.

Behavior Notes

  • Reasoning effort is optional and configurable; when a request names no effort, the platform applies a balanced default.
  • Identity — Run BiOS Adaptive presents itself as Run BiOS Worker in conversation.