Model library
Build with leading models
Every model we serve, priced per million tokens, behind one OpenAI-compatible API. Change the model id and nothing else in your code.
Run BiOS Adaptive
bios-adaptive
- Input
- $0.44 – $1.25
- Output
- $1.32 – $4.00
- Cached
- $0.14
1M context
Claude Opus 5
claude-opus-5
- Input
- $5.00
- Output
- $25.00
- Cached
- $0.50
1M context
Claude Sonnet 5
claude-sonnet-5
- Input
- $2.00
- Output
- $10.00
- Cached
- $0.20
1M context
DeepSeek V4.1 Flash
deepseek-v4.1-flash
- Input
$0.40$0.30- Output
$1.60$1.20- Cached
$0.08$0.06
1M context
DeepSeek V4 Pro
deepseek-v4-pro
- Input
- $1.40
- Output
- $3.40
- Cached
- $0.14
128K context
GLM-5.2
glm-5.2
- Input
- $1.40
- Output
- $3.40
- Cached
- $0.14
1M context
GLM-5.3-Flash
glm-5.3-flash
- Input
- $0.15
- Output
- $0.50
- Cached
- $0.03
1M context
Kimi K3
kimi-k3
- Input
- $3.00
- Output
- $15.00
- Cached
- $0.30
1M context
GLM-5.3
glm-5.3
- Input
- $1.40
- Output
- $4.40
- Cached
- $0.26
1M context
DeepSeek V4 Flash 0731
deepseek-v4-flash-0731
- Input
- $0.10
- Output
- $0.25
- Cached
- $0.01
128K context
DeepSeek V4 Flash
deepseek-v4-flash
- Input
- $0.10
- Output
- $0.25
- Cached
- $0.01
128K context
MiniMax M3
minimax-m3
- Input
- $0.30
- Output
- $1.20
- Cached
- $0.03
128K context
Qwen3.5 397B-A17B
qwen3.5-397b-a17b
- Input
- $0.50
- Output
- $3.40
- Cached
- $0.05
256K context
Kimi K2.7 Code
kimi-k2.7-code
- Input
- $0.80
- Output
- $3.40
- Cached
- $0.08
128K context
DeepSeek V4 Pro 0813
deepseek-v4-pro-0813
- Input
- $1.40
- Output
- $3.40
- Cached
- $0.14
128K context
Qwen3.8 2.4T-A95B
qwen3.8-2.4t-a95b
- Input
- $2.00
- Output
- $6.00
- Cached
- $0.20
1M context
USD per 1M tokens, read live from the Run BiOS API. Run BiOS Adaptive shows a range because its rate follows the model each request lands on; pinned models show their floor. Prompt caching is billed separately. The exact rate for the model you are about to call is shown in the dashboard before you send a request. Treat that as authoritative over this page.
Coming soon
Beyond text: image and voice
The next modalities on the same endpoint. Register now and we will tell you the moment each one ships.
Image generation
FLUX.2
Black Forest Labs
The open-weights quality bar for image generation, with fast klein variants for latency-sensitive work.
Billed per image · rates at launch
Qwen-Image-3.0
Alibaba
Text-heavy layouts -- posters, infographics, UI mockups -- from prompts up to 4.5K tokens.
Billed per image · rates at launch
Stable Diffusion 3.5 Large
Stability AI
The largest open ecosystem of fine-tunes, LoRAs and control tooling.
Billed per image · rates at launch
Text-to-speech
Fish Audio S2 Pro
Fish Audio
Open-weight quality leader with inline emotion control across ~50 languages.
Billed per 1M characters · rates at launch
Qwen3-TTS
Alibaba
Streaming speech in 10 languages, with voice cloning from a 3-second sample.
Billed per 1M characters · rates at launch
Dia2
Nari Labs
Streaming dialogue speech built for multi-speaker conversations.
Billed per 1M characters · rates at launch
Speech-to-text
Qwen3-ASR
Alibaba
Speech recognition across 52 languages, streaming and batch in one model.
Billed per audio minute · rates at launch
Voxtral Realtime
Mistral AI
Native streaming transcription at sub-second latency, Apache-2.0.
Billed per audio minute · rates at launch
Whisper large-v3 turbo
OpenAI
The most-deployed open transcription model, covering 99 languages.
Billed per audio minute · rates at launch
Voice-to-voice
PersonaPlex
NVIDIA
Full-duplex voice agents with persona and voice prompting; leads on task adherence.
Billed per audio minute · rates at launch
Moshi
Kyutai
The reference full-duplex speech model -- listens and speaks at once.
Billed per audio minute · rates at launch
Qwen3-Omni
Alibaba
One omni model taking text, image, audio and video in, streaming speech out.
Billed per audio minute · rates at launch
The lineup above is the open-weights release we are targeting per modality and can change before launch.