Model library
Serverless LLM APIs for leading models
Every model we serve, priced per million tokens, behind one OpenAI-compatible API. Change the model id and nothing else in your code.
Run BiOS Adaptive
runbios/bios-adaptive
- Input
- $0.44 to $1.25
- Output
- $1.32 to $4.00
- Cached
- $0.14
1M context
DeepSeek V4.1 Flash
deepseek/deepseek-v4.1-flash
- Input
- $0.30
- Output
- $1.20
- Cached
- $0.03
1M context
Claude Opus 5.5
anthropic/claude-opus-5-5
- Input
- $3.80
- Output
- $19.00
- Cached
- $0.19
1M context
Claude Sonnet 5.5
anthropic/claude-sonnet-5-5
- Input
- $1.90
- Output
- $9.50
- Cached
- $0.19
1M context
GLM-5.3-Flash
z-ai/glm-5.3-flash
- Input
- $0.15
- Output
- $0.50
- Cached
- $0.03
1M context
GLM-5.3
z-ai/glm-5.3
- Input
- $1.40
- Output
- $4.40
- Cached
- $0.26
1M context
DeepSeek V4 Pro 0813
deepseek/deepseek-v4-pro-0813
- Input
- $1.40
- Output
- $3.40
- Cached
- $0.14
128K context
Claude Opus 5
anthropic/claude-opus-5
- Input
- $4.75
- Output
- $23.75
- Cached
- $0.475
1M context
GPT-6 Astra
openai/gpt-6-astra
- Input
- $10.00
- Output
- $50.00
- Cached
- $1.00
266K context
GPT-6 Sol
openai/gpt-6-sol
- Input
- $2.00
- Output
- $10.00
- Cached
- $0.20
266K context
GPT-6 Luna
openai/gpt-6-luna
- Input
- $0.10
- Output
- $0.50
- Cached
- $0.01
266K context
Claude Sonnet 5
anthropic/claude-sonnet-5
- Input
- $1.90
- Output
- $9.50
- Cached
- $0.19
1M context
DeepSeek V4 Flash 0731
deepseek/deepseek-v4-flash-0731
- Input
- $0.10
- Output
- $0.25
- Cached
- $0.01
128K context
GPT-5.4 Mini
openai/gpt-5.4-mini
- Input
- $0.75
- Output
- $4.50
- Cached
- $0.08
266K context
MiniMax M3
minimax/minimax-m3
- Input
- $0.30
- Output
- $1.20
- Cached
- $0.03
128K context
Kimi K3
moonshotai/kimi-k3
- Input
- $3.00
- Output
- $15.00
- Cached
- $0.30
1M context
Qwen3.5 397B-A17B
qwen/qwen3.5-397b-a17b
- Input
- $0.50
- Output
- $3.40
- Cached
- $0.05
256K context
GLM-5.2
z-ai/glm-5.2
- Input
- $1.40
- Output
- $3.40
- Cached
- $0.14
1M context
DeepSeek V4 Pro
deepseek/deepseek-v4-pro
- Input
- $1.40
- Output
- $3.40
- Cached
- $0.14
128K context
DeepSeek V4 Flash
deepseek/deepseek-v4-flash
- Input
- $0.10
- Output
- $0.25
- Cached
- $0.01
128K context
GPT-5.6 Luna
openai/gpt-5.6-luna
- Input
- $0.20
- Output
- $1.20
- Cached
- $0.02
266K context
Kimi K2.7 Code
moonshotai/kimi-k2.7-code
- Input
- $0.80
- Output
- $3.40
- Cached
- $0.08
128K context
GPT-4.1 nano
openai/gpt-4.1-nano
- Input
- $0.10
- Output
- $0.40
- Cached
- $0.03
1M context
GPT-4.1 Mini
openai/gpt-4.1-mini
- Input
- $0.40
- Output
- $1.60
- Cached
- $0.10
1M context
GPT-4.1
openai/gpt-4.1
- Input
- $2.00
- Output
- $8.00
- Cached
- $0.50
1M context
GPT-5.5
openai/gpt-5.5
- Input
- $5.00
- Output
- $30.00
- Cached
- $0.50
266K context
Qwen3.8 2.4T-A95B
qwen/qwen3.8-2.4t-a95b
- Input
- $2.00
- Output
- $6.00
- Cached
- $0.20
1M context
GPT-4o Mini
openai/gpt-4o-mini
- Input
- $0.15
- Output
- $0.60
- Cached
- $0.08
125K context
GPT-5
openai/gpt-5
- Input
- $1.25
- Output
- $10.00
- Cached
- $0.13
266K context
GPT-4o
openai/gpt-4o
- Input
- $2.50
- Output
- $10.00
- Cached
- $1.25
125K context
GPT-5 Mini
openai/gpt-5-mini
- Input
- $0.25
- Output
- $2.00
- Cached
- $0.03
266K context
GPT-5.1
openai/gpt-5.1
- Input
- $1.25
- Output
- $10.00
- Cached
- $0.13
266K context
GPT-5.4
openai/gpt-5.4
- Input
- $2.50
- Output
- $15.00
- Cached
- $0.25
266K context
GPT-5 nano
openai/gpt-5-nano
- Input
- $0.05
- Output
- $0.40
- Cached
- $0.01
266K context
o3
openai/o3
- Input
- $2.00
- Output
- $8.00
- Cached
- $0.50
195K context
Chat Latest
openai/chat-latest
- Input
- $5.00
- Output
- $30.00
- Cached
- $0.50
266K context
GPT-5.6 Sol
openai/gpt-5.6-sol
- Input
- $4.00
- Output
- $20.00
- Cached
- $0.40
266K context
GPT-5.4 nano
openai/gpt-5.4-nano
- Input
- $0.20
- Output
- $1.25
- Cached
- $0.02
266K context
GPT-5.6 Terra
openai/gpt-5.6-terra
- Input
- $2.00
- Output
- $12.00
- Cached
- $0.20
266K context
o3-mini
openai/o3-mini
- Input
- $1.10
- Output
- $4.40
- Cached
- $0.55
195K context
o4-mini
openai/o4-mini
- Input
- $1.10
- Output
- $4.40
- Cached
- $0.28
195K context
USD per 1M tokens, read live from the Run BiOS API. Run BiOS Adaptive shows a range because its rate follows the model each request lands on. Pinned models show their floor. Prompt caching is billed separately. The exact rate for the model you are about to call is shown in the dashboard before you send a request. Treat that as authoritative over this page.
Coming soon
Beyond text: image and voice
The next modalities on the same endpoint. Register now and we will tell you the moment each one ships.
Image generation
FLUX.2
Black Forest Labs
The open-weights quality bar for image generation, with fast klein variants for latency-sensitive work.
Billed per image · rates at launch
Qwen-Image-3.0
Alibaba
Text-heavy layouts -- posters, infographics, UI mockups -- from prompts up to 4.5K tokens.
Billed per image · rates at launch
Stable Diffusion 3.5 Large
Stability AI
The largest open ecosystem of fine-tunes, LoRAs and control tooling.
Billed per image · rates at launch
Text-to-speech
Fish Audio S2 Pro
Fish Audio
Open-weight quality leader with inline emotion control across ~50 languages.
Billed per 1M characters · rates at launch
Qwen3-TTS
Alibaba
Streaming speech in 10 languages, with voice cloning from a 3-second sample.
Billed per 1M characters · rates at launch
Dia2
Nari Labs
Streaming dialogue speech built for multi-speaker conversations.
Billed per 1M characters · rates at launch
Speech-to-text
Qwen3-ASR
Alibaba
Speech recognition across 52 languages, streaming and batch in one model.
Billed per audio minute · rates at launch
Voxtral Realtime
Mistral AI
Native streaming transcription at sub-second latency, Apache-2.0.
Billed per audio minute · rates at launch
Whisper large-v3 turbo
OpenAI
The most-deployed open transcription model, covering 99 languages.
Billed per audio minute · rates at launch
Voice-to-voice
PersonaPlex
NVIDIA
Full-duplex voice agents with persona and voice prompting. Leads on task adherence.
Billed per audio minute · rates at launch
Moshi
Kyutai
The reference full-duplex speech model -- listens and speaks at once.
Billed per audio minute · rates at launch
Qwen3-Omni
Alibaba
One omni model taking text, image, audio and video in, streaming speech out.
Billed per audio minute · rates at launch
The lineup above is the open-weights release we are targeting per modality and can change before launch.