Qwen models
Every Qwen model we serve, priced per million tokens, behind one OpenAI-compatible API. Change the model id and nothing else in your code.
Qwen3.5 397B-A17B
qwen/qwen3.5-397b-a17b
- Input
- $0.50
- Output
- $3.40
- Cached
- $0.05
256K context
Qwen3.8 2.4T-A95B
qwen/qwen3.8-2.4t-a95b
- Input
- $2.00
- Output
- $6.00
- Cached
- $0.20
1M context
USD per 1M tokens, read live from the Run BiOS API. Run BiOS Adaptive shows a range because its rate follows the model each request lands on; pinned models show their floor. Prompt caching is billed separately. The exact rate for the model you are about to call is shown in the dashboard before you send a request. Treat that as authoritative over this page.
Coming soon
Beyond text: image and voice
The next modalities on the same endpoint. Register now and we will tell you the moment each one ships.
Image generation
FLUX.2
Black Forest Labs
The open-weights quality bar for image generation, with fast klein variants for latency-sensitive work.
Billed per image · rates at launch
Qwen-Image-3.0
Alibaba
Text-heavy layouts -- posters, infographics, UI mockups -- from prompts up to 4.5K tokens.
Billed per image · rates at launch
Stable Diffusion 3.5 Large
Stability AI
The largest open ecosystem of fine-tunes, LoRAs and control tooling.
Billed per image · rates at launch
Text-to-speech
Fish Audio S2 Pro
Fish Audio
Open-weight quality leader with inline emotion control across ~50 languages.
Billed per 1M characters · rates at launch
Qwen3-TTS
Alibaba
Streaming speech in 10 languages, with voice cloning from a 3-second sample.
Billed per 1M characters · rates at launch
Dia2
Nari Labs
Streaming dialogue speech built for multi-speaker conversations.
Billed per 1M characters · rates at launch
Speech-to-text
Qwen3-ASR
Alibaba
Speech recognition across 52 languages, streaming and batch in one model.
Billed per audio minute · rates at launch
Voxtral Realtime
Mistral AI
Native streaming transcription at sub-second latency, Apache-2.0.
Billed per audio minute · rates at launch
Whisper large-v3 turbo
OpenAI
The most-deployed open transcription model, covering 99 languages.
Billed per audio minute · rates at launch
Voice-to-voice
PersonaPlex
NVIDIA
Full-duplex voice agents with persona and voice prompting; leads on task adherence.
Billed per audio minute · rates at launch
Moshi
Kyutai
The reference full-duplex speech model -- listens and speaks at once.
Billed per audio minute · rates at launch
Qwen3-Omni
Alibaba
One omni model taking text, image, audio and video in, streaming speech out.
Billed per audio minute · rates at launch
The lineup above is the open-weights release we are targeting per modality and can change before launch.