Pricing calculator

See what it costs to run serverless inference on Run BiOS

Pick a model, set your volume, and see the bill per day, week or month — with what Fireworks, Together AI and Nebius would charge for the same work. Every rate is the committed rate card; rival figures are dated snapshots.

Save $137,500/mo

on deepseek-v4-pro vs Nebius

12%

vs Nebius

Pick the model you use most

Choose a period and volume, or type your own

tokens/mo

What share of your tokens are input rather than output

Input 70%Output 30%

Based on 500 billion tokens/mo · 350 billion input, 150 billion output

Input / 1MOutput / 1MMonthly
Run BiOS$1.40$3.40$1,000,000
Nebius$1.75$3.50$1,137,500

Get 50% extra on your first top-up

Inference billed per token at published rates.

Per-million-token rates

ModelContextInput / 1MCached / 1MOutput / 1M
bios-adaptive1M$0.44 – $1.25$0.14$1.32 – $4.00
chat-latest266K$5.00$0.50$30.00
claude-opus-51M$5.00$0.50$25.00
claude-opus-5-51M$4.00$0.20$20.00
claude-sonnet-51M$2.00$0.20$10.00
claude-sonnet-5-51M$2.00$0.20$10.00
deepseek-v4-flash128K$0.10$0.01$0.25
deepseek-v4-flash-0731128K$0.10$0.01$0.25
deepseek-v4-pro128K$1.40$0.14$3.40
deepseek-v4-pro-0813128K$1.40$0.14$3.40
deepseek-v4.1-flash1M$0.40$0.04$1.60
glm-5.21M$1.40$0.14$3.40
glm-5.31M$1.40$0.26$4.40
glm-5.3-flash1M$0.15$0.03$0.50
gpt-4.11M$2.00$0.50$8.00
gpt-4.1-mini1M$0.40$0.10$1.60
gpt-4.1-nano1M$0.10$0.03$0.40
gpt-4o125K$2.50$1.25$10.00
gpt-4o-mini125K$0.15$0.08$0.60
gpt-5266K$1.25$0.13$10.00
gpt-5-mini266K$0.25$0.03$2.00
gpt-5-nano266K$0.05$0.01$0.40
gpt-5.1266K$1.25$0.13$10.00
gpt-5.4266K$2.50$0.25$15.00
gpt-5.4-mini266K$0.75$0.08$4.50
gpt-5.4-nano266K$0.20$0.02$1.25
gpt-5.5266K$5.00$0.50$30.00
gpt-5.6-luna266K$0.20$0.02$1.20
gpt-5.6-sol266K$4.00$0.40$20.00
gpt-5.6-terra266K$2.00$0.20$12.00
gpt-6-astra266K$10.00$1.00$50.00
gpt-6-luna266K$0.10$0.01$0.50
gpt-6-sol266K$2.00$0.20$10.00
kimi-k2.7-code128K$0.80$0.08$3.40
kimi-k31M$3.00$0.30$15.00
minimax-m3128K$0.30$0.03$1.20
o3195K$2.00$0.50$8.00
o3-mini195K$1.10$0.55$4.40
o4-mini195K$1.10$0.28$4.40
qwen3.5-397b-a17b256K$0.50$0.05$3.40
qwen3.8-2.4t-a95b256K$2.00$0.20$6.00

Scroll inside the table to see all models. USD per 1M tokens, as of 30 September 2026. Run BiOS Adaptive shows a range because its rate follows the model each request lands on; pinned models show their floor. Prompt caching is billed separately. The exact rate for the model you are about to call is shown in the dashboard before you send a request — treat that as authoritative over this table.

Neither product can run up a debt. Both draw from one pre-paid wallet. If the balance reaches zero you get a grace warning, and a running endpoint pauses rather than accruing charges. Top up and resume.