Pricing calculator

See what it costs to run serverless inference on Run BiOS

Pick a model, set your volume, and see the bill per day, week or month — with what Fireworks, Together AI and Nebius would charge for the same work. Every rate is the committed rate card; rival figures are dated snapshots.

Save $137,500/mo

on deepseek-v4-pro vs Nebius

12%

vs Fireworks

12%

vs Together AI

12%

vs Nebius

Pick the model you use most

Choose a period, then a volume — or type your own

tokens/mo

What share of your tokens are input rather than output

Input 70%Output 30%

Based on 500 billion tokens/mo · 350 billion input, 150 billion output

Input / 1MOutput / 1MMonthly
Run BiOS$1.40$3.40$1,000,000
Fireworks$1.74$3.48$1,131,000
Together AI$1.74$3.48$1,131,000
Nebius$1.75$3.50$1,137,500

Fireworks Standard tier, as published 29 August 2026 · Together AI serverless, as published 29 August 2026.

Start with $10 in credits

No credit card required. Billed per second of GPU time.

Per-million-token rates

ModelContextInput / 1MCached / 1MOutput / 1M
bios-adaptive1M$0.44 – $1.25$0.14$1.32 – $3.95
claude-opus-51M$5.00$0.50$25.00
claude-sonnet-51M$2.00$0.20$10.00
deepseek-v4-flash128K$0.10$0.01$0.25
deepseek-v4-flash-0731128K$0.10$0.01$0.25
deepseek-v4-pro128K$1.40$0.14$3.40
deepseek-v4-pro-0813128K$1.40$0.14$3.40
glm-5.21M$1.40$0.14$3.40
glm-5.31M$1.40$0.26$4.40
glm-5.3-flash1M$0.15$0.03$0.50
kimi-k2.7-code128K$0.80$0.08$3.40
kimi-k31M$3.00$0.30$15.00
minimax-m3128K$0.30$0.03$1.20
qwen3.5-397b-a17b256K$0.50$0.05$3.40
qwen3.8-2.4t-a95b256K$2.00$0.20$6.00

USD per 1M tokens, as of 29 August 2026. Run BiOS Adaptive shows a range because its rate follows the model each request lands on; pinned models show their floor. Prompt caching is billed separately. The exact rate for the model you are about to call is shown in the dashboard before you send a request — treat that as authoritative over this table.

Neither product can run up a debt. Both draw from one pre-paid wallet. If the balance reaches zero you get a grace warning, and a running endpoint pauses rather than accruing charges. Top up and resume.