Save up to 70% on AI cost
We don't retain or store your data. Prompts and responses are processed in memory and discarded the moment the request completes.
Prompts and responses live in memory for the duration of the request. No request logs, no content store, no archive.
We never train on your data, and there is no stored copy to breach, subpoena, or misuse. What was never kept cannot be lost.
We count tokens to bill you accurately. That count is all that exists afterwards -- never what you sent or what came back.
What it costs
Save $17,398/mo
on deepseek-v4-pro vs Fireworks
25%
vs Fireworks
25%
vs Together AI
Pick the model you use most
Choose a period, then a volume — or type your own
What share of your tokens are input rather than output
Based on 1 billion tokens/day · 700 million input, 300 million output
| Input / 1M | Output / 1M | Daily | |
|---|---|---|---|
| Run BiOS | $1.30 | $2.60 | $1,690.00 |
| Fireworks | $1.74 | $3.48 | $2,262.00 |
| Together AI | $1.74 | $3.48 | $2,262.00 |
Over an average month Run BiOS is $51,404 — $17,398 saved against Fireworks, and $17,398 saved against Together AI.
Fireworks Standard tier, as published 26 July 2026 · Together AI serverless, as published 26 July 2026.
Not sure which model? Run BiOS Adaptive chooses for you
Long context, priced per million tokens. Change the model id and nothing else in your code.
| Model | Context | Input / 1M | Cached / 1M | Output / 1M |
|---|---|---|---|---|
| bios-adaptive | 1M | $0.14 – $1.25 | $0.03 | $0.28 – $3.95 |
| claude-opus-5 | 1M | $5.00 | $0.50 | $25.00 |
| claude-sonnet-5 | 1M | $3.00 | $0.30 | $15.00 |
| deepseek-v4-flash | 128K | $0.14 | $0.01 | $0.25 |
| deepseek-v4-pro | 128K | $1.30 | $0.10 | $2.60 |
| glm-5.2 | 1M | $1.40 | $0.14 | $3.00 |
| kimi-k2.7-code | 128K | $0.75 | $0.08 | $3.40 |
| minimax-m3 | 128K | $0.30 | $0.03 | $1.20 |
| qwen3.5-397b-a17b | 256K | $0.45 | $0.05 | $3.00 |
USD per 1M tokens, as of 12 August 2026. Run BiOS Adaptive shows a range because its rate follows the model each request lands on; pinned models show their floor. Prompt caching is billed separately. The exact rate for the model you are about to call is shown in the dashboard before you send a request — treat that as authoritative over this table.
Run BiOS Adaptive
Adaptive routes each request to the model that fits — quality, speed, and budget in balance — on per-token pricing with a published ceiling.
Custom models
Fine-tuning produces weights you own. A finished checkpoint can go straight onto a dedicated endpoint without leaving the platform.
Reached through the same OpenAI-compatible API as everything else. This is part of fine-tuning rather than the serverless product — different billing, and a different reason to reach for it.
A LoRA or QLoRA adapter is served with its base model automatically, with no manual merge step.
Chat, completion, embedding or reranker — each exposes the matching OpenAI route.
Light fits on the minimum that will hold the model; heavy adds GPUs for concurrency.
bf16 or fp8, with KV cache compression to fit more concurrent requests on the same card.
Which one you need
Fine-tuning costs more than a serverless call and takes longer to get right. Plenty of workloads should start on serverless and stay there. Here is how to tell which side you are on.
A general-purpose model is already doing the job — or you do not yet know exactly what the job is.
A general-purpose model gets close, but is consistently wrong in a way you can describe.
Call a production model through the serverless API, or train your own on dedicated GPUs. Two products, one account, no commitment on either.
Serverless per million tokens · Fine-tuning from $0.42/hr · Your weights stay yours