Enterprise AI inference. Zero logs. Zero data retention.Take back control of your enterprise AI spend.

Tokens used AI spend Zero data retention

Save up to 70% on AI cost

0 bytes retained
Day 1Day 15Day 30

Build with leading models

Zero-Log by Design

We don't retain or store your data. Prompts and responses are processed in memory and discarded the moment the request completes.

Nothing Written Down

Prompts and responses live in memory for the duration of the request. No request logs, no content store, no archive.

Nothing to Leak

We never train on your data, and there is no stored copy to breach, subpoena, or misuse. What was never kept cannot be lost.

Metering, Not Logging

We count tokens to bill you accurately. That count is all that exists afterwards -- never what you sent or what came back.

What it costs

See what it costs to run serverless inference on Run BiOS

Save $17,398/mo

on deepseek-v4-pro vs Fireworks

25%

vs Fireworks

25%

vs Together AI

Pick the model you use most

Choose a period, then a volume — or type your own

tokens/day

What share of your tokens are input rather than output

Input 70%Output 30%

Based on 1 billion tokens/day · 700 million input, 300 million output

Input / 1MOutput / 1MDaily
Run BiOS$1.30$2.60$1,690.00
Fireworks$1.74$3.48$2,262.00
Together AI$1.74$3.48$2,262.00

Over an average month Run BiOS is $51,404$17,398 saved against Fireworks, and $17,398 saved against Together AI.

Fireworks Standard tier, as published 26 July 2026 · Together AI serverless, as published 26 July 2026.

Start with $10 in credits

No credit card required. Billed per second of GPU time.

Not sure which model? Run BiOS Adaptive chooses for you

Frontier models, ready to call

Long context, priced per million tokens. Change the model id and nothing else in your code.

Six families
Claude · DeepSeek · GLM · Kimi · MiniMax · Qwen
1M tokens
Longest context window
$0.14
Lowest input, per 1M tokens
Text + vision
Models that accept image input

USD per 1M tokens, as of 12 August 2026. Run BiOS Adaptive shows a range because its rate follows the model each request lands on; pinned models show their floor. Prompt caching is billed separately. The exact rate for the model you are about to call is shown in the dashboard before you send a request — treat that as authoritative over this table.

Run BiOS Adaptive

One endpoint. Every frontier model. Zero model-ops.

Adaptive routes each request to the model that fits — quality, speed, and budget in balance — on per-token pricing with a published ceiling.

Custom models

Serve a model you trained, on a GPU of its own

Fine-tuning produces weights you own. A finished checkpoint can go straight onto a dedicated endpoint without leaving the platform.

Custom training, end to end

Comes with fine-tuning

Custom Model Endpoints

Reached through the same OpenAI-compatible API as everything else. This is part of fine-tuning rather than the serverless product — different billing, and a different reason to reach for it.

Per second of GPU time
The same rates as training — from $0.42/hr across 12 GPU types

Adapters serve themselves

A LoRA or QLoRA adapter is served with its base model automatically, with no manual merge step.

The route matches the model

Chat, completion, embedding or reranker — each exposes the matching OpenAI route.

Sized to expected traffic

Light fits on the minimum that will hold the model; heavy adds GPUs for concurrency.

Memory is yours to tune

bf16 or fp8, with KV cache compression to fit more concurrent requests on the same card.

Which one you need

When to fine-tune, and when not to

Fine-tuning costs more than a serverless call and takes longer to get right. Plenty of workloads should start on serverless and stay there. Here is how to tell which side you are on.

Stay on serverless inference

A general-purpose model is already doing the job — or you do not yet know exactly what the job is.

  • You are validating an idea or shipping a first version
  • Your prompts still change week to week
  • Traffic is spiky, seasonal, or low volume
  • An open model in the catalog already hits the accuracy you need
  • You have no labelled examples of the output you want

Fine-tune your own model

A general-purpose model gets close, but is consistently wrong in a way you can describe.

  • You have examples of the output you want — hundreds or thousands, not millions
  • You need a smaller, cheaper model to do one narrow job well
  • Your domain language is not in the base model — clinical, legal, internal jargon
  • Format or tone has to be right on every single call
  • You need the weights to belong to you

Start with either one

Call a production model through the serverless API, or train your own on dedicated GPUs. Two products, one account, no commitment on either.

Serverless per million tokens · Fine-tuning from $0.42/hr · Your weights stay yours