Model library

Build with leading models

Every model we serve, priced per million tokens, behind one OpenAI-compatible API. Change the model id and nothing else in your code.

USD per 1M tokens, as of 17 August 2026. Run BiOS Adaptive shows a range because its rate follows the model each request lands on; pinned models show their floor. Prompt caching is billed separately. The exact rate for the model you are about to call is shown in the dashboard before you send a request — treat that as authoritative over this page.

Coming soon

Beyond text: image and voice

The next modalities on the same endpoint. Register now and we will tell you the moment each one ships.

Image generation

FComing soon

FLUX.2

Black Forest Labs

The open-weights quality bar for image generation, with fast klein variants for latency-sensitive work.

Billed per image · rates at launch

Coming soon

Qwen-Image-3.0

Alibaba

Text-heavy layouts -- posters, infographics, UI mockups -- from prompts up to 4.5K tokens.

Billed per image · rates at launch

SComing soon

Stable Diffusion 3.5 Large

Stability AI

The largest open ecosystem of fine-tunes, LoRAs and control tooling.

Billed per image · rates at launch

Text-to-speech

FComing soon

Fish Audio S2 Pro

Fish Audio

Open-weight quality leader with inline emotion control across ~50 languages.

Billed per 1M characters · rates at launch

Coming soon

Qwen3-TTS

Alibaba

Streaming speech in 10 languages, with voice cloning from a 3-second sample.

Billed per 1M characters · rates at launch

DComing soon

Dia2

Nari Labs

Streaming dialogue speech built for multi-speaker conversations.

Billed per 1M characters · rates at launch

Speech-to-text

Coming soon

Qwen3-ASR

Alibaba

Speech recognition across 52 languages, streaming and batch in one model.

Billed per audio minute · rates at launch

Coming soon

Voxtral Realtime

Mistral AI

Native streaming transcription at sub-second latency, Apache-2.0.

Billed per audio minute · rates at launch

WComing soon

Whisper large-v3 turbo

OpenAI

The most-deployed open transcription model, covering 99 languages.

Billed per audio minute · rates at launch

Voice-to-voice

Coming soon

PersonaPlex

NVIDIA

Full-duplex voice agents with persona and voice prompting; leads on task adherence.

Billed per audio minute · rates at launch

MComing soon

Moshi

Kyutai

The reference full-duplex speech model -- listens and speaks at once.

Billed per audio minute · rates at launch

Coming soon

Qwen3-Omni

Alibaba

One omni model taking text, image, audio and video in, streaming speech out.

Billed per audio minute · rates at launch

The lineup above is the open-weights release we are targeting per modality and can change before launch.