Chat Latest on Run BiOS
chat-latest
Top up $20 and we add 50%. That gives you about 2.0M input tokens extra on Chat Latest, on us.
Create an accountPricing
USD per 1M tokens| Input | $5.00 |
| Output | $30.00 |
| Cached input | $0.5090% below fresh input |
Billed per token with no minimum and no monthly fee. Price your workload
Capabilities and limits
- Context window
- 266K tokens
- Max output
- 125K tokens
- Modality
- multimodal
- Tool calling
- Supported
- Vision
- Supported
- Reasoning
- Always on
- Reasoning effort
- medium (default: medium)
- Prompt caching
- Supported (automatic)
Measured performance
7d- Time to first token (p50)
- 800 ms
- Time to first token (p95)
- 800 ms
- Throughput (avg)
- 84.9 tok/s
- Throughput (p50)
- 50.0 tok/s
Measured on real Run BiOS traffic over the trailing 7d. Your figures will vary with prompt shape and region.
curl https://api.runbios.ai/v1/chat/completions \
-H "Authorization: Bearer $BIOS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "chat-latest",
"messages": [{ "role": "user", "content": "Hello" }]
}'Existing OpenAI SDK code works by changing the base URL and the key. API overview
Model card
Curated by Run BiOS · benchmark figures, where present, are the card’s own, not our measurements
Chat Latest points to the latest Instant model currently used in ChatGPT. The underlying model snapshot is regularly updated by OpenAI, so its behaviour can change over time. For production API usage, OpenAI recommends GPT-6 Astra.
Model details
| Property | Value |
|---|---|
| Developer | OpenAI |
| Model ID | chat-latest |
| OpenAI default snapshot | chat-latest |
| Input | Text, image |
| Output | Text |
| Context window | 272,000 tokens on Run BiOS (OpenAI's full window is 400,000 tokens) |
| Max output tokens | 128,000 |
| Knowledge cutoff | Aug 31, 2025 |
| Reasoning | Yes |
Details above are from OpenAI's official model page; context window and input types are as served on Run BiOS.
On Run BiOS
- Always reasons. Reasoning levels
medium; the default ismedium(its only level). This model cannot turn reasoning off: a request for no reasoning runs at the lowest available level. - Reasoning summaries. When reasoning is on, OpenAI may return a summary of the model's reasoning: in
reasoning_contenton the OpenAI API, and in thinking blocks on the Anthropic Messages API whenthinking.displayissummarized. OpenAI does not return a summary for every request. - Tool calling. Function tools with
tool_choiceauto, required, a named function or none, including parallel tool calls, and together with reasoning. - Structured outputs.
response_formatwithjson_schemaorjson_object. - Image input. Text and image input, text output.
- Sampling parameters.
temperature,top_p, penalties,stopandseedare not applied to this model on Run BiOS, because it always reasons; the request is still served. - Prompt caching. Automatic for repeated prompt prefixes; the minimum cacheable length varies with the request. Cache reads bill at the cached-input rate, and there is no separate charge for cache writes. Send a stable
prompt_cache_keyto improve cache hits. - Two APIs. Works with the OpenAI Chat Completions API (
/v1/chat/completions) and the Anthropic Messages API (/v1/messages).