GPT-5.6 Luna on Run BiOS
gpt-5.6-luna
+50%First top-up bonus
Top up $20 and we add 50%. That gives you about 50.0M input tokens extra on GPT-5.6 Luna, on us.
Create an account266K contextVisionTool callingReasoningPrompt caching
Pricing
USD per 1M tokens| Input | $0.20 |
| Output | $1.20 |
| Cached input | $0.0290% below fresh input |
| Cache write (5m) | $0.25 |
Billed per token with no minimum and no monthly fee. Price your workload
Capabilities and limits
- Context window
- 266K tokens
- Max output
- 125K tokens
- Modality
- multimodal
- Tool calling
- Supported
- Vision
- Supported
- Reasoning
- Optional
- Reasoning effort
- low, medium, high, xhigh, max (default: medium)
- Prompt caching
- Supported (automatic)
- Cache minimum
- 1K tokens
Measured performance
7d- Time to first token (p50)
- 3.0 s
- Time to first token (p95)
- 6.0 s
- Throughput (avg)
- 190.0 tok/s
- Throughput (p50)
- 200.0 tok/s
Measured on real Run BiOS traffic over the trailing 7d. Your figures will vary with prompt shape and region.
Call it now: OpenAI-compatible
curl https://api.runbios.ai/v1/chat/completions \
-H "Authorization: Bearer $BIOS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.6-luna",
"messages": [{ "role": "user", "content": "Hello" }]
}'Existing OpenAI SDK code works by changing the base URL and the key. API overview
Model card
Curated by Run BiOS · benchmark figures, where present, are the card’s own, not our measurements
GPT-5.6 Luna is designed for cost-sensitive, high-volume workloads. It roughly corresponds to the nano model tier used in earlier GPT-5 families.
Model details
| Property | Value |
|---|---|
| Developer | OpenAI |
| Model ID | gpt-5.6-luna |
| OpenAI default snapshot | gpt-5.6-luna |
| Input | Text, image |
| Output | Text |
| Context window | 272,000 tokens on Run BiOS (OpenAI's full window is 1,050,000 tokens) |
| Max output tokens | 128,000 |
| Knowledge cutoff | Feb 16, 2026 |
| Reasoning | Yes |
Details above are from OpenAI's official model page; context window and input types are as served on Run BiOS.
On Run BiOS
- Configurable reasoning. Reasoning levels
low,medium,high,xhigh,max; the default ismedium. Turn reasoning off withreasoning_effort: "none"(OpenAI API) orthinking: {"type": "disabled"}(Anthropic Messages API). - Reasoning summaries. When reasoning is on, OpenAI may return a summary of the model's reasoning: in
reasoning_contenton the OpenAI API, and in thinking blocks on the Anthropic Messages API whenthinking.displayissummarized. OpenAI does not return a summary for every request. - Tool calling. Function tools with
tool_choiceauto, required, a named function or none, including parallel tool calls, and together with reasoning. - Structured outputs.
response_formatwithjson_schemaorjson_object. - Image input. Text and image input, text output.
- Sampling parameters. With reasoning off,
temperature,top_pandseedare supported. With reasoning on, sampling parameters are not applied; the request is still served.stopis not supported by this model. - Prompt caching. Automatic for repeated prompt prefixes of at least 1,024 tokens. Cache reads bill at 0.1× the input rate and cache writes at 1.25×, as OpenAI charges.
- Two APIs. Works with the OpenAI Chat Completions API (
/v1/chat/completions) and the Anthropic Messages API (/v1/messages).
More from GPT-5.6
GPT-6 Astra$10.00/$50.00GPT-6 Sol$2.00/$10.00GPT-6 Luna$0.10/$0.50GPT-5$1.25/$10.00GPT-4o Mini$0.15/$0.60GPT-5.6 Sol$4.00/$20.00GPT-5 Mini$0.25/$2.00GPT-4.1 Mini$0.40/$1.60GPT-5.4 Mini$0.75/$4.50GPT-5.1$1.25/$10.00GPT-5.4$2.50/$15.00GPT-5.4 nano$0.20/$1.25GPT-5.5$5.00/$30.00GPT-5.6 Terra$2.00/$12.00GPT-4.1 nano$0.10/$0.40GPT-4o$2.50/$10.00GPT-5 nano$0.05/$0.40GPT-4.1$2.00/$8.00