1M contextTool callingPrompt caching

Pricing

USD per 1M tokens
Input$0.10
Output$0.40
Cached input$0.0375% below fresh input

Billed per token with no minimum and no monthly fee. Price your workload

Capabilities and limits

Context window
1M tokens
Max output
32K tokens
Modality
text
Tool calling
Supported
Vision
Not supported
Reasoning
Not supported
Prompt caching
Supported (automatic)

Measured performance

7d
Time to first token (p50)
400 ms
Time to first token (p95)
800 ms
Throughput (avg)
128.3 tok/s
Throughput (p50)
150.0 tok/s

Measured on real Run BiOS traffic over the trailing 7d. Your figures will vary with prompt shape and region.

Call it now: OpenAI-compatible
curl https://api.runbios.ai/v1/chat/completions \
  -H "Authorization: Bearer $BIOS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4.1-nano",
    "messages": [{ "role": "user", "content": "Hello" }]
  }'

Existing OpenAI SDK code works by changing the base URL and the key. API overview

Model card

Curated by Run BiOS · benchmark figures, where present, are the card’s own, not our measurements

GPT-4.1 nano is the fastest, most cost-efficient version of GPT-4.1. It excels at instruction following and tool calling, and features a 1M-token context window and low latency without a reasoning step. For more complex tasks, OpenAI recommends starting with GPT-5 nano.

Model details

Property Value
Developer OpenAI
Model ID gpt-4.1-nano
OpenAI default snapshot gpt-4.1-nano-2025-04-14
Input Text
Output Text
Context window 1,047,576 tokens
Max output tokens 32,768
Knowledge cutoff Jun 01, 2024
Reasoning No (instruction model, no reasoning step)

Details above are from OpenAI's official model page; context window and input types are as served on Run BiOS.

On Run BiOS

  • No reasoning step. A reasoning_effort setting is accepted and ignored, so requests written for reasoning models still work.
  • Tool calling. Function tools with tool_choice auto, required, a named function or none, including parallel tool calls.
  • Structured outputs. response_format with json_schema or json_object.
  • Text input only on Run BiOS.
  • Sampling parameters. temperature, top_p, frequency_penalty, presence_penalty, stop and seed are supported.
  • Prompt caching. Automatic for repeated prompt prefixes; the minimum cacheable length varies with the request. Cache reads bill at the cached-input rate, and there is no separate charge for cache writes. Send a stable prompt_cache_key to improve cache hits.
  • Two APIs. Works with the OpenAI Chat Completions API (/v1/chat/completions) and the Anthropic Messages API (/v1/messages).