125K contextVisionTool callingPrompt caching

Pricing

USD per 1M tokens
Input$0.15
Output$0.60
Cached input$0.0850% below fresh input

Billed per token with no minimum and no monthly fee. Price your workload

Capabilities and limits

Context window
125K tokens
Max output
16K tokens
Modality
multimodal
Tool calling
Supported
Vision
Supported
Reasoning
Not supported
Prompt caching
Supported (automatic)

Measured performance

7d
Time to first token (p50)
800 ms
Time to first token (p95)
800 ms
Throughput (avg)
83.1 tok/s
Throughput (p50)
100.0 tok/s

Measured on real Run BiOS traffic over the trailing 7d. Your figures will vary with prompt shape and region.

Call it now: OpenAI-compatible
curl https://api.runbios.ai/v1/chat/completions \
  -H "Authorization: Bearer $BIOS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o-mini",
    "messages": [{ "role": "user", "content": "Hello" }]
  }'

Existing OpenAI SDK code works by changing the base URL and the key. API overview

Model card

Curated by Run BiOS · benchmark figures, where present, are the card’s own, not our measurements

GPT-4o Mini ("o" for "omni") is a fast, affordable small model for focused tasks. It accepts both text and image inputs, and produces text outputs (including Structured Outputs).

Model details

Property Value
Developer OpenAI
Model ID gpt-4o-mini
OpenAI default snapshot gpt-4o-mini-2024-07-18
Input Text, image
Output Text
Context window 128,000 tokens
Max output tokens 16,384
Knowledge cutoff Oct 01, 2023
Reasoning No (instruction model, no reasoning step)

Details above are from OpenAI's official model page; context window and input types are as served on Run BiOS.

On Run BiOS

  • No reasoning step. A reasoning_effort setting is accepted and ignored, so requests written for reasoning models still work.
  • Tool calling. Function tools with tool_choice auto, required, a named function or none, including parallel tool calls.
  • Structured outputs. response_format with json_schema or json_object.
  • Image input. Text and image input, text output.
  • Sampling parameters. temperature, top_p, frequency_penalty, presence_penalty, stop and seed are supported.
  • Prompt caching. Automatic for repeated prompt prefixes; the minimum cacheable length varies with the request. Cache reads bill at the cached-input rate, and there is no separate charge for cache writes. Send a stable prompt_cache_key to improve cache hits.
  • Two APIs. Works with the OpenAI Chat Completions API (/v1/chat/completions) and the Anthropic Messages API (/v1/messages).