1M contextVisionTool callingReasoning

Pricing

USD per 1M tokens · as of 29 August 2026
Input$0.15
Output$0.50
Cached input$0.03

Flat rate — every request bills the same. Full catalog and comparisons on the pricing comparison page.

Where it fits

Not sure? BiOS Adaptive routes to the best model per request, including this one.

Call it now — OpenAI-compatible
curl https://api.runbios.ai/v1/chat/completions \
  -H "Authorization: Bearer $BIOS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "glm-5.3-flash",
    "messages": [{ "role": "user", "content": "Hello" }]
  }'

Existing OpenAI SDK code works by changing the base URL to https://api.runbios.ai/v1 and the key. API overview