1M contextVisionTool callingReasoningPrompt caching

Pricing

USD per 1M tokens
Input$0.40
Output$1.60
Cached input$0.0880% below fresh input

Billed per token with no minimum and no monthly fee. Price your workload

Capabilities and limits

Context window
1M tokens
Max output
375K tokens
Modality
multimodal
Tool calling
Supported
Vision
Supported
Reasoning
Optional
Reasoning effort
low, high, max (default: high)
Prompt caching
Supported (automatic)

Measured performance

7d
Time to first token (p50)
3.0 s
Time to first token (p95)
6.0 s
Throughput (avg)
29.0 tok/s
Throughput (p50)
50.0 tok/s

Measured on real Run BiOS traffic over the trailing 7d. Your figures will vary with prompt shape and region.

Call it now — OpenAI-compatible
curl https://api.runbios.ai/v1/chat/completions \
  -H "Authorization: Bearer $BIOS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v4.1-flash",
    "messages": [{ "role": "user", "content": "Hello" }]
  }'

Existing OpenAI SDK code works by changing the base URL and the key. API overview

Model card

Curated by Run BiOS · benchmark figures, where present, are the card’s own, not our measurements

DeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts (MoE) model built by DeepSeek (DeepSeek-AI) and described in the technical report "DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression" (2026). It has 552B backbone parameters, natively processes images and text, generates text autoregressively, and supports contexts of up to one million tokens.

Per DeepSeek's official model card, the headline of this release is efficiency rather than scale: DeepSeek-V4.1-Flash activates only 8B parameters per token during prefill and 16B during decode, and reduces the global KV cache footprint to 890 bytes per token — roughly a quarter of DeepSeek-V4-Flash and, per DeepSeek's Figure 1(b), about a 437-fold reduction relative to DeepSeek-V1. DeepSeek positions this as "substantially improving cost efficiency for input-heavy agentic workloads."

On the benchmarks DeepSeek publishes, V4.1-Flash scores higher than the much larger DeepSeek-V4-Pro on all twelve agentic benchmarks in the comparison table below, and higher than DeepSeek-V4-Flash on every one of them where V4-Flash has a published number. Against the full frontier field it takes the top score outright on Terminal-Bench 2.1, DeepSWE v1.1, CyberGym, AutomationBench, Agent's Last Exam, HLE-with-tools and Codeforces rating, and ties Kimi K3 for the top score on MathArena Apex. It does not lead on pure reasoning: Opus-5.0 and GPT-5.6 Sol score higher on GPQA Diamond and HLE.

Architecture

DeepSeek-V4.1-Flash adopts a Causal Encoder-Decoder (CED) architecture: a 40-layer Transformer organised as a 20-layer causal encoder followed by a 20-layer decoder. Under CED, the decoder's global KV cache is projected from the final encoder hidden states rather than derived from each decoder layer's own hidden states — the change that yields the 8B-prefill / 16B-decode activation profile. SWA Bounded Replay reconstructs missing sliding-window-attention KV states by replaying only the most recent n_win tokens, avoiding any need to persist SWA KV to SSD and reducing the persistent KV cache footprint to roughly 1/8 of DeepSeek-V4-Flash.

Compressed Sparse Attention 2 (CSA2). Each attention layer is assigned one of three static modes — Full, Reindex or Reuse — to share main KV and indexer K across layers and reuse Top-K sparse-attention indices. In the decoder, a Hierarchical Sparse Indexer restricts later indexing layers to a candidate pool built by the first Full Mode layer, bounding deeper indexer cost independently of context length. Together with FP4 main KV caching (E2M1 format, one E4M3 scale per 16 channels), these designs produce the 890 bytes-per-token global KV cache.

Further published components:

  • Single-Pass mHC — revised residual-stream mixing with an efficient Mega-mHC kernel.
  • Engram conditional memory — 196B parameters, sparsely accessed via token-based lookup.
  • DSpark speculative decoding — semi-autoregressive draft generation with confidence-scheduled verification.
  • MoE layers use 1 shared expert and 384 routed experts, activating 6 routed experts per token.

Multimodal path. A vision encoder (DeepSeek-ViT, trained from scratch with 2D-RoPE and 3x3 pixel-unshuffle downsampling) and a two-layer MLP projector convert images into visual embeddings, which are processed jointly with text embeddings from the start of language-model pre-training — images are native to the model, not bolted on afterwards.

Training

Per the technical report, DeepSeek-V4.1-Flash is trained from scratch on a multimodal corpus of 45T tokens, with sparse attention trained at a sequence length of 64K and context extended to 1M tokens at 34T tokens.

Post-training follows the standard SFT → RL → on-policy distillation (OPD) paradigm "without algorithmic modifications"; DeepSeek states that all substantive changes lie instead in the data pipeline — large-scale automated synthesis of agent tasks and environments, with progressive scaling of data, tasks and rollouts.

Reasoning Effort

DeepSeek exposes a continuously controllable reasoning effort on this model: an integer from 1 to 100 that trades inference cost for accuracy. Every instruct benchmark DeepSeek published uses the maximum setting (reasoning_effort=100).

On Run BiOS this model is served through the platform's standard effort ladder — low, high and max — with high applied when a request does not specify one, and reasoning can be switched off per request. DeepSeek's recommended sampling parameters are temperature = 1.0 with top_p = 0.95 (or 1.0), a 1M-token context window, and max_tokens of at least 256K.

Benchmarks

Comparison with frontier models (max reasoning effort)

Reproduced from DeepSeek's official model card. All numbers are as published; evaluations use temperature=1.0, top_p=0.95.

Benchmark (Metric) Opus-5.0 GPT-5.6 Sol K3 GLM-5.3 DS-V4-Pro DS-V4-Flash DS-V4.1-Flash
Reasoning
GPQA Diamond (Pass@1) 93.4 94.1 92.9 88.1 92.4 89.9 90.9
HLE (Pass@1) 56.3 44.5 43.5 42.0† 42.7† 37.8† 36.8 (39.1†)
Codeforces (Rating) 3348 3289 3471
MathArena Apex (Pass@1) 65.6 65.3 58.6 65.6
Agentic
Terminal-Bench 2.1 (Pass@1) 89.1 88.8 88.3 88.2 87.9 82.7 90.6
Terminal-Bench 3.0 (Pass@1) 43.3 34.4 17.7 28.3 11.8 7.6 30.0
Terminal-Bench 4.0 (Pass@1) 51.8 39.9 12.6 37.9 12.4 7.0 31.2
DeepSWE v1.1 (Resolved) 74.0 73.0 67.5 66.9 62.7 54.4 74.2
ProgramBench (Almost@1) 37.0 23.0 17.5 19.0 15.5 20.3
NL2Repo-Bench (Score) 75.3 56.8 58.0 58.0 61.5 54.2 64.0
CyberGym (Pass@1) 84.5 80.0 84.5 83.3 76.7 88.1
SEC-Bench Pro (Pass@1) 74.3 56.4 30.9 62.8
ExploitGym (Pass@1) 22.1 33.7 15.0 5.4 1.8 15.3
HLE w/ tools (Pass@1) 63.6 59.8 62.5 60.0 51.5 63.9
AutomationBench (Pass@1) 50.3 45.8 46.7 48.8 43.2 37.7 54.8
Agent's Last Exam (Pass@1) 28.6 26.7 27.6 28.5 25.7 25.2 31.8
Chartography w/ tools (Pass@1) 84.0 79.9 68.1 78.9
BabyVision w/ tools (Pass@1) 94.1 88.9 85.7 89.6
ZeroBench-main w/ tools (Pass@5) 52.0 53.0 41.0 49.0

† Text-only subset of HLE.

Code-agent benchmarks (Terminal-Bench 2.1/3.0/4.0, DeepSWE v1.1, NL2Repo-Bench, ProgramBench) are evaluated with the Minimal mode of DeepSeek Harness at a 1M-token context window; the mini-SWE harness is used for DeepSWE v1.1 and the Claude Code harness for SEC-Bench Pro, to match those benchmarks' official setups. Visual agent benchmarks (Chartography, BabyVision, ZeroBench) use the Claude Code harness at a 512K-token context window. Agent's Last Exam and AutomationBench use their official scaffolds.

Performance across agent scaffolds (max reasoning effort)

DeepSeek also publishes how the model holds up across harnesses — N=8 samples per task on DeepSWE v1.1, N=3 on Terminal-Bench 2.1, Linux containers, 1M-token context limit, max_steps=500, Terminal-Bench 2.1 without network access.

Benchmark (Metric) Claude Code Codex OpenCode Pi mini-SWE DSH Minimal DSH Standard DSH PTC
DeepSWE v1.1 (Resolved) 69.8 65.6 65.5 66.2 74.2 72.6 70.5 67.6
Terminal-Bench 2.1 (Pass@1) 88.0 84.1 85.0 86.1 90.3 90.6 85.8 85.8

Base-model evaluation

All base models are evaluated in DeepSeek's internal framework under the same settings; scores within 0.3 of each other are considered equivalent.

Benchmark (Metric) # Shots DeepSeek-V4-Flash-Base DeepSeek-V4-Pro-Base DeepSeek-V4.1-Flash-Base
Architecture MoE MoE MoE
# Backbone Params 284B 1.6T 552B
# Activated Params 13B 49B 8B / 16B
World Knowledge
AGIEval (EM) 3–5-shot 83.9 84.4 83.4
MMLU-Pro (EM) 5-shot 68.3 73.5 74.1
C-Eval (EM) 5-shot 92.1 93.1 92.1
MultiLoKo (LLM-Judge) 5-shot 42.6 50.9 45.5
SimpleQA-Verified (EM) 25-shot 30.1 55.2 42.3
SuperGPQA (EM) 5-shot 46.5 53.9 53.1
Language & Reasoning
BBH (EM) 3-shot 86.9 87.5 86.1
BBEH (EM) 1-shot 25.4 29.8 27.2
DROP (F1) 1-shot 88.6 88.7 87.9
HellaSwag (EM) 0-shot 85.7 88.0 87.2
Code & Math
BigCodeBench (Pass@1) 3-shot 56.8 59.2 60.6
HumanEval (Pass@1) 0-shot 69.5 76.8 79.4
GSM8K (EM) 8-shot 90.8 92.6 93.0
MATH (EM) 4-shot 57.4 64.5 61.1
MGSM (EM) 8-shot 85.7 84.4 80.2
Long Context
LongBench-V2 (EM) 1-shot 44.7 51.5 45.2
Multimodal
MMMU-Pro (EM) 4-shot 56.5
CVBench (EM) 4-shot 77.9
DocVQA (LLM-Judge) 4-shot 95.6
RefCOCO-avg (Acc@0.5) 0-shot 86.0

Prompt Encoding

This release ships no Jinja chat template. DeepSeek instead publishes a self-contained Python reference encoder, plus deepseek-recipe — Rust libraries with Python bindings that convert Messages, Chat Completions and Responses API requests into DeepSeek V4/V4.1 prompts and parse model output back, covering thinking, tool calls, images and generation settings. Calling this model through the Run BiOS API means none of that applies to you: send an ordinary OpenAI-shaped Chat Completions request and the platform handles the encoding.

Safety Notes

DeepSeek's official model card and technical report for DeepSeek-V4.1-Flash do not publish a dedicated safety evaluation, red-teaming or responsible-use section. The model is released under the MIT License. As with any general-purpose model, outputs should be reviewed before use in high-stakes applications.

On Run BiOS

On the Run BiOS platform DeepSeek-V4.1-Flash is available as the serverless model deepseek-v4.1-flash with a 1,048,576-token context window, a 384,000-token maximum output, native image input, function calling, optional reasoning at effort low, high or max, and automatic prompt caching for repeated prefixes.