What happened?

On 10 September 2026 DeepSeek put DeepSeek-V4.1-Flash on the API as deepseek-flash, with weights on Hugging Face under MIT. The lab blog is dated 9 Sep; the X thread, changelog, and new rate card landed the next UTC morning. Pricing moved at 04:00 UTC on 10 Sep.

The pitch is not a smaller brain. It is a cheaper memory for agents. The model is a 552B backbone MoE plus a 196B Engram lookup table. Hugging Face lists 763B parameters for the checkpoint. Active compute is asymmetric: 8B per token on prefill, 16B on decode. Context is 1M; the API caps output at 384K. Native vision is in from pre-training, not a bolted-on experimental SKU.

The number that organizes the rest: global KV is 890 bytes per token, always in HBM. That is about 1/4 of DeepSeek-V4-Flash and about 437× smaller than V1. Persistent KV (SSD / host) is about 1/8 of V4-Flash, via SWA Bounded Replay: they stop writing sliding-window KV to disk and rebuild the last window on a miss. Cache-hit charges are a large slice of agent bills. Compressing the cache is how Flash undercuts Pro.

The architecture is Causal Encoder-Decoder (CED): 40 layers, 20 + 20. Decoder global KV is projected from the last encoder hidden state, so most of the prompt skips the top half. Compressed Sparse Attention 2 (CSA2) shares main KV and indexer keys across layers (Full / Reindex / Reuse). Main KV is FP4. Decode FLOPs stay almost flat with length: stretching context 256× (4K to 1M) raises decode FLOPs by about 1/4.

DeepSeek is retiring the previous Flash line and, shortly, Pro. deepseek-v4-flash and deepseek-v4-flash-vision-exp already route to V4.1-Flash. From 04:00 UTC on 14 Sep (12:00 Beijing), deepseek-v4-pro does too, at Flash rates, until a V4.1-Pro exists. WorkBuddy (CodeBuddy) and OpenCode are named as day-one partners.

Off-peak Flash is $0.003 / $0.15 / $0.60 per 1M tokens (cache hit / miss / output). Peak is double: $0.006 / $0.30 / $1.20. Off-peak is every hour that is not 01:00-04:00 or 06:00-10:00 UTC, Monday-Friday. Pro's listed card is still several times that until the alias flips.

Why this is interesting

  • The SKU inversion is the product: Flash is billed as the smallest model in a new family, then DeepSeek's own instruct table puts it ahead of V4-Pro on the agent suite they care about. Terminal-Bench 2.1 90.6 vs Pro 87.9 vs Opus-5.0 89.1. DeepSWE v1.1 74.2 vs Pro 62.7 vs Opus 74.0 vs GPT-5.6 Sol 73.0. CyberGym 88.1 vs Pro 83.3. AutomationBench 54.8 vs Opus 50.3. That is their harness, max reasoning effort, 1M window. Do not read it as an independent board.

  • Same metric, different scaffold: The 74.2 DeepSWE number is mini-SWE. On Claude Code it is 69.8; Codex 65.6; OpenCode 65.5. Terminal-Bench 2.1 is 90.6 on DeepSeek Harness Minimal and 88.0 on Claude Code. Print both. The model is less harness-locked than a single headline implies, and the headline is still the friendliest scaffold.

  • The hard benches did not invert: Terminal-Bench 3.0 30.0 vs Opus 43.3. Terminal-Bench 4.0 31.2 vs Opus 51.8. HLE 36.8 (39.1 text-only) vs Opus 56.3 vs Pro 42.7. GPQA Diamond 90.9 vs Sol 94.1. DeepSeek says the remaining gap is expert-domain / science agents, not "can it run a terminal." That split is the honest one.

  • Cache is now a first-class cost center: 890 bytes × 1M tokens is about 890 MB of global KV. Agent loops are prefill-heavy and cache-hit-heavy. CED cuts prefill activation in half; CSA2 + FP4 + bounded replay cut what you store and ship. The API card follows: Flash miss/output is a fraction of listed Pro, and they are routing Pro traffic onto that card in four days. The "multiple parties" claim that Flash already beats Pro on cost, speed, and total runtime is DeepSeek's, unnamed. Treat it as a routing justification, not a third-party eval.

  • Open weights, not a laptop: MIT, recipe repo, they will talk 2,000 GPUs plus a storage cluster. Unsloth's reply is the local read: the 196B Engram is sparse lookup (mmap / SSD), which helps accessibility, and they still asked for smaller dense models. 8B/16B active is the serving story. The backbone is still hundreds of billions. Do not confuse "Flash" with Qwen3.8-27B on a Studio.

  • Effort is a dial, max is a trap: Integer 1-100. API presets map low/high/max to 50 / 75 / 100. Their plot: effort 25→100 lifts DeepSWE 66.0→74.2 and Terminal-Bench 2.1 82.4→90.6 at about 2.5× output tokens. Most of the accuracy is in 60-80; 100 is 1.6-1.8× longer agent traces for a thin gain. If you leave the knob on max, you bought the expensive row of their own chart.

What it is not

Not a proof that open weights closed the frontier. Opus still owns Terminal-Bench 3/4 and HLE. Not an independent bake-off. Not V4.1-Pro. Not a drop-in that keeps Pro quality on science tasks after 14 Sep - it keeps Flash rates. Not a 27B you run at home. The paper's "over 95% of real-world tasks" line is a lab slogan; ignore it.

Bottom line

V4.1-Flash is DeepSeek arguing that the bottleneck for agents is KV, not another trillion parameters. 890-byte global cache, 8B-in / 16B-out, MIT weights, and a rate card that makes Pro the expensive alias they are about to delete. On their agent benches Flash already sits on Opus's shoulder. On the harder terminal and HLE sets it does not. If your workload is long-horizon tools with fat prefixes, this is the default deepseek-flash switch. If your workload is expert science agents, wait for Pro's replacement, or keep a closed frontier model on those jobs.