DeepSeek-V4.1-Flash-AZ-Focus
DeepSeek's V4.1 Flash, with the second-guessing taken out. It reasons without the hesitation: about half the thinking, agentic coding on par in our SWE-bench run, and a drop-in replacement for the original. Competition maths is the exception: see the results.
It stops second-guessing itself
DeepSeek-V4.1-Flash is a strong reasoner, but a lot of its thinking is hesitation: "Wait…", "Hmm…","Let me double-check…", re-deriving answers it already had. We optimised the model's weights against that habit, while leaving alone the checking it does as an agent (running tests, inspecting output before it submits). It still reasons step by step. It just doesn't loop.
Illustration: each block is about 10 thinking tokens. The lengths match our measured averages on reasoning benchmarks (about half the thinking tokens) with accuracy within about three points. Fewer thinking tokens means less to generate, so answers generally arrive sooner.
Less thinking at every effort setting
DeepSeek-V4.1-Flash has a built-in reasoning-effort dial from 1 to 100. So the fair question is: why not just turn it down? We measured both models at the same settings. At every setting, AZ-Focus uses about half the thinking tokens, within about three points on accuracy, and at the lowest setting it goes where the dial alone can't: a third to a half less thinking than stock's own minimum.
Useful for both models: DeepSeek's default (75) spends more thinking than these tasks need. Stock at effort 50 gave the single highest MMLU-Pro score (89.2%); if you need every last point, that's the setting to use. If you want the most answers per GPU, AZ-Focus at a low effort setting is the most efficient point we measured.
Across benchmarks, at the default setting
DeepSeek-V4.1-Flash vs AZ-Focus, both served the same way (SGLang, thinking on, DeepSeek's recommended sampling). Scores are compared problem by problem; none of these benchmarks was used to build the model.
| Benchmark | What | Original | AZ-Focus | Thinking, original | Thinking, AZ-Focus | Change |
|---|---|---|---|---|---|---|
| Reasoning & knowledge | ||||||
| MMLU-Pro | 500 questions, 14 subjects | 87.2% | 87.4% | 1,692 | 758 | 55% less |
| BIG-Bench Hard | 540 logic & reasoning tasks | 86.9% | 84.3% | 888 | 393 | 56% less |
| GPQA Diamond | 198 graduate-level science questions | 80.8% | 86.7% | 5,658 | 2,872 | 49% less |
| Coding | ||||||
| SWE-bench Verified | first 100 tasks, agentic (mini-SWE-agent) | 99/100 | 98.5/100 | 13,575 | 11,911 | 12% less |
| HumanEval | 164 problems | 99.4% | 98.8% | 941 | 429 | 54% less |
| Maths | ||||||
| MATH-500 + GSM8K | 450 held-out problems | 96.2% | 95.9% | 751 | 546 | 27% less |
| AIME 2025 | 30 competition problems | 93.3% | 85.0% | 7,984 | 6,246 | 22% less |
Tuned for real work
We deliberately chose a lighter edit than our first version. A stronger one cut more thinking on reasoning benchmarks, but cost accuracy on agentic coding and competition maths. This release keeps coding at the original's level.
Answers aren't shortened
The change is to deliberation, not to the answers. With thinking switched off, output length on code, JSON, formatting and writing is unchanged.
Sooner, not slower
Speed per token is unchanged, so fewer tokens means earlier answers and more work from the same hardware.
A drop-in replacement
Same architecture, tokenizer, chat template, file layout and size as DeepSeek-V4.1-Flash, in the same MXFP4/FP8 formats. Serve it with the same engine and settings: no adapter, patch or extra code. Use DeepSeek's recommended sampling (temperature 1.0, top-p 0.95); the difference shows in thinking mode.
Limitations
- Measured on reasoning, knowledge, science, code, agentic coding and maths benchmarks. Long multi-hour agent sessions, other languages and very long contexts haven't been measured separately yet.
- Hard competition maths is the clear trade-off: AIME 2025 scored about 8 points lower in our runs (4 runs each). If olympiad-style maths is your main use, use the original.
- If you need the model's full self-verification trail, use the original.
MIT licence, as for the original. Based on DeepSeek-V4.1-Flash by DeepSeek-AI.
Want to build on this?
We're an independent AI lab, looking for a few partners to build with at the frontier.