← AZ-Focus releasesAZ-Focus · Open model · MIT licence · v2.0

DeepSeek-V4.1-Flash-AZ-Focus

DeepSeek's V4.1 Flash, with the second-guessing taken out. It reasons without the hesitation: about half the thinking, agentic coding on par in our SWE-bench run, and a drop-in replacement for the original. Competition maths is the exception: see the results.

2.2×
less thinking on MMLU-Pro
2.6×
less thinking on BIG-Bench Hard
2.0×
less than stock at its lowest effort setting (MMLU-Pro)
≤3
points below the original at most on knowledge & reasoning; lower on competition maths
What changes

It stops second-guessing itself

DeepSeek-V4.1-Flash is a strong reasoner, but a lot of its thinking is hesitation: "Wait…", "Hmm…","Let me double-check…", re-deriving answers it already had. We optimised the model's weights against that habit, while leaving alone the checking it does as an agent (running tests, inspecting output before it submits). It still reasons step by step. It just doesn't loop.

QuestionA train leaves at 3:40 pm and arrives at 6:15 pm. How long is the journey?
Original model0 tokens · 0.0s
 
AZ-Focus0 tokens · 0.0s
 
reasoning AZ-Focus reasoning scrap: "wait", "hmm", "let me double-check…"

Illustration: each block is about 10 thinking tokens. The lengths match our measured averages on reasoning benchmarks (about half the thinking tokens) with accuracy within about three points. Fewer thinking tokens means less to generate, so answers generally arrive sooner.

Beyond the dial

Less thinking at every effort setting

DeepSeek-V4.1-Flash has a built-in reasoning-effort dial from 1 to 100. So the fair question is: why not just turn it down? We measured both models at the same settings. At every setting, AZ-Focus uses about half the thinking tokens, within about three points on accuracy, and at the lowest setting it goes where the dial alone can't: a third to a half less thinking than stock's own minimum.

Fig. 2: Thinking tokens per item at each effort settingDeepSeek-V4.1-Flash AZ-Focus
BIG-Bench Hard
Effort 75default
888 · 86.9%
393 · 84.3%−56%
Effort 50
514 · 87.8%
238 · 86.7%−54%
Effort 35
391 · 88.0%
197 · 86.1%−50%
Effort 25
298 · 86.3%
189 · 87.0%−37%
↑ DeepSeek-V4.1-Flash's lowest setting
MMLU-Pro
Effort 75default
1,692 · 87.2%
758 · 87.4%−55%
Effort 50
1,175 · 89.2%
527 · 86.8%−55%
Effort 35
981 · 88.4%
409 · 86.2%−58%
Effort 25
699 · 88.4%
356 · 85.8%−49%
↑ DeepSeek-V4.1-Flash's lowest setting
Bars: mean thinking tokens per item. Labels: tokens · accuracy, and the change in thinking. Same questions, same serving setup, one run each; differences of about a point in accuracy are within run-to-run noise. Bars left of the dashed line use less thinking than DeepSeek-V4.1-Flash can at any setting.

Useful for both models: DeepSeek's default (75) spends more thinking than these tasks need. Stock at effort 50 gave the single highest MMLU-Pro score (89.2%); if you need every last point, that's the setting to use. If you want the most answers per GPU, AZ-Focus at a low effort setting is the most efficient point we measured.

Results

Across benchmarks, at the default setting

DeepSeek-V4.1-Flash vs AZ-Focus, both served the same way (SGLang, thinking on, DeepSeek's recommended sampling). Scores are compared problem by problem; none of these benchmarks was used to build the model.

BenchmarkWhatOriginalAZ-FocusThinking, originalThinking, AZ-FocusChange
Reasoning & knowledge
MMLU-Pro500 questions, 14 subjects87.2%87.4%1,69275855% less
BIG-Bench Hard540 logic & reasoning tasks86.9%84.3%88839356% less
GPQA Diamond198 graduate-level science questions80.8%86.7%5,6582,87249% less
Coding
SWE-bench Verifiedfirst 100 tasks, agentic (mini-SWE-agent)99/10098.5/10013,57511,91112% less
HumanEval164 problems99.4%98.8%94142954% less
Maths
MATH-500 + GSM8K450 held-out problems96.2%95.9%75154627% less
AIME 202530 competition problems93.3%85.0%7,9846,24622% less

Tuned for real work

We deliberately chose a lighter edit than our first version. A stronger one cut more thinking on reasoning benchmarks, but cost accuracy on agentic coding and competition maths. This release keeps coding at the original's level.

Answers aren't shortened

The change is to deliberation, not to the answers. With thinking switched off, output length on code, JSON, formatting and writing is unchanged.

Sooner, not slower

Speed per token is unchanged, so fewer tokens means earlier answers and more work from the same hardware.

Using it

A drop-in replacement

Same architecture, tokenizer, chat template, file layout and size as DeepSeek-V4.1-Flash, in the same MXFP4/FP8 formats. Serve it with the same engine and settings: no adapter, patch or extra code. Use DeepSeek's recommended sampling (temperature 1.0, top-p 0.95); the difference shows in thinking mode.

Good to know

Limitations

  • Measured on reasoning, knowledge, science, code, agentic coding and maths benchmarks. Long multi-hour agent sessions, other languages and very long contexts haven't been measured separately yet.
  • Hard competition maths is the clear trade-off: AIME 2025 scored about 8 points lower in our runs (4 runs each). If olympiad-style maths is your main use, use the original.
  • If you need the model's full self-verification trail, use the original.

MIT licence, as for the original. Based on DeepSeek-V4.1-Flash by DeepSeek-AI.

Want to build on this?

We're an independent AI lab, looking for a few partners to build with at the frontier.

Work with us