← AZ-Focus releasesAZ-Focus · Open model · MIT licence · FP8 · EXL3

GLM-5.3-Flash-AZ-Focus

Z.ai's GLM-5.3-Flash, with the second-guessing turned down. About 40% less thinking at max effort, with reasoning, maths and code accuracy level with or above the original.

39%
less thinking on MMLU-Pro
49%
less thinking on BIG-Bench Hard
+6.7
points on AIME 2025, with 32% less thinking
Same
or better accuracy on reasoning, science and code
What changes

It stops second-guessing itself

GLM-5.3-Flash thinks carefully, but a lot of that thinking is hesitation: "Wait…","Let me double-check…", re-deriving answers it already had. We optimised the model's weights against that habit, while leaving alone the checking it does as an agent (running tests, inspecting output before it submits). It still reasons step by step. It just loops less.

QuestionA train leaves at 3:40 pm and arrives at 6:15 pm. How long is the journey?
Original model0 tokens · 0.0s
 
AZ-Focus0 tokens · 0.0s
 
reasoning AZ-Focus reasoning scrap: "wait", "hmm", "let me double-check…"

Illustration: each block is about 10 thinking tokens. The lengths match our measured averages on reasoning benchmarks (about half the thinking tokens) with accuracy within about three points. Fewer thinking tokens means less to generate, so answers generally arrive sooner.

At max effort

Less thinking, same or better answers

At GLM-5.3-Flash's default (max) effort, AZ-Focus cuts thinking by 20% to 50% and scores the same or higher on reasoning, knowledge, maths and science.

Fig. 2: Thinking per item at max effortGLM-5.3-Flash AZ-Focus
MMLU-Pro500 questions
2,114 tokens · 86.2%
1,289 · 86.8%−39%
BIG-Bench Hard540 tasks
634 tokens · 83.7%
326 · 84.1%−49%
AIME 2025†30 problems, 2 runs
16.5k tokens · 75.0%
11.2k · 81.7%−32%
GPQA Diamond†198 questions
6,285 tokens · 80.8%
5,000 · 82.3%−20%
Each row is scaled to the original's thinking. Labels: mean thinking tokens per item · accuracy (GPQA: total output per question). Same questions, same serving setup; differences of about a point are within run-to-run noise. † Measured on an earlier build of the same edit.
Results

Across benchmarks, at max effort

GLM-5.3-Flash vs AZ-Focus, both served the same way (thinking on, Z.ai's recommended sampling). None of these benchmarks was used to build the model. Measured on the 4-bit EXL3 build; FP8 results will follow.

BenchmarkWhatOriginalAZ-FocusThinking, originalThinking, AZ-FocusChange
Reasoning & knowledge
MMLU-Pro500 questions, 14 subjects86.2%86.8%2,1141,28939% less
BIG-Bench Hard540 logic & reasoning tasks83.7%84.1%63432649% less
GPQA Diamond†198 graduate-level science questions80.8%82.3%6,2855,00020% less
Maths
AIME 2025†30 competition problems, 2 runs75.0%81.7%16,50011,20032% less
Coding
HumanEval+ & MBPP+†EvalPlus, 542 problems85.8%86.9%
SWE-bench Verifiedfirst 100 tasks, agentic (mini-SWE-agent), 2 runs95/10092/100
Agents
τ²-bench airline†tool-using customer-service agent, 2 runs0.890.88
τ²-bench retail†tool-using customer-service agent, 2 runs0.9190.905

SWE-bench: 93 and 91 vs 95 and 95. GPQA thinking is total output per question. † Measured on an earlier build of the same edit.

Best at max and high

That's where most of the hesitation is. At high effort AZ-Focus halves MMLU-Pro thinking and scores higher; at low effort the original already thinks very little, so use the original there.

Agents keep checking

Tool-using agents stay within about a point on τ²-bench. Agentic coding is about 3 points below the original on SWE-bench Verified: the one place the original's extra checking still pays.

Weights only

The change is in the weights: no adapter, runtime patch or extra code. Answers aren't shortened; with thinking off, output length is unchanged.

Effort setting

Best at max and high

GLM-5.3-Flash has three effort settings. At max and high, AZ-Focus cuts MMLU-Pro thinking by 39–47% with accuracy up. At low, the original already thinks very little; use the original there.

Fig. 3: Thinking tokens per item at each effort settingGLM-5.3-Flash AZ-Focus
BIG-Bench Hard
Effort maxdefault
634 · 83.7%
326 · 84.1%−49%
Effort high
117 · 83.5%
87 · 81.5%−26%
Effort low
89 · 83.5%
69 · 80.0%−22%
MMLU-Pro
Effort maxdefault
2,114 · 86.2%
1,289 · 86.8%−39%
Effort high
413 · 85.0%
219 · 86.2%−47%
Effort low
96 · 84.2%
89 · 83.0%−7%
Bars: mean thinking tokens per item. Labels: tokens · accuracy, and the change in thinking. Same questions, same serving setup, one run each.

Same questions, one run each. High and low were measured on an earlier build of the same edit.

Using it

Two formats

FP8: the same format and file layout as Z.ai's official FP8 release, a drop-in replacement for vLLM and SGLang. EXL3: routed experts at 4 bits and everything else BF16, quantized by us directly from Z.ai's BF16 release; load it with an engine that supports GLM-5.3-Flash's EXL3 layout. Same architecture, tokenizer and chat template as GLM-5.3-Flash; use Z.ai's recommended sampling at max or high effort.

Good to know

Limitations

  • Agentic coding is about 3 points below the original on SWE-bench Verified. If long agent runs are your main use, compare both on your own tasks.
  • At low effort the edit costs accuracy without saving much thinking. Use the original there.
  • Other languages, very long contexts and multi-hour agent sessions haven't been measured separately yet.

MIT licence, as for the original. Based on GLM-5.3-Flash by Z.ai.

Want to build on this?

We're an independent AI lab, looking for a few partners to build with at the frontier.

Work with us