GLM-5.3-Flash-AZ-Focus
Z.ai's GLM-5.3-Flash, with the second-guessing turned down. About 40% less thinking at max effort, with reasoning, maths and code accuracy level with or above the original.
It stops second-guessing itself
GLM-5.3-Flash thinks carefully, but a lot of that thinking is hesitation: "Wait…","Let me double-check…", re-deriving answers it already had. We optimised the model's weights against that habit, while leaving alone the checking it does as an agent (running tests, inspecting output before it submits). It still reasons step by step. It just loops less.
Illustration: each block is about 10 thinking tokens. The lengths match our measured averages on reasoning benchmarks (about half the thinking tokens) with accuracy within about three points. Fewer thinking tokens means less to generate, so answers generally arrive sooner.
Less thinking, same or better answers
At GLM-5.3-Flash's default (max) effort, AZ-Focus cuts thinking by 20% to 50% and scores the same or higher on reasoning, knowledge, maths and science.
Across benchmarks, at max effort
GLM-5.3-Flash vs AZ-Focus, both served the same way (thinking on, Z.ai's recommended sampling). None of these benchmarks was used to build the model. Measured on the 4-bit EXL3 build; FP8 results will follow.
| Benchmark | What | Original | AZ-Focus | Thinking, original | Thinking, AZ-Focus | Change |
|---|---|---|---|---|---|---|
| Reasoning & knowledge | ||||||
| MMLU-Pro | 500 questions, 14 subjects | 86.2% | 86.8% | 2,114 | 1,289 | 39% less |
| BIG-Bench Hard | 540 logic & reasoning tasks | 83.7% | 84.1% | 634 | 326 | 49% less |
| GPQA Diamond† | 198 graduate-level science questions | 80.8% | 82.3% | 6,285 | 5,000 | 20% less |
| Maths | ||||||
| AIME 2025† | 30 competition problems, 2 runs | 75.0% | 81.7% | 16,500 | 11,200 | 32% less |
| Coding | ||||||
| HumanEval+ & MBPP+† | EvalPlus, 542 problems | 85.8% | 86.9% | |||
| SWE-bench Verified | first 100 tasks, agentic (mini-SWE-agent), 2 runs | 95/100 | 92/100 | |||
| Agents | ||||||
| τ²-bench airline† | tool-using customer-service agent, 2 runs | 0.89 | 0.88 | |||
| τ²-bench retail† | tool-using customer-service agent, 2 runs | 0.919 | 0.905 | |||
SWE-bench: 93 and 91 vs 95 and 95. GPQA thinking is total output per question. † Measured on an earlier build of the same edit.
Best at max and high
That's where most of the hesitation is. At high effort AZ-Focus halves MMLU-Pro thinking and scores higher; at low effort the original already thinks very little, so use the original there.
Agents keep checking
Tool-using agents stay within about a point on τ²-bench. Agentic coding is about 3 points below the original on SWE-bench Verified: the one place the original's extra checking still pays.
Weights only
The change is in the weights: no adapter, runtime patch or extra code. Answers aren't shortened; with thinking off, output length is unchanged.
Best at max and high
GLM-5.3-Flash has three effort settings. At max and high, AZ-Focus cuts MMLU-Pro thinking by 39–47% with accuracy up. At low, the original already thinks very little; use the original there.
Same questions, one run each. High and low were measured on an earlier build of the same edit.
Two formats
FP8: the same format and file layout as Z.ai's official FP8 release, a drop-in replacement for vLLM and SGLang. EXL3: routed experts at 4 bits and everything else BF16, quantized by us directly from Z.ai's BF16 release; load it with an engine that supports GLM-5.3-Flash's EXL3 layout. Same architecture, tokenizer and chat template as GLM-5.3-Flash; use Z.ai's recommended sampling at max or high effort.
Limitations
- Agentic coding is about 3 points below the original on SWE-bench Verified. If long agent runs are your main use, compare both on your own tasks.
- At low effort the edit costs accuracy without saving much thinking. Use the original there.
- Other languages, very long contexts and multi-hour agent sessions haven't been measured separately yet.
MIT licence, as for the original. Based on GLM-5.3-Flash by Z.ai.
Want to build on this?
We're an independent AI lab, looking for a few partners to build with at the frontier.