╭────────────────────────────────────────────────────────────────────────╮ │ │ │ R I G M A R K // AGENT WORKLOAD RECEIPT │ │ BENCHMARKS LOCAL AI HOW CODING AGENTS ACTUALLY USE IT │ │ ● 9/9 BASIC OUTPUT GATES PASSED │ │ │ ├─ MODEL ────────────────────────────────────────────────────────────────┤ │ │ │ GLM-5.3-Flash-EXL3 │ │ │ ├─ SINGLE STREAM ────────────────────────────────────────────────────────┤ │ │ │ WORKLOAD tok/s Range tok/s Last (s) Checks │ │ PROSE 61.4 52.4–66.4 160.0 ✓ 3/3 │ │ CODE 64.6 64.1–66.3 115.2 ✓ 3/3 │ │ STRUCTURED* 70.8 69.7–71.0 14.5 ✓ 3/3 │ │ │ │ * Predictable JSON ceiling; not general agent performance. │ │ Decode medians are estimates; rates include streamed reasoning. │ │ │ ├─ PREFILL ──────────────────────────────────────────────────────────────┤ │ │ │ DEPTH Cold tok/s Cold TTFT (s) Replay TTFT (s) │ │ 8K 1,958 4.18 0.04 │ │ 32K 1,986 16.50 0.04 │ │ 64K 1,958 33.47 0.05 │ │ Medians; replay is an immediate repeat, not a proven hit. │ │ │ ├─ CAPPED CONCURRENT GENERATION ─────────────────────────────────────────┤ │ │ │ Aggregate tok/s: C1 48.2 | C2 52.3 | C4 51.9 │ │ C4: 0/12 normal stops; 0/12 with visible output. │ │ Capped throughput includes reasoning; not completed agent tasks. │ │ │ ├─ APPLIANCE ────────────────────────────────────────────────────────────┤ │ │ │ Hardware: 2x NVIDIA DGX Spark (GB10, 128 GB unified memory each) │ │ Topology: TP2 across 2 Sparks, RoCE over ConnectX-7 via switch │ │ Checkpoint: AxonZeta/GLM-5.3-Flash-AZ-Focus-EXL3 @ 13ba3ef94d88 │ │ Quantisation: EXL3 4 bpw routed experts (mcg codebook), BF16 elsewhere│ │ KV cache: fp8 (e4m3 DSA latent cache and indexer keys, Mia v1.10 │ │ KV=fp8 default); bf16 prompt activations │ │ Engine: TensorFold v0.6.0 + Mia GLM-5.3-Flash kit v1.10 (96 patches, │ │ commit 549fdc5) │ │ │ ├─ SETTINGS ─────────────────────────────────────────────────────────────┤ │ │ │ SUITE: CUSTOM SETTINGS │ │ Changed: --runs=3 │ │ Changed: --decode-tokens=32768 │ │ REQUEST: {"chat_template_kwargs": {"reasoning_effort": "max"}} │ │ Temperature 0.0 | top_p 1.0 | seed 20260905 | protocol 1.3.0 │ │ Decode: 3 runs, 32768-token cap | Prefill: 8K/32K/64K, 3 pairs │ │ Concurrency: C1/C2/C4, 3 rounds, 256-token cap; workload code │ │ Cache isolation: random-salt-per-pair │ │ Comparison ID: glm53flash-tp2-mtp-max │ │ │ ├─ RECEIPT ──────────────────────────────────────────────────────────────┤ │ │ │ SOURCE git:52c29c3fc1d7 • clean │ │ JSON sha256:4f2821e26aa481f2… │ │ SHARE THE CARD • LINK THE JSON RECEIPT • #RIGMARK │ │ github.com/alexellis/rigmark │ │ │ ╰────────────────────────────────────────────────────────────────────────╯