╭────────────────────────────────────────────────────────────────────────╮ │ │ │ R I G M A R K // AGENT WORKLOAD RECEIPT │ │ BENCHMARKS LOCAL AI HOW CODING AGENTS ACTUALLY USE IT │ │ ● 9/9 BASIC OUTPUT GATES PASSED │ │ │ ├─ MODEL ────────────────────────────────────────────────────────────────┤ │ │ │ GLM-5.3-Flash-EXL3 │ │ │ ├─ SINGLE STREAM ────────────────────────────────────────────────────────┤ │ │ │ WORKLOAD tok/s Range tok/s Last (s) Checks │ │ PROSE 48.6 46.8–49.7 60.6 ✓ 3/3 │ │ CODE 71.8 66.1–75.8 175.8 ✓ 3/3 │ │ STRUCTURED* 102.4 101.9–102.5 10.5 ✓ 3/3 │ │ │ │ * Predictable JSON ceiling; not general agent performance. │ │ Decode medians are estimates; rates include streamed reasoning. │ │ │ ├─ PREFILL ──────────────────────────────────────────────────────────────┤ │ │ │ DEPTH Cold tok/s Cold TTFT (s) Replay TTFT (s) │ │ 8K 1,996 4.10 0.03 │ │ 32K 2,027 16.17 0.06 │ │ 64K 2,003 32.72 0.06 │ │ Medians; replay is an immediate repeat, not a proven hit. │ │ │ ├─ CAPPED CONCURRENT GENERATION ─────────────────────────────────────────┤ │ │ │ Aggregate tok/s: C1 61.0 | C2 84.5 | C4 114.7 │ │ C4: 0/12 normal stops; 0/12 with visible output. │ │ Capped throughput includes reasoning; not completed agent tasks. │ │ │ ├─ APPLIANCE ────────────────────────────────────────────────────────────┤ │ │ │ Hardware: 2x NVIDIA DGX Spark (GB10, 128 GB unified memory each) │ │ Topology: TP2 across 2 Sparks, RoCE over ConnectX-7 via switch │ │ Checkpoint: AxonZeta/GLM-5.3-Flash-AZ-Focus-EXL3 @ 13ba3ef94d88 │ │ Quantisation: EXL3 4 bpw routed experts (mcg codebook), BF16 elsewhere│ │ KV cache: fp8 (e4m3 DSA latent cache and indexer keys, Mia v1.10 │ │ KV=fp8 default); bf16 prompt activations │ │ Engine: TensorFold v0.6.0 + Mia GLM-5.3-Flash kit v1.10 (96 patches, │ │ commit 549fdc5) │ │ │ ├─ SETTINGS ─────────────────────────────────────────────────────────────┤ │ │ │ SUITE: CUSTOM SETTINGS │ │ Changed: --runs=3 │ │ Changed: --decode-tokens=32768 │ │ REQUEST: {"chat_template_kwargs": {"reasoning_effort": "max"}} │ │ Temperature 0.0 | top_p 1.0 | seed 20260905 | protocol 1.3.0 │ │ Decode: 3 runs, 32768-token cap | Prefill: 8K/32K/64K, 3 pairs │ │ Concurrency: C1/C2/C4, 3 rounds, 256-token cap; workload code │ │ Cache isolation: random-salt-per-pair │ │ Comparison ID: glm53flash-tp2-dflash2-max │ │ │ ├─ RECEIPT ──────────────────────────────────────────────────────────────┤ │ │ │ SOURCE git:52c29c3fc1d7 • clean │ │ JSON sha256:e068abb87faa777b… │ │ SHARE THE CARD • LINK THE JSON RECEIPT • #RIGMARK │ │ github.com/alexellis/rigmark │ │ │ ╰────────────────────────────────────────────────────────────────────────╯