╭────────────────────────────────────────────────────────────────────────╮ │ │ │ R I G M A R K // AGENT WORKLOAD RECEIPT │ │ BENCHMARKS LOCAL AI HOW CODING AGENTS ACTUALLY USE IT │ │ ● 15/15 BASIC OUTPUT GATES PASSED │ │ │ ├─ MODEL ────────────────────────────────────────────────────────────────┤ │ │ │ GLM-5.3-Flash-EXL3 │ │ │ ├─ SINGLE STREAM ────────────────────────────────────────────────────────┤ │ │ │ WORKLOAD tok/s Range tok/s Last (s) Checks │ │ PROSE 44.8 44.0–45.0 22.8 ✓ 5/5 │ │ CODE 64.4 63.7–65.1 31.7 ✓ 5/5 │ │ STRUCTURED* 70.6 69.5–72.2 6.8 ✓ 5/5 │ │ │ │ * Predictable JSON ceiling; not general agent performance. │ │ Decode medians are estimates; rates include streamed reasoning. │ │ │ ├─ PREFILL ──────────────────────────────────────────────────────────────┤ │ │ │ DEPTH Cold tok/s Cold TTFT (s) Replay TTFT (s) │ │ 8K 1,961 4.18 0.04 │ │ 32K 1,991 16.46 0.04 │ │ 64K 1,964 33.37 0.05 │ │ Medians; replay is an immediate repeat, not a proven hit. │ │ │ ├─ CAPPED CONCURRENT GENERATION ─────────────────────────────────────────┤ │ │ │ Aggregate tok/s: C1 53.3 | C2 53.6 | C4 53.0 │ │ C4: 0/12 normal stops; 12/12 with visible output. │ │ Capped throughput includes reasoning; not completed agent tasks. │ │ │ ├─ APPLIANCE ────────────────────────────────────────────────────────────┤ │ │ │ Hardware: 2x NVIDIA DGX Spark (GB10, 128 GB unified memory each) │ │ Topology: TP2 across 2 Sparks, RoCE over ConnectX-7 via switch │ │ Checkpoint: Mia-AiLab/GLM-5.3-Flash-EXL3-4bpw-TensorFold @ │ │ 078455ffe647 │ │ Quantisation: EXL3 4 bpw routed experts (mcg codebook), BF16 elsewhere│ │ KV cache: fp8 (e4m3 DSA latent cache and indexer keys, Mia v1.10 │ │ KV=fp8 default); bf16 prompt activations │ │ Engine: TensorFold v0.6.0 + Mia GLM-5.3-Flash kit v1.10 (96 patches, │ │ commit 549fdc5) │ │ │ ├─ SETTINGS ─────────────────────────────────────────────────────────────┤ │ │ │ SUITE: DEFAULT SETTINGS │ │ REQUEST: {"chat_template_kwargs": {"reasoning_effort": "low"}} │ │ Temperature 0.0 | top_p 1.0 | seed 20260905 | protocol 1.3.0 │ │ Decode: 5 runs, 8192-token cap | Prefill: 8K/32K/64K, 3 pairs │ │ Concurrency: C1/C2/C4, 3 rounds, 256-token cap; workload code │ │ Cache isolation: random-salt-per-pair │ │ Comparison ID: glm53flash-tp2-mtp-low │ │ │ ├─ RECEIPT ──────────────────────────────────────────────────────────────┤ │ │ │ SOURCE git:52c29c3fc1d7 • clean │ │ JSON sha256:07a76d76d942fdcf… │ │ SHARE THE CARD • LINK THE JSON RECEIPT • #RIGMARK │ │ github.com/alexellis/rigmark │ │ │ ╰────────────────────────────────────────────────────────────────────────╯