One row per measured result. Every row links to the public artifact that backs it. Updating this page means adding a row — nothing here is ever rewritten.
2026-08-25 llama.cpp launch 1.79× faster than q8_0 KV at 128K, 92% of fp16 at ~2.45× less memory; 36 vs 28 users. → 2026-08-25 Second model, zero code changes Nemo-12B: 32 vs 24 users, losses disclosed. → 2026-08-25 Spec-decode compat 1.04× at max concurrency, receipted. → 2026-08-25 Distribution ~18K downloads across 14 public repos; ~12K in the last 30 days. → 2026-08 Quant lane (Hi-Fi) Qwen3-4B −55.9% KLD · 27B −45.3% · E2B −36% — all at byte parity, all needle-gated (Qwen3-4B through full 262K). → 2026-08-14 vLLM flagship 9 concurrent ≈128K users, one A100, 134.1 tok/s aggregate, 9/9 retrieval. → standing Repro Every number = run ID + receipt JSON + one-command kit (~$12 of rented A100). →Deep tables, method fine print and boundaries: /benchmark · historical record: /benchmark-archive · contact: contact@fraqtl.ai