receipt · 2026-08-14 · measured in vLLM

Don’t take the number. Watch it run.

Nine 128K users on one A100 at 134 tok/s aggregate — 2× fp16’s total throughput — where fp16 fits two and FP8’s best receipt-valid rung is four. Same card, same weights, no retraining.

For teams serving long-context open models on vLLM: more concurrent sessions per GPU you already own.

LATEST · AUG 14 — 9 users × 128K, retrieval-gated  ship log ↓

9/4/2
134.1 agg · 14.9 tok/s/user · fp16: ~33 for its 2
fraqtl_repro_receipts.py · Qwen3-4B-Instruct-2507 · 1× A100-80GB
KV pool · each block = one user at ≈128K
Drop-in

One kernel, three architectures, zero per-model changes. A new model’s sidecar builds in under an hour and the weights are never touched.

Retrieval-verified

405/405 passkey cells recovered exactly — three architectures, every arm, every run. Exact match or it doesn’t ship.

Reproducible

Every cell above reruns from the public sidecar repos on a rented A100. One script, run IDs on request.

MODEL MATRIX 3 ARCHITECTURES · ONE KERNEL

One kernel, every model — zero per-model changes.

Each model gets a sidecar from the factory (<1h, weights untouched) and runs on the identical membrane kernel. Same command, same recipe — three architectures, one pattern: ~2.4× fp16 capacity, weight-floor parity at short context, and the speed win grows with context.

MODEL KV CAPACITY VS FP16 VS FP8 8K DECODE 32K DECODE 128K DECODE NEEDLE CELLS
Qwen3-4B-Instruct-2507
2025 · native 256K
2.39× 1.21× 92% fp16 97% fp16 1.34× fp16 · 95% fp8 189/189
Llama-3.1-8B-Instruct
native 128K
2.35× 1.20× 96% fp16 97% fp16 1.23× fp16 · 1.00× fp8 90/90
Mistral-7B-v0.3
native 32K
2.42× 1.21× 95% fp16 99% fp16 n/a — collapses at 128K in all arms incl. fp16 126/126
Needle cells = exact-match passkey grids, 7 depths × 3 keys per context (Llama 5 × 3), all three arms. Capacity = measured KV pool tokens, 8K config. Cells are receipts; blanks are runs not done, not failures.
All receipts & benchmarks → python fraqtl_repro_receipts.py  # reproduce every cell on a rented A100
PILOT INTAKE

Send us your model. We'll tell you if fraQtl can help.

Free Proof of Concept: See a before-and-after analysis of your model in 7 days. 30-Day Pilot: Deploy on your actual workload with guaranteed success criteria — 100% credited toward your production license.

Direct contact: contact@fraqtl.ai
Replies within 1 business day.