Nine 128K users on one A100 at 134 tok/s aggregate — 2× fp16’s total throughput — where fp16 fits two and FP8’s best receipt-valid rung is four. Same card, same weights, no retraining.
For teams serving long-context open models on vLLM: more concurrent sessions per GPU you already own.
LATEST · AUG 14 — 9 users × 128K, retrieval-gated ship log ↓
One kernel, three architectures, zero per-model changes. A new model’s sidecar builds in under an hour and the weights are never touched.
405/405 passkey cells recovered exactly — three architectures, every arm, every run. Exact match or it doesn’t ship.
Every cell above reruns from the public sidecar repos on a rented A100. One script, run IDs on request.
Each model gets a sidecar from the factory (<1h, weights untouched) and runs on the identical membrane kernel. Same command, same recipe — three architectures, one pattern: ~2.4× fp16 capacity, weight-floor parity at short context, and the speed win grows with context.
| MODEL | KV CAPACITY VS FP16 | VS FP8 | 8K DECODE | 32K DECODE | 128K DECODE | NEEDLE CELLS |
|---|---|---|---|---|---|---|
Qwen3-4B-Instruct-2507 2025 · native 256K |
2.39× | 1.21× | 92% fp16 | 97% fp16 | 1.34× fp16 · 95% fp8 | 189/189 |
Llama-3.1-8B-Instruct native 128K |
2.35× | 1.20× | 96% fp16 | 97% fp16 | 1.23× fp16 · 1.00× fp8 | 90/90 |
Mistral-7B-v0.3 native 32K |
2.42× | 1.21× | 95% fp16 | 99% fp16 | n/a — collapses at 128K in all arms incl. fp16 | 126/126 |
python fraqtl_repro_receipts.py # reproduce every cell on a rented A100
Free Proof of Concept: See a before-and-after analysis of your model in 7 days. 30-Day Pilot: Deploy on your actual workload with guaranteed success criteria — 100% credited toward your production license.