FP8 beyond 4 users: higher rungs did not pass the receipt gate (r2 audit) — no claim either way.
256K: capacity verified, quality gate still open — no claims yet.
H100 / consumer GPUs: runtime is SM80 (A100) only today.
GPTQ-on-Phi-3, AWQ 3/5-bit, MoE matched-bits vs AWQ/GPTQ: not attempted — no implicit claim. Details in the archive.
Blanks in any table are runs not done, not failures.
THE GATE
Low-bit KV buys memory by losing the answer. Ours doesn't.
Mistral-7B at 128K in llama.cpp. Same model, same prompts — only the KV representation changes. Five needles buried at five depths; each dot is one retrieved.
fp16 KV
22,657 MiB
5 / 5
Q8 KV
15,437 MiB · −31.9%
1 / 5
Q4 KV
11,287 MiB · −50.2%
0 / 5
fraQtl D1
13,261 MiB · −41.5%
5 / 5
Mistral-7B-Instruct · 128K context · llama.cpp · NIAH = 5-needle retrieval · live VRAM is measured resident KV footprint. fraQtl D1 sits 14.1% below Q8 on memory and keeps 5/5 where Q8 keeps 1/5. Lower memory and intact retrieval at once.
WHY IT HOLDS
Quantize the wrong basis and you quantize the answer. So we change basis first.
Per-token integer quantization spends its bits uniformly across a basis where the signal is not uniform — which is why retrieval is the first thing to go. fraQtl rotates each head's K and V into their own eigenbasis, protects the leading directions, and quantizes the rest. The packed result is read fragment-native into the tensor cores: no dequantization pass, no reconstruction buffer.
ONE HEAD, ONE PAGE · WHAT THE MEMBRANE DOES TO IT
01 · RAW KV, fp16
Signal spread across every direction. Nothing to cut cheaply.
→
02 · EIGENBASIS ROTATION
Per-head rotation, calibrated in 0.3s. Variance collapses onto a few axes.
→
03 · PACKED MEMBRANE
Protected rows kept at rank. The tail goes to INT3 with sign correction.
k = 16
protected directions per head, V+K INT3
0.3 s
calibration per model — weights untouched
0
reconstruction passes — read fragment-native
σ = 0
bit-identical across seeds 7 / 123 / 2024
The V/K theorem is validated at architecture level on Mistral 7B GQA-4, Qwen 2.5 3B GQA-2, Llama 3.2 3B and Phi-3 — the same recipe transfers because the rotation is derived per head, not per model. Eviction methods (TOVA, SnapKV, StreamingLLM, H2O) drop tokens; per-token integer methods (KIVI, KVQuant) keep every token in the original basis. fraQtl keeps every token and changes the basis.
R1 · VLLM RUNTIMEFLAGSHIP
Alone it matches fp16. Under load it pulls away.
The runtime reads packed KV pages straight into the tensor cores — no reconstruction step. At one user, weights dominate and everyone lands within a few percent. Add users and fp16 runs out of KV memory first, fp8 second. Two models, one kernel, zero per-model changes.
CONCURRENT USERS · 128K CONTEXT EACH · ONE A100-80GB
Qwen3-4B-2507 · bit-plane V2 · Aug 14
fraQtl
134.1 tok/s
9/9 needles
fp8 KV
117.6 tok/s
peak at 4 users
fp16
66.6 tok/s
ceiling at 2 users
Each block is one user holding a full 128K context (130,048-token prompt + 1,024-token decode) with an independent passkey at a distinct depth. Prefill at 128K: ~3,950 tok/s, at parity with fp16.
KV POOL CAPACITY · MISTRAL-7B
fraQtl997,200
fp8 KV825,104
fp16412,544
Measured pool tokens, 8K config, one A100-80GB. 2.42× fp16.
DECODE SPEED VS CONTEXT
% of fp16, same hardware — the speed win grows with context
Exact-match passkey grids (7 depths × 3 keys per context) run on every arm of every receipt; a single missed passkey discards the result. Qwen3-4B-2507: 189/189 cells across 8K/32K/128K. Mistral-7B-v0.3: 126/126 across 8K/32K. Llama-3.1-8B: 90/90 — 405/405 total, plus the Aug 14 nine-user run's 9/9 per-user passkeys at depths spanning 0–100%. Same kernel for all three models; a new model's sidecar builds in under an hour.
METHOD & CAVEATS — full conditions, known bugs, what is not claimedexpand ▾
BF16/FP16 is the quality reference; public Q4 is the practical size baseline teams already deploy. Matched Q4_K_M bakeoff in progress — results published when locked. Method: vLLM real paged serving, CUDA graphs on, prefix caching off, two-pass subtractive decode timing, temperature 0, exact-match passkey retrieval per arm. fp8 = fp16 weights + fp8-E4M3 KV (FlashInfer), the strong baseline; weights identical across all arms. fraQtl = rank-protected eigenbasis KV read fragment-native into the tensor cores, zero reconstruction. 128K receipts run with chunked prefill disabled (known perf bug under chunking, fix in progress; stock arms unaffected). A batch-routing bug that dropped needles at batch ≥2 was found 07-03, root-caused, fixed, and fully re-receipted — pre-fix numbers were never published. No 256K claims yet: capacity verified, quality gate open. Qwen3-4B KV pool: 845,136 tokens vs fp16 361,776 (fp8 709,488) — 2.3× capacity. Aug 14 nine-user cells use bit-plane V2 KV tiering (+45% KV pool vs V1); the July V1-layout receipt remains the per-layout bandwidth roof — V2 trades ~5% aggregate for 3 more users. Per-user decode at 9 users is 14.9 tok/s (fp16: ~33, for its 2 users). All cells reproducible from the public Hugging Face sidecar repos with fraqtl_repro_receipts.py; run IDs available on request.
R2 · MODEL MATRIX
One kernel, every model, zero per-model changes.
Each model gets a sidecar from the factory in under an hour, weights untouched, and runs on the identical membrane kernel. Same command, same recipe, three architectures, one pattern.
Qwen3-4B-Instruct-2507
GQA · native 256K
189 / 189 NEEDLE CELLS
Llama-3.1-8B-Instruct
GQA-8 · native 128K
90 / 90 NEEDLE CELLS
Mistral-7B-v0.3
GQA-4 · native 32K
126 / 126 NEEDLE CELLS
Model
KV vs fp16
vs fp8
8K
32K
128K
Qwen3-4B-Instruct-2507
2025 · native 256K
2.39×
1.21×
92%
97%
1.34×
Llama-3.1-8B-Instruct
native 128K
2.35×
1.20×
96%
97%
1.23×
Mistral-7B-v0.3
native 32K
2.42×
1.21×
95%
99%
n/a — collapses in all arms incl. fp16
Decode columns are % of fp16 at that context. Needle cells = exact-match passkey grids, 7 depths × 3 keys per context (Llama 5 × 3), all three arms. Capacity = measured KV pool tokens, 8K config. Cells are receipts; blanks are runs not done, not failures.
R4 · QUALITY & FOOTPRINTHI-FI WEIGHTS
fp16 OOMs at 64K. We run 128K with 28 GB spare.
The public Qwen 3.6 35B-A3B artifact on a single A100-80GB. Same model, same hardware, eight times the context — the difference between needing one GPU and needing two.
VRAM FOOTPRINT · QWEN 3.6 35B-A3B · ONE A100-80GB
—— 80 GB CEILING
16K CONTEXTboth fit on one GPU
fp16 ~71 GB
25.6 GB
64K CONTEXTfp16 needs 2 GPUs · fraQtl 1
82.9 GB → OOM
36.8 GB
128K CONTEXT28 GB still free
85+ GB → 2 GPUs
51.7 GB
fp16 at 64K = 82.86 GB measured (OOMs on the 80 GB ceiling). 16K and 128K fp16 figures are model + KV extrapolations, conservative lower bounds. fraQtl figures measured with the public artifact at huggingface.co/fraQtl/Qwen3.6-35B-A3B-compressed.
RESEARCH BACKING — partial-stack bake-off vs KIVI / TOVA / SnapKV / H2Oexpand ▾
A separate research comparison, not the customer claim. Partial-stack, Mistral / Llama, sub-4-bit, C44 bake-off.
Metric
fraQtl
Nearest published
Difference
NIAH retention 1080 trials
98.5%
97.8% · TOVA
+0.7 pp
PPL Δ at 3.5× compression
+0.012
+0.214 · SnapKV
~18× tighter
NIAH at 128K Llama 3.1 8B
100%
0% · KIVI-2
+100 pp
GPUs needed at 128K 35B MoE
1× A100-80GB
2× A100-80GB
half the hardware
Honest note: at matched 4-bit, fraQtl ties KIVI-4 / KVQuant-4 within sampling noise. The wins above are at sub-4-bit, at 128K, against eviction methods, and on hardware footprint. Sources: Mistral-7B-Instruct C44 bake-off (1080 NIAH trials, 8K–31K, 3 needle types); Llama 3.1 8B C44d/e (128K, 3 needles × 5 trials per depth).
QUALITY EVIDENCE · QUOTED VERBATIM
“On the 2026-06-11 hardened paired packet, D2 located 63/63 NIAH needles (fp16-parity); observed failures are rare exact-ID transcription noise (~5%) and synthetic 3-hop state-tracking flips (~17%, symmetric), both root-caused; real multi-hop QA (LongBench) shows no F1 gap.”
D2-family recipe lineage · direct evidence for the current recipe: the 9/9 grid above
PUBLIC ARTIFACTS
Everything above is downloadable. Nothing is a demo.
Every receipt traces to a public repo on huggingface.co/fraQtl — KV sidecar kits with the runtime and repro script, and calibrated Hi-Fi GGUFs with KLD receipts.
ARCHIVE — research bake-offs, cross-architecture weight matrix, 2026-H1 receiptsexpand ▾
Superseded numbers preserved for the record: the matched-4-bit weight-compression matrix across MHA + GQA-2 + GQA-3 + GQA-4, the 9/9 matched-bits peer wins, per-row scripts and commit hashes, and the H1 cross-architecture tables. Restyled to this system in the production file — kept short here so the mock stays readable.
Run it against your own workload.
The runtime wheel and repro script are public. A rented A100 is enough to check every number on this page.