Pushing LLMs below 4 bits has never been about whether compression is possible; it has been about whether the model can still run fast. Today's two industry routes—Hadamard-rotation-based incoherence processing (QTIP, SpinQuant, etc.) and trellis-coded quantization (TCQ) with structured codebooks—each solve half the problem. Hadamard makes the weight distribution more uniform, so low-bit quantizers can approximate better, but it cannot be absorbed into neighboring operators and therefore has to run online in the inference path, eating throughput. Trellis quantization gives high-dimensional shaping via Viterbi search and state-dependent mapping, but every step requires state-keyed dequantization; if each dequant needs a bunch of logic/arithmetic/data-movement operations, GPU weight-consumption bandwidth cannot keep up, so the bandwidth dividend is eaten up.
The XOR-Trellis paper (arXiv:2610.00432, 2026-09-30) from Arm removes both walls at once. Instead of treating RHT as a mandatory pre-step, it re-explains why Hadamard makes quantization better: RHT is effective primarily because it smooths out non-uniform weight sensitivity, which makes conventional isotropic Euclidean distance a reasonable proxy for a model-aware objective. In other words, what is missing is not RHT itself, but "search objective must reflect model sensitivity". The paper therefore proposes two independent but complementary designs:
- XOR-Trellis FP4 decoder: uses four carefully designed local reconstruction palettes (P0 = [-6,-1.5,0,2], P1 = [-4,-1,0.5,3], P2 = [-3,-0.5,1,4], P3 = [-2,0,1.5,6]) combined with a 16-bit sliding state. Four masked-parity computations (0x1089/0x2112/0x0A24/0x2441) XOR out p0, p1, q0, q1; the first pair selects the palette, the second permutes how the four branch identities map to palette entries. Dequantization collapses from "state-to-value lookup" to "XOR + small FP4 palette lookup"—no dynamic query, no expensive computed function, only operations that are highly parallel on a GPU/NPU. Appendix C.2 proves the four palettes are pointwise optimal for the 4-way reconstruction problem under fixed FP4 elements, uniform use of four disjoint 4-entry palettes, and scalar squared-error distortion.
- Curvature-aware Viterbi objective: rather than rotating the weights, it feeds the diagonal factor D from the LDL decomposition H = LDL⊤ directly into the additive distortion cost of Viterbi. Large diagonal entries in D represent dimensions where quantization error is more costly to the model loss, so the search gives those positions higher weight. The whole objective is computed in the original coordinate space, fully offline, without modifying the compressed representation, with zero overhead added to the runtime dequantizer.
What the numbers say
On Llama-3.1-8B WikiText-2 perplexity (PPL, lower is better), as reported in the paper:
- Unrotated baseline trellis: 8.69
- Hadamard-enabled QTIP-3INST: 8.48
- XOR-Trellis (curvature-aware, no Hadamard): 8.00
That is, the paper's method beats the "add RHT" tradition and completely cuts the online Hadamard path.
Cross-model extension (full data in Appendix B Tables 4–7):
- Llama-3.2-1B-Instruct: BF16 13.16 → XOR-Trellis (no RHT + D-weighted + energy ordering) 21.38 PPL, zero-shot average accuracy lifted from baseline 46.74% to 51.43%.
- Llama-3.2-3B-Instruct: zero-shot average lifted from 54.03% to 58.18%.
- Llama-2-13B: 5.47 PPL / 65.38% average, slightly above the Hadamard-enabled QTIP-3INST (5.64 / 65.51%).
- Mistral-7B: 5.67 PPL / 67.95% average, well ahead of QTIP-3INST's 6.04 / 66.80%.
- Qwen3-8B: 10.25 PPL / 67.75% average, also beating the RHT-enabled QTIP-3INST (11.35 / 67.52%).
The consistency across models shows this is not "lucky tuning" of one model but a real structural improvement.
So what
XOR-Trellis's real selling point is not the 0.5 PPL bump, but the explicit argument that "Hadamard is a non-essential step". Over the past year, the industry has poured effort into RHT and learned rotation (QuaRot, SpinQuant) under the premise that "weight sensitivity must be smoothed by some rotation before compression can work". This paper removes that premise with "diagonal Hessian weighting + XOR local palette":
- For hardware vendors, the dequant path becomes tractable: XOR + 4-entry lookup is SIMD-friendly on a GPU, no longer tied to dynamic computed functions.
- For model-side researchers, the implicit "Hadamard is necessary" assumption can be retired; engineering effort can shift to harder and more underserved areas (e.g., activation quantization, MoE quantization).
- For the NVIDIA Blackwell/FP4 GEMM and Rubin block-scaled FP4 hardware roadmap, the paper provides software-layer proof of Hadamard-free viability, meaning future FP4 inference stacks can shorten from "general weights → rotation → quantization" to "general weights → quantization".
One more detail on the grouping format: the paper uses E2M1 elements + E5M3 block scales, aligning with the NVIDIA Blackwell/Rubin block-scaled FP4 roadmap—so this is not a theoretical toy but a battle-tested version aligned with hardware.
The LLM quantization track has run from GPTQ (2022) to today, with the core question always being "how to recover the accuracy loss from compression without a big cost". XOR-Trellis's contribution: the cost of that recovery edge is much lower than the industry assumes, and it can be done entirely on the quantization side, without RHT preprocessing that spans the entire inference graph. The next step is whether Arm/this SMFP4 can be integrated into the main paths of mainstream LLM inference frameworks (vLLM, TensorRT-LLM, llama.cpp). Reference: arXiv:2610.00432 full text and Appendix B Tables 4–7.