Enerzai announced a 1.58-bit weight quantization breakthrough on June 3, achieving 77% memory reduction over the FP16 baseline. The method uses a learned, layer-wise mixed-precision scheme that puts most weights at 1.58-bit while keeping critical layers at higher precision. On standard LLM benchmarks, quality loss is under 1 point, and inference can run on consumer GPUs that previously couldn't host these models.