A 164MB download averaging 5.21% word error rate — that is the scorecard Fermion Research published for Phonon-2, its newly open-sourced English speech recognition model. The base is not trained from scratch: it is NVIDIA's Parakeet TDT 0.6B v3, a 2.5GB full-precision model. Closing a 15x size gap to near-lossless comes down to one design choice: a five-level quantization scheme that keeps each encoder weight at roughly 2.1 bits of information.
Seven benchmarks, digit by digit
Per the official model card, across the Open ASR Leaderboard's seven English sets Phonon-2 averages 5.21% WER against the full-precision teacher's 4.96% — an average gap of 0.25 points. Set by set, wins go both ways: on the AMI meeting corpus it scores 9.37 versus 9.42, and on parliamentary speech (VoxPopuli) 2.46 versus 3.19, both beating the teacher; on the other five sets (LS clean, LS other, Earnings-22, GigaSpeech, SPGISpeech) the teacher still leads. The card states the model reaches 100.8% of the teacher's word accuracy on parliamentary speech and wins outright on meetings — consistent with the per-set numbers.
The same table includes a direct rival built on the same foundation: Parakeet Redux from moondream, a ternary-quantized 178MB model. Phonon-2 leads on average, 5.21% versus 5.69%, and set by set everywhere except AMI (9.37 versus 9.16). Two independent compression routes now meet head-on in the sub-200MB weight class.
How you store 2.1 bits
"Five levels" means every weight maps to one of five learned levels, about 2.1 bits on average; the codec implementation sits in the repo's quint5_codec.py. The model card calls it the most accurate open English speech recognition model under 900MB, with every open model scoring better being at least 5.8 times its size — an official, self-reported claim with no third-party replication yet. On speed, the published numbers: an hour of audio transcribed in about 20 seconds on an M5 MacBook Air (174x realtime), 143x on eight Zen 5 CPU cores, and 6,680x on a single H100 at batch 128.
More than a weights file
The GitHub repository ships a complete engine matrix: MLX on Apple Silicon, official Docker images for Linux and Windows CPUs plus NVIDIA CUDA; the command line handles file transcription and live microphone capture, and can serve an OpenAI-compatible endpoint. On the family tree, Phonon-1 (415MB) is built on Qwen3-ASR-0.6B and released under Apache-2.0, while Phonon-2's weights follow the upstream Parakeet license, CC-BY-4.0, with changes listed in a NOTICE file. Phonon-2 is also the engine inside Detta, a Mac dictation app — commercial closure and open weights are not in conflict.
The on-device ASR arms race
Within a single week, Parakeet Redux and Phonon-2 both squeezed a Parakeet foundation under 200MB: one bets on the generality of ternary quantization, the other on the accuracy ceiling of five-level quantization. The real story of Phonon-2 is not any single benchmark number — it is that 2.1-bit extreme quantization managed to beat a full-precision teacher on individual sets. Once quantization no longer automatically means an accuracy ceiling, keeping audio on the machine stops being a privacy compromise and becomes the engineering default.
So: next time you pick a transcription engine for a Mac app, check what 164MB buys you before paying that cloud API bill.