Speech-to-text has always assumed a GPU: models run over a gigabyte, and throughput is bought with graphics cards. Moondream's parakeet-redux breaks that assumption — NVIDIA's open ASR model parakeet-tdt-0.6b-v3 compressed to 178MB, transcribing at 113x realtime on an ordinary x86 CPU without touching a GPU.
Ternary weights: the most aggressive quantization
parakeet-redux applies 1.58-bit ternary quantization to parakeet-tdt-0.6b-v3: same architecture, same tokenizer, but every encoder weight is now -1, 0, or +1. Weights drop from 1.2GB to 178MB — under 15% of the original, small enough for any edge device. The companion Photon runtime reads the packed weights directly: AVX-512 VNNI on x86, NEON on ARM, Metal on Apple GPUs.
Speed: the fastest Parakeet CPU setup on the same hardware
On 8 physical cores of an AMD EPYC 9575F (Zen 5), parakeet-redux with Photon reaches 113x realtime — 2.5x the fastest other Parakeet CPU runtime measured. parakeet.cpp at q8_0 manages 45x, sherpa-onnx 42x, onnx-asr 28x. On an Apple M2 it also leads: 38x on CPU and 43x on GPU, versus 38x for parakeet.cpp's Metal path and 37x for fp32 parakeet-mlx. A ternary-weight model outpaces natively-tuned runtimes on a MacBook Air.
Accuracy: multilingual and long-form wins, noise is the weak flank
On the seven English test sets of the Open ASR Leaderboard, Redux averages 6.55 WER versus 6.26 for the original — within 0.3, with AMI and Earnings-22 slightly ahead. The counterintuitive result is 25-language FLEURS: 10.56 average, beating the original's 11.62 — Estonian 9.15 vs 13.23, Latvian 12.80 vs 21.38, Slovene 16.21 vs 21.76. TED-LIUM long-form (10-20 minute talks) also flips: 2.51 vs 2.71. The weak flank is noise: 9.04 average across nine MUSAN conditions, well behind the original's 6.72. The model card's own explanation: the ternary encoder has a thinner acoustic margin, substituting similar-sounding words more often at low SNR.
So what
For teams building local voice interfaces or privacy-sensitive transcription, this is a clear cost path: a 178MB model plus one runtime gives realtime transcription on laptops and edge boxes, with no cloud round-trip. Two caveats: Moondream ships it under a custom license — read the terms before commercial use — and the noise regression means ternary quantization is not free; stress-test with your own real audio. The de-GPU-ification of ASR is becoming its own lane: ternary quantization, rewritten runtimes, stripped Python dependencies — different methods, one direction: moving inference cost from the cloud to the endpoint.