Seven hobbyist boards now run a 0.5B-parameter language model end to end on about 1.53 watts. That is the headline of ESP32s3-LLM-Cluster, an MIT-licensed project by developer Low-Zi-Hong that pulled 149 points and more than 30 comments on Hacker News this week, with firmware and tooling fully public.
How a 0.5B model fits on microcontrollers
The base is Qwen2-0.5B: a 24-layer Transformer with hidden size 896, an MLP intermediate dimension of 4,864, and grouped-query attention with 14 Q-heads over 2 KV-heads. The cluster is seven ESP32-S3 boards: one master runs the BPE tokenizer, the INT4 token embedding (about 14MB in flash), the final RMSNorm and the LM head; the other six compute nodes each carry 4 layers, covering all 24. Nodes are chained over SPI in a daisy chain — each board has two SPI channels, one receiving the FP32 hidden state from the previous node and one sending it to the next, with dedicated reset and ready lines between master and nodes.
Quantization is BitNet-style 1.58-bit ternary weights: linear layers hold only -1, 0 and 1, packed at 4 weights per byte, while embeddings go INT4. The vocabulary is cropped from 151K to 32K — the maximum the ESP32-S3's 16MB flash can carry. That works out to roughly 3.82MB per layer and 15.3MB per node, just inside the flash partition. For speed, the author hand-wrote assembly MAC kernels for the ternary matmuls and added lookup tables.
The costs: 1.3 seconds per node, near-random output
The other side of the ledger is documented just as clearly. The workflow guide reports each node takes about 1.3 seconds per inference step, growing linearly as boards are added — 100 boards could run a 400-layer model, at proportionally worse latency. Power figures are author-measured: about 1.17W idle at 5V 0.23A, and 1.53W while inferring.
More candid still is the training note: the QAT fine-tuning script only partially trains, with loss stuck around 8.0, "just spitting out random tokens". The troubleshooting guide even has an entry for output stuck repeating the same word — a known behavior of a heavily under-trained model plus greedy sampling, not necessarily a code bug. In other words, the project currently proves that a 1.58-bit model can complete a full forward pass across a microcontroller cluster, not that it can hold a conversation.
HN pours cold water
The sharpest comments converge on one point: this is not a path to practical edge inference. One argues the bottleneck for LLM inference remains memory bandwidth — "one big GPU with twice the VRAM will always perform significantly better than two GPUs with half the VRAM each" — and that at GDDR7 speeds signals travel only about 10mm per clock cycle, so spreading a system across dozens of small chips wastes most of its resources on data movement. Another questions whether daisy-chained SPI can scale far at all. Some commenters dream of massively parallel RISC-V clusters, but nobody disputes the economics.
So what
The value here is not a benchmark. It is that the full chain — extreme quantization, model slicing, inter-board pipelining — has been walked end to end on seven development boards, and every piece (vocabulary pruning, ternary weight packing, assembly MAC kernels, KV cache in PSRAM) is reusable engineering. Whoever reruns the QAT with real compute could turn the same skeleton from "random tokens" into "actual speech". For edge-deployment teams, this open implementation is more concrete than most paper appendices.