[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-esp32s3-bitnet-llm-cluster":3,"topics-all":38,"news-related-3e28cfaf-74aa-4521-8eba-37332fe93901":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"3e28cfaf-74aa-4521-8eba-37332fe93901","7块ESP32跑1.58-bit Qwen:功耗1.53瓦","开发者 Low-Zi-Hong 把 Qwen2-0.5B 量化到 1.58-bit,切成 24 层、用 SPI 菊花链把 7 块 ESP32-S3 单片机串成流水线推理集群,每块板扛 4 层,整簇功耗约 1.53 瓦。作者自曝 QAT 只做了部分训练、loss 约 8.0,模型输出仍接近随机。","用 7 块十几块钱的开发板,把一个 0.5B 参数的语言模型真正\"跑\"起来,整簇功耗只要约 1.53 瓦——这是开发者 Low-Zi-Hong 开源的 ESP32s3-LLM-Cluster 交出的答案。项目在 Hacker News 拿下 149 点热度、30 多条讨论,MIT 协议,固件与工具链全部公开。\n\n## 怎么把 0.5B 模型塞进单片机\n\n项目基座是 Qwen2-0.5B:24 层 Transformer,hidden size 896,MLP 中间维 4864,注意力是 14 个 Q 头配 2 个 KV 头的 GQA。整个集群由 7 块 ESP32-S3 组成:1 块 master 负责 BPE 分词、token embedding(INT4 量化,约 14MB 放 flash)、最后的 RMSNorm 和输出头;剩下 6 块计算节点每块扛 4 层,恰好分完 24 层。节点间用 SPI 菊花链串联,每块板两个 SPI 通道,一个收上一节点传来的隐向量(FP32),一个发给下一节点,主从之间还有专门的复位与就绪信号线。\n\n量化方案是 BitNet 式 1.58-bit 三值权重——线性层权重只有 -1、0、1 三种取值,4 个权重打包进 1 个字节;embedding 单独走 INT4。词表从原始 151K 砍到 32K,因为 32K 已是 ESP32-S3 的 16MB flash 能装下的上限。算下来每层约 3.82MB,每块计算节点约 15.3MB,正好卡在 flash 分区内。为了榨速度,作者给三值矩阵乘写了汇编级 MAC 优化,并配上查找表。\n\n## 代价:每节点 1.3 秒,输出接近随机\n\n这份账本的另一面写得很清楚。workflow 文档自报:每个节点推理耗时约 1.3 秒,且随节点数线性增长——100 块板理论上能跑 400 层的模型,但延迟同步涨上去。功耗数字同样来自作者实测:待机 5V 0.23A 约 1.17W,推理时约 1.53W。\n\n更诚实的是训练侧的自曝:QAT 微调脚本只做了部分训练,loss 停在 8.0 左右,\"基本就是在吐随机 token\"。排错指南里专门写了一条:输出卡在重复同一个词,是欠训练模型叠加贪心采样的已知行为,不一定是代码 bug。换句话说,这个项目目前证明的是 1.58-bit 模型可以在单片机集群上完整跑完一遍前向计算,而不是\"能对话\"。\n\n## HN 的两盆冷水\n\n评论区最硬的两条批评指向同一件事:这不是通往实用边缘推理的路。一条指出 LLM 推理的瓶颈始终是内存带宽,\"一块两倍显存的大 GPU 永远显著快过两块各一半显存的 GPU\",GDDR7 时代信号一个时钟周期只能走约 10mm,把系统摊到几十块小芯片上,大部分资源都浪费在数据搬运上。另一条直接点名菊花链 SPI 的扩展性存疑。也有人畅想 RISC-V 大规模并行集群,但没人反驳\"经济上不划算\"这个判断。\n\n## 所以呢\n\n这个项目的价值不在跑分,而在把\"极端量化 + 模型切片 + 板间流水线\"这条完整链路在 7 块开发板上走通了,每一环——词表裁剪、三值权重打包、汇编 MAC、KV cache 放 PSRAM——都是可复用的工程积木。等有人用充足算力把 QAT 做完整,同一套骨架就能从\"吐随机 token\"变成\"能说话\"。对做边缘部署的团队,这份开源实现比大多数论文附录实在。","https:\u002F\u002Fgithub.com\u002FLow-Zi-Hong\u002FESP32s3-LLM-Cluster","998df6db-96e6-4b8e-8be1-cfa00a6cd177",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":19,"name":20,"slug":20,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b49648f9-963e-4082-8684-3d085b7358fe","quantization",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"8bbde89b-f970-48a5-b608-fde472878fbb","en","7 ESP32 boards run a 1.58-bit Qwen on 1.53 watts","Seven ESP32-S3 boards chained over SPI run a 1.58-bit Qwen2-0.5B cluster on 1.53 watts; the author admits the partly trained model still outputs random tokens.","Seven hobbyist boards now run a 0.5B-parameter language model end to end on about 1.53 watts. That is the headline of ESP32s3-LLM-Cluster, an MIT-licensed project by developer Low-Zi-Hong that pulled 149 points and more than 30 comments on Hacker News this week, with firmware and tooling fully public.\n\n## How a 0.5B model fits on microcontrollers\n\nThe base is Qwen2-0.5B: a 24-layer Transformer with hidden size 896, an MLP intermediate dimension of 4,864, and grouped-query attention with 14 Q-heads over 2 KV-heads. The cluster is seven ESP32-S3 boards: one master runs the BPE tokenizer, the INT4 token embedding (about 14MB in flash), the final RMSNorm and the LM head; the other six compute nodes each carry 4 layers, covering all 24. Nodes are chained over SPI in a daisy chain — each board has two SPI channels, one receiving the FP32 hidden state from the previous node and one sending it to the next, with dedicated reset and ready lines between master and nodes.\n\nQuantization is BitNet-style 1.58-bit ternary weights: linear layers hold only -1, 0 and 1, packed at 4 weights per byte, while embeddings go INT4. The vocabulary is cropped from 151K to 32K — the maximum the ESP32-S3's 16MB flash can carry. That works out to roughly 3.82MB per layer and 15.3MB per node, just inside the flash partition. For speed, the author hand-wrote assembly MAC kernels for the ternary matmuls and added lookup tables.\n\n## The costs: 1.3 seconds per node, near-random output\n\nThe other side of the ledger is documented just as clearly. The workflow guide reports each node takes about 1.3 seconds per inference step, growing linearly as boards are added — 100 boards could run a 400-layer model, at proportionally worse latency. Power figures are author-measured: about 1.17W idle at 5V 0.23A, and 1.53W while inferring.\n\nMore candid still is the training note: the QAT fine-tuning script only partially trains, with loss stuck around 8.0, \"just spitting out random tokens\". The troubleshooting guide even has an entry for output stuck repeating the same word — a known behavior of a heavily under-trained model plus greedy sampling, not necessarily a code bug. In other words, the project currently proves that a 1.58-bit model can complete a full forward pass across a microcontroller cluster, not that it can hold a conversation.\n\n## HN pours cold water\n\nThe sharpest comments converge on one point: this is not a path to practical edge inference. One argues the bottleneck for LLM inference remains memory bandwidth — \"one big GPU with twice the VRAM will always perform significantly better than two GPUs with half the VRAM each\" — and that at GDDR7 speeds signals travel only about 10mm per clock cycle, so spreading a system across dozens of small chips wastes most of its resources on data movement. Another questions whether daisy-chained SPI can scale far at all. Some commenters dream of massively parallel RISC-V clusters, but nobody disputes the economics.\n\n## So what\n\nThe value here is not a benchmark. It is that the full chain — extreme quantization, model slicing, inter-board pipelining — has been walked end to end on seven development boards, and every piece (vocabulary pruning, ternary weight packing, assembly MAC kernels, KV cache in PSRAM) is reusable engineering. Whoever reruns the QAT with real compute could turn the same skeleton from \"random tokens\" into \"actual speech\". For edge-deployment teams, this open implementation is more concrete than most paper appendices.","esp32s3-bitnet-llm-cluster","2026-09-30T13:11:01Z","2026-09-30T13:11:15.058824Z","2026-09-30T13:11:15.058835Z",true,"agent",295,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"6d60f9ba-4866-4520-9f12-d955e37f8472","Gated DeltaNet 全压 4-bit 没掉点:一篇论文拆掉混合 LLM 的量化禁忌","gated-deltanet-nvfp4-full-4bit","2026-09-04T15:08:02+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"48e1c261-a40a-4c71-9cba-450a459e6ad3","4-bit 模型反超全精度:QAH 把量化从性能税变成第二次蒸馏","quantization-aware-healing-hypernova-60b","2026-08-25T17:20:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"b2d34fba-93ee-469b-9f64-6a7568f89745","Liquid AI 用量化感知蒸馏,把 LFM2.5 4-bit 精度拉回 97%","lfm25-qad-quantization-aware-distillation-edge","2026-08-20T11:00:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"c197245e-0a8a-4028-9ad2-6547bdc01be5","BitNet 团队把 1.58-bit 量化推进到 embedding:检索向量也可以\"训练时就压\"","bitnet-1-58bit-embedding","2026-07-18T03:00:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"93981299-c75e-4140-ad2c-ae68fb5d27ce","CAT-Q：512 样本把 235B LLM 压到 1.58-bit，成本降 10 万倍","cat-q-1-58-bit-512-samples-icml-oral","2026-06-26T20:25:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"f3b52d77-bfad-4289-adab-777d19a79bcb","Parakeet Redux把ASR压到178MB:纯CPU跑出113倍实时","parakeet-redux-ternary-cpu-asr","2026-10-01T21:08:38+00:00"]