[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-lut-llm-fpga-lut-6x-energy":3,"topics-all":36,"news-related-4ef890c9-f223-46ca-b230-d0968dfc7306":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"4ef890c9-f223-46ca-b230-d0968dfc7306","LUT-LLM：FPGA上用「查表」替代计算，让大模型推理能效提升6倍","大模型推理的能效瓶颈，正在催生一种另辟蹊径的硬件方案。\n\n传统上，LLM推理依赖GPU的算术运算完成矩阵乘法等核心操作。但FPGA的强项并非算力，而是大量分布式片上存储与灵活的数据通路控制能力——这一点在传统算术范式下被白白浪费了。\n\nFCCM 2026会议上提出的LUT-LLM，首次在FPGA上实现以内存查表（memory-based computation）替代算术运算来运行十亿参数级语言模型。核心思路：将LLM的矩阵运算结果量化到离散编码表，推理时按索引查表取结果，大幅减少乘加操作。该方案采用激活-权重联合量化，最小化量化误差。\n\n工程层面有三个关键优化：带宽感知的并行质心搜索降低解码延迟；高效二维表查找减少表访问开销；时空混合架构削减数据缓存次数，提升吞吐量。\n\n在AMD V80 FPGA上以Qwen 3 1.7B实测，算术操作减少4倍，生成速度提升1.1~3.3倍，能效达到GPU的3~6.6倍。对于追求低功耗、低延迟推理的场景，这一方向值得关注。不过，当前方案针对特定量化策略做了深度适配，迁移至其他模型家族尚需进一步泛化研究。\n\nLLM推理的硬件竞争正从「拼算力」走向「拼架构」。查表取代计算，或许只是开始。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2511.06174","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b49648f9-963e-4082-8684-3d085b7358fe","quantization",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"eef59265-49ba-445d-a7f5-462c11c9ccc2","en","LUT-LLM: FPGA lookup tables replace compute, 6x efficiency","The energy-efficiency bottleneck of large-model inference is spawning an alternative hardware approach.\n\nTraditionally, LLM inference relies on GPUs' arithmetic operations to perform core operations like matrix multiplication. But FPGA's strength isn't raw compute — it's a large amount of distributed on-chip storage and flexible datapath control, capability that is wasted under the traditional arithmetic paradigm.\n\nLUT-LLM, presented at the FCCM 2026 conference, is the first to implement memory-based computation on FPGA to replace arithmetic operations for running billion-parameter language models. The core idea: quantize the LLM's matrix-operation results into discrete encoding tables, and look up results by index at inference, sharply reducing multiply-accumulate operations. The scheme adopts activation-weight joint quantization to minimize quantization error.\n\nOn the engineering side there are three key optimizations: bandwidth-aware parallel centroid search reduces decoding latency; efficient 2D table lookup cuts table-access overhead; a space-time hybrid architecture reduces data cache passes and improves throughput.\n\nBenchmarks on AMD V80 FPGA with Qwen 3 1.7B show 4× fewer arithmetic operations, 1.1-3.3× faster generation, and 3-6.6× the energy efficiency of GPU. For scenarios pursuing low-power, low-latency inference, this direction is worth attention. The current scheme, however, is deeply adapted to a specific quantization strategy, and generalization to other model families still requires further research.\n\nThe hardware competition for LLM inference is shifting from \"compete on compute\" to \"compete on architecture.\" Lookup tables replacing computation may be just the beginning.","lut-llm-fpga-lut-6x-energy","2026-05-20T02:00:00Z","2026-05-20T10:05:31.031636Z","2026-08-19T02:08:40.142862Z",true,"agent",185,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"3cecce90-70b9-4bb3-b9b7-93e6b0c05105","D-Quant 用熵编码压 KV:2.26bit 近无损","d-quant-entropy-coding-kv-cache","2026-09-20T17:10:42+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"178aa5e5-2a4f-4a87-a97c-0da16295d96f","EMNLP 2026 OCGQuant:用通道配对治 NVFP4 陪葬误差,Qwen3-1.7B 接近 FP16","ocgquant-nvfp4-outlier-companion-grouping","2026-09-10T09:15:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"4147f71b-eaa7-4d39-91cf-c2c572105e7f","FlashPrefill V2:128K 长文本 prefill 提速 47 倍,块稀疏注意力走进生产框架","flashprefill-v2-block-sparse-prefill","2026-08-21T19:10:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"c32d3160-4e07-4128-890f-4e135aac2cce","CompactifAI 把 Llama 3.3 70B 砍到一半:Multiverse 在 Intel Xeon 6 上跑出 1.9 倍吞吐","compactifai-llama-3-3-70b-intel-xeon","2026-07-26T04:00:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"cf7f8e7e-4efc-451e-9595-706b0be911ba","PolyQ:把\"3-bit LLM 跑在 CPU\"做成一件可预测的事","polyq-3bit-llm-cpu","2026-07-17T10:00:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"007a83c4-faed-459a-ab44-17b915e05fb5","TriRoute 把 MoE + MoD + KV 量化做成一个控制器：三条条件计算路径第一次协同","triroute-moe-mod-kv-quantization","2026-07-09T18:02:00+00:00"]