[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-magnitude-self-tuning-local-agent-inference-engine":3,"topics-all":38,"news-related-9a77e0d7-8164-494e-984d-54cb57a7a0dd":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"9a77e0d7-8164-494e-984d-54cb57a7a0dd","Magnitude 开源:本地 Agent 专用推理引擎","YC S25 团队开源 Magnitude 推理引擎:内核在你自己的设备上编译调优,专为本机 Agent 会话设计。官方自报对比 llama.cpp,Metal 上解码快 92%、单 Agent 内存省 28%,Apache 2.0,Rust 编写。数字全部为厂商自测,留一份谨慎。","本地跑 Agent 的开发者长期面对一个结构性的尴尬:vLLM、SGLang 这类引擎为数据中心批处理而生,单会话性能是被牺牲的那一头;llama.cpp、Ollama 走通用兼容路线,不为任何具体硬件压榨极限;oMLX、ds4 这类专精引擎又缺生态完整性。换句话说,几乎没有引擎认真为「本机同时挂几个长会话 Agent」这个场景做过设计。9 月 30 日,一支 YC S25 团队在 Hacker News 发布了开源推理引擎 Magnitude,思路是把性能调优这件事搬到你自己的设备上做。\n\n## 它做了什么不同\n\nMagnitude 用 Rust 写成,自带 GPU kernel 运行时与自动调优器,以 Apache 2.0 协议开源。官方给出的技术路线有三条:第一,内核带弹性参数,模型运行前先在你的实际设备上编译调优,让通用内核摸到专用内核的性能天花板;第二,内存动态分配,启动只预留权重所需,Agent 会话变长内存堆才增长,会话结束即释放,机器同时还能干别的;第三,混合分页注意力,借鉴 SGLang 的 radix cache 思路让并发会话共享前缀缓存,同时按内存邻接优化排布,避免单会话性能被并发拖垮。README 列出的技术灵感包括 FlashAttention、FlashInfer、TurboQuant。\n\n## 官方自报的数字\n\n统一配置为 Qwen 3.6 35B A3B(4-bit 量化)、64k 上下文、关闭投机解码,对比对象 llama.cpp:Mac M4 Pro 48GB 上解码快 92%(30→57 tok\u002Fs),预填快 9%(466→507 tok\u002Fs),单 Agent 内存省 28%;DGX Spark 上解码快 19%(49→58 tok\u002Fs),预填快 23%(2033→2507 tok\u002Fs),内存省 27%。产品形态是桌面应用,一键接入 Pi、OpenCode、Hermes、Codex、Claude Code 等客户端,其余工具走 OpenAI 兼容接口;Agent 需要时自动拉起模型,闲置后自动关闭。发布约一天,HN 帖 187 分 88 评,仓库 star 约 6k。\n\n## 自报数字要打折听\n\n值得肯定的是切入角度:单会话延迟与多 Agent 并存,是本地推理的真实痛点——数据中心引擎不屑于优化,通用引擎的目标函数又不是这个。但全部跑分都是厂商自测:单一模型、单一量化、关闭投机解码,恰好是自家最优的配置区间;「快 2 倍」的宣传语取的是 Metal 解码的最好数字,CUDA 侧只有 19%。第三方复现出现之前,更稳妥的读法是把它当成一个有诚意的方向验证,而不是已定论的胜负。路线图里的专家流式加载(把 MoE 专家放进内存\u002F磁盘、按需搬进 GPU)值得盯——那才是小显存跑大模型的钥匙。\n\n对每天在本地挂 Agent 的开发者,这多了一个免费且开源的选项;对推理引擎格局,「Agent 时代的本机引擎」这个生态位正在被快速填补。问题留给读者:当 Agent 真的常驻你的本机,你现在的推理引擎,是为这个场景优化的吗?\n\n参考:官方 HN 发布帖 https:\u002F\u002Fnews.ycombinator.com\u002Fitem?id=49911995 ;GitHub 仓库 https:\u002F\u002Fgithub.com\u002Fmagnitudedev\u002Fmagnitude","https:\u002F\u002Fgithub.com\u002Fmagnitudedev\u002Fmagnitude","998df6db-96e6-4b8e-8be1-cfa00a6cd177",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"ed492f63-eb22-45ab-9dec-3fdb49ef6a0b","en","Magnitude: Self-Tuning Inference Engine for Local Agents","YC S25's Magnitude is an open-source Rust engine tuning kernels on-device for local agents; vendor tests: 92% faster decode than llama.cpp on Metal.","Developers running agents locally have long faced a structural gap. Engines like vLLM and SGLang are built for datacenter batch inference, sacrificing single-session performance; llama.cpp and Ollama pursue broad compatibility instead of squeezing the maximum out of any specific chip; specialized engines like oMLX and ds4 lack ecosystem completeness. In other words, almost no engine has been seriously designed for the scenario of running several long agent sessions on your own machine at once. On September 30, a YC S25 team released Magnitude, an open-source inference engine, on Hacker News. Its idea: move performance tuning onto your own device.\n\n## What it does differently\n\nMagnitude is written in Rust, ships its own GPU kernel runtime and autotuner, and is open source under Apache 2.0. The team lays out three technical moves. First, kernels carry flexible parameters and are compiled and tuned on your actual device before a model runs, letting generic kernels reach the performance ceiling of hardware-specific ones. Second, memory allocation is dynamic: at startup only enough memory for model weights is reserved; the heap grows as agent sessions grow and frees itself when agents stop, so the machine stays usable for other work. Third, hybrid paged attention borrows the radix-cache idea from SGLang so concurrent sessions share prefix caches, while placement is optimized for memory adjacency so single-session performance does not collapse under concurrency. The README lists FlashAttention, FlashInfer, and TurboQuant among its inspirations.\n\n## The vendor-reported numbers\n\nAgainst llama.cpp, with a unified configuration of Qwen 3.6 35B A3B (4-bit), 64k context, and speculative decoding disabled: on a Mac M4 Pro 48GB, decode is 92% faster (30 to 57 tok\u002Fs), prefill 9% faster (466 to 507 tok\u002Fs), and per-agent memory 28% lower; on a DGX Spark, decode is 19% faster (49 to 58 tok\u002Fs), prefill 23% faster (2033 to 2507 tok\u002Fs), and memory 27% lower. The product ships as a desktop app that connects Pi, OpenCode, Hermes, Codex, Claude Code and other clients with one click, with an OpenAI-compatible API for everything else; models spin up on demand when agents need them and shut down after inactivity. About a day after launch, the HN post sat at 187 points with 88 comments, and the repository at roughly 6k stars.\n\n## Discount the self-reported numbers accordingly\n\nThe angle deserves credit: single-session latency and multiple co-resident agents are real pain points of local inference — datacenter engines do not bother optimizing for them, and generalist engines are not targeting that objective. But every benchmark here is vendor-tested: a single model, a single quantization, speculative decoding off — exactly the configuration interval that favors the vendor. The up-to-2x tagline takes the best Metal decode number, while the CUDA side is 19%. Until third-party replication appears, the safer reading is a well-intentioned directional validation, not a settled verdict. The roadmap's expert streaming — parking MoE experts in RAM or on disk and loading them into the GPU just in time — is worth watching; that is the real key to running large models on small VRAM.\n\nFor developers who keep agents running locally, this adds a free and open-source option; for the inference-engine landscape, the local-engine-for-the-agent-era niche is being filled fast. A question to leave with you: when agents truly live on your machine, is your current inference engine actually optimized for that?\n\nReferences: official launch post https:\u002F\u002Fnews.ycombinator.com\u002Fitem?id=49911995 ; GitHub repository https:\u002F\u002Fgithub.com\u002Fmagnitudedev\u002Fmagnitude","magnitude-self-tuning-local-agent-inference-engine","2026-10-01T17:08:35Z","2026-10-01T17:08:44.816417Z","2026-10-01T17:08:44.816427Z",true,"agent",1041,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"1edd86ea-3eb0-4424-9f57-add15c08d891","HF Hub 接入 SkyPilot：Xet 去重让 20+ 云共享模型数据","hf-skypilot-hf-storage","2026-07-12T06:30:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"f76e6f4b-1ae5-442a-abfa-823e7226ae81","Hugging Face 把 transformers 跑出 vLLM 原生速度：单 flag 让 235B MoE 直接吃到 EP 红利","hf-vllm-transformers-backend","2026-07-09T04:03:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"4d436945-18e9-4d69-a4c8-c1e3e975ab33","MiniMax M3发布：稀疏注意力打通百万token上下文，开源模型编程能力逼近闭源前沿","minimax-m3-sparse-attn-million-token-msa","2026-06-04T01:00:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"86c3465b-27f0-4f8c-83d1-9a8e1fa171c8","Compute Aligned Training：让模型学会协同推理的新训练范式","compute-aligned-training-pass-n-vote-michigan","2026-05-22T13:10:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"068bcc7a-d901-4d3d-924c-eefa7ced4467","TokenSpeed 开源推理引擎发布：剑指 Agentic Workloads 的高效推理","tokenspeed-agentic-inference-engine","2026-05-20T19:10:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"530e5aa5-2026-4c30-b4e6-421caca907b2","Transformer 提前罢工:13 个基座模型跟不住引用链,一个 rank-8 LoRA 修好","tiny-lora-frozen-transformer-chain-relay","2026-10-03T21:05:16+00:00"]