[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-basert-metal-kernel":3,"topics-all":36,"news-related-2f4423e8-d5f9-448d-b0da-a16cf7d627a7":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"2f4423e8-d5f9-448d-b0da-a16cf7d627a7","BaseRT：Apple Silicon LLM 推理第一，llama.cpp 1.56×","过去一年，「Apple Silicon 跑 LLM」几乎等同于 MLX 与 llama.cpp 两条路线。行业默认：要么走 llama.cpp 的成熟生态，要么走 MLX 借助 unified memory 的 Python 友好设计，剩下的空间似乎已被充分挖掘。basecompute 团队 7 月 1 日挂上 arXiv 的 BaseRT（2607.00501），直接把这个前提翻了个底朝天。\n\nBaseRT 的方法论不复杂、也不取巧：放弃跨平台抽象层，完全在 Metal API 上手写针对 M 系列芯片特性的 kernel fusion、统一内存感知调度、以及一套自定义的 dispatch 路径。换句话说，它把 llama.cpp 跨硬件抽象带来的固定开销、以及 MLX 把 Metal 嵌进 Python 解释器带来的运行时开销，一并烧掉了。\n\n数字上，M3 \u002F M4 Pro 上跑 Qwen3、Llama 3.2、Gemma 4（Q4 \u002F Q8 量化），BaseRT 的 decode 吞吐比 llama.cpp 高 1.56×、比 MLX 高 1.35×。更值得注意的是 MoE 模型的 prefill：架构抽象一旦被吃掉，权重 dispatch 与 expert 路由的固定成本被放大，论文里这部分加速比 decode 还要夸张。它支持 Q2 至 FP16 共 8 种量化格式，覆盖从 sub-1B 到 30B 全谱系模型，代码已在 GitHub 开源（github.com\u002Fbasecompute\u002FbaseRT）。\n\n这件事的意义不只在性能。它直接打掉了一个行业惯性——「M 系列做严肃 LLM 推理只是玩具」。随着隐私要求、延迟约束和云端成本压力把推理推向端侧，端上 LLM 不再是 demo，而是承担生产负载的备选层。BaseRT 用 1.56× 的数字钉死了这个转折点，也意味着接下来一两年，跨平台框架的「硬件抽象税」会被越来越多人重新估量。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.00501","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"e0d31e94-ce47-4c8f-831c-d3d2926d42f3","hardware",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"eb603380-1819-44b1-be30-6dd6bdd42b56","en","BaseRT: #1 LLM inference on Apple Silicon, llama.cpp at 1.56x","Over the past year, \"Apple Silicon running LLM\" has been almost synonymous with the two paths of MLX and llama.cpp. Industry default: either take llama.cpp's mature ecosystem, or take MLX's Python-friendly design leveraging unified memory, the remaining space seems already fully mined. The BaseRT (2607.00501) that basecompute's team posted to arXiv on July 1 directly flipped this premise. BaseRT's methodology is not complex nor clever: abandon the cross-platform abstraction layer, completely hand-write kernel fusion targeting M-series chip characteristics, unified memory-aware scheduling, and a custom dispatch path on the Metal API. In other words, it burns away the fixed overhead llama.cpp brings from cross-hardware abstraction, and the runtime overhead MLX brings from embedding Metal into the Python interpreter. On the numbers, with Qwen3, Llama 3.2, Gemma 4 (Q4\u002FQ8 quantization) running on M3\u002FM4 Pro, BaseRT's decode throughput is 1.56× higher than llama.cpp, 1.35× higher than MLX. More noteworthy is the prefill of MoE models: once architectural abstraction is eaten away, the fixed cost of weight dispatch and expert routing is amplified, the paper says this part of the speedup is even more dramatic than decode. It supports 8 quantization formats from Q2 to FP16, covering the full spectrum from sub-1B to 30B models, with code open-sourced on GitHub (github.com\u002Fbasecompute\u002FbaseRT). The significance of this isn't just in performance. It directly takes down an industry inertia — \"M-series doing serious LLM inference is just a toy\". As privacy requirements, latency constraints, and cloud cost pressure push inference to the device side, on-device LLM is no longer a demo, but an alternative layer taking on production workloads. BaseRT uses the 1.56× number to nail down this turning point, and also means that in the next year or two, the \"hardware abstraction tax\" of cross-platform frameworks will be reassessed by more and more people.","basert-metal-kernel","2026-07-03T20:05:00Z","2026-07-03T20:12:55.127504Z","2026-08-19T02:08:40.142862Z",true,"agent",297,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"183fb3be-e062-47e7-9591-7c2372e116c1","LLM 蒸馏的显存瓶颈不只在教师模型：离线 Top-K 与分块 KL 把长上下文训练装回单卡","llm-distillation-offline-top-k-chunked-kl","2026-08-05T20:08:13+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"cdc8e3ce-b1aa-4348-9436-04763179af9c","AMD MI455X：Transformers 99.5% 通过率，432GB HBM4","amd-mi455x-huggingface-99-5","2026-07-27T10:30:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"beec1ff3-22af-4657-b58a-90cb0797c3b1","PyroDash 让小模型「借力」大模型推理：把 LLM 调用砍到 1.9%，成本从 $49 降到 $1.78","pyrodash-small-large-routing","2026-07-24T00:00:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"7f723663-8405-43aa-b31e-73efd714fa97","KV-Cache Grafting：冻结权重，Gemma-4-12B AIME 80%→93.3%","byte-exact-kv-cache-grafting","2026-07-17T06:20:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"ed39ed38-b5fa-4f58-92cf-d05233ab998b","Speculate with Memory：LLM Agent 无损加速 2.5×，准确率涨 39pp","speculate-with-memory-llm-agent-acceleration","2026-07-15T08:15:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"3af7d9f7-9cb3-43a3-a338-00d716c8053e","JoLT 用 Tucker + JL 残差把 KV 缓存压到 1\u002F3：让长上下文 LLM 推理不再被显存卡脖子","jolt-tucker-jl-kv-cache","2026-07-15T02:18:00+00:00"]