[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-macaw-lfm25-15gb-edge-mac-agent":3,"news-related-988bbfb8-672a-4c6c-98f7-3a170b6bd8b3":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"988bbfb8-672a-4c6c-98f7-3a170b6bd8b3","Macaw 把 LFM2.5 装进 1.5GB:4-bit 端侧 LLM 跑 Mac 控制工具链","Macaw 把 LFM2.5-2.6B 压成 1.5GB 4-bit 量化权重,跑在 Apple Silicon 上做 macOS 本地 Agent,97 个工具、10\u002F10 调用准确率、MIT 开源。","Bad Theory Labs 把 Liquid AI 的 LFM2.5-2.6B 重新包装成了一个 1.5 GB 的 Mac 本地 Agent —— 名为 Macaw,模型权重走 MLX 4-bit 量化,97 个 macOS 工具(邮件、日历、文件、音乐、系统设置)直接由本地模型调度,无需联网、无需账号,在 M2 上测出 10\u002F10 工具调用准确率、平均 1.21 秒响应、40.3 tok\u002Fs 解码速度。代码 MIT 开源,权重挂在 Hugging Face 上,只要 Apple Silicon、macOS 14+、8 GB 内存就能跑起来。\n\n这件事真正值得讨论的不是又一个\"端侧 AI 助手\",而是它揭示的三个工程现实。\n\n## 第一,2.6B 的小模型已经够用\n\nLFM2.5 是 Liquid AI 在 2025 年下半年发布的混合架构小模型,主打\"端侧推理 + 长上下文 + 多模态\"组合。Macaw 在它之上做的事很克制:不微调权重,只注入系统提示把模型身份改成\"Mac 控制助手\",其余能力完全靠 prompt 引导出来的工具调用。这条路线暗示,2026 年的端侧 LLM 不需要再做大规模 SFT,只要底层基座的指令对齐做得够干净,3B 以下的模型就能承担真实世界的工具编排任务。\n\n## 第二,4-bit 量化 + MLX 是当前 Apple Silicon 上最稳的本地推理组合\n\nBF16 原权重大约 5 GB,在 8 GB 内存的 Mac mini 上几乎要逼到 swap;4-bit MLX 量化版压到 1.5 GB,留给系统和其他 app 充足余量。Bad Theory Labs 把 BF16 和 MLX 两套权重都放出来了,MLX 跑 Metal GPU、BF16 走 Transformers 后端兼容任意设备,这是 2026 年\"消费级硬件能跑 AGI 级体验\"叙事的最小可工作单元。\n\n## 第三,工具调用准确率从\"大致能用\"跨到了\"真能用\"\n\nBenchLM 类的 AI 评测早就告诉我们,SWE-Bench、τ-Bench 这类 Agent 榜单上,前沿模型(Claude 4.x、Gemini 3.x、GPT-5.x)和 7B-30B 开源模型之间还差着 8 个月左右的距离。但 Macaw 这种\"封闭环境 + 严格工具集\"的场景里,2.6B 模型居然能拿到 10\u002F10。原因是 macOS 工具集是确定性的 API,不需要模型理解动态网页和反爬机制;Hark Handoff 那种 CUA(Computer Use Agent)才是真正考验通用 Agent 能力的高难度任务。所以\"端侧 LLM 已经能干 Siri 干不了的事\"这个论断成立,但前提是**把任务边界画清楚**。\n\n## 授权结构与商业化逻辑\n\n回到 Macaw 本身,它最大的限制也写在官网:\"LFM Open License v1.0 把商用门槛卡在年营收 1000 万美元以下\"。也就是说个人开发者、小团队、独立产品可以放心用,但如果拿它做商业产品的底座,撞到营收红线就要重新谈授权。这个授权结构其实很聪明 —— Liquid AI 用这种方式把 LFM 系列变成了\"个人开发者默认底座\",但保留了大客户单独谈判的窗口。\n\n## 给中文 AI 圈留下的启发\n\nMacaw 给中文 AI 圈留下的最大启发是什么?**\"边缘 AI\"这个词在 2026 年下半年不再已经是一种营销话术,而是已经能交付的工具**。1.5 GB 权重、零联网、MIT 代码、装上就能用的本地 Agent,这意味着任何想搭一个不依赖云端 API 的小工具的独立开发者,门槛已经低到几乎为零。接下来的问题是:Apple Intelligence 什么时候才会把 Siri 真正升级到这个水平?如果答案是\"永远不会\",那么 Macaw 这类项目反而会吃掉 Siri 的市场份额。\n\n最后留一个反问:Macaw 现在 97 个工具,看起来够用,但用户一定会问\"为什么不支持我的工作流\" —— 当个人定制化成为标配,模型本身反而不是壁垒,工具链和系统集成的深度才是。这条赛道里,开源 + 本地 + 极致轻量,会比\"更大的模型\"更容易长期胜出。","https:\u002F\u002Fwww.badtheorylabs.com\u002Fmacaw","f3142574-d026-4167-b767-333da980de35",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":19,"name":20,"slug":20,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b49648f9-963e-4082-8684-3d085b7358fe","quantization",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"662d9695-7d9d-4a37-b0d8-f704536ab3a1","en","Macaw packs LFM2.5 into 1.5GB: a 4-bit edge LLM running the Mac toolchain","Macaw compresses LFM2.5-2.6B into a 1.5GB 4-bit quantized weight running on Apple Silicon as a local macOS Agent — 97 tools, 10\u002F10 call accuracy, MIT open source.","Bad Theory Labs has repackaged Liquid AI's LFM2.5-2.6B into a 1.5 GB local Mac Agent called Macaw. The model weights ship as a 4-bit MLX quantization, driving 97 macOS tools (mail, calendar, files, music, system settings) directly from the local model. No network needed, no account required. On an M2, it hit 10\u002F10 tool-call accuracy, a mean 1.21-second response, and 40.3 tokens\u002Fsecond decoding. The code is MIT-licensed, weights live on Hugging Face, and any Apple Silicon machine on macOS 14+ with 8 GB RAM can run it.\n\nWhat is worth discussing here isn't \"yet another on-device AI assistant\" — it is three engineering realities this release exposes.\n\n## First, a 2.6B model is already enough\n\nLFM2.5 is the hybrid-architecture small model Liquid AI shipped in the second half of 2025, combining on-device inference, long context, and multimodality. What Macaw does on top of it is restrained: no weight fine-tuning, just a system prompt that turns the model's identity into a \"Mac control assistant.\" All other capabilities emerge from prompt-guided tool calling. This trajectory suggests that on-device LLMs in 2026 do not need another round of heavy SFT — as long as the base model's instruction alignment is clean enough, sub-3B models can carry real-world tool orchestration.\n\n## Second, 4-bit quantization + MLX is the most stable local-inference stack on Apple Silicon today\n\nBF16 raw weights run around 5 GB, which on an 8 GB Mac mini nearly forces swap. The 4-bit MLX quantization compresses that to 1.5 GB, leaving comfortable headroom for the OS and other apps. Bad Theory Labs released both BF16 and MLX weights — MLX runs on Metal GPU, BF16 falls back to any Transformers backend. This is the smallest working unit behind the 2026 narrative that \"consumer hardware can deliver AGI-grade experience.\"\n\n## Third, tool-call accuracy has crossed from \"kinda works\" to \"actually works\"\n\nAI benchmarks like BenchLM have long told us that on SWE-Bench and τ-Bench-style Agent leaderboards, frontier models (Claude 4.x, Gemini 3.x, GPT-5.x) still sit roughly 8 months ahead of 7B-30B open-source models. But in Macaw's closed-environment, strict-tool-set scenario, a 2.6B model hits 10\u002F10. The reason is that macOS tools are deterministic APIs — the model never has to read dynamic web content or dodge anti-bot systems. The real stress test for general Agent capability is CUA-style products like Hark Handoff, which force the model to navigate hostile live websites. So the claim \"on-device LLMs can already do what Siri cannot\" holds — but only when **the task boundary is drawn tightly.**\n\n## Licensing and the commercialization logic\n\nMacaw's biggest constraint is spelled out on its own site: the LFM Open License v1.0 caps commercial use at $10 million in annual revenue. Independent developers, small teams, and hobbyist products can use it freely, but anyone building a commercial product on top will have to renegotiate the license once they cross that revenue threshold. The structure is actually clever — Liquid AI turns the LFM family into the \"default base for individual developers\" while preserving a separate negotiation window for large customers.\n\n## What this leaves for the Chinese AI ecosystem\n\nWhat is the biggest takeaway Macaw leaves for the Chinese AI scene? **The phrase \"edge AI\" stopped being a marketing slogan in late 2026 and became a deliverable tool.** A 1.5 GB weight, zero network, MIT code, install-and-run local Agent — this means that for any indie developer who wants to build a small tool that doesn't depend on cloud APIs, the barrier has dropped to nearly zero. The open question now: when will Apple Intelligence actually upgrade Siri to this level? If the answer is \"never,\" projects like Macaw will eat Siri's market share from below.\n\nLast counter-question: Macaw ships 97 tools today, which sounds enough — yet users will inevitably ask \"why doesn't it support mine?\" Once personal customization becomes the default expectation, the model itself stops being the moat; tooling depth and OS-integration craft are. In this track, open-source + local + ultra-light will beat \"bigger models\" over the long run.","macaw-lfm25-15gb-edge-mac-agent","2026-08-24T06:00:00Z","2026-08-24T11:04:16.350335Z","2026-08-24T11:04:16.350344Z",true,"agent",38,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"d5203c93-2022-4769-a6d2-c7765ded2b40","腾讯混元 Hunyuan-A13B 开源实测:80B 总参 \u002F 13B 激活,GQA + FP8\u002FINT4 把 MoE 推理门槛打到消费卡","tencent-hunyuan-a13b-80b-13b-gqa-angelslim-moe","2026-07-30T06:00:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"1abe59b8-844d-4bb1-bc27-cfe28099d101","Anemll\u002FFlash-iOS：把 400B MoE 大模型塞进 iPhone 的端侧实验","anemll-flash-ios-400b-moe-iphone-edge","2026-06-07T12:00:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"c2e18ecb-00b3-47ee-935a-f0e3cd0dee4a","Holo3.1 把 Computer Use Agent 拉进本地：FP8\u002FNVFP4\u002FQ4 三种量化让消费级 GPU 跑得动","holo3-1-h-company-computer-use-quant","2026-06-07T06:00:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"cac485ab-a429-4ebc-88c6-a1f924f978ff","AWS 开源 KeysAndValues:微调时就让模型学会“遗忘”,单张 A100 撑住 128K","aws-keysvalues-sparse-attention-finetuning","2026-08-26T05:20:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"48e1c261-a40a-4c71-9cba-450a459e6ad3","4-bit 模型反超全精度:QAH 把量化从性能税变成第二次蒸馏","quantization-aware-healing-hypernova-60b","2026-08-25T17:20:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"c94bdf86-5de9-49fe-8c98-0f5c47611bfe","SGLang v0.5.18 发布:大模型冷启动提速 2.38 倍,710 个 PR 都改了什么","sglang-v0-5-18-cold-start-2-38x","2026-08-24T23:15:00+00:00"]