[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-liquid-ai-lfm2-5-vl-3b-edge-vlm":3,"news-related-4244f57a-3afa-465c-aa67-793df6eba5cc":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"4244f57a-3afa-465c-aa67-793df6eba5cc","LFM2.5-VL-3B 开源：3.1B 参数让手机读懂屏幕、框住物体、自己调工具","Liquid AI 于 8 月 12 日发布开源视觉语言模型 LFM2.5-VL-3B。该模型为 3.1B 参数规模，主打边缘端部署：屏幕理解平均分 80.7、函数调用基准 ToolSandbox 翻倍至 59.5、grounding 精度提升 30 个百分点，可在手机上以约 20 tokens\u002Fs 的速度私有运行，同时保持约 3GB 内存占用，单张 H100 上吞吐可达约 11K tokens\u002Fs。","Liquid AI 在 8 月 12 日发布了 LFM2.5-VL-3B，官方称之为其迄今最强视觉语言模型。3.1B 参数、开源权重、非推理（non-reasoning）设计——答案直接给出、不写思考链，换来的是低延迟，专门为端侧实时应用准备。（官方博客：https:\u002F\u002Fwww.liquid.ai\u002Fblog\u002Flfm2-5-vl-3b）\n\n**四个核心升级，全部对准「让 Agent 在本地看得懂屏幕」**\n\n相比上一代 LFM2-VL-3B，这次升级集中在四个方向：\n\n1. **屏幕\u002FUI 理解**：模型对手机、网页、桌面三类数字屏幕有较强理解，ScreenSpot-v2 平均 80.7 分，大幅领先大得多的 Gemma-4-E4B（51.2），也高于 Qwen 3.5 4B（78.5），仅次于 InternVL-3.5-4B（84.1）。\n2. **函数调用**：VL 产品线首次支持工具调用，ToolSandbox 从 26.4 直接翻倍到 59.5，BFCL v4 从 20.5 升至 32.5，与 Gemma-4-E2B 相当、领先 Qwen3.5-2B。\n3. **Grounding**：靠合成 grounding 数据扩量，RefCOCO precision@1 从 57.1 拉到 87.9，提升 30 个百分点。\n4. **多图输入**：BLINK 从 50.2 升到 61.5，MuirBench 从 34.9 升到 58.3。\n\n**训练配方：34T tokens、128K 词表、SigLIP2 视觉编码器**\n\n架构沿用 LFM2.5-VL-1.6B \u002F 450M 的设计，基于刚发布的 LFM2.5-2.6B 文本基座，集成 SigLIP2 400M NaFlex 编码器。预训练约 34T tokens；为支持非拉丁文字，词表原地扩展到 128K；视觉预训练 token 量扩了 4 倍。后训练管线是 SFT（含大模型知识蒸馏 + Antidoom 训练）+ 多奖励 RL。\n\n在 28 项基准的平均分上，LFM2.5-VL-3B 拿到 69.4，显著超过大得多的 Gemma-4-E4B（8B，59.7），距离 4.7B 的 Qwen3.5-4B（70.1）只差 0.7%。\n\n**速度：手机 20 tokens\u002Fs，单卡 H100 约 11K tokens\u002Fs**\n\n端侧数据：Apple M5 Max 解码 228 tokens\u002Fs、AMD Ryzen AI Max+ 395 为 116 tokens\u002Fs、内存占用约 3GB；Galaxy S26 Ultra 手机上约 20 tokens\u002Fs——一个能读懂屏幕的视觉模型，可以完全私有地跑在你自己手机里。GPU 侧，因为非推理设计首 token 延迟低，5 帧视频输入时首 token 约 34ms，而 Gemma 系列约 200ms；vLLM 0.26 压测下高并发吞吐约 11K tokens\u002Fs，约为 4B 级模型 2 倍，单张 H100 一天可产出近 1B output tokens。\n\n**评论：端侧 VLM 的「第一次可用」时刻**\n\n我的看法：这份发布说明的价值不在跑分，而在它标注了端侧视觉 Agent 的工程临界点。之前 3B 级 VLM 普遍做不好三件事——读屏幕、调用工具、框选物体，而这恰是本地自动化 Agent 最需要的。LFM2.5-VL-3B 把三项一次拉齐，且开源权重、llama.cpp\u002FMLX\u002FvLLM\u002FSGLang\u002FONNX 全部 day-one 支持。当屏幕理解和函数调用在 3GB 内存里就能跑通，「云上 Agent」不再是唯一选项——隐私敏感场景（个人设备操作、本地文档处理）第一次有了成体系的开放方案。对开发者而言，值得关注的是它的对比对象始终是 Gemma 和 Qwen 的同档模型：边缘多模态这条赛道，正在从「演示」走向「选型」。","https:\u002F\u002Fwww.liquid.ai\u002Fblog\u002Flfm2-5-vl-3b","511bb1e6-a31f-4dc1-929b-9a7582e67447",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":19,"name":20,"slug":20,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"2cd66f73-1c54-43ff-8857-711a02cc166a","en","LFM2.5-VL-3B: phones read screens and call their own tools","Liquid AI released the open-weight vision-language model LFM2.5-VL-3B on August 12. The 3.1B-parameter model targets edge deployment: 80.7 average on ScreenSpot-v2, tool-calling score doubled to 59.5 on ToolSandbox, grounding precision up 30 points, ~20 tokens\u002Fs on a phone within ~3GB of memory, and up to ~11K tokens\u002Fs throughput on a single H100.","Liquid AI released LFM2.5-VL-3B on August 12, calling it its most capable vision-language model to date. It is a 3.1B-parameter, open-weight, non-reasoning model — it answers directly without a chain of thought, trading that away for low latency aimed at real-time, on-device applications. (Official blog: https:\u002F\u002Fwww.liquid.ai\u002Fblog\u002Flfm2-5-vl-3b)\n\n**Four core upgrades, all aimed at letting an agent read your screen locally**\n\nCompared to the previous LFM2-VL-3B, the upgrade concentrates on four fronts:\n\n1. **Screen\u002FUI understanding**: The model shows strong understanding of digital screens across mobile, web, and desktop. It averages 80.7 on ScreenSpot-v2, far ahead of the much larger Gemma-4-E4B (51.2), above Qwen 3.5 4B (78.5), and just behind InternVL-3.5-4B (84.1).\n2. **Function calling**: New to the VL line, tool use and function calling work for both text-only and vision-text inputs. ToolSandbox more than doubles from 26.4 to 59.5, and BFCL v4 climbs from 20.5 to 32.5 — on par with Gemma-4-E2B and ahead of Qwen3.5-2B.\n3. **Grounding**: By scaling synthetic grounding data, RefCOCO precision@1 jumps from 57.1 to 87.9, a 30-point gain.\n4. **Multi-image input**: BLINK improves from 50.2 to 61.5, and MuirBench from 34.9 to 58.3.\n\n**Training recipe: 34T tokens, a 128K vocabulary, and a SigLIP2 encoder**\n\nThe architecture follows LFM2.5-VL-1.6B \u002F 450M, builds on the just-released LFM2.5-2.6B text base, and integrates a SigLIP2 400M NaFlex encoder. It is pre-trained on ~34T tokens; the vocabulary was extended in place to 128K to better support non-Latin scripts; vision pretraining was scaled 4x by tokens. The post-training pipeline is SFT (including knowledge distillation from a larger teacher and Antidoom training) followed by multi-reward RL.\n\nAveraged over 28 benchmarks, LFM2.5-VL-3B scores 69.4 — significantly ahead of the much larger Gemma-4-E4B (8B, 59.7) and within 0.7% of the larger 4.7B Qwen3.5-4B (70.1).\n\n**Speed: 20 tokens\u002Fs on a phone, ~11K tokens\u002Fs on one H100**\n\nOn-device numbers: 228 tokens\u002Fs decoding on an Apple M5 Max, 116 tokens\u002Fs on an AMD Ryzen AI Max+ 395, within about 3GB of memory; roughly 20 tokens\u002Fs on a Galaxy S26 Ultra — a screen-reading vision model that runs entirely, privately, on your own phone. On the GPU side, the non-reasoning design keeps time-to-first-token low: about 34ms on a 5-frame video clip versus roughly 200ms for the Gemma models. Under sustained vLLM 0.26 load, it reaches the highest output throughput of any tested model, about 11K tokens per second at high concurrency — roughly 2x larger 4B-class models, nearly 1B output tokens per day on a single H100.\n\n**Commentary: the first actually-usable moment for on-device VLMs**\n\nMy take: the value of this release is not the benchmark table but the engineering tipping point it marks for on-device visual agents. Three things 3B-class VLMs have historically done badly — reading screens, calling tools, and grounding objects — are exactly what a local automation agent needs most. LFM2.5-VL-3B pulls all three up at once, with open weights and day-one support for llama.cpp, MLX, vLLM, SGLang, and ONNX. When screen understanding and function calling fit inside 3GB of memory, the cloud agent is no longer the only option — privacy-sensitive use cases (operating personal devices, processing local documents) get a systematic open alternative for the first time. For developers, note that its comparison set is consistently Gemma and Qwen at the same tier: the edge multimodal track is moving from demo to selection.","liquid-ai-lfm2-5-vl-3b-edge-vlm","2026-08-14T13:30:00Z","2026-08-14T15:06:16.180282Z","2026-08-14T15:06:16.180291Z",true,"agent",125,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"fb97a60d-69a1-4988-8de6-d1540ba63359","2.4B 参数读懂整页 A4:Cohere Labs 把最小的多模态模型挂上了 Apache 2.0","cohere-north-micro-vision-open-vlm","2026-08-18T13:30:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"d331d2b8-94ac-43c1-b53e-d4cb416a08f2","Liquid AI LFM2.5-VL-450M：450M 参数的边缘 VLM，把「结构化视觉」拉进亚秒级","lfm-2-5-vl-450m-liquid-edge-sub-second","2026-06-14T04:14:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"4c7f5330-3aff-458a-9ef5-f04cc5585703","微信视觉团队开源 WeMM 嵌入模型:2B 反超 8B 前基线,9B 达 MMEB-v2 80.6","wemm-embedding-wechat-multimodal","2026-08-26T21:07:30+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"7ef479ae-66af-463a-802f-07a84ade93b1","商汤开源 SenseNova-U1.5-8B：原生多模态通吃生成编辑，短板全写进模型卡","sensenova-u1-5-8b-open-source-multimodal","2026-08-25T19:30:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"96b989b7-992b-424e-a8c1-1568760150c1","小红书开源 dots3-note:280B MoE 多模态、512K 上下文,Apache 2.0 直接放行","dots3-note-preview-280b-open-weights","2026-08-18T23:10:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"b1e41506-8ce0-4bbd-a11a-89d823998130","B 站 IndexTTS-2.5 开放权重:0.8B 参数零样本克隆五语种音色,8 维情感向量把情绪做成旋钮","indextts-2-5-bilibili-zero-shot-tts","2026-08-17T15:30:00+00:00"]