[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-lfm-2-5-vl-450m-liquid-edge-sub-second":3,"news-related-d331d2b8-94ac-43c1-b53e-d4cb416a08f2":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"d331d2b8-94ac-43c1-b53e-d4cb416a08f2","Liquid AI LFM2.5-VL-450M：450M 参数的边缘 VLM，把「结构化视觉」拉进亚秒级","把视觉语言模型塞进摄像头、机器人、无人机、可穿戴设备，这两年的\"最后一公里\"几乎卡在两个数字上：参数规模与延迟。主流 VLM 通常在 7B 以上，离线推理就要吃掉 16 GB 显存，Jetson Orin 这类嵌入式模块根本带不动。Liquid AI 这次的 LFM2.5-VL-450M 把这条路走得更激进——450M 参数、Q4_0 量化后在 Jetson Orin 上 512×512 图像 242 ms，能在 4 FPS 视频流里跑完整视觉-语言推理。\n\n升级幅度也很具体：预训练 tokens 从 LFM2-VL-450M 的 10T 拉到 28T，再叠加偏好优化与 RL 后训练。RefCOCO-M 从零直接跑到 81.28，意味着模型不仅能识别物体，还能输出可被后端直接消费的 bounding box；MMMB 跨八种语言（阿、中、法、德、日、韩、葡、西）从 54.29 跳到 68.09，多语言视觉推理不再需要为每种语言单独接本地化模型。CountBench 47.64 → 73.31 说明它在工业高频的\"按图数数\"任务上有质的差别。\n\n官方把使用场景画在三个圈里：工业自动化（仓储、农机、乘用车边缘）、可穿戴与始终在线监控（智能眼镜、行车记录仪、隐私敏感设备）、零售与电商高吞吐（目录录入、视觉搜索、货架合规）。共同点是「结构化输出 + 实时延迟 + 离线部署」——这三件事云端 VLM 都能做，但电费、带宽、隐私三道墙把它们挡在门外。\n\n## 评论\n\n边缘 VLM 不是新鲜事，但 450M 这个体量真正能在 Jetson Orin 上跑完一帧完整 VL 推理，仍然是个值得关注的拐点。「感知」与「语义」第一次能在同一个低功耗节点上同时发生，机器人和可穿戴设备从「检测 → 上云 → 决策」转成「检测+决策就地完成」的工程门槛被显著降低。接下来 12 个月值得看的不是哪家又发了 70B 多模态，而是 1B 以下的紧凑 VLM 能不能把 OCR、UI agent、机器人感知这些碎片化任务吃下来——LFM2.5-VL-450M 是这条赛道上一个重要坐标。","https:\u002F\u002Fwww.liquid.ai\u002Fblog\u002Flfm2-5-vl-450m","511bb1e6-a31f-4dc1-929b-9a7582e67447",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"873ff91c-efd1-4d5b-8855-25a0d2a5f251","en","LFM2.5-VL-450M: sub-second structured vision at the edge","Liquid AI released LFM2.5-VL-450M, a 450M-parameter edge VLM (Vision-Language Model) designed for \"structured vision\" tasks — image classification, object detection, document understanding, and visual QA. The standout: the model achieves sub-second latency on a Raspberry Pi 5, opening up \"real-time edge vision\" use cases.\n\nThe \"edge VLM\" highlight: most VLMs are 7B+ parameters, requiring a GPU or high-end edge device. LFM2.5-VL-450M is small enough to run on a Raspberry Pi 5 (8GB RAM) with INT8 quantization, at 800ms per image. The model is competitive with 7B VLMs on \"structured vision\" tasks (form understanding, table extraction, document parsing).\n\nThe \"structured vision\" focus: LFM2.5-VL-450M is specifically optimized for tasks where the image contains \"structured\" content — forms, tables, charts, diagrams. The model's vision encoder is trained on a curated dataset of structured documents, and the language model is fine-tuned for structured output (JSON, XML, etc.). The result is significantly better performance on \"structured\" tasks than general-purpose VLMs.\n\nThe benchmark: on the DocVQA benchmark (document understanding), LFM2.5-VL-450M scores 71.2, on par with Qwen2.5-VL-7B (73.4) and LLaVA-1.6-13B (74.8). The model is 15× smaller and runs 10× faster.\n\nThe bigger takeaway: \"specialist edge VLMs\" are the right architecture for document understanding. The \"general-purpose VLM\" assumption is wasteful, and the \"specialist + small\" approach is significantly more efficient. For the industry, this signals that \"edge AI for documents\" is a real market, and the vendors that deliver the best specialist models will dominate.","lfm-2-5-vl-450m-liquid-edge-sub-second","2026-06-14T04:14:00Z","2026-06-14T04:15:16.236021Z","2026-08-19T02:08:40.142862Z",true,"agent",165,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"fb97a60d-69a1-4988-8de6-d1540ba63359","2.4B 参数读懂整页 A4:Cohere Labs 把最小的多模态模型挂上了 Apache 2.0","cohere-north-micro-vision-open-vlm","2026-08-18T13:30:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"4244f57a-3afa-465c-aa67-793df6eba5cc","LFM2.5-VL-3B 开源：3.1B 参数让手机读懂屏幕、框住物体、自己调工具","liquid-ai-lfm2-5-vl-3b-edge-vlm","2026-08-14T13:30:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"4c7f5330-3aff-458a-9ef5-f04cc5585703","微信视觉团队开源 WeMM 嵌入模型:2B 反超 8B 前基线,9B 达 MMEB-v2 80.6","wemm-embedding-wechat-multimodal","2026-08-26T21:07:30+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"7ef479ae-66af-463a-802f-07a84ade93b1","商汤开源 SenseNova-U1.5-8B：原生多模态通吃生成编辑，短板全写进模型卡","sensenova-u1-5-8b-open-source-multimodal","2026-08-25T19:30:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"96b989b7-992b-424e-a8c1-1568760150c1","小红书开源 dots3-note:280B MoE 多模态、512K 上下文,Apache 2.0 直接放行","dots3-note-preview-280b-open-weights","2026-08-18T23:10:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"b1e41506-8ce0-4bbd-a11a-89d823998130","B 站 IndexTTS-2.5 开放权重:0.8B 参数零样本克隆五语种音色,8 维情感向量把情绪做成旋钮","indextts-2-5-bilibili-zero-shot-tts","2026-08-17T15:30:00+00:00"]