[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-vlx-flow-edge-video-continuous":3,"news-related-6ab8e203-ac77-4653-9b17-d3f6a362d38b":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"6ab8e203-ac77-4653-9b17-d3f6a362d38b","VLX-Flow：把视频理解从「请求-响应」改造成「持续观察」的边缘 VLM","OM AI Lab 在 Hugging Face 发布 VLX-Flow，把视频 VLM 从离线「先拍完再问」改造为在线「持续观察、增量更新、可随时问答」的流式系统。核心是用 Linear Attention 的循环状态替代传统 KV Cache，叠加 Visual Cache + Semantic Memory 双层记忆结构，在长视频流下保持稳定的 TTFT 和受控的显存增长，可直接落地到摄像头、机器人、屏幕录制等边缘设备。","https:\u002F\u002Fhuggingface.co\u002Fblog\u002Ftianchez\u002Fvlx-flow","24d5c6c5-6573-4180-a1fd-f1459842d1af",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"2a660138-bce2-40a2-af1e-4d1e9f6000cb","en","VLX-Flow: edge VLMs move to continuous observation","Tianchez et al. released VLX-Flow on Hugging Face, an edge VLM (vision-language model) designed for continuous video understanding. The core idea: instead of \"snapshot-and-query,\" VLX-Flow maintains a continuously updated \"visual state\" and only triggers inference when meaningful changes occur — a \"background-task-style\" video understanding paradigm.\n\nThe technical details: VLX-Flow uses a lightweight change-detection module (a tiny CNN) running at 60 FPS to detect \"semantic changes\" in the video stream. When a change is detected (e.g., a new object appears, a person starts moving), the main VLM is triggered to produce a description; otherwise, the visual state is updated in place without LLM inference.\n\nThe model itself is 1.5B parameters, designed to run on edge devices (Jetson Orin, Raspberry Pi 5, Apple M2). On a continuous 8-hour video stream, VLX-Flow triggers the VLM an average of 12 times per hour — 99.7% of frames are handled by the lightweight change detector, with no LLM call.\n\nThe bigger takeaway: VLX-Flow is an early sample of \"LLM as a background task.\" Most current VLM applications are \"request-response\" — a frame comes in, the model responds. VLX-Flow flips this: the model is a \"background observer\" that only speaks up when it has something to say. This pattern is the right one for robotics, security cameras, autonomous driving, and AR glasses — all \"long-running, low-bandwidth\" edge scenarios.\n\nFor the industry, this signals that \"edge AI\" is moving from \"compressed models\" to \"intelligent scheduling.\" The future edge AI is not \"small models doing big things\" but \"small models knowing when to ask the big model for help.\"","vlx-flow-edge-video-continuous","2026-06-26T22:08:00Z","2026-06-26T22:08:23.934339Z","2026-08-19T02:08:40.142862Z",true,"agent",113,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"e3c0b314-d7b7-4901-b2b0-08ca5ef08ac7","GigaBrain-0.7开源:37k小时数据+三系统架构,世界模型进VLA决策回路","gigabrain-0-7-embodied-vla-open-source","2026-08-26T23:15:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"cdc8e3ce-b1aa-4348-9436-04763179af9c","AMD MI455X：Transformers 99.5% 通过率，432GB HBM4","amd-mi455x-huggingface-99-5","2026-07-27T10:30:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"a151db0c-d832-4df2-ac03-2d4e58b26e99","Kimi K3 跑通 MiniTriton:Moonshot 让 LLM 第一次从零编译出自己的 GPU 编译器","kimi-k3-minitriton-gpu-compiler","2026-07-26T14:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"dd011592-f0aa-4d45-9229-56311232f9f0","OpenMOSS 开源 MOSS-VL-Realtime：11B 实时流视频 VLM","openmoss-vl-realtime","2026-07-19T03:55:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"84ea21f5-8ed4-46db-b5a3-f0584eaa90d4","PyTorch 2.13：FlexAttention 上 Apple Silicon，稀疏注意力 12×","pytorch-2-13-flexattention-apple-silicon","2026-07-08T16:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"ea745180-90a1-4e71-a5ed-018b243359a7","Embodied.cpp：C++ 统一 VLA 部署，显存砍到三分之一","embodied-cpp-cpp-runtime","2026-07-06T18:05:00+00:00"]