[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-apple-lensvlm-9b-weights-huggingface":3,"topics-all":38,"news-related-0d357a0b-42da-40af-8038-e8c035cb9810":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"0d357a0b-42da-40af-8038-e8c035cb9810","Apple 开源 LensVLM-9B:先扫压缩图,再读原页","Apple 把 5 月的 LensVLM 论文落成可用权重:9B 视觉语言模型上架 Hugging Face,基座 Qwen3.5-9B。它把文档渲染成 5-15 倍压缩页面图,模型先扫缩略图定位相关页,再调工具取回全文,长文档问答不必把所有文本塞进上下文。配套代码同步开源,社区已跟进 9 个量化版本。","9 月 21 日,Apple 把今年 5 月发表的论文 [LensVLM: Selective Context Expansion for Compressed Visual Representation of Text](https:\u002F\u002Farxiv.org\u002Fabs\u002F2605.07019) 正式落成了可用工件:9B 参数的视觉语言模型 LensVLM-9B 上架 [Hugging Face](https:\u002F\u002Fhuggingface.co\u002Fapple\u002FLensVLM-9B),官方代码仓库 [ml-lensvlm](https:\u002F\u002Fgithub.com\u002Fapple-aiml-research\u002Fml-lensvlm) 同步开放。论文与权重之间隔了四个多月,这次的新闻不是\"又发了一个模型\",而是一条研究路线兑现了可复现的交付物。\n\n## 它解决什么问题\n\n长文档问答的标准做法是把全文塞进上下文窗口,代价随页数线性上涨。LensVLM 换了一个思路:把整份文档渲染成压缩后的页面图像——提供 5x、10x、15x 三档压缩——模型先\"扫\"这些低分辨率缩略图,判断哪几页与问题相关,再通过一个学出来的 read_page 工具,只把相关页面按原始分辨率展开读取。\n\n官方 demo 给了一条完整轨迹:输入 15 页压缩图像,问题问\"在电影 Kiss and Tell 中饰演 Corliss Archer 的女演员后来担任过什么政府职务\"。模型在 \u003Cthink> 推理里判断第 10 页提到\"美国女演员、歌手\",调用 read_page(10) 取回全文,串起 Shirley Temple → 美国首席礼宾官(Chief of Protocol of the United States),两轮内闭环。\n\n## 规格与实现\n\n模型是 dense 结构,9B 参数,BF16 精度,基座是 Qwen\u002FQwen3.5-9B-Base——大厂把强开源文本骨干适配成多模态,而不是从零训练,这已经是通行模式。论文作者阵容 10 人(Roy Xie、Dan Friedman、Donghan Yu 等)。\n\n评测管线也一并开源:从 HotpotQA、Natural Questions、Musique 构造带干扰段的长文档评测集,推理后用 LLM-as-judge 打分;论文里的 judge 用的是 Qwen3.5-397B-A17B-FP8,跑在 8 卡 B200 节点上。也就是说,复现这条研究路线不需要 Apple 私有设施。\n\n## 生态与许可\n\nHugging Face 页面显示,模型上架后最近一个月下载 1,432 次,社区已跟进 9 个量化版本,2 个 Spaces 在用。注意许可条款:模型权重走 Apple Machine Learning Research Model License,代码走 Apple Sample Code License——都不是 Apache\u002FMIT 级别的宽松许可,商用前需要读条款。这是 Apple 一贯的研究开源姿态:工件可复现,边界画得清楚。\n\n## 值得留意的点\n\n把\"压缩视觉表示 + 选择性展开\"这条路线和当前两股潮流对照着看更有意思:一边是上下文窗口军备竞赛,百万 token 成为旗舰标配;另一边是 KV cache 压缩、上下文蒸馏等\"省着用\"的工程派。LensVLM 属于后者,但它把省上下文的动作从推理层挪到了输入表示层——文档根本不以全文形式进入模型,而是以图像缩略图的形式进入,再按需取回。对于 PDF、扫描件这类本来就是视觉形态的输入,这条路绕开了 OCR 依赖。\n\n对开发者的\"所以呢\":如果你的场景是几百页文档里的偶发查询,先扫缩略图、再精读少量页面的检索策略,可能比无脑开百万上下文更便宜。9B 的体量单卡可跑,量化版本已经就位,门槛不高——但先确认研究许可覆盖你的用途。","https:\u002F\u002Fhuggingface.co\u002Fapple\u002FLensVLM-9B","a2e6145a-2a88-4c51-8d09-c4375b2a833b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"2d9c2fb0-2be5-4ad1-aedb-e9747addf355","compression",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":19,"name":20,"slug":20,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"5647897c-25ae-471c-bec7-a9e08b1a839f","en","Apple Open-Sources LensVLM-9B: Scan Thumbnails, Read Pages","Apple ships LensVLM as weights: a 9B VLM on Qwen3.5-9B, live on Hugging Face. It scans compressed page thumbnails, then expands relevant pages only.","On September 21, Apple turned its May paper \"[LensVLM: Selective Context Expansion for Compressed Visual Representation of Text](https:\u002F\u002Farxiv.org\u002Fabs\u002F2605.07019)\" into a usable artifact: the 9B-parameter vision-language model LensVLM-9B is now live on [Hugging Face](https:\u002F\u002Fhuggingface.co\u002Fapple\u002FLensVLM-9B), with the official code repository [ml-lensvlm](https:\u002F\u002Fgithub.com\u002Fapple-aiml-research\u002Fml-lensvlm) released alongside it. Four-plus months separate the paper from the weights — the story here is not \"another model dropped\" but a research line delivering a reproducible artifact.\n\n## The Problem It Targets\n\nThe standard approach to long-document QA is stuffing the full text into the context window, with costs climbing linearly in pages. LensVLM takes a different route: render the entire document as compressed page images — at 5x, 10x, or 15x compression — let the model \"scan\" those low-resolution thumbnails first, decide which pages matter for the question, and then use a learned read_page tool to expand only the relevant pages back to full resolution.\n\nThe official demo shows a full trajectory: given 15 compressed page images and the question \"What government position was held by the woman who portrayed Corliss Archer in the film Kiss and Tell?\", the model reasons inside \u003Cthink> tags that Page 10 mentions \"American actress, singer\", calls read_page(10) to retrieve the full text, and chains Shirley Temple → Chief of Protocol of the United States in two turns.\n\n## Specs and Implementation\n\nThe model is a dense 9B-parameter model in BF16, built on Qwen\u002FQwen3.5-9B-Base — the now-standard pattern of adapting a strong open text backbone for multimodal use rather than training from scratch. The paper lists ten authors (Roy Xie, Dan Friedman, Donghan Yu, and colleagues).\n\nThe evaluation pipeline ships too: long-document eval sets built from HotpotQA, Natural Questions, and Musique with distractor augmentation, scored by an LLM-as-judge after inference. The paper's judge is Qwen3.5-397B-A17B-FP8 served on an 8x B200 node — meaning this research line can be reproduced without any Apple-internal infrastructure.\n\n## Ecosystem and Licensing\n\nThe Hugging Face page shows 1,432 downloads in the past month since listing, with 9 community quantizations already available and 2 Spaces using the model. Note the licensing: model weights fall under the Apple Machine Learning Research Model License, while the code uses the Apple Sample Code License — neither is an Apache\u002FMIT-style permissive license, so read the terms before commercial use. This is Apple's consistent research-open posture: reproducible artifacts, clearly drawn boundaries.\n\n## Why It Matters\n\nIt is more interesting to read the \"compressed visual representation + selective expansion\" line against the two current trends: on one side, the context-window arms race with million-token flagships; on the other, the engineering school of KV-cache compression and context distillation. LensVLM belongs to the latter camp, but it moves the context-saving action from the inference layer to the input-representation layer — the document never enters the model as full text at all. It enters as image thumbnails, retrieved on demand. For PDFs and scanned documents that are natively visual anyway, this path sidesteps OCR entirely.\n\nThe \"so what\" for developers: if your workload is occasional queries over hundred-page documents, a scan-thumbnails-first, read-few-pages-on-demand retrieval policy may be cheaper than reflexively opening a million-token context. A 9B model runs on a single GPU, quantized variants are already up, and the barrier is low — just make sure the research license covers your use case.","apple-lensvlm-9b-weights-huggingface","2026-09-26T15:20:00Z","2026-09-26T15:12:10.511284Z","2026-09-26T15:12:10.511298Z",true,"agent",370,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"36c71e4e-e398-4583-8710-1732aefff06a","OneStreamer:4B 流式模型先记再答,八榜最佳","onestreamer-4b-streaming-video-memory","2026-10-02T15:07:48+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"bbe8d55a-1069-42fa-a342-d944d53fdb4b","操作电脑的 27B 开源权重模型 Holo4:最强版禁商用","holo4-open-weight-computer-use","2026-10-02T13:12:01+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"dc61debd-1d77-4d5a-9d59-5b23c3da07de","蚂蚁开源Realtime-Venus：9B全双工模型边说边干活，三项续聊指标超GPT-4o","ant-realtime-venus-full-duplex-delegation","2026-09-30T23:10:53+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"d056f67b-7e0d-4e44-8d39-e31ea50deeae","Bonsai 2 27B 三元压缩:Qwen3.8 压到 5.9 GB,benchmark 留存 98.2%","bonsai-2-27b-ternary-qwen3-8-compression","2026-09-17T15:47:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"d055ddb8-4d82-4523-99b7-39c5f77e2ff7","PhysBrain 1.5 开源：8B 具身基座 28 项评测均分 72.5，官方称追平 GPT-6-Astra","physbrain-1-5-open-embodied-base","2026-09-16T21:07:24+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"51c13e24-8072-404c-a8d4-75c40cff05ee","Ling-3.0-flash-VL 开源：124B MoE 只激活 5.5B，视觉塞进 Agent 闭环","ling-3-0-flash-vl-open-weights","2026-09-15T13:18:00+00:00"]