[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-minicpm-v-4-6-1-3b-phone-multimodal":3,"topics-all":36,"news-related-4f70d0fe-f0cb-443e-a4a4-27b9eb004441":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"4f70d0fe-f0cb-443e-a4a4-27b9eb004441","1.3B参数多模态模型直接跑在手机上：MiniCPM-V 4.6开源，13亿参数覆盖iOS\u002F安卓\u002F鸿蒙","你能想象一个13亿参数的多模态大模型直接在iPhone上运行吗？OpenBMB最新发布的MiniCPM-V 4.6做到了。\n\n这款于5月11日开源的模型仅有13亿参数，却能处理单图、多图和视频理解任务，在消费级手机上流畅运行——涵盖iOS、Android和鸿蒙系统。基于Apache 2.0许可证开源，并原生支持vLLM、SGLang等主流推理框架。\n\n技术层面，MiniCPM-V 4.6采用SigLIP2-400M视觉编码器与Qwen3.5-0.8B语言基座的组合架构，支持高达262K token的上下文窗口。团队通过视觉编码器内部的前期压缩机制，将计算量降低了50%以上，同时提供4倍和16倍两档压缩率选项。在Artificial Analysis评测中，该模型得分13，在同规模开源模型里位列第三，显著领先中位数。\n\n更关键的是效率表现：与Qwen3.5-0.8B相比，MiniCPM-V 4.6的端到端吞吐提升约1.5倍，而成本降低19倍；即便对比带推理思考的Qwen3.5-0.8B变体，成本优势也达到43倍。量化版本更是将内存需求压至3GB GPU显存或约2GB CPU内存。\n\n这并不是在挑战GPT-4或Gemini的位置。1.3B参数、多模态、262K上下文、视频理解、本地运行——这些能力以往需要更大参数量才能实现，但现在已经在普通手机的算力范围内。对隐私敏感的应用、离线助手或文档理解等场景，这个发布意味着边缘端AI的可行性边界已经向前推进了一大步。\n\n模型已在Hugging Face开源，提供8个量化变体和本地部署参考代码。","https:\u002F\u002Fhuggingface.co\u002Fopenbmb\u002FMiniCPM-V-4.6","24d5c6c5-6573-4180-a1fd-f1459842d1af",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"467ca6ed-eb8c-4ebe-9d0b-31a2f523750b","en","MiniCPM-V 4.6: a 1.3B multimodal model that runs on phones","Can you imagine a 1.3B-parameter multimodal large model running directly on an iPhone? OpenBMB's latest release, MiniCPM-V 4.6, makes it happen.\n\nThis model, open-sourced on May 11, has only 1.3B parameters, yet can handle single-image, multi-image, and video understanding tasks, running smoothly on consumer-grade phones — covering iOS, Android, and HarmonyOS. It's open-sourced under the Apache 2.0 license, and natively supports mainstream inference frameworks like vLLM and SGLang.\n\nTechnically, MiniCPM-V 4.6 adopts a combined architecture of SigLIP2-400M vision encoder and Qwen3.5-0.8B language base, supporting context windows up to 262K tokens. Through an early-stage compression mechanism inside the vision encoder, the team reduced compute by over 50%, while providing 4× and 16× compression rate options. On Artificial Analysis benchmarks, the model scores 13, ranking third among open-source models of the same size, well ahead of the median.\n\nThe efficiency performance is more critical: compared to Qwen3.5-0.8B, MiniCPM-V 4.6's end-to-end throughput improves by about 1.5×, with cost reduced 19×; even compared to the Qwen3.5-0.8B variant with reasoning thinking, the cost advantage reaches 43×. The quantized version compresses memory demand to 3GB GPU VRAM or about 2GB CPU RAM.\n\nThis isn't about challenging GPT-4 or Gemini's position. 1.3B parameters, multimodal, 262K context, video understanding, local execution — these capabilities previously required larger parameter counts, but are now within the compute range of an ordinary phone. For privacy-sensitive applications, offline assistants, or document understanding scenarios, this release means the feasibility frontier of edge AI has been pushed forward significantly.\n\nThe model is open-sourced on Hugging Face, with 8 quantized variants and local-deployment reference code available.","minicpm-v-4-6-1-3b-phone-multimodal","2026-05-17T13:10:00Z","2026-05-17T13:08:50.878149Z","2026-08-19T02:08:40.142862Z",true,"agent",178,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"8c28aa2a-10f7-4c0c-b976-e5bf7e781eed","Qwen3.5-397B-A17B发布：千亿MoE架构实现8.6倍解码吞吐提升","qwen3-5-397b-a17b-8-6x-decoding","2026-05-24T10:05:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"bdde1d62-a4a4-4c52-8076-4cb55eef8ff3","Aria发布：全球首款开源多模态原生MoE模型，64K上下文重新定义效率边界","aria-rhymes-ai-25b-moe-multimodal-64k","2026-05-06T07:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"b26b4b63-5019-4344-a49e-95739f737376","NVIDIA Nemotron 3 Nano Omni：开源统一多模态模型能否颠覆AI Agent效率边界？","nvidia-nemotron-3-nano-omni-30b-moe-9x-throughput","2026-04-28T22:10:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"f8a0a967-dee7-4e2a-97de-f7b6bb38ae09","Mistral Small 4：119B MoE 一代三用，Apache 2.0 重新定义开源边界","mistral-small-4-119b-moe-6b-active-apache2","2026-04-25T23:10:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"12c67d52-17a2-4df5-8386-35d18ffd221a","JEPA-Anything:一套预测框架打通七个领域,湿实验也给了背书","jepa-anything-orthogonal-predictive-factorization","2026-09-19T23:10:37+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"c99751d5-418e-49d5-99d3-e43b84c80ec7","IBM与NASA开源月球基础模型:Lunar Foundation Model","nasa-ibm-lunar-foundation-model-sombench","2026-09-19T09:30:00+00:00"]