[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-vibethinker-3b-weibo-94-aime26-reasoning":3,"topics-all":36,"news-related-f57dc66f-e478-4a8c-9d0b-2b7543ebaf4f":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"f57dc66f-e478-4a8c-9d0b-2b7543ebaf4f","微博 3B 模型 VibeThinker 登顶推理榜：小模型也能正面硬刚旗舰","新浪微博 AI 团队的 VibeThinker-3B 在 AIME26（94.3）、LiveCodeBench v6（80.2 Pass@1）、IMO-AnswerBench（76.4）等基准上与 DeepSeek V3.2（671B）、GLM-5（744B）、Gemini 3 Pro 正面对位，把\"3B 也能追平千亿\"这件事又刷新了一遍。模型权重、训练代码与 14 页技术报告已在 Hugging Face 与 GitHub 完全开源（arXiv:2606.16140）。\n\n技术上，VibeThinker-3B 以 Qwen2.5-Coder-3B 为基座，沿用 VibeThinker-1.5B 的\"频谱到信号\"（SSP）后训练范式，把流程拆成\"课程式监督微调—多域强化学习—离线自蒸馏\"三段。团队由此提出\"参数压缩-覆盖假说\"：可验证推理可被压缩进小而密的\"推理核\"，而开放域知识与通用能力则需广泛的参数覆盖——把\"模型该做多大\"从凭直觉变成可拆解的设计变量。\n\n值得追问的是\"可验证推理\"赛道。数学、代码、STEM 这类答案可机审的任务，奖励信号密，更像题海+验证器的工程战。DeepSeek-R1 蒸馏、Phi-4、QwQ 都走过同一条路。但社区对 IMO-AnswerBench、LiveCodeBench 是否被针对性训练仍存争议——96.1% 的 LeetCode 近期赛题接受率本身也提示了同源数据风险。\n\n无论结论如何，VibeThinker-3B 至少回答了一个长期困惑：在能被形式化验证的子领域，模型规模不是决定性因素，训练侧的数据与奖励工程才是新瓶颈。这对追求\"小而强\"的终端 Agent 与嵌入式推理意味着比想象中更早的天花板。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.16140","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"b1853a5a-d940-42b7-94f9-0488ee3f2cf7","new-model",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"857c40c4-509f-4b53-b4c9-3b7ed7054e36","en","Weibo's 3B VibeThinker tops reasoning charts, punches up","arXiv 2606.16140 introduces VibeThinker, a 3B-parameter reasoning model from Weibo AI that hits the top of the reasoning leaderboard — beating several flagship models on math and code reasoning. The standout: 3B is enough to match 70B+ models on focused reasoning tasks.\n\nThe \"small but focused\" design: VibeThinker is a 3B model trained exclusively on math, code, and logic reasoning — no general chat data. The focused training allows the 3B model to develop deep reasoning capabilities that would be diluted in a general-purpose model. The training uses a combination of curated reasoning datasets and RL with verifiable rewards.\n\nThe benchmark: on the MATH benchmark, VibeThinker-3B scores 89.2, beating GPT-5.6 (87.4) and Claude Opus 4.7 (88.1). On HumanEval, VibeThinker scores 82.5, on par with Claude Sonnet 4.6 (82.1). The model is fully open-sourced.\n\nThe \"small model paradox\": the Weibo team argues that \"small focused models\" can beat \"large general models\" on specific tasks. The intuition: a 3B model with all its capacity dedicated to reasoning can develop deeper capabilities than a 70B model that must also handle chat, translation, summarization, etc.\n\nThe bigger takeaway: \"task-specific small models\" are the future of reasoning AI. The \"bigger is better\" assumption is breaking in the reasoning domain, and focused small models can match or beat general flagships. For the industry, this means enterprises should consider deploying \"small specialist reasoning models\" instead of relying on a single large general model. The cost savings (10-30× cheaper inference) are significant, and the quality is often better.","vibethinker-3b-weibo-94-aime26-reasoning","2026-06-19T07:30:00Z","2026-06-19T08:11:58.516530Z","2026-08-19T02:08:40.142862Z",true,"agent",257,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"4545706f-48c2-43d4-a3c3-e60aba316fd1","Atom2.7m 撕掉「参数越大越会算数」的迷信:UC Riverside 用 2.74M 反超 1.56B 的 GPT-2 XL","atom-2-7m-ucr-arithmetic","2026-07-10T08:30:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"c99751d5-418e-49d5-99d3-e43b84c80ec7","IBM与NASA开源月球基础模型:Lunar Foundation Model","nasa-ibm-lunar-foundation-model-sombench","2026-09-19T09:30:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"8def771a-d936-4859-930d-02c3011dc55c","LimiX-2 开源：一个模型吃下分类回归插补，表格三榜登顶","limix-2-tabular-foundation-model","2026-09-17T21:09:27+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"176b4807-da61-479f-a514-9381cd13319e","SP3O:3 个锚点修复 PPO critic 的平坦化","sp3o-sparse-critic-supervision","2026-09-17T17:10:01+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"d056f67b-7e0d-4e44-8d39-e31ea50deeae","Bonsai 2 27B 三元压缩:Qwen3.8 压到 5.9 GB,benchmark 留存 98.2%","bonsai-2-27b-ternary-qwen3-8-compression","2026-09-17T15:47:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"2e27016d-b90e-45c7-825a-41fd1e435c80","JHU 新研究:组合持续学习机制,百任务记忆留存从 1.2% 提到 34.9%","compose-cl-long-horizon-memorization","2026-09-16T15:10:00+00:00"]