[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-nanbeige-4-2-3b-looped-transformer-agentic-3b-beats-qwen3-5-9b":3,"news-related-80315de0-7eb3-491a-b2e6-103a691a8bd7":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"80315de0-7eb3-491a-b2e6-103a691a8bd7","Nanbeige4.2-3B 用 Looped Transformer 在 11 项基准上跑赢 Qwen3.5-9B","Nanbeige \u002F 看准(BOSS直聘旗下)推出 4B 总参、3B 激活的紧凑 agentic LLM Nanbeige4.2-3B,基于 Looped Transformer 把 transformer 层复用多次来拉容量,绕过\"参数膨胀\"的常规路径。在 11 项 Agent \u002F 代码 \u002F 推理评测上它全部跑赢 Qwen3.5-9B 与 Gemma4-12B —— SWE-bench Verified 63.6、GPQA-Diamond 87.4、Terminal-Bench 2.0 44.1 —— SFT 与 RL 阶段同时把 process reward 与 outcome reward 摆上,tool-call 稳定性远胜同尺寸对手。GGUF 量化后 MacBook 本地就能跑,五栈部署齐活。","**一句话**:Nanbeige \u002F 看准(BOSS直聘旗下)推 4B 总参 \u002F 3B 激活的紧凑 Agentic 模型,在 11 项评测上把 Qwen3.5-9B 和 Gemma4-12B 摁在身下。\n\n技术背景:2026 年的 open-weight LLM 战场有一条清晰的\"小模型反超大模型\"叙事 —— GLM-5.2、Kimi K3 用 MoE 把激活参数做小,Nanbeige 选的是另一条路:Looped Transformer,即同一组 transformer 层在网络里被复用多次,把容量挤出来的同时参数总数不涨。Nanbeige4.2-3B 是这条路线上目前最强的开源 checkpoint,3B 非词嵌入参数、256K 上下文、自家训练栈 SFT + outcome\u002Fprocess 双奖励 RL。\n\n核心数据(11 项评测全部领先 9B 对标):\n- SWE-bench Verified **63.6** vs Qwen3.5-9B 53.1、Gemma4-12B 44.2\n- SWE-bench Pro **46.9** vs Qwen3.5-9B 33.8、Gemma4-12B 21.9\n- Terminal-Bench 2.0 **44.1** vs Qwen3.5-9B 29.2、Gemma4-12B 21.1\n- GPQA-Diamond **87.4**(3B 模型比大多数博士还高)\n- HMMT-Feb-2026 **82.8**\n- Office-QA-Pro **21.1**(OpenClaw 框架 + 真实 PDF 上下文)\n- MCP-Atlas **57.8**(MCP 工具调用榜单前列)\n\n机制核心:LoopSplit 把循环层切开做轻量推理、mHC with depth attention 在循环内部加多头压缩、concat n-gram embeddings 让相邻 token 信息流贯穿 —— 三件套联手让 4B 参数装下了原本 9B 起步的能力。RL 阶段把 SFT 的轨迹在第二轮按 outcome + process reward 重新打分,显著收紧工具调用的稳定性。\n\n部署门槛:SGLang \u002F vLLM \u002F llama.cpp \u002F Ollama \u002F MLX 五栈全支持。GGUF Q4_K_M 量化后能在 MacBook 本地跑,OpenClaw 框架实测日常任务 \u002F 办公 \u002F 深研究 6 项基准全超 Qwen3.5-9B。这意味着 3B 模型在 Apple Silicon 上已经够用,云 API 调用不再是必选项。\n\n个人评论:Nanbeige 的发布信息密度很高 —— 模型卡写清了 Looped Transformer 在循环维度和注意力维度的具体收益,但缺一个对比 DeepSeek V4 MLA \u002F Hunyuan Hy3 KDA \u002F Qwen3.5 GQA 的横向 subsection。\"小模型 > 大模型\"这条叙事走到 2026 下半年,真正的胜负手已经不在参数总量,而在(1)循环复用带来的有效深度 + (2)RL 阶段是否给出 process-level reward + (3)agent 工具栈的连续训练 —— Nanbeige 这条路线,前两条都做了。Nanbeige4.5 也已经在路上,Lora-NB(老板直聘)把招聘垂直场景的 agent 能力下放给 3B 模型,这是 SaaS 公司把模型当基础设施用的典型打法。\n\n所以呢:不是参数越大越强 —— 看架构 + 训练管线 + 部署链路的整体工程深度。如果你想跑一个能在 MacBook 上帮你处理文档、查资料、写代码的本地 agent,3B 已经不是\"够用就行\",而是\"超额交付\"。\n","https:\u002F\u002Fhuggingface.co\u002FNanbeige\u002FNanbeige4.2-3B","24d5c6c5-6573-4180-a1fd-f1459842d1af",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":19,"name":20,"slug":20,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"e9c844b7-13f6-4b74-800e-012d78b78dfe","en","Nanbeige4.2-3B: Looped Transformer tops Qwen3.5-9B on 11","Nanbeige (under BOSS Zhipin) released a compact 4B-total \u002F 3B-active agentic LLM that beats Qwen3.5-9B and Gemma4-12B across 11 benchmarks — using a Looped Transformer that reuses transformer layers to amplify effective capacity without growing parameters. SWE-bench Verified 63.6, GPQA-Diamond 87.4, SGLang\u002FvLLM\u002Fllama.cpp\u002FOllama\u002FMLX fully supported, GGUF quantization runs locally on a MacBook.","Bottom line: Nanbeige (the AI lab under BOSS Zhipin \u002F Kanzhun) released a compact 4B-total \u002F 3B-active agentic LLM that beats Qwen3.5-9B and Gemma4-12B on 11 benchmarks — and runs locally on a MacBook.\n\nBackground: by mid-2026 the \"small model beats big model\" narrative is one of the clearest storylines in open-weight LLMs. GLM-5.2 and Kimi K3 cut activated parameter counts via MoE; Nanbeige takes a different route — a Looped Transformer architecture that reuses the same set of transformer layers multiple times, squeezing capacity out without growing the parameter count. Nanbeige4.2-3B is the strongest open-weight checkpoint on this track so far: 3B non-embedding parameters, 256K context, and an SFT + outcome\u002Fprocess dual-reward RL pipeline that they wrote themselves.\n\nThe numbers (11 benchmarks, all wins against the 9B \u002F 12B comparators):\n\n- SWE-bench Verified 63.6 vs Qwen3.5-9B 53.1 \u002F Gemma4-12B 44.2\n- SWE-bench Pro 46.9 vs 33.8 \u002F 21.9\n- Terminal-Bench 2.0 44.1 vs 29.2 \u002F 21.1\n- GPQA-Diamond 87.4 (a 3B model outscoring most PhDs)\n- HMMT-Feb-2026 82.8\n- Office-QA-Pro 21.1 (using the OpenClaw framework with real PDF context)\n- MCP-Atlas 57.8 (top tier on the MCP tool-use leaderboard)\n\nThe mechanism: three architectural tricks — LoopSplit slices the looped layers for cheaper inference at the front, mHC with depth attention adds multi-head compression inside the loop, and concatenated n-gram embeddings let adjacent-token information flow across the reused layers. Together they make a 4B parameter model hold what would otherwise require 9B+ parameters. The RL stage then re-scores SFT trajectories with combined outcome + process rewards, dramatically tightening tool-call reliability.\n\nDeployment: SGLang, vLLM, llama.cpp, Ollama, and MLX are all supported. A GGUF Q4_K_M quant runs natively on a MacBook, and OpenClaw framework tests beat Qwen3.5-9B on 6 daily-assistant \u002F office \u002F deep-research benchmarks. A 3B model being usable locally — not just \"barely adequate\" — means the cloud API call is no longer mandatory.\n\nMy take: Nanbeige's model card is dense with information — they wrote up exactly which architectural dimensions (loop depth, attention) buy which gains, although a side-by-side comparison against DeepSeek V4 MLA \u002F Hunyuan Hy3 KDA \u002F Qwen3.5 GQA would have been welcome. By H2 2026 the \"small beats big\" race is no longer about total parameters; it's about (1) effective depth via loop reuse, (2) process-level rewards in RL, and (3) a contiguous agent tool-stack training corpus. Nanbeige delivers on the first two. With Nanbeige4.5 already on its way, BOSS Zhipin is positioning vertical-domain (recruiting) agent capability on a 3B model — that is a textbook SaaS-company move of treating the model as infrastructure.\n\nSo what: bigger isn't better — architecture, training pipeline, and deployment chain together determine what a model actually delivers. If you want a local agent that handles documents, research, and coding on a MacBook without sending data to the cloud, 3B is no longer \"good enough\" — it is now over-delivering.","nanbeige-4-2-3b-looped-transformer-agentic-3b-beats-qwen3-5-9b","2026-07-30T10:30:00Z","2026-07-30T06:07:50.005791Z","2026-07-30T06:07:50.005801Z",true,"agent",198,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"f9cf9f03-6aca-4d29-94d3-5c6acfeaf435","匿名模型 OX Alpha 短暂登顶 OpenRouter 编码榜:研究者推测底座指向智谱 GLM-5.x","ox-alpha-stealth-openrouter-glm-5-zhipu","2026-08-24T03:00:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"d4fa7e14-8fbd-4940-93a6-3dd6f0a3991d","DeepSeek V4 Pro 正式版：1.6T MoE，1M 上下文","deepseek-v4-pro-0813-ga-1m-context-moe","2026-08-13T02:00:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"c94766df-827e-4e4e-a006-b6639ec76722","DeepSeek V4-Flash-0731 转正观察:权重不动,后训练把 Agent 分数打到 V4-Pro 之上","deepseek-v4-flash-0731-agent-benchmark-official-aug2026","2026-08-01T02:00:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"5bfdf32b-44eb-4eb5-a98b-39e921168182","九天内连发五款前沿模型:7 月的大模型军备赛,真正决胜负的不再是 benchmark","july-2026-five-frontier-models","2026-07-23T12:00:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"6f1f105b-8e80-4b2c-b88c-b392556952aa","2026年本地LLM深度评测：开源模型性能全解析","local-llm-2026-deep-eval-swe-bench-aime","2026-04-25T11:15:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"f6e4aab0-7693-4c2c-bb66-c1641fc2cc3e","Ox Alpha 谜底揭晓:智谱 GLM-5.3-Flash,MIT 开源 320B MoE","ox-alpha-glm-5-3-flash-reveal","2026-08-27T13:30:00+00:00"]