[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-bytedance-seed-2-1-agent-first":3,"news-related-6ba58314-305f-4255-83c5-87bdd1123b49":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"6ba58314-305f-4255-83c5-87bdd1123b49","字节 Seed 2.1 押注「Agent-first」：模型自己参与训练，多模态重夺 SOTA","6月24日，字节 Seed 团队正式上线 Seed 2.1 模型家族（标准版与 Pro），通过豆包与火山引擎开放。与单纯刷榜不同，这次更清晰的信号是：模型评估体系正从「静态基准」转向「真实工作流」。\n\n核心改进集中在三点。**通用 Agent 能力**：Seed 2.1 Pro 在衡量经济任务完成质量的 GDPVal 拿下最高分，在 MobileWorld 移动 GUI 任务上同样排名第一，平均完成任务所需的步数减少 16%。**端到端 Coding 能力**：模型能跨文件理解代码架构、依赖关系和业务逻辑，产出可维护的工程级交付；前端开发场景在 Code Arena 排名第八，得分 1539。**多模态与长上下文**：在 CharXiv-RQ、MeasureBench、ERQA 等视觉理解基准，以及 TVBench、TOMATO、OVBench 等视频理解基准上均刷新 SOTA，并在 MMLongBench-128K 长上下文测试中表现稳定。\n\n更具想象空间的是「Seed for Seed」计划——模型不再只是被训练的对象，而是直接参与评估系统搭建、能力诊断、SFT 数据合成和 RL 训练框架优化等环节，多个 Agent 以执行、评估、诊断、优化等不同角色协同，把 R&D 流水线变成可持续的闭环。\n\n关键意义不在某项跑分，而在「Agent-first」的产品定位：把模型从对话工具拉回工作流执行者。当模型能稳定完成多步跨工具任务、交付可维护的工程代码时，企业级部署意愿才真正被点燃。字节这一步的对手是 OpenAI 的 GPT-5.6 与 Anthropic 的 Claude 系列，国内则有 GLM-5.2、Kimi K2.7 Code 等开源力量贴身竞争——2026 下半年的 LLM 主战场，已从「刷榜」转向「能不能干活」。","https:\u002F\u002Fseed.bytedance.com\u002Fen\u002Fblog\u002Fseed2-1-officially-released-advancing-ai-productivity","d4a24db7-b6c2-410c-ba99-c16625c61305",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"e82b2d09-81b2-43d1-977e-e018443b3c14","coding-agent",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"6479a4fc-634f-4336-9691-c881b87c29dd","en","Seed 2.1 goes agent-first: models join their own training","ByteDance Seed officially released Seed 2.1, its next-generation foundation model. The biggest change from Seed 1.x: a clear \"Agent-first\" design philosophy — the model is no longer just a passive responder but actively participates in its own training loop.\n\nThe technical details: Seed 2.1 introduces a \"self-distillation with self-generated trajectories\" mechanism — the model generates its own Agent trajectories, evaluates which paths are most valuable, and uses those paths as its own training data. This is a \"self-rewarding + self-improving\" closed loop, breaking the traditional \"human-annotated data ceiling.\"\n\nThe multimodal side is the highlight: Seed 2.1 unifies text, image, video, and audio into a single architecture, and the multimodal benchmark refreshes SOTA on multiple items — visual reasoning, long video understanding, audio-visual alignment, and cross-modal generation all see significant gains. Particularly impressive is the long-video understanding, which surpasses GPT-5.6 on hour-long videos.\n\nThe bigger takeaway: \"Agent-first\" is not just a slogan. Seed 2.1's training pipeline already uses the model itself as the data generator and evaluator, blurring the line between \"training\" and \"inference.\" For the industry, this means the next generation of foundation models will not just be \"bigger\" but \"more autonomous in their own improvement\" — and the gap between closed and open models may widen further.\n\nThe challenge: \"self-distillation\" has obvious feedback-loop risks. Whether Seed 2.1 can avoid \"model collapse\" will be a long-term question, and ByteDance says it has added external-curation and human-spot-check mechanisms to mitigate this.","bytedance-seed-2-1-agent-first","2026-06-27T15:30:00Z","2026-06-27T16:10:10.612337Z","2026-08-19T02:08:40.142862Z",true,"agent",87,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"c40263bb-c193-46e8-ba63-76499bb1c2af","豆包 2.1 Pro 抢跑 Agent 时代：180T 日均 token 背后的 MaaS 规模战","doubao-2-1-pro-180t-tokens-maas","2026-06-23T04:30:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"f6e4aab0-7693-4c2c-bb66-c1641fc2cc3e","Ox Alpha 谜底揭晓:智谱 GLM-5.3-Flash,MIT 开源 320B MoE","ox-alpha-glm-5-3-flash-reveal","2026-08-27T13:30:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"804ab59a-a8d6-4b61-bf74-8f6f2bdae83c","智谱把 Flash 做成一件正经事:一次说清 GLM-5.3-Flash 的架构和 benchmark 真相","glm-5-3-flash-hybrid-attention-architecture","2026-08-27T08:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"a64d03b9-1d07-404b-9231-d434c65c44ce","OX Alpha 免费一周:模型页说不训练,EULA 却保留训练权","ox-alpha-stealth-eula-retention-conflict","2026-08-23T13:10:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"89a79f9a-bfd2-4ebe-8f03-92fa74a3a34f","Ornith-1.5 开源：模型自己出题、自己搭考场，397B 到 9B 三档齐发","ornith-1-5-self-improvement-open-models","2026-08-20T13:30:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"d1e8997e-bb60-453d-9ef8-71b8bdde5386","Harvey 首个自研法律模型 Tenet 曝光:底座没选 GPT 和 Claude,选了 Kimi K3","harvey-tenet-kimi-k3-legal-model","2026-08-18T17:30:00+00:00"]