[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-thomson-1-0-small-continual-learning":3,"news-related-3d36921f-3b84-4663-97a0-fee7d4eff795":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"3d36921f-3b84-4663-97a0-fee7d4eff795","汤森路透开源 Thomson-1.0-Small:持续学习改造 Qwen,3B 激活的 35B MoE","Thomson Reuters 开源 Thomson-1.0-Small:基于 Qwen3.6-35B-A3B 持续学习,聚焦法律\u002F税务\u002F新闻,训练量约 1.63×10²³ FLOP。法律 Agent 基准 73.4、总体均分 74.6 领先 Gemma 4-31B 和 Haiku 4.5,编码与多语言偏弱。","汤森路透把 Thomson-1.0 家族的开源权重成员挂上了 HuggingFace:Thomson-1.0-Small,35B 总参数、3B 激活的 MoE,原生 262,144 上下文,底座是 Qwen3.6-35B-A3B。技术报告《Thomson: Continual Learning of Frontier Models for SovereignAI》8 月 27 日提交 arXiv(编号 2608.27147),权重以 BF16 格式提供,兼容 Transformers、vLLM、SGLang 等推理栈。\n\n## 算力账:前沿模型到底要花多少钱\n\n论文的核心论点很直接:前沿性能不是少数巨额融资玩家的专利。整个流水线消耗约 1.63×10²³ FLOP,折合 35,207 个 B200 GPU 小时;中期训练语料从超过 19T tokens 的池子里筛出 200B tokens,由专有文档、这些文档的同义改写、通用能力回放数据三部分大致均分。合作方包括帝国理工学院(共同完成价值对齐)、DatologyAI(数据策展)与 Lambda。\n\n## 三段式流水线\n\n第一阶段做价值对齐,用 Constitutional DPO 把模型对齐到公开可修改的 Public AI Constitution,而不是某家公司的私有价值观;第二阶段持续预训练,吸收汤森路透数十年积累的新闻、合同、监管文件、判例、法规与实务指引,同时用模型合并保护通用能力;第三阶段后训练混合 DPO 与强化学习,偏好数据直接派生自专家撰写的内容,还包括基于 IRAC 案例分析框架的本体驱动偏好数据,以及 Deep Research 代理式 harness 的奖励设计——后者的奖励结构专门激励忠实的工具使用和准确的引用习惯。\n\n## 跑分真相:赢在哪、输在哪\n\n总体均分 74.6,高于同底座 Qwen3.6-35B-A3B 和中间检查点 Snowdon-1.1-Small 的 71.7,也压过 Gemma 4-31B(71.2)和 Haiku 4.5(68.2)。亮点集中在专业场景:Harvey 法律 Agent 基准 73.4,高于底座的 69.5,Gemma 4-31B 在该项甚至没过 35 分;税务 Deep Research 78.6;通用 Agent 85.8;政治中立性 98.5。短板同样明显:编码 37.4 反而低于底座的 39.8;MBE 律师资格考试 83.4 输给 Gemma 的 88.8;多语言 71.9 距 Haiku 的 85.8 差距不小。更值得注意的是论文所说的「π 形」增益——非目标能力不降反升:AIME 2026 从底座的 86.7 提到 90.0,MMLU-Pro、GPQA-Diamond 均小幅上涨,基本消除了窄域适配常见的灾难性遗忘。\n\n## 所以呢\n\n这篇论文真正想论证的是 SovereignAI:一家内容公司拿着开源底座、3.5 万个 B200 GPU 小时的预算和自家专有语料,就能在法律、税务这类高价值垂域做出有竞争力的模型,并把模型权重、工具链、价值观对齐整条栈握在自己手里。对行业观察者而言,Qwen 开放权重生态正在成为全球机构做「主权 AI」的事实底座——这次是伦敦的出版商,下一次可能是银行或律所。技术报告:arxiv.org\u002Fabs\u002F2608.27147;权重:huggingface.co\u002Fthomsonreuters\u002FThomson-1.0-Small。","https:\u002F\u002Fhuggingface.co\u002Fthomsonreuters\u002FThomson-1.0-Small","44f4a577-4568-406d-a37e-c8fd651ffcc3",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":19,"name":20,"slug":20,"description":14,"color":14},"d11f0044-8aef-487c-bebe-89ce4683a4a3","moe",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"138eb733-c01d-4340-a258-dc226610dd52","en","Thomson Reuters open-sources Thomson-1.0-Small, 35B MoE built on Qwen","Thomson Reuters open-sources Thomson-1.0-Small: a 35B\u002F3B-active MoE built on Qwen3.6-35B-A3B for legal and tax; overall 74.6, legal agent 73.4, coding lags.","Thomson Reuters has put the open-weight member of its Thomson-1.0 family on HuggingFace: Thomson-1.0-Small, a Mixture-of-Experts model with 35B total parameters and 3B activated, a native context length of 262,144, and Qwen3.6-35B-A3B underneath. The technical report, \"Thomson: Continual Learning of Frontier Models for SovereignAI\", hit arXiv on August 27 (2608.27147), with weights shipped in BF16 and compatible with Transformers, vLLM, SGLang and other serving stacks.\n\n## The compute bill\n\nThe paper's central claim is blunt: frontier performance is not the exclusive remit of heavily funded players. The full pipeline consumed roughly 1.63 x 10^23 FLOP — 35,207 B200 GPU-hours. The mid-training corpus was 200B tokens curated from a pool of over 19T tokens, split roughly equally between curated proprietary documents, synthetic rephrasings of those documents, and general-capability replay data. Partners include Imperial College London (value re-alignment), DatologyAI (data curation) and Lambda.\n\n## A three-stage pipeline\n\nStage one handles values: Constitutional DPO aligns the model to the openly developed, freely modifiable Public AI Constitution rather than a proprietary value system. Stage two is data-centric continual pre-training on decades of Thomson Reuters material — news, contracts, regulatory filings, case law, statutes and practitioner guidance — with model merging protecting general capabilities as domain knowledge sinks in. Stage three combines DPO with reinforcement learning; preference data derives from expert-authored material, ontology-driven preference data built on schemas like IRAC for case law, and an agentic Deep Research harness whose reward structure explicitly rewards faithful tool use and accurate citation.\n\n## What the benchmarks actually say\n\nThe overall average is 74.6 — above the 71.7 of both its base Qwen3.6-35B-A3B and the intermediate Snowdon-1.1-Small checkpoint, ahead of Gemma 4-31B (71.2) and Haiku 4.5 (68.2). The wins concentrate in professional work: Harvey legal agent benchmark 73.4 versus the base's 69.5 (Gemma 4-31B scored under 35 there), tax Deep Research 78.6, general agent 85.8, political neutrality 98.5. The weak flanks are equally clear: coding 37.4 is below the base's 39.8; the MBE bar exam 83.4 loses to Gemma's 88.8; multilingualism 71.9 trails Haiku's 85.8 by a wide margin. What the paper calls the pi-shaped gain is the more striking part: non-targeted capabilities went up — AIME 2026 rose from the base's 86.7 to 90.0, with MMLU-Pro and GPQA-Diamond also ticking up — while the catastrophic forgetting common to narrow domain adaptation was almost eliminated.\n\n## So what\n\nWhat the paper really argues for is SovereignAI: a content company, an open-weight base, 35 thousand B200 GPU-hours and proprietary corpus can produce a competitive model in high-value verticals like legal and tax — owning the full stack of weights, tooling and value alignment. For anyone watching the industry, the Qwen open-weight ecosystem is becoming the de facto foundation for institutions building \"sovereign AI\" — this time a London publisher; next time, perhaps a bank or a law firm. Technical report: arxiv.org\u002Fabs\u002F2608.27147; weights: huggingface.co\u002Fthomsonreuters\u002FThomson-1.0-Small.","thomson-1-0-small-continual-learning","2026-08-28T19:10:00Z","2026-08-28T19:11:02.587236Z","2026-08-28T19:11:02.587251Z",true,"agent",63,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"33f3b08b-c8a2-43ec-81cf-85e2b918f913","腾讯开源 Hy4 preview:770B MoE、1M 上下文,模型首次参与自身训练","tencent-hy4-preview-770b-moe","2026-08-29T15:00:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"804ab59a-a8d6-4b61-bf74-8f6f2bdae83c","智谱把 Flash 做成一件正经事:一次说清 GLM-5.3-Flash 的架构和 benchmark 真相","glm-5-3-flash-hybrid-attention-architecture","2026-08-27T08:00:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"7958a2f1-028c-4b4e-b134-0d5de9afc1c1","Motif 3 收官:韩国 314B MoE 改用 MIT 许可,从零起步架构首次面向商用","motif-3-mit-license-sovereign-ai","2026-08-24T00:00:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"491f4904-c854-4925-b3e3-e34b8afd5e50","KDA+MLA 混合栈下沉到 1.3B 激活:Ling-3.0-tiny 把 MoE 端侧化,INT4 跑出 115 tok\u002Fs","ling-3-tiny-kda-mla-edge-deployment","2026-08-18T00:00:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"637f84e0-e6dc-490a-bba1-879f6527bdd5","Qwen3.8-Max 2.4T 开源:Gated DeltaNet 把长上下文成本砍到 1\u002F8","qwen3-8-max-2-4t-open-weights-gated-deltanet","2026-08-30T03:00:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"61de017b-bdd6-44b3-9f45-d4fb233bd24d","PhoneLLM 开源:30B MoE 电话客服模型,自称比 GPT-5.6 Terra 便宜 94%","phonellm-alpha-1-voice-agent-open-model","2026-08-29T21:10:00+00:00"]