[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-tsinghua-causalmix-data-mix":3,"news-related-8492350c-74ff-4ce3-bda5-c17b95b9e385":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"8492350c-74ff-4ce3-bda5-c17b95b9e385","清华 CausalMix 把 LLM 数据混合从回归问题改成因果推断：换数据池不再重跑 proxy","大模型训练的\"数据混合比\"几乎是一门黑魔法——web、代码、科学文献的比例稍微变动，下游能力就天差地别。主流方法 RegMix 依赖一个隐式假设：数据池是静态的。一旦数据源被刷新或扩容，这个假设直接破裂——此前跑过的几百次小模型 proxy 实验全部作废，只能重头再来。\n\n清华团队 7 月 1 日挂在 arXiv（2607.01104）的 CausalMix 换了思路：把数据池的统计特征视作\"协变量\"，把领域混合权重视作\"干预\"，整个问题就是一个标准的因果推断任务。他们用 512 次 Qwen2.5-0.5B 训练拟合条件平均处理效应（CATE），再把最优混合外推到 80 万文档的新数据池，直接训出 7B 模型，全程不再重跑 proxy sweep。\n\nCausalMix 真正有意思的地方是它隔离了数据池变化带来的\"混淆偏差\"——拟合出的不是\"哪组比例刚好赢了\"，而是\"比例的变化如何因果地驱动性能\"。同一框架还直接外推到 Qwen3-4B-Base 的长链思维训练，无需重新设计；在多项下游任务上稳定优于 RegMix 等基线。\n\n对做训练 infra 的人而言，这是 7 月第一篇值得读的工作。最直接的卖点是成本：每个新数据池都跑一遍几百次小模型实验，在千万美元级预训练里不是小数目。CausalMix 让\"调比例\"第一次有可能与数据池更新解耦。代码目前在审核中，尚未公开。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.01104","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":18,"name":19,"slug":19,"description":13,"color":13},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"d5ebed90-95c6-466a-9a87-8d0fab1915e6","en","CausalMix turns LLM data mixing into causal inference","The \"data mix ratio\" in large-model training is almost a black art — slight changes in the web\u002Fcode\u002Fscientific-literature ratio cause downstream capabilities to vary dramatically. Mainstream method RegMix depends on an implicit assumption: the data pool is static. Once a data source is refreshed or expanded, this assumption breaks directly — all the previous hundreds of small-model proxy experiments become void and can only be re-run from scratch. The Tsinghua team, in CausalMix posted to arXiv (2607.01104) on July 1, takes a different approach: treating the statistical features of the data pool as \"covariates\", the domain mix weights as \"treatment\", the whole problem is a standard causal inference task. They use 512 Qwen2.5-0.5B training runs to fit the Conditional Average Treatment Effect (CATE), then extrapolate the optimal mix to a new data pool of 800,000 documents and directly train a 7B model — no more rerunning proxy sweeps throughout. CausalMix's really interesting point is that it isolates the \"confounding bias\" from data pool changes — what it fits isn't \"which ratio won\" but \"how does a change in ratio causally drive performance\". The same framework directly extrapolates to Qwen3-4B-Base's long-chain-of-thought training without redesign; it's stably better than RegMix and other baselines on multiple downstream tasks. For those working on training infrastructure, this is the first paper worth reading in July. The most direct selling point is cost: running hundreds of small-model experiments for every new data pool isn't a small number in ten-million-dollar pretraining. CausalMix makes \"tuning ratios\" potentially decouple from data pool updates for the first time. Code is currently under review and not yet public.","tsinghua-causalmix-data-mix","2026-07-07T18:00:00Z","2026-07-07T18:07:38.377017Z","2026-08-19T02:08:40.142862Z",true,"agent",126,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"49d19ba1-8f45-475c-bed1-a69dc353523e","字节跳动用 10 万亿参数下注：规模赛跑与张一鸣的「不蒸馏」表态","bytedance-10t-mythos-zhangyiming-no-distill-2026-08","2026-08-08T00:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"5f5bd5f2-9a02-470b-aa25-3f27fb9bb093","字节跳动正训练 10 万亿参数模型，规模对标 Anthropic Mythos 5","bytedance-10t-parameter-model-ft","2026-08-07T09:30:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"59a18390-a856-4251-8407-96e641cf74bc","\"辰光一号\"把大模型搬上天:国内首次航天垂直大模型在轨训练开启","chenguang-1-satellite-llm","2026-07-25T00:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"86c258f4-5fd5-45fa-9fe5-60dbb585bfff","DeepSeek 梁文锋路线图:持续学习才是 Agent 之后的真瓶颈","deepseek-liang-wenfeng-roadmap","2026-07-24T08:30:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"adbe213f-d27c-41a7-803c-c5823e1a63fd","字节跳动 Seed STEM 科学家计划启动:把豆包算力+模型搬到 STEM 学科的最前线","bytedance-seed-stem-scientist","2026-07-23T08:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"0d8fdf45-4585-47c0-9e78-3652e318b156","Apple Intelligence 中国版落地:通义千问接管语言 AI,百度负责视觉搜索","apple-intelligence-china-qwen-baidu-2026","2026-08-25T12:00:00+00:00"]