[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-hplt-multisynt-multilingual-dataset":3,"topics-all":36,"news-related-ab2d6e9e-8890-4ae7-b6ca-8febc831a279":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"ab2d6e9e-8890-4ae7-b6ca-8febc831a279","HPLT MultiSynt\u002FMT：4.8 万亿 token 多语种数据集","由 HPLT 联盟（图尔库大学、赫尔辛基大学等机构）联合发布的 MultiSynt\u002FMT，是当前规模最大的开放多语种预训练语料之一：覆盖 36 种欧洲语言、合计约 4.8 万亿 token 的目标语料，由 1000 亿 token 高质量 Nemotron-CC 文本经过 Tower+ 与 OPUS-MT\u002FHPLT-MT 系统翻译生成。对许多中低资源欧洲语言（如冰岛语、爱尔兰语、加利西亚语）来说，这是它们能直接拿到的最大开源预训练资源——过去这些语言要么依赖零散爬取的 web 语料，要么干脆进不了主流预训练流水线。\n\n真正让业内侧目的是训练效率的实证。在多语基准评测套件上，使用 MultiSynt\u002FMT 训练的参考 LLM 仅需 HPLT 2.0（原生爬取数据基线）约 72% 的预训练 token 即可追平最终分数——换算下来，预算可以直接压缩到 28%。而在把预算锁在固定的 1000 亿 token 时，相对于原生基线还有约 15% 的相对提升。它把\"机器翻译合成语料是否真能替代原生语料\"这件吵了几年的事，推进到了量化实证阶段。\n\n论文还顺手做了一件对评测社区很重要的事：用 LLM-as-judge（流畅度敏感的判别器）复测时，标准多选题基准几乎抹平了不同 MT 系统翻译质量的差异，而基于流畅度的判别评测却把这层信号重新拉了回来，并证实问题不在 MultiSynt 本身。同时也坦承挪威语等文化语境任务仍然更适合用原生数据——这是对\"token 总量神话\"的一次冷静校准。对想在欧洲多语场景下做下一轮预训练实验的团队，这份 4.8 万亿 token 的开放语料几乎是裸送的礼物。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.00890v1","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"2763ec52-31aa-47e0-a538-cd5ec6801d82","en","HPLT MultiSynt\u002FMT: 4.8-trillion-token multilingual corpus","MultiSynt\u002FMT, jointly released by the HPLT consortium (University of Turku, University of Helsinki, and other institutions), is one of the largest open multilingual pretraining corpora currently: covering 36 European languages, with a target corpus of about 4.8 trillion tokens, generated from 100 billion tokens of high-quality Nemotron-CC text through Tower+ and OPUS-MT\u002FHPLT-MT system translation. For many low-to-medium-resource European languages (like Icelandic, Irish, Galician), this is the largest open-source pretraining resource they can directly access — in the past these languages either relied on scattered web crawling or simply couldn't enter mainstream pretraining pipelines. What really catches the industry's attention is the empirical evidence of training efficiency. On the multilingual benchmark suite, the reference LLM trained with MultiSynt\u002FMT only needs about 72% of HPLT 2.0's (the native-crawl-data baseline) pretraining tokens to match the final score — converted, the budget can be directly compressed to 28%. And when the budget is locked at a fixed 100 billion tokens, there's still about a 15% relative improvement over the native baseline. It pushes the years-long debate over \"can machine-translated synthetic corpora really replace native corpora\" to the quantitative-empirical stage. The paper also did a favor for the evaluation community: when re-tested with LLM-as-judge (a fluency-sensitive discriminator), the standard multiple-choice benchmarks almost flatten out the quality differences between different MT system translations, while the fluency-based discriminator evaluation pulls this layer of signal back and confirms the issue isn't in MultiSynt itself. At the same time, it candidly admits that cultural-context tasks like Norwegian are still better served with native data — this is a calm calibration of the \"token total mythology\". For teams wanting to do the next round of pretraining experiments in European multilingual scenarios, this 4.8-trillion-token open corpus is almost a freebie.","hplt-multisynt-multilingual-dataset","2026-07-07T16:01:00Z","2026-07-07T16:11:56.369057Z","2026-08-19T02:08:40.142862Z",true,"agent",176,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"054e060c-e182-42a9-b3ed-229feb8ac0ac","2026 年的蒸馏长什么样:Hugging Face 拆解前沿模型三大范式","2026-distillation-three-paradigms","2026-07-09T12:00:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"93dfc6f4-e4a9-47a3-aa66-b4b9cee864e7","Beyond LoRA 不只是口号：HF 给 40+ PEFT 方法拍下公平基准，OFT 在图像任务上反超 LoRA","beyond-lora-hf-peft-benchmark-oft","2026-06-18T12:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"28c41f06-d20f-481c-b133-cd109af3aed1","答对之后停不下来:微软团队揪出在线蒸馏的 EOS 错配元凶","eos-mismatch-opd-length-inflation","2026-09-18T21:09:06+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"8def771a-d936-4859-930d-02c3011dc55c","LimiX-2 开源：一个模型吃下分类回归插补，表格三榜登顶","limix-2-tabular-foundation-model","2026-09-17T21:09:27+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"176b4807-da61-479f-a514-9381cd13319e","SP3O:3 个锚点修复 PPO critic 的平坦化","sp3o-sparse-critic-supervision","2026-09-17T17:10:01+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"7bae3d71-a5c2-4588-95e7-b5d4b5c7085a","开源模型 4.4 个月追上闭源前沿:Hugging Face 被 NVIDIA 129 亿美元收编","nvidia-acquires-hugging-face-open-source-ai","2026-09-17T08:00:00+00:00"]