[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-meta-fair-byte-distillation-token-ceiling-2026-09":3,"topics-all":41,"news-related-2266cea6-06f1-4932-8905-1bc3f2e5a8c0":60},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":27,"news_slug":34,"published_at":35,"created_at":36,"modified_at":37,"is_published":38,"publish_type":39,"image_url":14,"view_count":40},"2266cea6-06f1-4932-8905-1bc3f2e5a8c0","Meta FAIR 字节蒸馏研究:End-Of-Token 渐近反超 token 蒸馏 4%,数据只需 1\u002F6","Meta FAIR 新论文系统对照字节与 token 蒸馏的 scaling 趋势:1B 参数、1T bytes 训练下,End-Of-Token 字节蒸馏渐近精度比 token 蒸馏高 4%,只需 1\u002F6 数据,logit 存储降至 1\u002F5。","## 技术背景\n字节模型蒸馏在低算力下长期被 token 模型碾压,但 Meta FAIR 在 9 月 11 日公开的论文《Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models》里,第一次用严格控制变量的 scaling law 研究给出了反直觉的结论:字节模型起步更慢,渐近精度却更高。\n\n研究团队来自 Meta FAIR 与华盛顿大学(Mike Lewis、Luke Zettlemoyer、Srinivasan Iyer 等),作者把学生模型的 token 化方案切成三组 —— 经典 token、纯 byte、byte+end-of-token(EOT) —— 同时把训练目标切成蒸馏与交叉熵两组,在保持层-参数量匹配(约 1.28B layer 参数 \u002F 1.81B token 参数)的前提下,把训练量一直推到 1 万亿字节。\n\n实验的关键设计是\"先把教师的 token logits 转换成 byte logits\"。为此论文提出两种一次性转换:Marginalize-It(近似,把首字节之后未匹配的尾部概率重新归一化)与 End-Of-Token(精确,在词表中加 `\u003Ceot>` 标记 token 边界)。前者是 baseline,后者保留教师分布的全部概率质量。\n\n把六组实验在八项 benchmark(多选 QA、语言生成、机器翻译)上跑出来的 scaling 曲线叠加后,趋势出乎意料。低算力阶段,Token-1B 蒸馏模型几乎在所有任务上领先;但随着 FLOPs 增长,token 曲线快速饱和,字节曲线反而保持陡峭爬升。论文的渐近预测是:End-Of-Token 蒸馏比 Token 蒸馏渐近准确率高约 4%,比 Marginalize-It 蒸馏高约 1.9%。\n\n更值得工程圈关注的两个数字:其一,End-Of-Token 蒸馏仅用 1\u002F6 的训练数据就能追平 Token 蒸馏的最终性能,这意味着同样的算力预算下,字节学生模型在数据采购和训练时间上的成本压到极致;其二,把词表从约 10 万 token 砍到 256 字节,蒸馏时不需要 top-k 截断,logit 存储成本降到约 1\u002F5,这对蒸馏流水线的工程开销是数量级影响。\n\n论文还把渐近预测与开源权重模型对照:End-Of-Token 蒸馏 1B 模型渐近平均精度可望比 Llama 3.2-1B、Gemma-3-1B-pt、Gemma 2B 分别高 6.5%、8.1%、2.1%。这一结论对应的是外推的 scaling law,并非当下真实 checkpoint 已超过这些模型——读者别把\"渐近潜力\"误读成\"今日成绩\"。\n\n我自己更看重方法论上的两点提示。第一,BPB(Bits-Per-Byte)这个常见 validation 指标,在跨 token 化方案、跨训练目标时并不能直接对比——同一 BPB 对应的下游准确率可以差出几十个百分点。后续选学生模型时,光盯 validation loss 不够,必须配套看下游任务的 scaling law。第二,蒸馏 logit 的存储成本长期被低估。10 万 token 的词表,top-100 截断下每条样本仍要存 100×float;而 256 字节词表几乎可以全量存。Meta FAIR 这一刀切下去,本质上把\"是否能跑大蒸馏\"的瓶颈,从算力挪到了 logit I\u002FO。\n\n所以这次研究更像是给字节路径\"正名\",而不是宣告 token 模型终结。Token 在中低算力预算下仍是更快的选择,字节模型的价值会随数据\u002F算力放大才显现。对国内做小模型蒸馏、字节级多模态预训练(比如想统一文本\u002F图像\u002F音频的字节表示)的团队,这是一份直接可借鉴的 scaling law 模板。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.12303v1","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21,24],{"id":12,"name":13,"slug":13,"description":14,"color":14},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b49648f9-963e-4082-8684-3d085b7358fe","quantization",{"id":25,"name":26,"slug":26,"description":14,"color":14},"4f214978-cac1-4f39-aa4b-f92a0d0934b7","transformer",[28],{"id":29,"lang":30,"title":31,"summary":32,"content":33},"bfe4ee6d-edc2-45b4-9d5f-cc3730383fba","en","Meta FAIR byte-distillation study: End-Of-Token asymptotically beats token distillation by 4%, using only 1\u002F6 of the data","A new Meta FAIR paper systematically compares byte-level and token-level distillation scaling: at 1B parameters and 1T bytes of training, End-Of-Token byte distillation asymptotically exceeds token distillation by 4% on averaged downstream accuracy, needs only one-sixth of the data, and cuts logit storage to about one-fifth.","## Why byte distillation looks slow but wins asymptotically\n\nByte-level distillation has long been outshone by token-level distillation at modest compute budgets. A new paper from Meta FAIR and the University of Washington, *Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models* (arXiv:2609.12303), flips the script: at scale, the byte student can outperform its token counterpart.\n\nThe team — led by Kalyani Marathe, Artidoro Pagnoni, Tomasz Limisiewicz, Margaret Li, Mike Lewis, Luke Zettlemoyer and Srinivasan Iyer — fixed layer-parameter counts (about 1.28B layer params for byte variants, 1.81B for token variants) and swept training up to roughly one trillion bytes. They crossed two axes simultaneously: tokenization (tokens, bytes, bytes with end-of-token) and objective (distillation vs. cross-entropy). To compare apples to apples they had to convert the teacher's token logits into byte logits in a single forward pass, introducing two variants: Marginalize-It (approximate, redistributes dropped probability mass) and End-Of-Token (exact, appends an `\u003Ceot>` marker so no mass is lost).\n\nAcross eight benchmarks spanning multiple-choice QA (ARC-Easy, ARC-Challenge, HellaSwag, PIQA), language generation (MBPP, Natural Questions) and machine translation (Flores), the token-1B curves lead at low FLOPs but saturate quickly. The byte curves start worse but climb at a steeper slope. Fitting power laws on downstream error vs. validation BPB, the authors extrapolate that distilled End-Of-Token-1B beats distilled Token-1B asymptotically by up to 4% on averaged downstream accuracy, and beats Marginalize-It distillation by about 1.9%.\n\nTwo numbers matter most for production pipelines. First, End-Of-Token distillation matches the final Token-distilled performance using only one-sixth of the training data — a data-cost win on top of a compute win. Second, by shrinking the vocabulary from roughly 100K tokens to about 256 bytes, the distillation pipeline no longer needs top-k truncation when dumping logits, dropping logit storage cost to roughly one-fifth.\n\nThe paper also benchmarks the asymptotic prediction against open-weight models: distilled End-Of-Token-1B would exceed Llama 3.2-1B by up to 6.5%, Gemma-3-1B-pt by 8.1%, and Gemma 2B by 2.1% on averaged tasks. These are extrapolation ceilings from fitted scaling laws, not numbers measured on today's checkpoints, so read them as \"what the trend predicts at convergence\" rather than \"what the model already does.\"\n\nThe methodological takeaway is sharp. Validation BPB, the lingua franca of pretraining, becomes misleading across tokenization schemes and objectives: the same BPB maps to very different downstream accuracy depending on whether the student operates on tokens or bytes. Picking a student on validation loss alone is no longer safe. Pair it with a downstream-error scaling law before signing off.\n\nThe second lesson is engineering-flavored. Logit-storage cost in distillation is widely underpriced. With a 100K-token vocabulary, even top-100 truncation stores 100 floats per position; with a 256-byte vocabulary you can keep the full distribution. Meta FAIR's result is in effect saying the bottleneck for \"can we run a really big distillation run\" is shifting from FLOPs to logit I\u002FO.\n\nFor Chinese teams training small distilled models or doing byte-level multimodal pretraining (e.g. unifying text, image, audio at the byte level), this paper is a directly usable scaling-law template. Byte models are not winning the next quarter, but the ceiling is higher and the data bill is smaller — a combination that compounds.","meta-fair-byte-distillation-token-ceiling-2026-09","2026-09-15T02:00:00Z","2026-09-15T03:06:58.941373Z","2026-09-15T03:06:58.941381Z",true,"agent",31,[42,51],{"slug":43,"tag_slug":43,"title_zh":44,"title_en":45,"intro_zh":46,"intro_en":47,"id":48,"is_active":38,"created_at":49,"modified_at":50},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":52,"tag_slug":52,"title_zh":53,"title_en":54,"intro_zh":55,"intro_en":56,"id":57,"is_active":38,"created_at":58,"modified_at":59},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":61},[62,67,72,77,82,87],{"id":63,"title":64,"news_slug":65,"published_at":66},"3d922c00-afcb-4f1c-a6d5-8f9d6c10c642","从 Kimi Linear 到 Kimi K3:MoE 推理效率战里被忽略的架构升级","kimi-k3-latentmoe-kda-attnres-nope","2026-07-30T00:30:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"b5638cab-a4d6-44ac-9230-32ed0a4cba9d","ARMT 把「记忆」焊进 Transformer:用恒定显存换无限上下文","armt-associative-recurrent-memory-transformer","2026-07-23T00:00:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"d2844cbb-b70e-469c-95b8-71cee8d735a6","给 Transformer 装上「CNN 鼻子」:用 0.01% 的参数量换 benchmark 普涨","transformer-cnn-nose-0-01-percent","2026-07-22T12:10:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"067a3f68-9bc8-4486-9715-5a391e537909","ETH 统一评测循环线性注意力：Kimi Delta Attention 损失最低","eth-zurich-clvr-kda","2026-07-12T04:00:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"29774f38-c361-4dca-b11a-c14df2fc84d9","HiLS 把\"无限上下文\"从口号变成数学:让稀疏注意力首次跑赢 Full Attention","hils-hierarchical-landmark-sparse","2026-07-07T14:00:00+00:00",{"id":88,"title":89,"news_slug":90,"published_at":91},"5c53c383-9727-4880-95e9-9fa752132b01","把混合注意力推到 head 级：HydraHead 用 7:1 LA\u002FFA 比实现 3:1 层混的长上下文性能","hydrahead-7-to-1-la-fa-head-mixed-attention","2026-06-20T16:14:00+00:00"]