[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-glm-flash-single-token-decisions":3,"topics-all":38,"news-related-523d3ae4-9c60-4bcb-8e6c-ea98e58c7f71":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"523d3ae4-9c60-4bcb-8e6c-ea98e58c7f71","一次前向一个决策:GLM-5.3-Flash 平替 Jev","Privatemode 开源免训练方法:prompt 预填答案位,读 GLM-5.3-Flash 首个 token 在各选项上的 logprob,一次前向返回带置信度的决策。28 个文本数据集上与专用模型 Jev 统计打平,还支持图像,百万次约 62 欧元。","\"这条工单派哪个团队?\"\"这条评论正面还是负面?\"——软件问大模型的多数问题不是要文章,而是要一个决策:选项固定、答案落在集合里,最好带置信度。Privatemode(Edgeless Systems)上周发布的方法把这个场景压到极致:GLM-5.3-Flash 一次前向、一个 token,吐出完整决策。HN 上这条帖子已攒 133 分、58 条评论。\n\n## 三个动作,零微调\n\n做法拆开只有三步。第一,把状态、问题和编号选项打包进 prompt,要求模型用选项序号回答;第二,让 prompt 结尾停在答案位——直接预填 `choice_index:`,模型下一个 token 必然是某个序号;第三,不读模型生成的内容,改读它在这一位置给所有选项序号分配的概率,归一化后取最大者。\n\n关键洞察:既然 JSON 的形状是我们自己定的,让模型生成整个 JSON 就是浪费——要的只是那个被约束住的\"类型化判断\"。全程不微调,用的就是官方原版 GLM-5.3-Flash。Jev(TypeSafe)和 Laya(Convai)是为决策专门训练的 System One 模型;Privatemode 的赌注是:通用 LLM 配上合适的采样姿势,不重训也能站上同一量级。\n\n## 29 个数据集的战报\n\n基准覆盖 29 个公开数据集,2 到 151 个选项,意图路由、情感、分类、蕴含、法律文本、扫描文档都有,英德双语。28 个文本数据集上,GLM-5.3-Flash 与 Jev 各赢 10 个、8 个差距在 1 个百分点内,中位差 0.7 个百分点偏向 Jev,统计不显著(p=0.64)。温度 0 下两次运行仍有最多 3.5% 漂移(批处理与浮点所致),更小差距都当噪声。\n\n选项数量的影响比选哪个系统更大。TREC 上从 6 选项细化到 42 选项,Jev 从 92.1% 掉到 85.6%,GLM-5.3-Flash 从 91.2% 掉到 79.6%,Laya 从 88.4% 掉到 51.2%。\n\n成本与延迟是诚实的:百万次决策 GLM-5.3-Flash 约 62 欧元、Jev 约 16 欧元,便宜四倍的是 Jev;从德国实测,前者 180 ms、后者 264 ms。边界也自曝:151 选项的 CLINC150 上它用两次请求拼出 87.5%(Jev 78.4%),但要 719 ms(Jev 249 ms);上限是 191 个选项——tokenizer 会把更大的数字拆成多 token。图像是专属能力:RVL-CDIP 1600 份扫描商务文档 16 分类,70.2%,Jev 与 Laya 都是纯文本模型无法作答。\n\n## 标签噪声与推理的代价\n\n两个发现值得单独说。其一,部分错误其实在标签里:banking77 约 17% 的样本两个答案都说得通(如 get_physical_card 与 order_physical_card),任何系统在该数据集的理论上限约 85% 而非 100%。其二,推理确实更准但按 token 计价:让同一个模型先推理再答,准确率全区间更高(89.9% vs 85.5% 到 82.0% vs 79.2%),代价是每次几百 token、百万次约 350 欧元——对比 62 欧元的单 token 方案,贵的不是模型,是\"想一下\"本身。\n\n## 所以呢\n\n这套方法的真正读者,是被专用决策模型挡在门外的人:一个 MIT 协议的 Python 库,配任何 vLLM 后端就能跑,换模型不用改代码。可复现性也罕见:基准方法学、冻结数据集、harness、聚合脚本全开源,博客里每个数字都能重算。下一个被\"通用模型 + 推理工程\"拉平的专用赛道会是哪个?评论区聊聊。","https:\u002F\u002Fwww.privatemode.ai\u002Fblog\u002Fsystem-one-from-glm-flash","3a432e87-47dc-491d-aec4-31ce455c9416",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"fca9258a-9430-455a-b95d-b9fae5e373a8","ai-inference",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"319a2646-6075-4b00-b7fb-fd4c77ca002a","en","One Token, One Decision: GLM-5.3-Flash as a Jev Alternative","Skip fine-tuning, read one token's logprobs: GLM-5.3-Flash ties dedicated decision model Jev on 28 text datasets, one forward pass per decision.","\"Which team should handle this ticket?\" \"Is this review positive or negative?\" Most of what software asks an LLM is not an essay but a typed decision: options fixed, answer confined to a set, ideally with a confidence attached. Privatemode (Edgeless Systems) pushed this scenario to its limit last week: GLM-5.3-Flash returns a full decision in a single forward pass and a single token. On Hacker News the post has gathered 133 points and 58 comments.\n\n## Three moves, zero fine-tuning\n\nThe recipe has three steps. First, pack state, question, and numbered options into the prompt as JSON, and ask the model to answer with an option index. Second, end the prompt inside the answer: prefill `choice_index:`, so the model's next token must be an index. Third, instead of reading the token the model emits, read the probabilities it assigned to all option indexes at that position, normalize, and take the max.\n\nThe core insight: since we already know the shape of the JSON, having the LLM predict the whole object is waste — what we want is its typed judgement under constraint. No fine-tuning, no distillation; the model runs exactly as it ships. The comparison targets, Jev (TypeSafe) and Laya (Convai), are purpose-trained \"System One\" decision models. Privatemode's bet: a general LLM with the right sampling posture can reach the same tier without retraining.\n\n## Results across 29 datasets\n\nThe benchmark covers 29 public labeled datasets, 2 to 151 options, spanning intent routing, sentiment, topic classification, moderation, entailment, QA, legal text, and scanned documents, in English and German. Across the 28 text datasets, GLM-5.3-Flash and Jev each win 10; 8 are within one percentage point; the median gap is 0.7 points in Jev's favor, not statistically significant (p=0.64). The authors also report that even at temperature 0, identical runs drift on up to 3.5% of answers, caused by batching and floating-point arithmetic — smaller differences are treated as noise.\n\nThe number of options matters more than the choice of system. On TREC, going from 6 to 42 options, Jev drops from 92.1% to 85.6%, GLM-5.3-Flash from 91.2% to 79.6%, Laya from 88.4% to 51.2%.\n\nCost and latency are reported honestly. A million decisions cost about EUR 62 with GLM-5.3-Flash versus EUR 16 with Jev — the dedicated model is four times cheaper; measured from Germany, latency was 180 ms versus 264 ms. The limits are self-disclosed too: on CLINC150 with 151 options, GLM-5.3-Flash splits the query into two requests and reaches 87.5% accuracy (Jev 78.4%), but takes 719 ms (Jev 249 ms); the ceiling is 191 option indexes, because GLM-5.3-Flash's tokenizer splits larger numbers into multiple tokens. Images are the exclusive capability: on RVL-CDIP, 1,600 scanned business documents in 16 classes, it scores 70.2%, while Jev and Laya are text-only and cannot answer at all.\n\n## Label noise and the price of reasoning\n\nTwo findings deserve their own section. First, some of the remaining errors are in the labels: about 17% of banking77 examples admit two defensible answers (e.g. get_physical_card vs order_physical_card), so the theoretical ceiling there is about 85%, not 100%. Second, reasoning is more accurate but priced per token: letting the same GLM-5.3-Flash reason before answering lifts accuracy in every band, 89.9% vs 85.5% (two options) up to 82.0% vs 79.2% (21-80 options), at the cost of hundreds of tokens per decision and about EUR 350 per million — against EUR 62 for the single-token setup. The expensive part is not the model; it is the act of thinking itself.\n\n## So what\n\nThe real audience for this method is anyone locked out by dedicated decision models. Jev is a closed API with a fixed model; here is an MIT-licensed Python library that runs against any vLLM backend, fetches token IDs from the server, and switches models without code changes. Reproducibility is rare in its class: the benchmark methodology, frozen dataset specs, test harness, and aggregation scripts are all open, and every number in the post can be recomputed. Which specialized niche will \"general model + inference engineering\" flatten next? Drop your pick in the comments.","glm-flash-single-token-decisions","2026-09-28T13:11:55Z","2026-09-28T13:11:57.884987Z","2026-09-28T13:11:57.884997Z",true,"agent",1,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"c6dc2edc-1a4a-46b9-85e6-f1c8ef32faa6","DeepSeek DSpark 跑进 Apple Silicon：mlx-dspark 给出首个原生 MLX 移植,逐字节保持原模型输出","mlx-dspark-apple-silicon","2026-07-04T12:00:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"62e17707-e36f-45f6-8749-0d0370382cbd","llm-d：混合 GPU 集群 3-5 倍加速，KV Cache 感知路由","llm-d-mixed-gpu-kv-cache-aware-routing","2026-06-23T22:00:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"267a9244-2ed7-4034-86cb-be4cbd196a08","NVIDIA 开源 Nemotron 3 Super：Latent MoE 如何让 120B 模型「省着跑」","nvidia-nemotron-3-super-120b-latent-moe","2026-05-18T04:05:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"2dbc7c0f-083a-47d4-ba9a-8a7d66b22ae0","亚马逊八阶段配方:后训练让 GLM-4.5-Air 反超官方版","amazon-rufus-air-post-training-recipe","2026-09-25T23:12:24+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"e4b3903e-c65f-46cb-90ae-81502eb8cdd9","CliffCompaction开源:只删不改的会话压缩,长程Agent成本砍半","cliffcompaction-truncate-only-agent-compaction","2026-09-23T17:10:56+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"3559e613-9558-48e1-ab20-f53b62796363","让每个 token 用上全部专家:高德 IntBMoE 解耦参与度、计算与显存,60ms 服务数亿用户","intbmoe-full-participation-block-moe","2026-09-21T13:01:54+00:00"]