[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-jeff-open-decision-models-0-8b":3,"topics-all":38,"news-related-ccce6dfe-776f-4cd8-9605-9163daea4627":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"ccce6dfe-776f-4cd8-9605-9163daea4627","Jeff 开源决策模型:2B 追平 Jev","独立项目 Jeff 把 Qwen3.5 与 Gemma 4 微调成三个 Apache 2.0 决策模型,单次前向输出选项概率,0.8B 版 22 毫秒一次决策;五基准综合 83.1 追平 Jev,分类任务反超,单卡半小时微调把语音导航准确率从 31.7% 拉到 95.8%。","开源圈这两天出现了一个很容易被忽略、但值得仔细看的方向:把大模型从一个\"什么都会聊两句\"的通用助手,收窄成一个\"一次前向、一次决策\"的分类器。继 TypeSafe 的 Jev 之后,独立开发者项目 Jeff 把 Qwen3.5 和 Gemma 4 微调成三个小型决策模型,权重以 Apache 2.0 放出,单个 0.8B 模型做一次决策只要约 22 毫秒,在家用显卡和 Apple M4 Max 上都能跑。\n\n## 它到底做什么\n\nJeff 的思路是:你用自然语言描述一个情境,列出候选项,模型从**单次前向传播**直接输出每个选项的校准概率。没有生成文本、没有解析,回答的是\"该选哪个\"而不是\"为什么\"。官方对它的定位说得很直白:这是一个分类器,不是规划器。推理放在代码里,判断交给模型。\n\n这个格式与 TypeSafe 的 Jev 完全兼容,但项目方明确声明与 TypeSafe 没有任何隶属关系,训练代码基于 MIT 协议的开源方案 AutoJev 改造。\n\n## 数字层面:小模型打赢了大模型(部分)\n\n先说最有说服力的部分。在五个公开基准共 4599 道题上,Jeff-Qwen3.5-2B 综合得分 83.1,追平了 Jev 公布的 83.0;0.8B 版本也有 79.1。更细的拆解很有信息量:在 Financial PhraseBank 上 Jeff 三个版本全部超过 96,压过 Jev 的 77.0;RAGTruth 上 Jeff-2B 的 88.9 与 AutoJev-27B 并列最高。但在推理密集的 BBH、JudgeBench、JevBench hard 上,小模型差距明显——JevBench hard 上 Jev 拿到 73.3,Jeff-2B 只有 53.3。\n\n这个结果很符合直觉:分类、接地的活儿,微调过的小模型足够用;需要多步推理的活儿,参数量就是硬约束。\n\n## 全程本地硬件,没有云 GPU\n\n整个项目在一张 RTX PRO 6000 工作站显卡上完成:0.8B 训练约 2 小时,2B 约 3.5 小时。合成训练数据由开源模型 Qwen3.8-Flash-Next 在两台 DGX Sparks 上生成,闭环模型只用于抽查合成数据质量,训练数据里没有闭源模型的输出。训练细节也开源在仓库里:全参数微调、单 epoch、批 256、对选项字母做交叉熵,再用一个拟合温度做校准。\n\n## 微调半小时,准确率从 31.7% 到 95.8%\n\n对实际使用者最有价值的一条数据:如果零样本效果不够,在自建数据上短微调一轮,收益极大。官方的语音导航微调只用了约 1.1 万条应用内样本、单卡半小时,留存准确率从 31.7% 拉到 95.8%,单次决策延迟约 40 毫秒。另一个极端案例:用 60 万条 Lichess 棋局(Stockfish 标注)微调 0.8B 模型 3.5 小时,在 1000 道留存棋题上解题率从 15.5% 提升到 55.8%,一块 GPU 能同时下约 600 盘人类快棋。仓库还提供了 Doom、Frogger、Pac-Man 三个零样本游戏测试,0.8B 在 Doom 里追平了手写规则机器人。\n\n## 局限与\"所以呢\"\n\n项目自己列的边界很诚实:每个问题最多 26 个选项(位置 27 以后基本不会被选中)、只支持英文和文本、小模型不做多步推理、2B 版玩游戏反而不如 0.8B。另外注意:基准成绩不预测游戏表现,基准分更低的 0.8B 是游戏甜点位。\n\n这件事的行业意义在于:Agent 框架里大量\"路由、意图识别、审核分类、指令抽取\"的环节,并不需要动用千亿参数的大模型。一张工作站显卡加几小时微调,就能得到一个 22 毫秒级、可校准、Apache 2.0 商用友好的决策单元。当推理成本成为 Agent 规模化的瓶颈时,\"大模型管规划、小模型管决策\"的分工正在从论文变成可复制的工程路径。下次你在 Agent 里写下一个 if-else 链,可以想一想:这个分支判断,是不是也可以交给一个 0.8B 的模型。\n\n(项目地址:[firelex\u002Fjeff](https:\u002F\u002Fgithub.com\u002Ffirelex\u002Fjeff),权重在 Hugging Face mstrasser 名下,代码 MIT、权重 Apache 2.0)","https:\u002F\u002Fgithub.com\u002Ffirelex\u002Fjeff","998df6db-96e6-4b8e-8be1-cfa00a6cd177",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"fca9258a-9430-455a-b95d-b9fae5e373a8","ai-inference",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"d6a8b36c-2544-49ea-a3e5-f4fa1a8e6f98","en","Jeff: Open Decision Models Tie Jev, 0.8B at 22ms","Jeff fine-tunes Qwen3.5\u002FGemma 4 into Apache 2.0 decision models: one forward pass, 0.8B at 22 ms, 2B ties Jev at 83.1; a half-hour tune hit 95.8%.","An easy-to-miss but worth-examining direction emerged in the open-source community: narrowing a general-purpose LLM assistant into a \"one forward pass, one decision\" classifier. Following TypeSafe's Jev, the independent project Jeff fine-tuned Qwen3.5 and Gemma 4 into three small decision models, released under Apache 2.0. A single 0.8B model makes a decision in about 22 ms, running on consumer GPUs and the Apple M4 Max.\n\n## What it actually does\n\nThe idea: you describe a situation in plain words and list candidate options; the model outputs a calibrated probability for each option from a single forward pass. No generated text, no parsing — the answer is \"which one\", not \"why\". The README is blunt about positioning: it is a classifier, not a planner. Reason in code, decide with the model.\n\nThe request format is fully compatible with TypeSafe's Jev, though the project explicitly states it has no affiliation with TypeSafe; the training code builds on the MIT-licensed open recipe AutoJev.\n\n## The numbers: small models beat the big one (partly)\n\nOn five public benchmarks totaling 4,599 questions, Jeff-Qwen3.5-2B scored 83.1 overall, matching Jev's published 83.0; the 0.8B reached 79.1. The breakdown is informative: on Financial PhraseBank all three Jeff variants scored above 96, beating Jev's 77.0; on RAGTruth, Jeff-2B's 88.9 tied AutoJev-27B for the best. But on reasoning-heavy BBH, JudgeBench and JevBench hard, the small models lag clearly — Jev scored 73.3 on JevBench hard while Jeff-2B got 53.3.\n\nThe pattern matches intuition: for classification and grounding, fine-tuned small models are enough; for multi-step reasoning, parameter count is a hard constraint.\n\n## All local hardware, no cloud GPUs\n\nThe whole project ran on one RTX PRO 6000 workstation GPU: the 0.8B trained in about 2 hours, the 2B in about 3.5. Synthetic training data was written by the open model Qwen3.8-Flash-Next on two DGX Sparks; a closed model was only used to spot-check sample quality, and no closed-model output is in the training data. Training details are open too: full-weight fine-tuning, one epoch, batch 256, cross-entropy over option letters, plus a fitted temperature for calibration.\n\n## Half an hour of fine-tuning: 31.7% to 95.8%\n\nThe most practical data point: when zero-shot is not enough, a short fine-tune on your own examples pays off enormously. The official voice-navigation fine-tune used about 11k in-app examples and half an hour on one GPU, moving held-out accuracy from 31.7% to 95.8%, at about 40 ms per decision. Another extreme case: fine-tuning the 0.8B on 600,000 Lichess positions labelled by Stockfish for 3.5 hours lifted puzzle-solving from 15.5% to 55.8% on 1,000 held-out puzzles, and one GPU keeps up with about 600 human blitz games at once. The repo also ships zero-shot game tests — Doom, Frogger, Pac-Man — where the 0.8B matched the hand-coded rule bot in Doom.\n\n## Limits, and the \"so what\"\n\nThe project is honest about boundaries: at most 26 options per question, English and text only, no multi-step reasoning at 0.8B–2B, and the 2B plays games worse than the 0.8B. Note that benchmark scores don't predict gameplay: the lower-scoring 0.8B is the sweet spot.\n\nThe industry takeaway: massive amounts of \"routing, intent recognition, moderation, command extraction\" steps inside agent frameworks don't need a hundred-billion-parameter model. One workstation GPU plus a few hours of fine-tuning yields a 22-ms, calibratable, Apache-2.0 commercial-friendly decision unit. As inference cost becomes the bottleneck of agent scaling, the division of labor — big models plan, small models decide — is turning from papers into a reproducible engineering path. Next time you write an if-else chain inside an agent, ask yourself: could this branch decision go to a 0.8B model?\n\n(Project: [firelex\u002Fjeff](https:\u002F\u002Fgithub.com\u002Ffirelex\u002Fjeff), weights under mstrasser on Hugging Face; code MIT, weights Apache 2.0)","jeff-open-decision-models-0-8b","2026-09-29T15:10:00Z","2026-09-29T15:09:44.321909Z","2026-09-29T15:09:44.321917Z",true,"agent",433,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"523d3ae4-9c60-4bcb-8e6c-ea98e58c7f71","一次前向一个决策:GLM-5.3-Flash 平替 Jev","glm-flash-single-token-decisions","2026-09-28T13:11:55+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"c6dc2edc-1a4a-46b9-85e6-f1c8ef32faa6","DeepSeek DSpark 跑进 Apple Silicon：mlx-dspark 给出首个原生 MLX 移植,逐字节保持原模型输出","mlx-dspark-apple-silicon","2026-07-04T12:00:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"62e17707-e36f-45f6-8749-0d0370382cbd","llm-d：混合 GPU 集群 3-5 倍加速，KV Cache 感知路由","llm-d-mixed-gpu-kv-cache-aware-routing","2026-06-23T22:00:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"267a9244-2ed7-4034-86cb-be4cbd196a08","NVIDIA 开源 Nemotron 3 Super：Latent MoE 如何让 120B 模型「省着跑」","nvidia-nemotron-3-super-120b-latent-moe","2026-05-18T04:05:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"9dc3bd1e-95da-4c07-b352-900688d1af44","TGRL把温度分组变成训练信号:RLVR训练提速36%","tgrl-temperature-grouped-rl","2026-10-01T19:02:31+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"3e28cfaf-74aa-4521-8eba-37332fe93901","7块ESP32跑1.58-bit Qwen:功耗1.53瓦","esp32s3-bitnet-llm-cluster","2026-09-30T13:11:01+00:00"]