[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-kalibench-cybersecurity-cli-runtime-rewards":3,"topics-all":42,"news-related-4977d1f4-c8f7-480c-aaa1-ec01d69f44d8":61},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":27,"news_slug":34,"published_at":35,"created_at":36,"modified_at":37,"is_published":38,"publish_type":39,"image_url":40,"view_count":41},"4977d1f4-c8f7-480c-aaa1-ec01d69f44d8","8B 拿 SFT+GRPO 打 685B MoE:KaliBench 把 LLM 网络安全工具调用拆到命令行级","网络安全里的 LLM 评估一直停留在选择题和端到端 Agent 上,真正能不能写出可执行的 Kali Linux 命令始终没人直接测。NeurIPS 2026 接收的 KaliBench(arXiv:2610.02206)把这个缺口填上 —— 8,504 条手工核验的查询-命令对、1,642 个工具、23 个能力维度、5 个安全阶段;24 个开源模型在无提示设定下没有任何一个超过 42% 的命令精确率。配套训练出的 8B 模型 RedSage-K,经 SFT+GRPO 把总分从 71.7 推到 79.2,距 685B 的 DeepSeek-V3.2 仅差 1 个点。","网络安全这条赛道,LLM 评估长期分两极:一边是 CTF 式知识题,测的是「知道不知道」;另一边是端到端 Agent 跑通率,测的是「整条链路能不能走完」。真正卡在中间、决定一名安全工程师能不能用上 AI 的关键能力 —— 把自然语言请求翻译成一行能在终端里执行正确的 nmap、sqlmap、burp 命令 —— 反而没人直接测过。原因不复杂:命令行对参数顺序、flag 别名、键值绑定极其敏感,微小的拼写错误就会让整条命令失效,而打分又必须脱离模型自描述、与真实工具行为对齐。\n\nKhalifa University 与西澳大学团队把这个缺口补上。10 月 1 日挂到 arXiv 的 2610.02206(NeurIPS 2026 Evaluations and Datasets Track)带来 KaliBench:一个面向 Kali Linux 的细粒度网络安全工具调用基准,数据来自官方工具手册,经 LLM 校验 + 沙盒终端执行 + 人工复核 + 语义去重四道工序,最终留下 8,504 条查询-命令对(3,504 训练 \u002F 5,000 测试),覆盖 1,642 个工具、23 个能力维度与 5 个安全阶段(侦察与初始访问、漏洞分析、利用与投递、后期渗透、防御分析与报告)。\n\n## 难的不是选哪个工具,而是写对参数\n\nKaliBench 把评估拆成四块:工具选择、可选参数 F1(按 alias-aware 匹配)、位置参数 F1(多重集匹配),以及严格意义上的「命令精确正确」(允许文档化的别名与可选 flag 重排,但位置参数顺序保持不变)。在无任何工具提示的「unrestricted」设定下,24 个开源权重模型跑出的最强成绩是 GLM-5.2(753B)的 41.3%,DeepSeek-V3.2(685B 总参 \u002F 37B 激活 MoE)33.0%,Qwen3-Coder-Next(80B \u002F 3B MoE)26.2%,中位区间大多在 20%-30% 之间 —— 没有任何一个模型超过 42% 的精确率。给 20 个候选工具的「restricted」设定把均值抬到 28.3%,而给到目标工具及文档的「hinted」设定直接抬到 73.1%,差距接近 50 个点。\n\n这条曲线说明一件事:工具选择并不是瓶颈。RedSage-K(论文中作者把自训模型叫 Kali-SFT+GRPO,这里沿用项目页的 RedSage-K 命名)在工具选择这一项拿到了 77.9% 的准确率,与 685B 的 DeepSeek-V3.2(75.8%)同档;真正的差距在可选参数 F1 与位置参数 F1 上 —— 8B 模型在可选参数 alias 还原、键值绑定、参数顺序保持这些细节上仍明显落后。这意味着对工程团队来说,KaliBench 给出的不是「网络安全能不能交给 LLM」的二元答案,而是「工具调用接口层的具体哪一档具体在哪一格」的可量化诊断。\n\n## 训练信号不用现成模型,而是 ground-truth 命令\n\nKaliBench 的另一层价值在于「runtime-free verifiable rewards」。数据集构建阶段已对每条样本做过命令正确性核验,在训练阶段,确定性打分可以同时给工具选择、可选参数、位置参数、命令级精确率、输出格式各维度的奖励,无需在训练循环中执行模型生成的命令。这与 RLVR(Reinforcement Learning with Verifiable Rewards)近一年成为热门训练范式的趋势是契合的,但 KaliBench 的设计把它推到了「网络安全工具调用」这个尚未被 RLVR 渗透的细分领域。\n\n作者团队用 RedSage-Ins 8B 作为基座,先做监督微调(SFT),再做 GRPO 强化学习,得到 RedSage-K 的三个变体(纯 SFT \u002F 纯 GRPO \u002F SFT+GRPO)。在三个评估设定的平均总分上,8B 基线 71.7 → 纯 SFT 77.4 → SFT+GRPO 79.2,而 685B \u002F 37B 激活的 DeepSeek-V3.2 是 80.2。1.0 分的差距,在「8B vs 685B \u002F 18× 参数量」的悬殊对照下,几乎是平起平坐。开源权重的 RedSage-K(SFT-GRPO)模型权重同步发在 Hugging Face(RISys-Lab\u002FRedSage-K-SFT-GRPO),训练与评估代码在 GitHub(RISys-Lab\u002FKaliBench),数据本身也在 HF 集合中公开。\n\n## 商业模型把天花板抬到哪里\n\n论文 Table 2 给出闭源系统的对照:GPT-5.6-Sol 在 5,000 条评测里答出 4,988 条,精确率 61.83%(覆盖全部样本的精确率 61.68%);Codex CLI(GPT-5.5, xhigh 推理强度)答 4,633 条,精确率 55.77%(全样本 51.68%);Claude Opus 5 答 3,675 条,精确率 59.89%(全样本 44.02%)。注意这里的全样本精确率把「拒绝回答 \u002F 未生成」也算错 —— 也就是说在网络安全这种可能触发安全护栏的领域,模型敢不敢答、答不答得动,本身就是分数的一部分。这一组数字与开源权重那段一起读,KaliBench 给出的图景是:闭源模型路线拿到 60%+ 的精确率,开源路线卡在 30-40%,但经过针对性训练后,8B 开源模型能把距离压到 1.0 个百分点。\n\n对做网络安全工具的工程团队,KaliBench 把「自动驾驶工具推理」这件事从一句口号变成了可量化、可训练、可对照的实验台 —— 8B 模型跑 79.2 分,DeepSeek-V3.2 跑 80.2 分,差 1 个点这件事本身,比「LLM 不能用」或「LLM 可以用」更接近真实工程现场。\n\n参考:arXiv:2610.02206(risys-lab.github.io\u002FKaliBench);GitHub 仓库 RISys-Lab\u002FKaliBench;HF 集合 RISys-Lab\u002Fkalibench-datasets-and-models;GPT-5.6-Sol 与 Codex CLI、Claude Opus 5 数字见论文 Table 2","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.02206","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21,24],{"id":12,"name":13,"slug":13,"description":14,"color":14},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"1fcfaaf2-67de-43d3-9e35-5784852fec60","ai-safety",{"id":19,"name":20,"slug":20,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":22,"name":23,"slug":23,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":25,"name":26,"slug":26,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[28],{"id":29,"lang":30,"title":31,"summary":32,"content":33},"c43adcb6-d502-4f94-a3b2-09876ebc7910","en","8B with SFT+GRPO vs 685B MoE: KaliBench drills LLM cybersecurity tool use down to the command line","Cybersecurity evaluation of LLMs has long been stuck on two extremes — CTF-style knowledge questions and end-to-end agentic run-throughs — with no direct measurement of whether the model can write an executable Kali Linux command. NeurIPS 2026-accepted KaliBench (arXiv:2610.02206) fills the gap: 8,504 manually-verified query-command pairs, 1,642 tools, 23 capability dimensions, 5 security phases. None of 24 evaluated open-weight models exceeds 42% exact-command accuracy in the unrestricted setting. The accompanying 8B RedSage-K, trained with SFT+GRPO, lifts its total score from 71.7 to 79.2 — landing within 1.0 point of DeepSeek-V3.2's 685B MoE.","Cybersecurity evaluation of LLMs has long been stuck on two extremes: CTF-style knowledge questions that ask \"does the model know about it\" and end-to-end agentic run-throughs that ask \"can the model finish the loop\". The capability sitting in the middle that actually decides whether a security engineer can put AI to work — translating a natural-language request into a single nmap, sqlmap, or burp command that actually executes correctly on the terminal — has not been measured directly. The reason is not subtle: command lines are intolerant of argument order, flag aliases, and key-value bindings, and grading has to escape the model's self-description and align with the real tool's behavior.\n\nA team from Khalifa University and the University of Western Australia closed that gap. Posted on arXiv on October 1 (2610.02206, NeurIPS 2026 Evaluations and Datasets Track), KaliBench is a fine-grained cybersecurity tool-use benchmark targeting Kali Linux. The data is grounded in official tool documentation, goes through LLM verification plus sandboxed terminal execution plus human review plus semantic deduplication, and ends with 8,504 query-command pairs (3,504 for training, 5,000 for evaluation) covering 1,642 tools, 23 capability dimensions, and 5 security phases (reconnaissance & initial access, vulnerability analysis, exploitation & payload delivery, post-exploitation & lateral movement, defensive analysis & reporting).\n\n## The hard part is not picking the tool, it is writing the right arguments\n\nKaliBench splits evaluation into four pieces: tool selection, optional-argument F1 (alias-aware matching), positional-argument F1 (multiset matching), and strict \"exact-command correctness\" (which permits documented aliases and reordered optional flags but preserves positional argument order). Under the unrestricted setting with no tool hints, the strongest of 24 open-weight models is GLM-5.2 (753B) at 41.3%; DeepSeek-V3.2 (685B total \u002F 37B active MoE) reaches 33.0%, Qwen3-Coder-Next (80B \u002F 3B MoE) 26.2%, and the median band sits mostly in the 20-30% range — no model exceeds 42% exact-command accuracy. The restricted setting that provides 20 candidate tools pushes the average up to 28.3%; the hinted setting that hands the model the target tool plus its documentation jumps straight to 73.1%, a spread of close to 50 percentage points.\n\nThat curve says one thing: tool selection is not the bottleneck. RedSage-K (the paper names its self-trained model Kali-SFT+GRPO; the project page calls it RedSage-K, the name used here) lands at 77.9% tool selection, in the same band as DeepSeek-V3.2's 75.8%; the real gap lives in optional-argument F1 and positional-argument F1 — the 8B model still falls behind on optional-parameter alias recall, key-value bindings, and argument order preservation. For engineering teams this means KaliBench delivers not a binary answer to \"can LLMs do cybersecurity\" but a quantifiable diagnostic of which exact row of the tool-call interface each model sits in.\n\n## Training signal from ground-truth commands, not from running models\n\nThe other value of KaliBench is \"runtime-free verifiable rewards\". The dataset construction stage has already verified command correctness for every sample; during training, deterministic scoring can return rewards for tool selection, optional arguments, positional arguments, command-level exact correctness, and output format — all without executing model-generated commands inside the training loop. This fits the trend of verifiable-reward RL (RLVR) becoming a popular training paradigm over the past year, but KaliBench pushes that design into a sub-area — cybersecurity tool calling — that RLVR has not yet penetrated.\n\nThe authors take RedSage-Ins 8B as the base, do supervised fine-tuning first, then GRPO reinforcement learning, to obtain three RedSage-K variants (SFT only, GRPO only, SFT+GRPO). On the average total score across the three evaluation settings, the 8B baseline 71.7 → SFT only 77.4 → SFT+GRPO 79.2, while DeepSeek-V3.2 at 685B \u002F 37B active sits at 80.2. A 1.0-point gap with a roughly 18× parameter-count spread is effectively a tie. The open-weight RedSage-K (SFT-GRPO) is released on Hugging Face (RISys-Lab\u002FRedSage-K-SFT-GRPO), the training and evaluation code is on GitHub (RISys-Lab\u002FKaliBench), and the data itself is published in the HF collection.\n\n## Where the proprietary systems push the ceiling\n\nPaper Table 2 supplies the closed-source comparison: GPT-5.6-Sol answers 4,988 of the 5,000 evaluation prompts with 61.83% exact-command accuracy on answered (61.68% over all 5,000); Codex CLI (GPT-5.5, xhigh reasoning) answers 4,633 with 55.77% on answered (51.68% over all); Claude Opus 5 answers 3,675 with 59.89% on answered (44.02% over all). Note that the all-5,000 number counts refusal-to-answer and non-generation as failures — in a cybersecurity domain that may trigger safety guardrails, whether the model dares to answer is itself part of the score. Read together with the open-weight block, KaliBench's picture is: the closed-source lane clears 60%+ exact-command accuracy, the open-weight lane is stuck in the 30-40% band, but targeted training lets an 8B open-weight model close the gap to 1.0 percentage point.\n\nFor teams building cybersecurity tooling, KaliBench turns \"autonomous tool reasoning\" from a slogan into a quantifiable, trainable, comparable experimental surface — an 8B model scoring 79.2 versus DeepSeek-V3.2 scoring 80.2, a 1-point gap, is closer to actual engineering reality than either \"LLMs can't do this\" or \"LLMs can do this\".\n\nReferences: arXiv:2610.02206 (risys-lab.github.io\u002FKaliBench); GitHub repository RISys-Lab\u002FKaliBench; Hugging Face collection RISys-Lab\u002Fkalibench-datasets-and-models; GPT-5.6-Sol, Codex CLI, and Claude Opus 5 numbers from paper Table 2.","kalibench-cybersecurity-cli-runtime-rewards","2026-10-03T05:00:00Z","2026-10-03T05:06:33.531819Z","2026-10-03T05:06:33.531831Z",true,"agent","https:\u002F\u002Frisys-lab.github.io\u002FKaliBench\u002Fassets\u002Fsocial-card.png",1,[43,52],{"slug":44,"tag_slug":44,"title_zh":45,"title_en":46,"intro_zh":47,"intro_en":48,"id":49,"is_active":38,"created_at":50,"modified_at":51},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":53,"tag_slug":53,"title_zh":54,"title_en":55,"intro_zh":56,"intro_en":57,"id":58,"is_active":38,"created_at":59,"modified_at":60},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":62},[63,68,73,78,83,88],{"id":64,"title":65,"news_slug":66,"published_at":67},"9ebb888c-dfe7-416a-9940-a913527d4f73","AI Agent 的失败比成功更值钱:5 万对错误诊断数据,修正通过率 18.4%→51.1%","agent-error-dataset","2026-10-01T15:11:08+00:00",{"id":69,"title":70,"news_slug":71,"published_at":72},"e98cf9a2-8348-4142-9a56-c11166774798","SpeakerMem-R1:多方对话记忆,分清谁说了什么","speakermem-r1-multi-party-memory","2026-09-24T19:05:00+00:00",{"id":74,"title":75,"news_slug":76,"published_at":77},"3bcb0e1d-99ea-4fae-9bdd-b6b625aabf10","代码 agent 8 成都在骗你:12 模型实测揭晓","overclaimbench-llm-agents","2026-09-21T07:00:00+00:00",{"id":79,"title":80,"news_slug":81,"published_at":82},"089195fb-7fe5-4ba9-a4bc-8e356fe5e923","BAAI把1000个GitHub仓库蒸馏成5000个技能,科研agent奖牌率31%冲到73%","baai-disco-repo-to-skill-library","2026-09-03T17:07:35+00:00",{"id":84,"title":85,"news_slug":86,"published_at":87},"e8965513-b56f-475b-b15f-22a5ea2d2a4e","Agent 取代人成为 HF Hub 一号用户:Claude Code 占 44.4%,还有一次 4.5 天未察觉的入侵","hf-hub-agent-user-claude-code-4-5-day-intrusion","2026-08-21T08:00:00+00:00",{"id":89,"title":90,"news_slug":91,"published_at":92},"c94766df-827e-4e4e-a006-b6639ec76722","DeepSeek V4-Flash-0731 转正观察:权重不动,后训练把 Agent 分数打到 V4-Pro 之上","deepseek-v4-flash-0731-agent-benchmark-official-aug2026","2026-08-01T02:00:00+00:00"]