[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-gcc-rejects-llm-contributions-15-line-threshold":3,"news-related-ed8ef087-9dc6-4295-bbdd-d0d2c92977d3":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"ed8ef087-9dc6-4295-bbdd-d0d2c92977d3","GCC 拒绝 LLM 生成的实质性贡献:开源基础设施开始为 AI 代码划红线","GCC 指导委员会正式采纳 AI 贡献政策,拒绝包含约 15 行以上 LLM 生成内容的代码或文本补丁,但允许用 LLM 做研究、审查与 bug 分析。这条规则首次把\"AI 生成代码是否算贡献\"从社区讨论推到编译器级别的硬约束。","## GCC 拒绝 LLM 生成的实质性贡献:开源基础设施开始为 AI 代码划红线\n\n刚刚,GCC 指导委员会正式宣布采纳 AI 政策工作组推荐的 AI 贡献政策,把\"AI 生成代码能不能算合法贡献\"从社区争论拉到了编译器级别的硬约束。\n\n### 政策的核心边界\n\n新规明确:GCC 将拒绝任何\"包含 LLM 生成的内容或源自 LLM 生成内容的具有法律意义的贡献\"。所谓\"具有法律意义\",在 GCC 的定义里门槛并不高——大约 **15 行代码或文本**就足以构成具有版权意义的贡献。低于这个门槛的微小修改、纯风格调整则不在限制范围内。\n\n值得注意的是,政策并非全面封杀 LLM。维护者仍可以**自主选择**接受由 LLM 生成的具有法律意义的测试用例,因为测试代码通常不承担\"作品\"角色,合规风险更可控。同时,使用 LLM 做研究、分析、Bug 发现与报告、补丁审查等工作流是被允许的——只要最终提交到仓库的成果里**不包含** LLM 输出。\n\n### 为什么这条规则重要\n\nGCC 是 Linux 内核、GNU 工具链、嵌入式与高性能计算生态的编译底层。把\"AI 生成代码\"挡在门外的决定,有三个层面的含义:\n\n**1. 法律层面的现实压力**。随着 Anthropic、OpenAI、Midjourney 等厂商被多次卷入版权诉讼,项目维护者越来越担心自己成为被告链的一环——上游贡献可能源自受版权争议训练的模型,下游商业用户使用编译器产物时,法律风险会沿调用链传导。GCC 选择\"宁严勿松\",是为了让贡献链的法律责任清晰可追溯。\n\n**2. 代码质量的工程判断**。LLM 生成的代码常带有\"看起来对、跑起来也像对、但维护起来是黑洞\"的特征:它能模仿开源风格、堆叠 API 调用,却难以承担 GCC 这种需要长期演进、对生成代码质量、ABI 兼容性、平台支持面负责的项目。15 行这条线划得保守,但反映了维护者对 LLM 在大尺度系统软件里可靠性的不信任。\n\n**3. 治理边界的示范效应**。Linux 内核社区在 2025 年就发起过类似的 RLAIF 政策讨论,围绕\"是否标注 AI 协助代码\"反复拉锯。GCC 作为 GNU 旗舰项目,这次给出的是**拒绝式**而非**标注式**的答案,后续可能影响 systemd、binutils、glibc 等依赖链上游的跟进态度。\n\n### 仍可用的 LLM 工作流\n\nGCC 并没有把开发者关在门外。**研究阶段、bug 排查、补丁 review、测试用例生成**这些场景,LLM 都是被默许甚至欢迎的工具。这意味着贡献者仍然可以用 AI 加速\"理解代码—定位问题—构思方案\"的环节,只是不能把 AI 的输出**直接当成提交物**。\n\n这是开源世界正在形成的共识:**AI 是放大器,不是合著者**。\n\n### 行业影响\n\n对国内做编译器、操作系统、数据库等基础设施的开源项目,GCC 的政策是一个值得参照的样板:\n\n- **法律风险**:接入 LLM 工具链的厂商,需要明确内部生成代码的\"贡献资格\"和\"署名规则\"。\n- **人才策略**:参与 GCC、Linux kernel、Rust 等上游的工程师,使用 Copilot\u002FClaude Code 等工具的方式需要重新培训。\n- **模型评估**:能输出\"看起来像合著者\"质量代码的模型,在企业场景里反而比\"工具型\"模型更危险——它会模糊人机责任边界。\n\n### 所以呢\n\nGCC 的政策不是反 AI,而是**给 AI 在关键基础设施中的角色定调**。它把 LLM 锁回了\"加速器\"的位置,不让它变成\"作者\"。这对整个开源生态来说是健康的:开源软件最珍贵的资产是**可追溯的责任链**,而 AI 生成代码的最大弱点恰好也是责任不清。\n\n未来一年,Kernel、Python、Curl、Kubernetes 这些头部项目,大概率会跟进或微调类似规则。对开发者来说,接受这个现实比抵抗它更划算——把 LLM 用在脑力劳动而非署名劳动,既是合规的,也是诚实的。","https:\u002F\u002Fwww.solidot.org\u002Fstory?sid=84958","d59894d3-308e-4fd8-8865-86dc1eeac4a2",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"1fcfaaf2-67de-43d3-9e35-5784852fec60","ai-safety",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"a100a6c9-be5d-445b-b3f1-f04a1644e250","en","GCC rejects substantive LLM contributions: a red line for AI code","The GCC steering committee has formally adopted an AI contribution policy, refusing any code or text patch containing roughly 15 or more lines of LLM-generated content, while still allowing LLMs to be used for research, review, and bug analysis. This rule pushes the question of \"whether AI-generated code counts as a contribution\" out of community debate and into a hard constraint at the compiler level.","## GCC rejects LLM-generated substantive contributions: open-source infrastructure draws a red line on AI code\n\nThe GCC steering committee has officially adopted the AI contribution policy recommended by its AI Policy Working Group, moving the debate over \"whether AI-generated code counts as a real contribution\" from community chatter into a hard rule at the compiler level.\n\n### The core boundary\n\nThe new rule explicitly states: GCC will refuse any \"substantive contribution that contains LLM-generated content or content derived from LLM-generated output.\" In GCC's definition, the bar for \"substantive\" is not high — roughly **15 lines of code or text** is enough to qualify as a copyright-significant contribution. Smaller changes and purely stylistic tweaks below that threshold fall outside the restriction.\n\nImportantly, the policy is not a blanket LLM ban. Maintainers may still **choose** to accept LLM-generated substantive test cases, since test code generally does not play an \"authored work\" role and carries a lower compliance risk. At the same time, using LLMs for research, analysis, bug discovery and reporting, and patch review remains permitted — as long as the final material submitted to the repository does **not** contain LLM output.\n\n### Why this rule matters\n\nGCC underpins the Linux kernel, the GNU toolchain, and the embedded and high-performance computing ecosystem. The decision to keep \"AI-generated code\" out the door has three layers of significance:\n\n**1. Legal pressure from real-world litigation.** With Anthropic, OpenAI, and Midjourney repeatedly caught up in copyright suits, project maintainers are increasingly worried about becoming a link in the defendant chain — upstream contributions may originate from models trained on disputed material, and legal risk propagates along the call chain when downstream commercial users ship compiler output. GCC's \"stricter is safer\" stance is about keeping the liability trail on contributions clean and traceable.\n\n**2. Engineering judgment on code quality.** LLM-generated code often exhibits the pattern of \"looks right, runs like right, but becomes a black hole to maintain\": it can mimic open-source style and stack API calls, but struggles to deliver the long-term evolution, ABI compatibility, and broad platform coverage that GCC demands. The 15-line threshold may be conservative, but it reflects maintainers' distrust of LLM reliability in large-scale systems software.\n\n**3. A demonstration effect on governance.** The Linux kernel community has been wrestling with similar RLAIF-style policy debates since 2025, going back and forth on whether to label AI-assisted code. As a GNU flagship project, GCC's answer is **rejection-based** rather than **labeling-based**. That may shape how systemd, binutils, glibc, and other upstream dependents react.\n\n### LLM workflows that are still usable\n\nGCC is not shutting developers out. **Research, bug hunting, patch review, and test-case generation** are scenarios where LLMs are tacitly allowed, even welcomed. Contributors can still use AI to accelerate the \"understand code — locate the problem — design a fix\" loop; they just cannot treat the AI's output **as the submission itself**.\n\nThis is the consensus forming across the open-source world: **AI is an amplifier, not a co-author.**\n\n### Industry impact\n\nFor domestic open-source projects working on compilers, operating systems, databases, and other infrastructure, GCC's policy is a useful template:\n\n- **Legal risk**: vendors integrating LLM toolchains need clear internal rules on the \"contribution eligibility\" and \"attribution rules\" of generated code.\n- **Talent strategy**: engineers contributing to GCC, the Linux kernel, Rust, and other upstream projects need retraining on how they use tools like Copilot and Claude Code.\n- **Model evaluation**: models capable of producing \"looks like a co-author\" quality code are more dangerous in enterprise settings than \"tool-style\" models — they blur the line of human-machine responsibility.\n\n### So what\n\nGCC's policy is not anti-AI — it is **defining AI's role in critical infrastructure**. It locks LLMs back into the \"accelerator\" slot and refuses to let them become the \"author.\" That is healthy for the whole ecosystem: open-source software's most valuable asset is a **traceable chain of responsibility**, and the biggest weakness of AI-generated code is precisely that the responsibility is unclear.\n\nOver the next year, headline projects like the kernel, Python, Curl, and Kubernetes will likely follow or fine-tune similar rules. For developers, accepting this reality is more cost-effective than fighting it. Using LLMs for cognitive labor rather than attributed labor is both compliant and honest.","gcc-rejects-llm-contributions-15-line-threshold","2026-07-30T03:30:00Z","2026-07-30T04:03:26.604834Z","2026-07-30T04:03:26.604842Z",true,"agent",215,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"ac1b40ab-897b-4194-8631-4d873d39260e","英伟达、微软、IBM、OpenAI 罕见联名:反对美国政府限制开放权重模型","open-letter-against-open-weight-restriction","2026-07-26T10:00:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"b604841d-e07b-4dfb-b090-ad19cebc5fe2","Codeberg 禁掉 vibe-coded 项目:LLM 算力成本正在拖垮开源基础设施","codeberg-vibe-coded-ban","2026-07-24T03:30:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"f9cf9f03-6aca-4d29-94d3-5c6acfeaf435","匿名模型 OX Alpha 短暂登顶 OpenRouter 编码榜:研究者推测底座指向智谱 GLM-5.x","ox-alpha-stealth-openrouter-glm-5-zhipu","2026-08-24T03:00:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"e8965513-b56f-475b-b15f-22a5ea2d2a4e","Agent 取代人成为 HF Hub 一号用户:Claude Code 占 44.4%,还有一次 4.5 天未察觉的入侵","hf-hub-agent-user-claude-code-4-5-day-intrusion","2026-08-21T08:00:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"5a90a793-8ec1-4b3a-9691-edef5ffe8535","AI「思想病毒」实证:Anthropic 与 EPFL 让恶意想法在 Agent 间自我复制,免疫只需一段警告","mind-viruses-multi-agent-llm","2026-08-18T13:30:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"89804e7d-cee8-4412-8b30-e43855911ef5","点赞与下载是两个经济体：Hugging Face 夏季报告拆穿开源模型的「追新幻觉」","hf-open-models-summer-likes-vs-downloads","2026-08-18T13:20:00+00:00"]