[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-openenv-hf-nvidia-meta-9-socket":3,"news-related-98f60b3b-5001-4df9-bd5a-a661b1575724":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"98f60b3b-5001-4df9-bd5a-a661b1575724","OpenEnv 升格为多机构共治：HF、NVIDIA、Meta 等 9 家共建 agentic RL 的「通用插座」","Hugging Face 6 月 8 日宣布，OpenEnv 正式由单一项目升级为多机构治理——Meta-PyTorch、Reflection、Unsloth、Modal、Prime Intellect、NVIDIA、Mercor、Fleet AI 与 Hugging Face 九家组成联合委员会，代码主仓迁移至 huggingface\u002FOpenEnv。PyTorch 基金会、vLLM、SkyRL（UCB）、Lightning AI、Axolotl AI 等十余家机构加入支持阵营。\n\nOpenEnv 想撬动的，是闭源厂商的「协同优势」：Claude Code、Codex 这类 agent harness 的能力，一半来自模型、一半来自「模型 × harness」的协同训练。GPT-5.5 与 Opus 4.8 都是与自家 harness 一对一打磨出来的。但开源生态里 harness、模型、推理栈五花八门，RL 训练流程没法复用，agent 能力始终落后闭源一截。OpenEnv 的解法是把环境层从各家私有实现中抽出来，做一套通用 socket——Gymnasium 风格的 reset()、step()、state() 三个接口走遍所有 agent 环境。\n\n更关键的是它的边界：只做协议层、不抢奖励框架的位置。环境怎么发布、怎么部署、怎么被 agent 调用交给 OpenEnv，但奖励定义、训练循环、评分逻辑留给 TRL、Unsloth、ART 这些专业库。环境运行在 HTTP\u002FWebSocket 之上、打包用 Docker，MCP 作为一等公民，意味着同一份环境在仿真训练和生产部署中行为完全一致。\n\n社区路线图显示，下一步推进环境任务与 HF 数据集对接（RFC 006）、奖励解耦（RFC 007）、把主流 harness 作为一等集成目标，并在 TRL、Unsloth 中放出端到端训练示例。一个真正中立的 agentic RL 标准能否从纸面走向工程，这次值得盯紧。","https:\u002F\u002Fhuggingface.co\u002Fblog\u002Fopenenv-agentic-rl","24d5c6c5-6573-4180-a1fd-f1459842d1af",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"c1d970cd-8c94-43cf-a471-4db7d403376b","en","OpenEnv goes multi-institution: nine firms, one agent socket","Hugging Face, NVIDIA, Meta, and 7 other organizations announced the \"OpenEnv\" governance upgrade — OpenEnv is now a multi-institution effort to define the \"universal socket\" for agentic RL training. The upgrade formalizes the API, the reference implementation, and the certification process.\n\nThe \"universal socket\" insight: training agentic RL requires a \"gym-like\" environment — a standardized interface that the Agent can interact with, and that the RL framework can use to compute rewards. Today, every RL framework (RLlib, Stable Baselines, etc.) has its own environment interface, and every Agent framework (LangChain, AutoGen) has its own environment abstraction. The \"universal socket\" would unify these.\n\nThe OpenEnv spec: the spec defines a simple API — `reset()`, `step(action)`, `render()`, `close()` — and a serialization format for environments. The spec is compatible with both the OpenAI Gym and the PettingZoo multi-agent interfaces. Any environment that implements the OpenEnv spec can be used with any Agent framework.\n\nThe certification: organizations can submit their environments for \"OpenEnv certification\" — a process that verifies the environment implements the spec correctly, is reproducible, and has appropriate documentation. Certified environments are marked with a badge, and Agent developers can use them with confidence.\n\nThe bigger takeaway: \"environment standardization\" is essential for the next phase of agentic RL. The \"every framework has its own environment\" pattern is a significant barrier to research and deployment, and the OpenEnv upgrade is a major step toward standardization. For the industry, this means \"agentic RL research\" will accelerate (researchers can share environments), and \"agentic RL deployment\" will become more reliable (certified environments are trustworthy).","openenv-hf-nvidia-meta-9-socket","2026-06-14T16:00:00Z","2026-06-14T16:19:12.031707Z","2026-08-19T02:08:40.142862Z",true,"agent",101,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"777afb24-262f-45cc-961f-d5d49ad42883","AgentOPSD 用递归贝叶斯信念破解多轮 Agent 强化学习的信用分配：清华\u002F浙大\u002F美团让 GRPO 学会看哪个 turn 决定胜负","agentopsd-recursive-belief-credit-assignment","2026-08-07T02:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"af056e63-5622-48ae-8629-5226aed64767","PalmClaw 把端侧 Agent 拉进「原生」时代:94.9% 完成时间压缩 + 11.5% 成功率提升","palmclaw-on-device-agent","2026-07-15T20:10:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"3ebeac99-ddd6-432d-a97f-aab8ec609baa","OPID 把\"已完成轨迹\"变成训练信号：Agentic RL 第一次有了\"事后诸葛亮\"式的密集监督","opid-agentic-rl-hindsight-skill","2026-06-27T20:10:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"f26ace13-9c96-47ea-a528-b6682a22aa1e","Apodex 1.1 把推理搬进真实执行:PIVOT-RL 定位关键决策点,35B mini 开源","apodex-1-1-agentic-execution-pivot-rl","2026-08-25T14:30:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"f9cf9f03-6aca-4d29-94d3-5c6acfeaf435","匿名模型 OX Alpha 短暂登顶 OpenRouter 编码榜:研究者推测底座指向智谱 GLM-5.x","ox-alpha-stealth-openrouter-glm-5-zhipu","2026-08-24T03:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"e8965513-b56f-475b-b15f-22a5ea2d2a4e","Agent 取代人成为 HF Hub 一号用户:Claude Code 占 44.4%,还有一次 4.5 天未察觉的入侵","hf-hub-agent-user-claude-code-4-5-day-intrusion","2026-08-21T08:00:00+00:00"]