[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-microsoft-token-budget-gpt-5-6":3,"topics-all":38,"news-related-35ab3cd2-938f-4ae1-9c09-472391e1a929":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"35ab3cd2-938f-4ae1-9c09-472391e1a929","当「用得越多越好」撞上 token 账单:微软给工程师的 AI 预算紧箍咒","微软执行副总裁 Jay Parikh 在内部邮件中明确「Tokenmaxxing 不是优化目标」,把 OpenAI GPT-5.6 设为内部默认模型,并自 2026 年 7 月起为各部门设 AI token 预算目标。这标志大模型企业落地从「拼命塞 AI」切到「算账」阶段。","# 当「用得越多越好」撞上 token 账单:微软给工程师的 AI 预算紧箍咒\n\n微软执行副总裁 Jay Parikh 最近给员工发了一封内部邮件,把一件事说得很直白:**「Tokenmaxxing is not what we are optimizing for.」**(最大化 token 消耗不是我们的优化目标)。这是科技巨头里第一个把「反对 tokenmaxxing」写到内部文件里的声音,而它来的时间点,正好卡在大模型企业级落地 18 个月、token 账单第一次肉眼可见地在 IT 预算里撕开一道口子的当口。\n\n## 背景:从「用 AI」到「被 AI 用」\n\n过去两年,大量企业把 AI 使用率写进了绩效考核表:用得多,加分;用得少,扣分。这原本是为了推工具,但副作用很快显现——员工为了 KPI 最大化 token 消耗,把能用一句话解决的事扩成三段 prompt,再把 ChatGPT 生成的内容贴回 ChatGPT 总结。Solidot 报道指出,不少公司发现 token 费用大幅超出预算后,开始反向调整:既然员工被 KPI 推着「烧 token」,那干脆换个 KPI,让大家主动「省 token」。404 Media 把这种现象叫做 **tokenmaxxing**,并将其与 2025 年那波「我们用 AI 替代了一千名员工」的企业叙事直接挂钩——前者是过度使用的故事,后者是过度承诺的故事,两者背后是同一本算不清的账。\n\n## 微软这次的三个动作\n\nParikh 的邮件没有停留在口号上,做了三件具体的事:\n\n1. **设默认模型**。微软把 **OpenAI 的 GPT-5.6** 设为内部 GitHub Copilot 等工具的默认模型,因为它「比其它模型更便宜」。在 GPT-5 系里,这个版本不是旗舰,而是面向性价比的那一档——选它作为默认,等于微软在内部「按价位做了分流」。\n\n2. **设预算目标**。自 **2026 年 7 月起**,微软各部门将设定「AI token 预算目标」,员工可以在仪表盘上追踪各自的 AI 支出。据 404 Media 转引的内部指引,目前没有公开统一的目标值,但数据显示很多工程师每月在 token 上花费「从几百美元到几千美元不等」——也就是说,即便在微软这种和 OpenAI 有深度合作、能拿到内部价的公司,单工程师月烧 token 也能上万人民币。\n\n3. **改指标,改叙事**。Parikh 明确说:**「We are not optimizing for fewer tokens. We are optimizing for more impact per token.」**(我们不是在优化 token 数量,而是在优化每个 token 的影响力)。这句话是整套策略的灵魂——它把「省」重新定义成「值」:你仍然可以用 AI,但要能讲清楚这笔钱换来了什么。\n\n## 为什么是现在?三个看不见的成本\n\n这次转向的真正驱动力不是「微软突然变抠」,而是三个账面上看不到的成本集中爆发:\n\n- **算力即电力**。GPT-5 系模型单次推理的能耗远高于 GPT-4 系列,把默认模型降到 GPT-5.6 后,在微软内部数十万工程师 × 每天上千次调用的规模下,数据中心侧的电费和冷却成本会被显著压低。\n- **效率反噬**。Slashdot 评论区有人调侃「AI 把 3 天的活干到 30 分钟,但你盯着它生成的代码、跑测试、修 bug 的时间加起来比 3 天还多」——这是当前企业最普遍的「AI 提效悖论」:单位任务更快,但任务量爆炸,净工时不降反升。\n- **KPI 失真**。当「AI 使用量」成为绩效,所有人都会去刷数据,直到刷出来的数字跟业务结果脱钩。微软这次等于公开承认,把 AI 使用率当 KPI 这件事本身做错了。\n\n## 评论:大模型企业落地的「第二阶段信号」\n\n如果把 2023–2025 年看成是「拼命塞 AI」的阶段,那么 2026 年下半年到 2027 年的主旋律,大概率是**「算账」**。三个观察值得行业跟进:\n\n第一,**默认模型会变成新的成本战场**。当一家公司的内部代码助手默认模型从旗舰版切到性价比版,等于它在替模型厂商做产品分层——这会反过来影响 OpenAI、Anthropic、Google 的 SKU 设计,以及自托管\u002F开源模型的商业空间。\n\n第二,**「每 token 影响」会取代「AI 使用率」,成为新一代 KPI**。这件事一旦被微软这样体量的公司立成标杆,其它企业的 HR 和 IT 部门会照搬。「AI 提效」从口号变成可量化指标,反过来又会催生一批做 token 成本归因、prompt 效率审计的工具——这是个新的 SaaS 切口。\n\n第三,**「tokenmaxxing」会进企业 IT 治理词典**。404 Media 已经把它作为一个专有名词使用;可以预见,在 CIO 季度报告里,「token 消耗治理」会和「模型选型」「数据合规」并列出现。\n\n## 所以呢?\n\n对开发者来说,这封信的实际意义很朴素:**用 AI 不会被卡,但被问到「这个 AI 调用到底解决了什么问题」时,得有答案**。对管理者来说,微软这次的转向提醒一件事——AI 不是越用越便宜,也不是越用越好,而是要把它放回 ROI 的框架里。一封内部邮件,标着的是大模型在企业里从「上半场」切到「下半场」的转折点。\n\n参考资料:\n- [Solidot 报道(2026-08-06)](https:\u002F\u002Fwww.solidot.org\u002Fstory?sid=85018)\n- [Slashdot 转载 404 Media 原文(2026-08-04)](https:\u002F\u002Fslashdot.org\u002Fstory\u002F26\u002F08\u002F04\u002F1833219\u002Fmicrosoft-tells-engineers-tokenmaxxing-is-not-what-we-are-optimizing-for)\n- [404 Media 原始报道(2026-08-04)](https:\u002F\u002Fwww.404media.co\u002Fmicrosoft-tells-engineers-tokenmaxxing-is-not-what-we-are-optimizing-for\u002F)","https:\u002F\u002Fwww.solidot.org\u002Fstory?sid=85018","e06537c4-1c62-46c4-a4ac-d28107bbca86",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"c33b1bbc-d6ce-4f61-9d5d-1a0704a6a09b","ai-policy",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",{"id":19,"name":20,"slug":20,"description":14,"color":14},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":22,"name":23,"slug":23,"description":14,"color":14},"42e59a88-7795-47dc-a334-ef1e72c24347","openai",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"a5f4acaf-5334-4277-9967-46ce8758a0e5","en","More usage meets the token bill: Microsoft's engineer budgets","Microsoft EVP Jay Parikh told engineers in an internal email that 'Tokenmaxxing is not what we are optimizing for,' switched the default internal model to OpenAI's GPT-5.6, and set AI token budget targets for every division starting July 2026. The move signals enterprise LLM rollouts shifting from the 'shove AI in everything' phase to the 'add up the bill' phase.","# When \"More is Better\" Meets the Token Bill: Microsoft's AI Budget Wake-Up Call for Engineers\n\nMicrosoft executive vice president Jay Parikh sent an internal email recently that said it plainly: **\"Tokenmaxxing is not what we are optimizing for.\"** (Maximizing token consumption is not our optimization target.) This is the first time a tech giant has put \"anti-tokenmaxxing\" into a written internal policy, and it lands at a moment when enterprise LLM rollouts are 18 months in and token bills are visibly tearing holes in IT budgets.\n\n## Background: From \"Use AI\" to \"Be Used by AI\"\n\nOver the past two years, many companies baked AI usage into performance reviews: use more, score higher; use less, score lower. The intent was tool adoption, but the side effect surfaced fast — employees maximized token spend to chase KPIs, expanded one-line asks into three-paragraph prompts, then pasted ChatGPT output back into ChatGPT for a summary. As Solidot reported, once the token bills visibly exceeded budgets, companies began reversing course: if KPI pressure drives \"burn tokens,\" then swap the KPI and let people actively \"save tokens.\" 404 Media dubbed this **tokenmaxxing** and tied it directly to the 2025 corporate narrative of \"we replaced a thousand employees with AI\" — the over-use story and the over-promise story sit on the same unbalanced ledger.\n\n## Microsoft's Three Concrete Moves\n\nParikh's email did not stop at slogans. Three things happened:\n\n1. **A default model switch.** Microsoft made **OpenAI's GPT-5.6** the default model for internal tools including GitHub Copilot, on the grounds that it is \"cheaper than other models.\" Within the GPT-5 family, this is not the flagship — it is the price-performance tier. Picking it as the default is Microsoft quietly running price-tier routing internally.\n\n2. **A budget target.** Starting **July 2026**, Microsoft divisions will set an \"AI token budget target,\" and employees can track their individual AI spend in a dashboard. According to 404 Media's reading of the internal guidance, no uniform target value has been published, but data shows many engineers spend \"anywhere from hundreds to thousands of dollars per month\" on tokens. Even at Microsoft — with deep OpenAI ties and internal pricing — per-engineer token burn runs into five figures RMB monthly.\n\n3. **A metric and narrative reset.** Parikh said it explicitly: **\"We are not optimizing for fewer tokens. We are optimizing for more impact per token.\"** This sentence is the soul of the whole strategy. It redefines \"save\" as \"worth it\" — you may still use AI, but you need to articulate what that money bought.\n\n## Why Now? Three Hidden Costs\n\nThe real driver is not \"Microsoft suddenly got cheap.\" It is three off-balance-sheet costs hitting at once:\n\n- **Compute is electricity.** Per-inference energy for the GPT-5 family is significantly higher than GPT-4 era models. Routing default calls to GPT-5.6 across hundreds of thousands of engineers × thousands of daily invocations meaningfully cuts data-center power and cooling spend.\n- **The efficiency paradox.** As one Slashdot commenter put it, \"AI turns 3 days of work into 30 minutes, but the time you spend staring at its code, running tests, and fixing bugs adds up to more than 3 days.\" This is the most common \"AI productivity paradox\" in enterprises today: per-task speed is up, but task volume explodes, net working hours trend up, not down.\n- **KPI distortion.** Once \"AI usage\" becomes performance review fuel, everyone games the data until the numbers decouple from business outcomes. Microsoft's move is a public admission that using AI usage as a KPI was the wrong design from the start.\n\n## Commentary: The \"Phase Two\" Signal for Enterprise LLM Rollout\n\nIf 2023–2025 was the \"shovel AI into everything\" phase, then H2 2026 to 2027 will likely be the **\"add up the bill\"** phase. Three observations worth tracking:\n\nFirst, **the default model slot becomes the new cost battlefield.** When a company moves its default coding assistant from flagship to price-performance, it is essentially running product-tier routing for the model vendor — which in turn reshapes OpenAI, Anthropic, and Google's SKU design, and reshapes the business case for self-hosted and open-weight alternatives.\n\nSecond, **\"impact per token\" replaces \"AI usage\" as the next-generation KPI.** Once Microsoft stakes this out, HR and IT departments elsewhere will copy the playbook. \"AI productivity\" moves from slogan to a measurable metric, which in turn spawns a new SaaS category for token cost attribution and prompt-efficiency auditing.\n\nThird, **\"tokenmaxxing\" enters the enterprise IT governance vocabulary.** 404 Media is already using it as a proper noun; expect \"token consumption governance\" to sit alongside \"model selection\" and \"data compliance\" in CIO quarterly reviews.\n\n## So What?\n\nFor developers, the practical meaning of this email is simple: **no one is banning AI use, but when someone asks \"what did this AI call actually solve?\" you need an answer.** For managers, Microsoft's turn is a reminder that AI is neither cheaper with more usage nor better with more usage — it has to be put back in an ROI frame. An internal email marks the inflection point where enterprise LLMs move from the first half to the second half of the cycle.\n\nReferences:\n- [Solidot coverage (2026-08-06)](https:\u002F\u002Fwww.solidot.org\u002Fstory?sid=85018)\n- [Slashdot digest of 404 Media (2026-08-04)](https:\u002F\u002Fslashdot.org\u002Fstory\u002F26\u002F08\u002F04\u002F1833219\u002Fmicrosoft-tells-engineers-tokenmaxxing-is-not-what-we-are-optimizing-for)\n- [Original 404 Media report (2026-08-04)](https:\u002F\u002Fwww.404media.co\u002Fmicrosoft-tells-engineers-tokenmaxxing-is-not-what-we-are-optimizing-for\u002F)","microsoft-token-budget-gpt-5-6","2026-08-16T02:00:00Z","2026-08-16T07:08:11.851713Z","2026-08-16T07:08:11.851727Z",true,"agent",131,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"f8d091db-ca3a-44f6-b004-0b5f8aba0bef","OpenAI 请来9位数学家,却管不住模型节奏","openai-math-advisory-group","2026-09-21T21:15:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"d157b4f9-537e-405c-b557-859f6d2cf18c","微软自家高管警告:抓新闻训 AI 是「人类史上最大规模劳动盗窃」","microsoft-ai-scraping-theft-of-labor","2026-09-19T00:11:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"2fc4f696-a697-498c-b9d5-28250bfeaa79","ChatGPT 进欧盟 VLOP 名单:OpenAI 第一次要为生成式 AI 内容负全责","chatgpt-eu-vlop-dsa-first-ai-platform-rules","2026-09-05T07:00:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"b3632f86-c054-49ba-af6c-a09337ee6b5d","菲尔兹奖得主联名警告:AI 偏离数学本质","fields-medal-mathematicians-ai-misalignment-declaration","2026-09-22T07:00:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"b0c434a5-4911-4297-b1ef-2c44cbc26653","蚂蚁新研究:19769 个代码仓库,炼出百万条 agent 技能","code2skill-agent-skill-synthesis","2026-09-21T19:06:31+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"b0c4e8d2-5662-4e3e-b489-6202eabbe97b","Dream-RSI 把历史当模拟器:162 倍杠杆重写 RSI 算力账本","dream-rsi-replay-simulator-162x","2026-09-16T06:00:00+00:00"]