[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-microsoft-internal-ramp-ai-spending-28000-28-days":3,"topics-all":38,"news-related-6c5f1bc4-d877-483c-a0e9-70aff0e30dbe":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"6c5f1bc4-d877-483c-a0e9-70aff0e30dbe","微软内部 Ramp 账单:一名工程师 28 天烧掉 2.8 万美元 AI 费","微软按部门追踪的 350 名员工 AI 账单:Customer and Partner Solutions 一名员工 28 天支出 2.8 万美元,多人破 1 万,中位数 300 美元;CoreAI 部门中位数 975 美元最高。","## 从 28 天 2.8 万美元开始的那封内部邮件\n\nOpenAI 的 GPT-5.6 Sol 在七月上线,被微软内部定为 GitHub Copilot 与相关工作流的默认模型——这条消息最先是 CNBC 与 404 Media 在八月初放出来的,执行副总裁 Jay Parikh 的原话是「Tokenmaxxing is not what we are optimizing for」。Solidot 八月三十日补了一块市场公开稿里没仔细展开的数字:那是员工自愿提交的内部账单——7 万家公司支出数据由支付服务集团 Ramp 收集,微软 223,000 名全球员工里抽样了 350 名美国员工,按 28 天一个观察窗口,只要看一眼分布就知道这封邮件为何会被写出来。\n\n## 分布两端\n\n- **极端高位**:Customer and Partner Solutions(客户服务与伙伴解决方案部门)有一名工程师 28 天内烧掉 2.8 万美元的 AI 费用;\n- **次高位**:多名员工月支出超过 1 万美元;\n- **中位水平**:全体样本 28 天 AI 支出中位数约 300 美元;\n- **部门冷热差**:CoreAI 部门中位数最高,达 975 美元;而少数部门某些人 28 天只花了几十美元;\n- **样本规模**:**350 名美国员工**,占微软 223,000 全球员工的极小部分,**这个分布是基于自愿上报的账单**(而非审计口径)。\n\n材料出处:**404 Media(Futurism 转引)**、**CNBC(直接引用 Parikh 备忘录全文)**、**Gadget Review(Ynet News 转引同源数据)**三家独立报道交叉给出 2.8 万、1 万+、300 美元中位、CoreAI 975 美元中位这四个核心数字,无版本冲突。\n\n## 这件事的技术含义不只在金额\n\n把同一份账单放在模型选型逻辑里读,可以拆出三层判断:\n\n**第一层:贵模型不再自动获胜**。Customer and Partner Solutions 那位 2.8 万美元的 case,显然不是用最便宜的模型跑通业务——更像是典型 tokenmaxxing:选择最强模型、不主动收敛上下文、不主动换路径,直到月底账单自己跳出来。Parikh 的邮件重点不是单价,是把「用最强模型随手解决」这条肌肉记忆打断。\n\n**第二层:GPT-5.6 Sol 切默认的真实理由**。CNBC 引用 Parikh 原话:「Internally, shifting more workloads to OpenAI models helps us get greater value from our token investment」。OpenAI 七月发布的 GPT-5.6 Sol 比上代 GPT-5.6 便宜一个数量级、但能力接近前沿。微软把内部默认模型切过去,本质是用更便宜的同档位模型吞掉一部分 tokenmaxxing 的需求,**而不是让工程师主动选小模型**——这是采购层面的杠杆,不是工程层面的自律。\n\n**第三层:CoreAI 部门中位数 975 美元**很反直觉。CoreAI 是写 Copilot、写 GitHub、写模型本身的部门,他们用 AI 是为了造 AI,**975 美元中位数最高反而合理**:他们工作在 Agent、长程任务、自动化工作流这些「天然吃 token」的边界上。**真正异常的反而是中位数 300 美元两侧那条长尾**:少数人烧 2.8 万、1 万+,意味着这部分员工不是在「用 AI 干业务」,而在内部排行榜上做 token 通胀。\n\n## 对企业的启发\n\n这件事给所有把 AI 工具大规模铺下去的公司的启示只有一句话:**别用平均数管理 token 账单**。一位中位数 300 美元的部门和一位尾部 2.8 万美元的同事,在同一个 G&A 报销口径里拉平均,数字会显得「还好」——但尾部那一两位就能吃掉几十人份的预算。Ramp 这类支付聚合数据最有用的不是平均值,是「每部门 P95\u002FP99 + 中位数」的并列视图。\n\n所以正确动作不是继续喊「别用太多 AI」,而是把模型路由做细:**默认用 GPT-5.6 Sol 这种中等价位 + 高能力的档位、给真正需要长上下文的 Agent 工作流开一个独立配额、为单次任务设 token 上限**。微软这封邮件其实是把这套流程的开关打开;Copilot 默认模型切到更便宜档位,是把这套流程的采购侧焊死。等下一次账单周期出来,我们才会看到尾部那条 2.8 万美元的曲线有没有缩回去——这才是 tokenmaxxing 话题真正有没有降温的判据。\n\n本文事实来源:404 Media 报道与 OpenAI 后续披露 + CNBC Tech 报道 + Gadget Review + Ynet News,Solidot 中文综述整理。","https:\u002F\u002Fwww.solidot.org\u002Fstory?sid=85216","d59894d3-308e-4fd8-8865-86dc1eeac4a2",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"baf131c1-687a-49f4-87f6-4dd87c1c692f","gpt",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"42e59a88-7795-47dc-a334-ef1e72c24347","openai",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"70f36e63-b2d5-4159-bb33-2bd3bde0fa14","en","Microsoft's Internal Ramp Bills: One Engineer Burned $28,000 on AI in 28 Days","Microsoft tracked 350 US employees' AI bills by department: a Customer and Partner Solutions engineer logged $28,000 in 28 days, several others crossed $10,000, median was $300, and CoreAI had the highest department median at $975.","## The Internal Memo That Started With $28,000 in 28 Days\n\nWhen OpenAI released GPT-5.6 Sol in July, Microsoft picked it as the default model for GitHub Copilot and related internal workflows — a story first broken by CNBC and 404 Media in early August, anchored by Executive Vice President Jay Parikh's line: \"Tokenmaxxing is not what we are optimizing for.\" On August 30, Solidot surfaced a block of numbers the English coverage had glossed over: a Microsoft-internal sample of 350 US employees (out of 223,000 globally) self-reported their AI bills over a rolling 28-day window, and looking at the distribution makes it obvious why that memo had to be written.\n\n## Both Ends of the Distribution\n\n- **The extreme high**: an engineer in Customer and Partner Solutions burned $28,000 on AI tooling over a 28-day window.\n- **The next tier**: several other employees ran their 28-day bills past $10,000.\n- **The middle**: across the 350-person sample, median AI spend was about $300 per 28-day period.\n- **Departmental spread**: CoreAI posted the highest median at $975; in some departments certain employees logged only tens of dollars over the same window.\n- **Sample scope**: **350 US-based Microsoft employees, voluntary self-reporting** — a thin sliver of a 223,000-person company, not an audited figure.\n\nSources cross-checked: 404 Media (re-reported by Futurism), CNBC (which quoted the full Parikh memo), Gadget Review (re-reported by Ynet News) all align on the $28,000 \u002F $10,000+ \u002F $300 median \u002F CoreAI $975 figures — no version conflict across the four.\n\n## What This Actually Means Beyond the Headline\n\nRead through the same ledger with a model-routing lens, three things fall out:\n\n**1. Expensive models stop winning by default.** The $28,000 Customer-and-Partner-Solutions case isn't someone using the cheapest model and still over-running — it's textbook tokenmaxxing: always pick the strongest model, never trim context, never switch paths, let the bill arrive at month-end. Parikh's memo isn't about per-token price; it's about breaking the muscle memory of \"strongest first.\"\n\n**2. The real reason GPT-5.6 Sol became the default.** CNBC quotes Parikh verbatim: \"Internally, shifting more workloads to OpenAI models helps us get greater value from our token investment.\" GPT-5.6 Sol, released by OpenAI in July, is roughly an order of magnitude cheaper than its predecessor while staying near-frontier. Microsoft flipped the internal default to absorb part of the tokenmaxxing demand through procurement, not through asking engineers to self-discipline — that's leverage on the buying side, not engineering discipline.\n\n**3. CoreAI's $975 median is the part that looks paradoxical but isn't.** CoreAI builds Copilot, GitHub, and the models themselves. They use AI to make AI. Their work sits naturally on the token-hungry frontier — Agent loops, long-horizon tasks, automated workflows. A $975 median there is exactly what you'd expect. **What is genuinely anomalous is the long tail around the $300 median: a handful of employees at $28,000 and $10,000+ means those individuals aren't \"running AI to do their job\" — they're inflating a token leaderboard internally.**\n\n## What Other Companies Should Take From This\n\nThe lesson for any enterprise that has broadly deployed AI tooling is one sentence: **don't manage token spend by mean**. Pull median, P95, and P99 next to each other — by department. A team at median $300 and a colleague at $28,000 get blended into \"the average looks fine\" if you only show the mean; the tail person alone eats dozens of headcount's worth of budget. Ramp-style payment-aggregator data is most useful as a distribution shape, not as a single monthly number.\n\nSo the right action isn't \"use less AI.\" It's making the routing fine-grained: **default to mid-tier, high-capability models like GPT-5.6 Sol; carve out an isolated budget for genuinely long-context Agent workflows; cap tokens-per-task at the workflow level.** Microsoft's memo is the on-switch for that workflow; the Copilot default flip is the procurement-side weld that locks it in. The next month's internal Ramp bill — whether that $28,000 tail pulls back or keeps recurring — is the only honest readout on whether \"tokenmaxxing\" actually went away.\n\nSources: 404 Media reporting and OpenAI follow-up disclosures; CNBC Tech; Gadget Review; Ynet News. Solidot compiled the Chinese-language aggregation.","microsoft-internal-ramp-ai-spending-28000-28-days","2026-08-31T03:00:00Z","2026-08-31T01:03:41.027961Z","2026-08-31T01:03:41.027971Z",true,"agent",120,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"c2ee2a09-d001-4740-9820-21fb672eee8b","Copilot 默认模型切到 GPT-5.6 Sol：tokenmaxxing 终结","microsoft-gpt5-6-default-token-budget","2026-08-08T08:00:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"e12d2e7d-35b7-42d2-b02f-bdcc0a547878","机械臂实测 GPT-6 Astra:19\u002F20 对 8\u002F20 完胜 Fable 5.1,精细插入却全员卡壳","gpt-6-astra-robot-arm-benchmark","2026-09-07T19:13:54+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"390c2437-4e4f-45ec-8270-67c5bfa4fa47","ChatGPT、Claude、Grok、Gemini 罕见同时下线,周四早晨全球 AI 集体失声","chatgpt-claude-grok-gemini-thursday-outage","2026-09-05T06:00:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"b5be4ce8-4a41-461c-9202-148e64fab329","GPT-6 Astra 系统卡:零日自用、对齐升 53%,CoT 可监控性反向下滑","gpt-6-astra-system-card-2026-monitorability","2026-09-04T03:30:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"5bf8fa2d-258e-41d3-bfb6-5c2053433cfd","GPT-6 Astra 正式上线:8 月因安全被暂停的旗舰回来了","gpt-6-astra-launch","2026-09-04T03:12:38+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"9e58d587-3c1b-44c5-ad36-daf23aeb42a2","微软叫停 tokenmaxxing:GitHub Copilot 默认切回 GPT-5.6 Sol,Parikh 设 token 预算","microsoft-token-budget-gpt-5-6-default","2026-09-03T00:30:00+00:00"]