[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-microsoft-internal-ai-bill-explode-default-model-swap":3,"topics-all":38,"news-related-7bb4f5ec-14e0-43b6-9913-07cad82a520b":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"7bb4f5ec-14e0-43b6-9913-07cad82a520b","微软内部 AI 账单失控:单员工月烧 2.8 万美元,倒逼默认模型换人","微软内部披露员工 AI 自报账单,一名 CPS 部门工程师 28 天花费高达 2.8 万美元,公司级中位数约 300 美元。CoreAI EVP Jay Parikh 发邮件叫停 tokenmaxxing 行为,并将内部默认模型从 Claude 切到 GPT-5.6 Sol。","微软最近在内部通讯里悄悄做了一件不太体面的事:把员工的「AI 自付账单」汇总公布,然后把默认模型从 Claude 换成了 GPT-5.6 Sol。这事本身不算新闻,但账单数字透露出的信号,值得所有正在大规模铺 AI 的公司认真读一下。\n\n## 账单里到底有什么\n\n根据 Gadget Review 拿到的内部数据,在自愿提交的约 350 名美国员工样本里,微软全球 22.3 万名员工中出现了相当极端的分布:Customer and Partner Solutions 部门一名工程师在 28 天里烧了 2.8 万美元的 AI 工具开销,多名员工超过 1 万美元,而公司级中位数只有约 300 美元。部门中位数也呈现明显梯度——CoreAI 约 975 美元、Security 约 526 美元、Microsoft AI 约 490 美元、Cloud + AI 约 325 美元。\n\n考虑到这只是「自愿样本」,真实人均支出大概率更高;但极端值与中位数之间上百倍的鸿沟,说明 AI 工具使用在公司内部远没有标准化——有的团队把 Copilot 当日常搭子,有的团队基本不碰。\n\n## Tokenmaxxing:当用量榜变成竞技场\n\n真正让管理层坐不住的是一种被内部称为「tokenmaxxing」的行为。Copilot 仪表盘把每个人的 AI 用量做成可见指标,于是有员工开始在低价值甚至毫无意义的查询上反复刷量冲榜。微软 CoreAI 执行副总裁 Jay Parikh 在 8 月初的内部邮件里直接喊停,原话是「Tokenmaxxing 不是我们要优化的方向」。同步落地的措施包括:部门级 AI token 预算硬指标、员工个人花费上墙、GitHub Copilot 默认模型切换——这才是事件里被忽视但分量更重的部分。\n\n## 默认模型切换的真实含义\n\n微软给出的官方理由是「从 token 投资中获得更大价值」。翻译一下:GPT-5.6 Sol 比 Claude 在单价上更便宜,叠加微软同时是 OpenAI 的大股东,这笔账从财务模型上就更好看。但更深层的信号是——即使是最紧密的合作伙伴,当账单开始失控,商业本能会优先压住成本,而不是守住模型多样性。这与 Ramp 这类企业支付数据正在揭示的趋势一致:企业买 AI 时,「旗舰是不是最强」已经让位于「够用且便宜」。\n\n## 行业层面的 FinOps 时刻\n\n云时代有过同样的剧本:「随便开机器」的鼓励文化撞上六位数的 AWS 月账单,然后 FinOps 工具应运而生。现在轮到 AI token 重演这个循环。可见的下一步会是模型路由、按 query 难度分配价位、硬性 token 配额、团队级 usage 报表。对所有正在做企业级 AI 落地的团队来说,微软这次的内部数据是一份提前寄到的预警信。\n\n需要注意的是,样本量约 350 人相对 22.3 万全球员工并不大,极端值可能被自选偏差放大;但账单量级的方向性信号——尤其是默认模型切换的成本驱动逻辑——是确定无疑的。来源:Gadget Review,2026 年 8 月 26 日(https:\u002F\u002Fwww.gadgetreview.com\u002Fone-microsoft-employee-spent-28000-on-ai-in-28-days-now-the-company-is-watching-everyone)。","https:\u002F\u002Fwww.gadgetreview.com\u002Fone-microsoft-employee-spent-28000-on-ai-in-28-days-now-the-company-is-watching-everyone","7ccbce8a-ef55-4429-89a6-d3781b72b2df",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"fca9258a-9430-455a-b95d-b9fae5e373a8","ai-inference",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"42e59a88-7795-47dc-a334-ef1e72c24347","openai",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"a9f0ad26-141f-4250-bf61-d25b431bb65f","en","Microsoft AI bills explode, 8K per engineer, default model swapped","Microsoft's internal disclosure of employee self-reported AI bills shows a Customer and Partner Solutions engineer racked up 8,000 in 28 days against a company-wide median of ~00. CoreAI EVP Jay Parikh ordered an end to 'tokenmaxxing' and quietly shifted the default model from Claude to OpenAI's GPT-5.6 Sol.","Microsoft recently did something rather undignified in its internal communications: it aggregated and published employees' self-reported AI bills, then switched the default model from Anthropic's Claude to OpenAI's GPT-5.6 Sol. The action itself is not the news — but the bill numbers reveal a signal worth reading carefully for any company rolling out AI at scale.\n\n## What the bills actually show\n\nAccording to internal data obtained by Gadget Review, across roughly 350 self-reporting US employees (out of Microsoft's 223,000 global headcount), the distribution is extremely lopsided: one engineer in Customer and Partner Solutions burned 8,000 in AI tool spend over a 28-day period; multiple employees cleared 0,000; while the company-wide median sat at around 00. Department-level medians form a clear gradient — CoreAI at ~75, Security at ~26, Microsoft AI at ~90, Cloud + AI at ~25.\n\nBecause this is a voluntary sample, true per-capita spend is likely higher, but the 100x gap between the extremes and the median means AI tool usage is far from standardized inside the company — some teams treat Copilot as a daily companion; others barely touch it.\n\n## Tokenmaxxing: when the usage leaderboard becomes an arena\n\nWhat actually spooked management was a practice internally branded \"tokenmaxxing.\" Copilot dashboards surfaced per-employee AI consumption as a visible number, so some employees started burning tokens on low-value or even meaningless queries just to climb the leaderboard. Microsoft CoreAI EVP Jay Parikh called it out directly in an early-August internal memo: \"Tokenmaxxing is not what we are optimizing for.\" Measures rolled out alongside: department-level AI token budget hard targets, individual spend surfaced on internal dashboards, and — heavier than the memo itself — the GitHub Copilot default model switch.\n\n## What the default-model switch really means\n\nMicrosoft's stated rationale is \"greater value from our token investment.\" Translated: GPT-5.6 Sol is cheaper per token than Claude, and Microsoft is simultaneously OpenAI's biggest backer — the math looks better on the financial model. But the deeper signal is this: even when Anthropic is one of your closest partners, once bills start spiraling, commercial instinct favors cost control over model diversity. This matches a trend enterprise-payment services such as Ramp are surfacing: when companies buy AI, \"is the flagship the strongest\" has yielded to \"good enough and cheaper.\"\n\n## The FinOps moment at the industry level\n\nThe cloud era ran this exact script. \"Spin up whatever you need\" culture collided with six-figure AWS monthly bills, then FinOps tooling emerged. AI tokens are now running the same loop. The next visible moves will be model routing (assigning different price tiers by query difficulty), hard token quotas, team-level usage dashboards, and limits on agent call counts. For every team shipping enterprise-grade AI today, Microsoft's internal data is an early warning letter.\n\nA caveat: the ~350-person sample is small relative to Microsoft's 223,000 global headcount, and the extremes may be inflated by response bias. But the directional signal in the bill magnitude — especially the cost-driven logic of the default-model switch — is unambiguous. Source: Gadget Review, Aug 26 2026 (https:\u002F\u002Fwww.gadgetreview.com\u002Fone-microsoft-employee-spent-28000-on-ai-in-28-days-now-the-company-is-watching-everyone).","microsoft-internal-ai-bill-explode-default-model-swap","2026-08-28T04:00:00Z","2026-08-28T09:05:15.671314Z","2026-08-28T09:05:15.671322Z",true,"agent",160,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"6c5f1bc4-d877-483c-a0e9-70aff0e30dbe","微软内部 Ramp 账单:一名工程师 28 天烧掉 2.8 万美元 AI 费","microsoft-internal-ramp-ai-spending-28000-28-days","2026-08-31T03:00:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"97eb895f-8294-47b9-a9e0-5f165a75cb5b","OpenAI Jalapeño 实测:自研推理芯片在 Hot Chips 上跑赢 Blackwell","openai-jalapeno-hot-chips-broadcom-blackwell","2026-08-26T08:00:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"4aa9534a-778e-4cd7-8194-fdf3097249b8","OpenAI Jalapeño Hot Chips 实测:峰值每瓦 1.9×,延迟压到 1 秒","openai-jalapeno-hot-chips-benchmark-2026","2026-08-26T02:00:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"6fa1bc74-c98e-476b-bc4c-9ae057105ffb","ParaTempo:免训练并行推理,延迟最高降 32%、token 省三成","paratempo-temporal-confidence-parallel-reasoning","2026-08-24T17:20:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"31f3215c-0892-419d-a610-fe815cc60bbe","GPT-5.6 降价 80% 把竞争拉进「同等智能成本」：DeepSeek V4 Flash 接招，国产模型卡出双线赛道","gpt-5-6-luna-price-cut-equal-intelligence-cost","2026-08-12T03:00:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"c2ee2a09-d001-4740-9820-21fb672eee8b","Copilot 默认模型切到 GPT-5.6 Sol：tokenmaxxing 终结","microsoft-gpt5-6-default-token-budget","2026-08-08T08:00:00+00:00"]