The Internal Memo That Started With $28,000 in 28 Days

When OpenAI released GPT-5.6 Sol in July, Microsoft picked it as the default model for GitHub Copilot and related internal workflows — a story first broken by CNBC and 404 Media in early August, anchored by Executive Vice President Jay Parikh's line: "Tokenmaxxing is not what we are optimizing for." On August 30, Solidot surfaced a block of numbers the English coverage had glossed over: a Microsoft-internal sample of 350 US employees (out of 223,000 globally) self-reported their AI bills over a rolling 28-day window, and looking at the distribution makes it obvious why that memo had to be written.

Both Ends of the Distribution

  • The extreme high: an engineer in Customer and Partner Solutions burned $28,000 on AI tooling over a 28-day window.
  • The next tier: several other employees ran their 28-day bills past $10,000.
  • The middle: across the 350-person sample, median AI spend was about $300 per 28-day period.
  • Departmental spread: CoreAI posted the highest median at $975; in some departments certain employees logged only tens of dollars over the same window.
  • Sample scope: 350 US-based Microsoft employees, voluntary self-reporting — a thin sliver of a 223,000-person company, not an audited figure.

Sources cross-checked: 404 Media (re-reported by Futurism), CNBC (which quoted the full Parikh memo), Gadget Review (re-reported by Ynet News) all align on the $28,000 / $10,000+ / $300 median / CoreAI $975 figures — no version conflict across the four.

What This Actually Means Beyond the Headline

Read through the same ledger with a model-routing lens, three things fall out:

1. Expensive models stop winning by default. The $28,000 Customer-and-Partner-Solutions case isn't someone using the cheapest model and still over-running — it's textbook tokenmaxxing: always pick the strongest model, never trim context, never switch paths, let the bill arrive at month-end. Parikh's memo isn't about per-token price; it's about breaking the muscle memory of "strongest first."

2. The real reason GPT-5.6 Sol became the default. CNBC quotes Parikh verbatim: "Internally, shifting more workloads to OpenAI models helps us get greater value from our token investment." GPT-5.6 Sol, released by OpenAI in July, is roughly an order of magnitude cheaper than its predecessor while staying near-frontier. Microsoft flipped the internal default to absorb part of the tokenmaxxing demand through procurement, not through asking engineers to self-discipline — that's leverage on the buying side, not engineering discipline.

3. CoreAI's $975 median is the part that looks paradoxical but isn't. CoreAI builds Copilot, GitHub, and the models themselves. They use AI to make AI. Their work sits naturally on the token-hungry frontier — Agent loops, long-horizon tasks, automated workflows. A $975 median there is exactly what you'd expect. What is genuinely anomalous is the long tail around the $300 median: a handful of employees at $28,000 and $10,000+ means those individuals aren't "running AI to do their job" — they're inflating a token leaderboard internally.

What Other Companies Should Take From This

The lesson for any enterprise that has broadly deployed AI tooling is one sentence: don't manage token spend by mean. Pull median, P95, and P99 next to each other — by department. A team at median $300 and a colleague at $28,000 get blended into "the average looks fine" if you only show the mean; the tail person alone eats dozens of headcount's worth of budget. Ramp-style payment-aggregator data is most useful as a distribution shape, not as a single monthly number.

So the right action isn't "use less AI." It's making the routing fine-grained: default to mid-tier, high-capability models like GPT-5.6 Sol; carve out an isolated budget for genuinely long-context Agent workflows; cap tokens-per-task at the workflow level. Microsoft's memo is the on-switch for that workflow; the Copilot default flip is the procurement-side weld that locks it in. The next month's internal Ramp bill — whether that $28,000 tail pulls back or keeps recurring — is the only honest readout on whether "tokenmaxxing" actually went away.

Sources: 404 Media reporting and OpenAI follow-up disclosures; CNBC Tech; Gadget Review; Ynet News. Solidot compiled the Chinese-language aggregation.