On September 17, 2026, GitHub distinguished engineer Stephen Toub published a 65-minute write-up detailing how GitHub rewrote the engine behind Copilot CLI, the Copilot app and the Copilot SDK from TypeScript into Rust. The numbers GitHub itself released tell the story: 14.5 weeks, 128 pull requests landing on main, roughly 1.3 PRs a day, about 435,000 lines of TypeScript replaced with 832,378 lines of Rust plus 468,689 lines of Rust unit tests. Total token spend was about 136.3 billion, of which 130.6 billion were cached input reads at a 96.22% prompt-cache hit rate, for a bill of roughly $120,000. The bulk of the work is credited to a single engineer, Stephen Toub.
The performance numbers are equally concrete. The same benchmark ran 7.55 single-turn lifecycles per second on the TypeScript runtime and 120 per second on the Rust build loaded in-process, a 15.9x throughput gain on that workload. A ten-client agent cluster used 1,383 MB of memory on TypeScript and 126 MB on Rust — about an order of magnitude less. GitHub rebuilt the runtime behind a 19-function C ABI so the SDKs in six languages can load the runtime in-process instead of spawning a Node subprocess and talking to it over JSON-RPC.
GitHub also published the tool-call shape of the migration: 630,423 shell calls, 590,988 file views, 281,783 ripgrep searches, 53,715 apply_patch operations, 40,591 edits and 13,080 task calls delegating to subagents. Toub's takeaway is that "the popular image of AI spewing code is almost backwards" — most of the time went into reading, searching and running diagnostics, not editing. For the hardest file, session.ts at over 30,000 lines, a parent session ran for 25 hours, made 222 shell calls, read 205 files, ran 197 ripgrep searches and spawned 15 child sessions across seven waves: 10 on GPT-5.6 Sol and 5 on Claude Opus 4.8, all in autopilot mode.
The subagent model mix was led by Claude Opus 4.8, GPT-5.6 Sol, Claude Haiku 4.5 and GPT-5.5, with Gemini 3.1 Pro and Claude Opus 5 also in the rotation. Part of that mix was determined by the subagent definitions themselves, not by the developer. A separate review pass had Opus 5, GPT-5.6 Sol and Grok 4.6 compare TypeScript and Rust behavior.
The cost of correctness was real. By September 14, GitHub had traced and fixed dozens of regressions, almost all stemming from too-thin end-to-end test coverage. GitHub's five-point operating manual reads like a checklist for anyone running agents on production code: state the full end state, treat end-to-end tests as an oracle that the agent cannot modify, physically separate the agent changing code from the agent guarding tests, translate before redesigning, and when a failure mode shows up twice, write it into standing instructions, a reusable skill or the harness itself.
Sources verified: official GitHub blog post at https://github.blog/ai-and-ml/generative-ai/migrating-the-github-copilot-runtime-to-rust-using-copilot/ ; detailed third-party record at https://cellcog.ai/blog/github-copilot-runtime-rust-rewrite/