Enterprise AI

Microsoft Memo Makes GPT-5.6 the Default, Cuts Anthropic’s Copilot Traffic

Microsoft's internal memo names GPT-5.6 the default AI model for engineers, replacing a GitHub Copilot router that quietly sent workloads to Anthropic's Claude.

LUMIEN4 min read
Microsoft Memo Makes GPT-5.6 the Default, Cuts Anthropic’s Copilot Traffic

Microsoft CoreAI executive vice president Jay Parikh sent an internal memo this week naming OpenAI's GPT-5.6 as the default AI model for the company's engineers. Reported first by 404 Media, the memo follows a May directive that told Windows, Office, and Surface teams to stop using Claude Code and move to GitHub Copilot CLI before the fiscal year ended June 30. The bigger financial hit for Anthropic, however, is the router change: GitHub Copilot's internal auto-router had been quietly sending the bulk of Microsoft's engineering traffic to Claude models by default.

What happened

Detail Fact
New default model OpenAI GPT-5.6
Memo author Jay Parikh, EVP, Microsoft CoreAI
Source of report 404 Media, reporting by Alex Heath
Claude Code wind-down deadline June 30, 2026 (fiscal year close)
Token budget tracking started July 2026
Individual engineer spend range cited A few hundred to a few thousand dollars per month

Parikh’s memo contains the phrase “tokenmaxxing is not what we are optimizing for,” and asks engineers to review their own AI spending. He was also clear it is not about cutting usage: “We are not optimizing for fewer tokens. We are optimizing for more impact per token.” Shifting workloads to OpenAI models, he wrote, gives Microsoft greater value from its token investment.

The switch matters because the router, not individual product choices, determines most of the token volume. According to Alex Heath’s reporting, GitHub Copilot’s auto-router inside Microsoft leaned on Anthropic’s models for the bulk of its work. Thousands of engineers were using Claude tokens simply because Copilot sent them there, not because they selected Claude. Changing the router removes that passive traffic entirely.

Why the router change hits harder than the Claude Code ban

Cancelling Claude Code was visible and deliberate. Engineers knew it was going away, and the decision looked like a licensing cost call ahead of the fiscal year close. The router change is structural. No individual engineer has to do anything differently; the volume just flows to GPT-5.6 now.

Every Microsoft division has carried an AI token budget target since July, and employees can look up their individual spend on an internal dashboard. Microsoft is not alone in discovering this problem late: Uber’s engineering organisation burned through its entire planned 2026 AI coding budget in just four months. Microsoft CEO Satya Nadella acknowledged at a live Hard Fork taping in June that tokenmaxxing was happening inside the company and called the habit addictive.

What Anthropic still has at Microsoft

The relationship is not over. Claude models remain reachable through Copilot CLI. Microsoft Foundry customers still get access to Sonnet, Opus, and Haiku under the deal signed in November. Microsoft has also continued to favour Anthropic inside Microsoft 365 Copilot for tasks where it outperforms OpenAI, and the companies worked together to bring Claude Cowork technology into Microsoft 365 Copilot.

But defaults are where volume lives, and Anthropic’s enterprise business has been built substantially on coding workloads. Rajesh Jha’s May memo framed Copilot CLI as a product Microsoft could shape directly with GitHub. Parikh’s memo completes that picture: Microsoft wanted a tool it controlled and a bill it could predict.

Our take

This is a useful reminder that “available” and “default” are not the same thing in enterprise software. Anthropic still has deals, still has model listings, and still beats OpenAI on some Microsoft 365 tasks. But the passive routing that probably drove a large share of its Microsoft token revenue is gone. For any business evaluating AI vendor risk, the lesson is clear: if you are building on top of a model that sits inside another company’s product, know where the traffic actually comes from and who controls it.

For teams managing their own AI integration strategy, this is also a good moment to audit which models your internal tools are actually calling and what they cost per task. The “auto-router sends it to the cheapest capable model” pattern is sensible in principle, but it creates cost surprises when the vendor changes the routing logic. Track token spend by model and by task type now, before an internal memo forces you to.

If you follow how the open-weight vs. closed-model balance shifts enterprise decisions, our earlier coverage on open-weight AI closing the capability gap adds useful context on why closed-model defaults are increasingly a negotiated position, not a technical inevitability.

What to do about it

  1. Audit which AI models your tools are actually routing to, not just which ones you think you selected.
  2. Set up spend tracking per model and per use case before your bill surprises you the way it surprised Uber and Microsoft.
  3. Define a “frontier vs. non-frontier” task split for your team, matching model cost to task complexity.
  4. If your AI setup relies on a vendor’s auto-router, read the documentation to understand what controls the routing and when it can change.

Source: Bing News · Anthropic

Frequently asked questions

Did Microsoft ban Claude AI for its engineers?

Microsoft told its Windows, Office, and Surface engineers to stop using Claude Code and switch to GitHub Copilot CLI before June 30, 2026. Claude models remain available through Copilot CLI and Microsoft Foundry, but GPT-5.6 is now the default model for internal use.

What is tokenmaxxing and why does Microsoft want to stop it?

Tokenmaxxing refers to using large, expensive AI models for tasks that simpler models could handle, running up token costs without proportional benefit. Microsoft's internal memo from Jay Parikh told engineers to optimize for impact per token, not raw token volume.

What is GPT-5.6 and why did Microsoft make it the default?

GPT-5.6 is an OpenAI model that Microsoft has designated as the default for its engineers in a July 2026 internal memo. The switch is partly about cost control and partly about consolidating workloads with OpenAI, where Microsoft has a major investment, rather than paying Anthropic for Claude tokens through Copilot's auto-router.

Does Microsoft still use Anthropic's Claude models?

Yes. Claude models remain available through Copilot CLI, Microsoft Foundry customers still get Sonnet, Opus, and Haiku under a November deal, and Microsoft continues to use Anthropic models inside Microsoft 365 Copilot for tasks where they outperform OpenAI. The change is to the default routing, not a full removal.

More from AI