The token vendor asked its own staff to burn fewer tokens
On July 29, Satya Nadella closed out Microsoft's fiscal 2026 fourth quarter with a line built for the moment: "We are advancing the frontier on the cost-to-outcome curve, ensuring every customer can turn tokens into business results." Quarterly revenue of $90.0 billion, Azure up 43%, Microsoft 365 Copilot past 30 million paid seats. It read exactly like a company that sells tokens for a living.
Six days later, Emanuel Maiberg at 404 Media published an internal Microsoft email pointing the other direction. The sender was Jay Parikh, the executive vice president who runs Microsoft's CoreAI organization. "Tokenmaxxing is not what we are optimizing for," he wrote. "I want all of us focused on maximizing outcomes that move the needle for our customers and our business."
Here's the deal on the vocabulary: "tokenmaxxing" hardened into a real term over the first half of 2026. It means burning as many AI tokens as possible and treating the volume itself as proof of productivity. Nvidia CEO Jensen Huang poured fuel on it by suggesting a $500,000-a-year employee ought to consume roughly $250,000 in tokens. Andrej Karpathy made it a meme when he admitted on a podcast that he gets "nervous when I have subscription left over," and rotates between Codex and Claude specifically to burn through his allotment. Token spend had graduated from a line item into a status signal.
That's the part that matters about who tapped the brakes. Microsoft is not a token buyer — it's a token seller. It hosts OpenAI models on Azure, resells inference through GitHub Copilot, holds a 27% stake in OpenAI, and disclosed $24.1 billion of fiscal 2026 revenue from commercial arrangements with OpenAI. No company on earth sits closer to the raw cost of a token. When that company tells its own engineers to optimize for impact per token, the same sentence lands considerably harder on every enterprise buying from it.
The memo does three concrete things. As of July 2026, every Microsoft division carries an "AI token budget target." Individual employees can see their own token spend on an internal dashboard. And the default model for internal GitHub Copilot use switched to OpenAI's GPT-5.6 Sol. Parikh added one more sentence that turns out to be the most revealing in the whole document: "Internally, shifting more workloads to OpenAI models helps us get greater value from our token investment." Hold onto that one — it doesn't mean what the headlines assumed.
Who Jay Parikh is, and how big the chair he sits in actually is
Parikh's résumé changes how you read the memo. He ran engineering at Meta — then Facebook — for more than a decade, which means he spent years as the person responsible for bending an infrastructure cost curve while user counts went vertical. He then served as CEO of the cloud security startup Lacework. On January 13, 2025, Nadella announced a new organization called CoreAI – Platform and Tools in an internal memo and named Parikh to lead it. He reports directly to Nadella, and reporting put roughly 10,000 employees under the new group.
The composition of that group is the point. Under Parikh sit Eric Boyd, who runs the AI platform; Jason Taylor, deputy CTO for AI infrastructure; Julia Liuson, president of Microsoft's developer division; and Tim Bozarth, head of developer infrastructure. So the people who build the Copilot and AI stack, the people who sell Visual Studio and developer tooling, and the people who operate Microsoft's own internal developer infrastructure all sit in one reporting line. The token supply side and the token demand side share a manager. That's why this memo reads less like a finance directive and more like a product owner redefining how his own product gets used at home.
The company-level numbers make the restraint look odd at first glance. Microsoft's fiscal 2026 revenue was $331.8 billion, up 18% — a $50.1 billion increase, the largest single-year gain in company history. In Q4 alone: revenue $90.0 billion (up 18%), operating income $40.6 billion (up 18%), net income $35.8 billion (up 31%), diluted EPS $4.81 (up 32%). Microsoft Cloud revenue hit $59.3 billion (up 27%), and Azure and other cloud services grew 43%, pushing Azure past $100 billion in annual revenue for the first time.
The spending column is where the pressure lives. Fourth-quarter capital expenditures came in around $41 billion, up roughly 70% year over year, and CFO Amy Hood said about two-thirds of that went into short-lived assets like CPUs and GPUs. Microsoft is planning roughly $190 billion of capex for calendar 2026, up 61% from 2025. In the same quarter, Hood also disclosed that Microsoft is extending the useful life of office and data center buildings from 15 years to 25 — an accounting change that pushes depreciation further out. Read the two together: the company is shoveling record money into infrastructure while actively managing how fast that infrastructure hits the income statement.
Headcount moved the other way for the first time in a decade. Microsoft's 10-K put employees at 223,000 as of June 30, 2026, down 5,000 from 228,000 a year earlier — the first annual decline since 2016, when it was winding down the Nokia phone business. Product R&D roles fell by 3,000 to 77,000, a second straight annual decline from the 2024 peak of 81,000. Fewer people, more GPUs, and until July, no ceiling on how much inference those remaining people could consume.
And there's an irony sitting right inside the same reporting line. In June 2025, Julia Liuson — who reports to Parikh — wrote an internal memo saying that "just like collaboration, data-driven thinking, and effective communication, using AI is no longer optional — it's core to every role and every level," and instructed managers to fold employees' AI use into holistic performance reflections. Business Insider reported it after viewing the document. Fourteen months later, the same org chart is sending the opposite signal: volume is not the goal. Asking employees to internalize "use it or it counts against you" and then "but not that much" inside the same career cycle is a genuinely hard management ask.
What the memo actually says: budget targets, a dashboard, and GPT-5.6 Sol
Reporting doesn't pin down the exact send date of the email, but the policy was already live from July and 404 Media's story ran on August 4. Here's the mechanism, item by item.
| Element | What the memo established | Nature of the control |
|---|---|---|
| Budget unit | Per-division "AI token budget target," effective July 2026 | A target, not a hard ceiling |
| Individual control | Internal dashboard showing each employee's token spend | Visibility only; no formal personal cap yet |
| Current usage | Internal data shows some engineers consuming hundreds to a few thousand dollars of tokens monthly | The evidence of the problem's size |
| Default model | Internal GitHub Copilot default switched to GPT-5.6 Sol | Model-routing policy |
| Stated goal | "We are not optimizing for fewer tokens. We are optimizing for more impact per token" | Framing |
| What's preserved | The "AI-first" strategy stays; policy keeps adjusting as models and products evolve | The defensive line against "retreat" |
Two design choices are worth pausing on. First, the budget lands on divisions rather than individuals. Give a person a hard personal cap and you immediately create a new game called "I haven't used my allowance yet" — which is tokenmaxxing with extra steps. A division-level target hands allocation authority and accountability to the same team lead. Second, at the individual level Microsoft shipped a dashboard instead of a limit. Make it visible, don't make it forbidden. That's the more sophisticated organizational move, and it also imports a fresh risk: visible spend becomes a ranking, and a ranking becomes a performance signal. Uber demonstrated that failure mode in exactly that order, which we'll get to.
Now the genuinely interesting part — the model switch. Several outlets summarized it as Microsoft moving to "a cheaper OpenAI model." Put the price lists side by side and that summary is only half right, because GPT-5.6 Sol is OpenAI's flagship. It is not the budget tier.
| Model | Input (per 1M tokens) | Output (per 1M tokens) | Position |
|---|---|---|---|
| GPT-5.6 Sol | $5.00 | $30.00 | OpenAI flagship |
| GPT-5.6 Terra | $2.00 | $12.00 | Balanced tier |
| GPT-5.6 Luna | $0.20 | $1.20 | Cheapest tier |
| Claude Fable 5 | $10.00 | $50.00 | Anthropic top end |
| Claude Opus 5 | $5.00 | $25.00 | Anthropic workhorse |
| Claude Sonnet 5 | $2.00 | $10.00 | Promotional through Aug 31, 2026; list $3/$15 |
Read the table literally and Sol's $30 output rate is more expensive than Claude Opus 5's $25. There are two cheaper models inside OpenAI's own GPT-5.6 lineup. If Microsoft were optimizing purely on sticker price, the internal default should have been Terra or Luna. The model Sol clearly beats on cost is Anthropic's top-end Claude Fable 5 at $10/$50 — half the input price and 40% less on output.
That's where the picture snaps into focus, and it depends on what Microsoft engineers were actually selecting inside Copilot. On June 1, 2026, GitHub replaced premium request units with "GitHub AI Credits," a unit derived from the tokens an interaction consumes, priced at each model's listed API rates across input, output, and cached tokens. Which means internal Copilot usage is now metered against public list prices just like a customer's would be. In that environment, moving the default off a Fable-5-class model and onto Sol produces real savings. And Parikh's line — shifting more workloads to OpenAI models "helps us get greater value from our token investment" — is not a statement about unit price. It's a statement about where the money lands. Dollars spent on OpenAI models come partly back to Microsoft, through OpenAI's Azure consumption and through the value of Microsoft's OpenAI stake. Dollars spent on Anthropic models do not return through the same loop as directly.
The scale of that loop is now on public record. In its fiscal 2026 Form 10-K, Microsoft disclosed for the first time that it recorded $24.1 billion of revenue from commercial arrangements with OpenAI, including revenue-sharing payments, plus $6 billion in accounts receivable from OpenAI as of June 30, 2026. The filing does not break that figure down between Azure consumption, revenue share, and other agreements. Outside analysis has estimated the $24.1 billion at roughly 70% of Microsoft's total AI revenue and about 7% of the $331.8 billion company total — an estimate, not a disclosure, but the direction is unambiguous. When internal engineers pick OpenAI models, a meaningful share of the spend circles back inside the house.
Who actually gets paid in this trade
Microsoft collects on three separate layers. The first is plain savings: if a substantial chunk of 77,000 product R&D employees uses Copilot and some slice of them is burning thousands of dollars of tokens monthly, changing one default produces a material annual number. The second is predictability — division-level targets matter less as an absolute cap than as the thing that makes quarterly budgeting possible at all. The third is the attribution effect above: at equal spend, a dollar routed to OpenAI is worth more to Microsoft's own P&L than a dollar routed elsewhere.
OpenAI quietly banks something too. Being the default model inside the internal developer tooling of the world's largest software company is worth more in practice than a leaderboard win. That said, the relationship loosened considerably during 2026. When OpenAI completed its conversion to a public benefit corporation on October 28, 2025, Microsoft took a 27% stake valued around $135 billion, kept technology access through 2032, and OpenAI committed to $250 billion of Azure spending. Then on April 27, 2026, the partnership was restructured again: Microsoft's cloud exclusivity ended, OpenAI became free to license its models on any cloud, and Microsoft's cut of OpenAI revenue moved to a fixed ceiling rather than an open-ended share. The equity and the Azure volume survived; the exclusivity didn't. Steering internal defaults to OpenAI is a way of squeezing value out of the parts of the deal that remain.
Anthropic's position is the strangest one on the board. Losing the internal default is a loss, but the underlying relationship with Microsoft has never been closer. On November 18, 2025, Microsoft and Nvidia announced investments of up to $5 billion and up to $10 billion respectively in Anthropic, while Anthropic committed to purchase $30 billion of Azure compute with the option to contract up to a gigawatt more. The round put Anthropic's valuation in the neighborhood of $350 billion. Microsoft booked a $3.2 billion gain on its Anthropic investment in fiscal Q4 2026. So Microsoft is simultaneously Anthropic's shareholder, its cloud provider, and a company steering its own engineers away from Anthropic's models. In this market the line between competitor and customer stopped meaning anything a while ago.
For engineers, the trade is mixed. The upside is a rule where there wasn't one. Until now the only signal was "not using AI counts against you," with no guidance on how much is appropriate. The downside is the dashboard. Gartner senior principal analyst Nitish Tyagi put the problem plainly: "There is no direct relation between the increase in token consumption and an increase in productivity gains." When a metric that doesn't track productivity becomes visible, it stops being a management tool and becomes a political one. Linear COO Cristina Cordova framed it as ranking engineers by token spend being equivalent to ranking who spent the most money; Khosla Ventures partner Jon Chu called the practice "an absolutely stupid policy." There's a defense too — Y Combinator CEO Garry Tan has said "we've been tokenmaxxing longer than most people" — and Cursor's Edwin Arbus compared token spend to BMI: a useful, fast proxy that's slightly flawed. The debate is genuinely unsettled.
GitHub is a structural winner here. The AI Credits change shipped on June 1 as a customer billing policy, but it's what made per-model, per-person internal metering possible in the first place. That's why a dashboard could even exist. Under the new scheme, Copilot Pro at $10/month includes $15 of credits and Pro+ at $39 includes $70, while existing Business and Enterprise customers get a larger credit allotment for the first three months of the transition, June 1 through September 1, 2026. The move from flat-rate to usage-based billing was the technical precondition for this internal policy.
Microsoft's enterprise customers get something as well: a reference implementation. AI cost control has mostly been a private CFO problem, and now the vendor itself has effectively published its playbook — division targets, plus individual visibility, plus a designated default model. For a procurement lead, that's both a negotiating artifact and an internal-persuasion artifact.
The Uber cautionary tale, and why FinOps is the closest success story
This movie has screened before, and the endings diverge sharply.
The freshest and most painful failure is Uber. Uber told employees to use AI "as much as possible" and ran internal leaderboards ranking usage competitively. The result: the entire 2026 annual AI budget consumed in four months. Some engineers were generating $500 to $2,000 in monthly token bills. On June 2, Uber imposed a $1,500 monthly cap per employee per agentic coding tool — Claude Code and Cursor among them — with usage tracked on a personal dashboard and exceptions available with approval. The more uncomfortable detail came from COO Andrew Macdonald, who acknowledged it remains hard to connect rising AI usage to tangible product improvement even though AI agents now generate roughly 10% of Uber's committed code. The lesson is compact: what you encourage with a leaderboard, you cannot then control with a leaderboard. Reading Microsoft's choice of division targets over individual caps as a direct response to Uber's experience is not a stretch.
The second precedent is Microsoft's own history. In early 2023, the Wall Street Journal reported that Microsoft was losing an average of more than $20 per user per month on the $10-a-month GitHub Copilot, with some heavy users costing as much as $80. Flat-rate pricing plus per-query inference cost was never a solvable equation, and the June 1, 2026 shift to usage-based billing was the three-year-late correction. The instructive part is the timeline. Microsoft learned earlier and more expensively than almost anyone that AI economics break under flat rates — and still ran its internal usage without a ceiling for three more years. Fixing someone else's invoice turned out to be easier than fixing your own.
Third: Accenture. On February 19, 2026, Accenture told senior staff that failing to use AI could affect their promotions — the same grammar as Microsoft's 2025 Liuson memo. Months later, Accenture's agentic AI strategy lead Justice Kwak looked at the data and said, "It's actually not our engineers that are driving the token consumption. It's a lot of the non-engineers" doing things like converting PDFs into presentation slides. Mandate, then spend explosion, then root-cause analysis, then control: that four-step cycle has become the industry's standard path. Microsoft is currently somewhere between steps three and four.
For the success case you have to change categories. When cloud spending went out of control in the mid-2010s, the industry's answer wasn't blanket reduction — it was FinOps. Tagging, showback and chargeback, reserved capacity, and above all the principle that the team spending the money sees the money. The organizations that won didn't use less cloud; they made unit economics visible. Parikh's framing — not fewer tokens, more impact per token — is close to a direct quote from page one of that playbook, and the fact that the approach has a working precedent is genuinely in Microsoft's favor.
There is one decisive difference, though. Cloud costs mapped reasonably cleanly onto countable outputs: instances, storage, requests. Tokens don't yet. Tyagi's other observation is the operative one here: "None of the vendors have incredible features when it comes to cost optimization." Measuring unit economics requires both a denominator and a numerator, and right now the industry only counts the denominator precisely.
The counter-play: the people selling tokens are on the other side of this
The awkward parties in all of this are the model vendors, whose revenue is token consumption. Customer cost discipline is a direct compression of their top line, and the counter-play is already visible in the price lists.
OpenAI cut GPT-5.6 Luna by 80% and Terra by 20% on July 30, 2026, and left Sol untouched at $5/$30. The message in that combination is clear enough: defend frontier pricing, absorb high-volume workloads into the cheaper tiers. When a customer feels cost pressure, the goal is to have them move down your own ladder rather than across to a competitor. Codex has a second structural advantage — coding usage is bundled into the ChatGPT subscription, so a developer doesn't see a separate API invoice. "Predictable monthly" is the best-selling feature in this market right now.
Anthropic is fighting from a different angle. Claude Opus 5 arrived on July 24 at $5/$25 — half the price of Anthropic's most capable model, Fable 5, at $10/$50. In other words, keep the premium tier intact and drop the price of the model people actually run all day, competing on cost-per-capability. Claude Sonnet 5 is sitting at a promotional $2/$10 through August 31, against a $3/$15 list rate, more than a third off. Whether that promotion's overlap with Microsoft's internal default switch is coincidence isn't something we can verify, but it shows exactly where enterprise default-model decisions are being contested.
The third front is the tooling layer. Cached input at 10% of list and the Batch API at 50% off both input and output are now standard across vendors, and community tools have piled on top. The symbol of the moment is the "caveman" plugin 404 Media covered on June 30 — it forces models to answer in extremely terse language instead of polite prose, purely to cut output tokens, and a senior OpenAI employee contributed the code adding Codex support. A vendor employee shipping code that reduces his employer's revenue is a fair summary of the summer of 2026.
Gartner's read is blunter still. Per Tyagi, vendors are leaning into tokenmaxxing to boost the consumption high, monthly AI coding bills are migrating from the $20–$100 per developer range toward $2,000–$5,000, and extreme cases hit $20,000 a month. Gartner's June 24 release reported that nearly a quarter of technology leaders already spend $200–$500 per developer per month on AI coding tokens and about 6% spend more than $2,000, and predicted that by 2028 AI coding costs will surpass the average developer's salary. In markets with lower wages — Gartner cited India — token costs already exceed the salaries of engineers with four to six years of experience.
So the real vendor counter-play probably won't be price at all. It'll be metering: partial retreats toward seat-based pricing, experiments with outcome-based billing, or at minimum a dashboard that explains why a task cost what it cost. Gartner's specific criticism is that vendors lack transparency into how token consumption is calculated and billed. Microsoft building that dashboard internally first is best read as a preview of the same feature arriving in GitHub Copilot for customers.
So what actually changes
For developers, the change is a new instinct: model choice is not free. Opening the day with the strongest available model pinned as your default now leaves a trace on an internal dashboard. Three practical responses. Split models by task type — planning a refactor or chasing a hard bug justifies a top-tier model, while generating test scaffolding or tidying docs often doesn't; GPT-5.6 Luna is one twenty-fifth of Sol's price. Restructure workflows around caching, because cached input at 10% of list fundamentally changes the economics of agent loops that resubmit the same context. And push anything non-urgent through the Batch API for 50% off both directions. Combine all three and the cost of the same output moves by an order of magnitude.
For enterprise AI leads, this memo is a blueprint you can copy nearly verbatim: division targets rather than individual hard caps, visibility rather than prohibition, and a designated default model. Of the three, the default model is the highest-leverage lever in practice — token spend distributions are usually brutally skewed, and changing one default drags the whole curve down. The thing to guard against is equally clear: the moment the dashboard feeds into performance reviews, you have rebuilt Uber's leaderboard. And as Accenture's data showed, the actual epicenter of waste may not be your engineers at all, so look at consumption by department and job function before you decide whom to constrain. On the vendor side, you now have grounds to demand billing transparency as a contract term — Microsoft just conceded that the feature is necessary.
For investors, this reads two ways. Bearishly: the largest beneficiary of the token economy trimming its own internal consumption is the vendor itself expressing doubt about the premise that token value exceeds token cost. Set next to $190 billion of calendar 2026 capex and $41 billion in a single quarter with two-thirds going into short-lived assets, that lands heavily. Bullishly: it's a discipline signal. With the 10-K now showing $24.1 billion of OpenAI-related revenue and $6 billion receivable from a single counterparty, aligning internal token spend toward OpenAI tightens a loop that's already disclosed as a concentration. The metrics to watch next are whether Azure holds anywhere near 43% growth and how Microsoft Cloud margins move from fiscal Q1 2027.
For everyday users, the direct impact is close to nil, but two indirect effects are predictable. One is pricing: the flat-rate-to-usage-based migration already happened at GitHub Copilot, and the same pressure sits on consumer AI products. The word "unlimited" attached to an AI subscription is going to get rarer. The other is response quality. As organizations shift defaults toward cheaper tiers, the AI features inside services you use may quietly move to smaller models — and in products that don't disclose which model answered, there's no way for you to check.
Compressed to one sentence: the industry consensus in the first half of 2026 was that burning more tokens equals more competitiveness, and the company that sold that consensus hardest revoked it internally by summer. If this really is a shift from "spend less" to "impact per token," as Parikh framed it, that's a healthy normalization. The catch is that nobody has a tool for measuring impact per token yet. With the denominator on a dashboard and the numerator still a matter of judgment, the thing that actually gets managed is the denominator — and that's the deepest weakness in this policy.
🥄 Three Things You're Probably Wondering
— So what does this mean for me? If you use AI tools on a company account, your spend is likely headed for a dashboard somewhere. The one thing worth building now is the habit of matching model tier to task, plus a record of what each expensive run actually produced. When controls arrive, the people who can explain their spend are the ones who get exceptions.
— Why is this happening now? Because GitHub started billing Copilot in token-derived AI Credits on June 1, which gave Microsoft the ability to measure internal consumption per person and per model for the first time. Before that there was no data to govern. Layer on Uber torching a full annual AI budget in four months and Gartner's projection that AI coding costs will pass the average developer's salary by 2028, and July became the moment to act.
— Is Microsoft retreating from AI? Hard to argue that. This is a company that spent $41 billion in capex in the same quarter and is holding a roughly $190 billion calendar-2026 plan, and Parikh explicitly kept the "AI-first" framing. What is worth flagging is that designating GPT-5.6 Sol as the default isn't pure unit-price optimization — on list rates Sol's output costs more than Claude Opus 5's, and it isn't even the cheapest model in OpenAI's own lineup. Whether this is cost cutting or partner alignment is too early to call without internal usage data nobody outside Microsoft has.
Sources
- 404 Media — Microsoft Tells Engineers 'Tokenmaxxing Is Not What We Are Optimizing For' (original report, August 4, 2026)
- CNBC — Microsoft makes OpenAI GPT-5.6 Sol default in GitHub Copilot for staff
- The Register — Microsoft tells engineers to curb their token-burning enthusiasm
- The Next Web — Microsoft tells employees to stop tokenmaxxing, sets division-level AI budgets
- Sources (Alex Heath) — Microsoft's tokenmaxxing crackdown (Parikh quote and CoreAI scope)
- Microsoft — Microsoft Cloud and AI strength fuels fourth quarter results (FY26 Q4 official release)
- Microsoft Investor Relations — FY26 Q4 Press Release & Webcast
- CNBC — Microsoft CEO Nadella forms new CoreAI group, led by Jay Parikh (January 13, 2025)
- GeekWire — Microsoft creates new AI platform and tools division, led by former Facebook engineering chief (org structure)
- GitHub Blog — GitHub Copilot is moving to usage-based billing (AI Credits, live June 1, 2026)
- TechCrunch — Uber caps employee AI spending after blowing through budget in four months
- Bloomberg — Uber Caps Employee Spending on AI Tools Like Claude Code to Manage Costs
- Fortune — Uber burned through its entire 2026 AI budget in four months (COO Andrew Macdonald)
- Gartner — AI Coding Costs Will Surpass Average Developer's Salary by 2028 as Token Consumption Surges (June 24, 2026)
- The Register — AI coding agents could soon cost more than the developers using them (Nitish Tyagi quotes)
- 404 Media — The Tokenpocalypse Is Here: Companies Are Scrambling To Stop Spending So Much on AI (Accenture's Justice Kwak)
- 404 Media — Companies Are Making Claude and Codex Talk Like Cavemen to Stop AI's Soaring Costs
- Microsoft Blog — Microsoft, NVIDIA and Anthropic announce strategic partnerships (November 18, 2025)
- Neowin — Microsoft reveals it generated $24.1 billion in revenue from OpenAI in fiscal 2026 (10-K disclosure)
- GeekWire — Microsoft product R&D jobs decline for second straight year, new filing shows (223,000 employees)
- Forbes — OpenAI And Microsoft End Exclusive Partnership And Revenue Sharing (April 27, 2026)
- Claude Platform Docs — Pricing (Anthropic list rates)
- Ramp — The $1 trillion AI spend blind spot (enterprise AI spend trend)
Numbers and criteria are as of announcement and may change. Investment calls are yours to make!



