The Company That Makes Llama Is Renting Someone Else's Models by the Trillion

Here's the deal: Bloomberg reported on August 20 that Meta Platforms spends hundreds of millions of dollars a year accessing AI models through Microsoft Azure — enough to rank among Microsoft's largest AI customers.

One number conveys the scale. Meta consumes trillions of tokens every week through that channel. Trillions weekly isn't experimentation; that's production workload.

What makes it news is who Meta is. This is the company that led the open-weight camp with the Llama series, that stood up a superintelligence lab and is investing heavily in frontier model development, and that has reportedly been planning its own cloud business. And it's spending nine figures annually renting a competitor's models on a competitor's cloud.

Bloomberg framed it within the industry's circular business dealings — the accumulating web of companies that are simultaneously each other's customers and competitors, which makes genuine external demand hard to separate from intra-industry churn. It's a criticism that keeps attaching to recent AI infrastructure deals.

The Stage: Azure AI Foundry

Azure AI Foundry is Microsoft's AI model marketplace. It serves OpenAI models and many others via API, letting enterprises select from within their existing Azure agreements. As of July, Foundry reportedly had roughly 100,000 customers.

The top-spending names among those 100,000 are the interesting part. Per reporting, ByteDance has generally been the largest spender, with Meta now joining the top tier. Other large customers named include Adobe, Perplexity, and Sierra.

That list reveals Foundry's character. These aren't companies that merely use AI — they're companies that build products with AI. ByteDance has its own models. Adobe has Firefly. Perplexity and Sierra are AI product companies outright. Firms with substantial in-house AI capability are simultaneously buying other people's models at volume.

Meta's specific uses were described too: supporting internal software development, and using OpenAI models through Foundry to evaluate the output of its own in-house models. The second one stands out — using a competitor's model as the judge of your own.

One more piece of context. OpenAI accounts for roughly 70% of Microsoft's total AI revenue. Everything else on Foundry combined is smaller than that single relationship. "One of the largest customers" is best read as a ranking within the remaining 30%.

Why Buy Someone Else's Model When You Build Your Own

There are several answers.

Different models are good at different things. Llama is strong in some areas and OpenAI models in others. For a tool internal developers use as a coding assistant, using whatever works best right now is rational — corporate allegiance isn't a factor. What ships in the product and what the staff uses internally are separate decisions.

Evaluation needs an independent yardstick. Scoring your own model's output with your own model bakes the same biases into the measurement. Using a model from a different lineage as judge escapes some of that. This is standard practice, not something unusual on Meta's part.

In-house infrastructure gets prioritized for training. Meta's GPUs need to go toward training the next generation. Spending that capacity on ancillary inference — internal tooling, evaluation runs — pushes training schedules back. Buying externally is often cheaper in total cost. That's why cloud exists in the first place.

Procurement speed. Turning on an API in an already-contracted cloud is far faster than standing up new serving infrastructure internally. At large companies the real bottleneck in AI adoption is frequently contracts and approvals rather than technology, and the difference is measured in months.

There's a fifth reason that's less visible: you only know where your model stands by using the competition. Benchmark scores don't capture everyday quality. Hundreds of internal developers using a rival model daily generate feedback more specific than any leaderboard. Part of that annual nine-figure spend can be read as paying for that information.

What Each Side Gets

Item Detail
Reported Bloomberg, 2026-08-20
Meta spend Hundreds of millions annually on AI model access via Azure
Usage Trillions of tokens weekly
Channel Azure AI Foundry
Foundry customers ~100,000 (as of July)
Top spenders ByteDance, Meta, Adobe, Perplexity, Sierra among them
OpenAI share of Microsoft AI revenue ~70%
Meta's stated uses Internal software development, evaluating own model output

Microsoft gets revenue diversification. Deriving 70% of AI revenue from one partner is a strength and a risk simultaneously — if that relationship wobbles, so does the number. Large customers like Meta, ByteDance and Adobe thickening the remainder reduces that concentration. And since these are companies with real in-house AI capability, their continued use of Foundry functions as a quality signal for the platform.

Meta gets speed and flexibility. Use exactly what you need when you need it, and switch the moment something better appears. Concentrating in-house infrastructure on training while sourcing ancillary workloads externally is sound resource allocation. It has a price, though: hundreds of millions books directly as a competitor's revenue, and the more internal workflows acclimate to external APIs, the higher the switching cost back to in-house models later.

For OpenAI it's awkward. Its models selling through Azure is revenue, but the buyer building Llama — and using the purchase partly to evaluate Llama — is a different matter. Your model is contributing to a competitor's improvement loop.

Other large customers like Adobe and Perplexity are part of the same story. Each holds its own AI assets while buying heavily on Foundry. The assumption that "we have our own model, so we don't buy others" simply doesn't hold in this industry. Multi-model operation is becoming the default, not the exception.

For the three major clouds collectively, this validates the marketplace strategy. Becoming the storefront for every model rather than only your own pulls competitors' customers into your cloud revenue. Google listing xAI's Grok 4.6 on Vertex AI on August 21 runs on the same logic.

This Relationship Shape Isn't New

Competitor-as-largest-customer is an old pattern in technology.

The canonical case is Samsung and Apple. They fought head-on in smartphones while Samsung remained one of Apple's largest suppliers of displays and memory. Component contracts held even while patent litigation ran in court. It worked because what each got from the other was unambiguous: Apple needed best-in-class components, Samsung needed volume.

Netflix and AWS is the other standard citation. While Amazon competed via Prime Video, Netflix remained a major AWS customer. Netflix's calculation was that renting infrastructure cost less than building it, and spending the difference on content would win. That judgment proved right — and Netflix also paid Amazon a great deal of money every year.

The closest thing to a failure case comes from the early 2010s, when many companies built services on competitors' platforms and got destabilized when the platform owner changed policy or pricing. The mass extinction of apps dependent on social platform APIs is the standard example. The lesson is clean: the biggest risk in running on a competitor's infrastructure isn't price — it's that the right to change the rules belongs to them.

There's also the Apple-Google search default arrangement — two competitors maintaining a long-running commercial agreement, which at sufficient scale attracts regulatory attention. As AI deals between major players grow, similar scrutiny becomes plausible.

For Meta, the platform risk looks relatively low. The usage sits in internal tooling and evaluation rather than a critical product path, and it's the kind of workload that can move to another cloud or in-house at will. But the larger it gets, the more that assumption needs testing.

How the Competitive Picture Moves

Amazon and Google respond with the same strategy — neither chose to sell only its own models. Bedrock and Vertex AI are both multi-vendor marketplaces competing to land large AI companies. If even a company with its own models has to buy externally somewhere, who books that contract becomes an axis of cloud share competition.

Meta's own cloud ambitions are a live variable. Reports of Meta preparing a cloud business have circulated, and the money it currently pays Azure is itself an argument for that business — if your internal consumption alone reaches this scale, insourcing pencils out. But a cloud business needs operations, support, and ecosystem, not just infrastructure, so entry is not simple.

Meta's open-weight strategy may also feel the effect. Part of the case for publishing Llama was widening the ecosystem into a de facto standard. If Meta itself is internally running large volumes of other models, that case weakens. Read alongside indicators showing Meta slipping behind Qwen in the open-weight field, this report reads as a signal that Meta's overall AI strategy is being recalibrated.

For OpenAI, Azure distribution is reach traded against control. Microsoft holding the sales channel widens coverage but makes the customer relationship indirect. That's precisely why OpenAI keeps reinforcing its own API and enterprise sales motion.

The Chinese model camp takes a different route. Open weights spread without passing through gateways like Azure or Bedrock. But in large-enterprise procurement, access via an approved cloud remains overwhelmingly more convenient, so marketplace listings still heavily determine enterprise revenue.

So What Actually Changes

If you design AI infrastructure, take one principle from Meta's choice: don't make training and inference compete for the same resource pool. Allocate owned GPUs to irreplaceable work like training, and buy the substitutable inference — internal tooling, evaluation — externally. That's often cheaper in total, and it applies regardless of company size.

If you evaluate models, consider using a different model lineage as judge. Grading your own model with your own model lets shared weaknesses pass through. Judge models carry their own biases, though, so this doesn't fully replace human review.

If you're negotiating a cloud contract, this report is leverage. Large AI customers concentrating on marketplaces like Foundry means cloud providers want that revenue. There's room to ask for both broad model access terms and usage-based discounts.

If you're an investor, look at AI revenue composition. 70% of Microsoft's AI revenue coming from OpenAI is concentration risk, and how much customers like Meta and ByteDance dilute it determines revenue quality. Factor in too that as circular dealings grow, distinguishing genuine external demand from intra-industry transactions gets harder.

If you run a startup, the practical lesson is where to draw the build-versus-buy line. Even a company Meta's size doesn't insource everything. Narrow what you build to what directly differentiates your product and buy the rest. Making that call early rather than at scale saves substantial switching cost later.

If you watch the industry, the implication is about the limits of self-sufficiency. Even companies building their own models mix several in practice. The picture of one model covering all workloads doesn't hold up, and multi-model operation is likely to remain the default.

🥄 Three Things You're Probably Wondering

— Does this mean Meta doesn't trust its own models? That's an over-read. Internal developer tooling and product-embedded models are separate decisions, and using a different lineage for evaluation is standard bias reduction. That said, if Llama were best at everything, there'd be no reason to spend nine figures annually. The accurate reading is that strengths vary by domain.

— Is this good news for Microsoft? On revenue, yes — the 70%-from-OpenAI concentration gets diluted. The caveat travels with it: a rising share of circular dealings. Revenue from selling within the industry and revenue from outside it have different durability, and the blurring is exactly what makes current AI revenue figures hard to read.

— If Meta builds its own cloud, does this spend disappear? Some of it, not all. Even with your own cloud, using OpenAI models means buying them from somewhere. Insourcing reduces infrastructure cost, not model licensing. And running a cloud business is a different undertaking from buying servers — that transition alone is a multi-year project.

Sources

Numbers and criteria are as of announcement and may change.