The line everyone skipped: Databricks runs its own company on Azure now
On July 23, Microsoft and Databricks said they're extending their partnership into the 2030s. Normally that's a headline you scroll past. "Strategic collaboration deepened," "accelerating enterprise AI" — press releases like that ship every week and mean almost nothing.
Here's the thing about this one, though. Buried in it is a commitment that Databricks will run its own core business operations and analytics on Azure Databricks. A data company running itself on its own product is unremarkable. The word doing the work is Azure. Databricks has spent a decade selling itself as the platform that runs identically on AWS, Azure, and Google Cloud — that neutrality is the pitch. And it just told the world its own back office sits on one of those three.
Judson Althoff, CEO of Microsoft Commercial Business, went straight at it: "Databricks' decision to run its own core business operations on Azure Databricks also gives customers confidence in a platform proven at enterprise scale." Yes, that's a sales line. But sales lines like that don't survive legal review unless something in the actual agreement backs them up.
Now for what's missing. There is no dollar figure anywhere in this announcement. No committed spend. No contract value. Not even a hard end date — just "into the 2030s," which could mean 2030 or 2039. Compare that to two months earlier, when Snowflake announced its expanded AWS collaboration as $6 billion over multiple years, number front and center. That asymmetry is half of this story.
Three parties at this table: Databricks, Microsoft, and a piece of Arm silicon
Start with Databricks. Founded in 2013 by the Apache Spark team out of UC Berkeley's AMPLab, it now says more than 20,000 organizations use its Data + AI Platform, including 70% of the Fortune 500. It's the company that made "lakehouse" a category — the idea that you stop maintaining a separate warehouse for structured analytics and a separate lake for everything else, and just run both on one governed storage layer.
What Databricks actually cares about right now is its IPO. Reporting says it closed a roughly $4 billion Series L in December 2025 at a $134 billion pre-money valuation, added about $1.8 billion in debt financing led by JPMorgan in January 2026, and by June was in discussions at something in the $165–175 billion range. Annualized revenue run rate figures floating around — high-$6-billions, growing 80%-plus year over year — trace back to conference remarks rather than filings, and the tallies genuinely differ by outlet. There's no S-1 yet, so there's no way to check. What is certain: for a company about to go public, "locked into a hyperscaler relationship through the 2030s" removes a line item from the risk section.
Then Microsoft. Its FY26 Q4 results, out July 29, explain why it keeps stamping out announcements like this. Quarterly revenue of $90.0 billion, up 18%. Microsoft Cloud at $59.3 billion, up 27%. Azure and other cloud services up 43%. Operating income $40.6 billion, GAAP diluted EPS $4.81 up 32%. Satya Nadella's framing: "This year, Azure revenue surpassed $100 billion for the first time, and Microsoft 365 Copilot reached over 30 million paid seats." Commercial remaining performance obligation hit $678 billion, up 84%.
And Microsoft is stacking partnerships at a pace that's genuinely hard to track. Two days before the Databricks news, on July 21, it announced an expanded strategic partnership with Mistral: a multibillion-dollar commitment to European AI infrastructure, use of Mistral's GPU capacity including thousands of Nvidia Vera Rubin GPUs, and Mistral Medium 3.5 plus OCR 4 landing in Microsoft Foundry with Medium 3.5 also going into Copilot Studio. Brad Smith's line was that "Europe should have access to the world's most capable AI without compromising control over their data, operations or digital future." Mistral on Monday, Databricks on Wednesday. Before that, Anthropic. Before that, the OpenAI renegotiation. There's a pattern here, and we'll come back to it.
The third party isn't a company, it's a chip. Azure Cobalt — Microsoft's in-house Arm server CPU line. Databricks currently runs on Cobalt 100 and, per the announcement, plans to adopt Cobalt 200. This is the quietly important piece. Everyone talks GPUs, but data pipelines, query engines, catalog services, and orchestration all run on CPU. Whoever shaves the CPU bill wins the margin fight, and Microsoft designs its own CPUs now.
What Cobalt 200 actually is: 132 cores, 3nm, and memory encryption you can't turn off
Cobalt 200 wasn't unveiled with this partnership — Microsoft announced it at Ignite. It's built on Arm's Neoverse CSS V3 platform, fabbed on TSMC's 3nm process, with 132 Arm Neoverse V3 cores per SoC, 3MB of L2 cache per core, 192MB of shared L3 system cache, and per-core DVFS for dynamic voltage and frequency scaling.
The performance numbers Microsoft published on its own Azure blog: up to 50% better CPU performance versus Cobalt 100, up to 135% improvement on cloud database workloads, up to 40% better web serving, 20% higher remote storage IOPS with NVMe, and 15% higher network bandwidth. Internally, Dataverse reported 60% performance gains and Azure SQL Database is adopting it too. Worth flagging plainly: these are first-party measurements, not independent benchmarks.
The security design is arguably the more interesting engineering story. Microsoft built a custom memory controller for this chip specifically so that memory encryption is on by default with what it describes as negligible performance impact. That same controller supports Arm's Confidential Compute Architecture (CCA), which is meant to give stronger tenant isolation at lower overhead than purely software-based encrypted-memory schemes. Compression and cryptography engines also moved into the SoC itself — the repetitive grunt work of a datacenter, pushed down into hardware.
On availability: Cobalt 200 VMs are in preview, scaling up to 128 vCPUs. General purpose comes as Dplsv7/Dpldsv7 at a 2:1 memory ratio and Dpsv7/Dpdsv7 at 4:1; memory optimized as Epsv7/Epdsv7 at 8:1; plus two new families — high-memory Mpsv4/Mpdsv4 at 16:1 running 1–84 vCPUs, and storage-optimized Lpsv5 with up to 23TB of local NVMe. Launch regions are West US3, East US2, Central US, Sweden Central, East US, West US2, Spain Central, and Indonesia Central, with more promised. Teradata, Elastic, Arm, and Canonical supplied public endorsements.
So here's the honest read on the infrastructure half of this deal. Databricks plans to adopt Cobalt 200, and the VMs are in preview across eight regions. That's a roadmap, not a completed migration. How much of Databricks' workload actually moves, and when, is not public information.
What was actually in the announcement
Enough narrative. Here are the facts.
| Item | Detail |
|---|---|
| Announced | July 23, 2026 |
| Term | Into the 2030s (specific end year undisclosed) |
| Contract value | Undisclosed |
| Databricks commitment | Run core business operations and analytics on Azure Databricks; build unified lakehouse |
| Infrastructure | Currently on Cobalt 100; plans to adopt Cobalt 200 |
| Cobalt 200 performance | Up to 50% better CPU performance vs Cobalt 100; memory encryption on by default |
| Microsoft commitment | Continue integrating Databricks Data + AI Platform across its products |
| Core integrated products | Genie, Genie Ontology, Unity AI Gateway |
| Integration surface | Entra, Azure Data Lake Storage, OneLake, Power BI, Purview, Foundry, Power Platform, Microsoft 365, Teams, Copilot |
| Named customers | Banco Bradesco, Cincinnati Reds, Electrolux, SMBC, Unilever |
| Databricks scale | 20,000+ organizations; 70% of Fortune 500 |
| Executives quoted | Ali Ghodsi (Databricks co-founder, CEO); Judson Althoff (CEO, Microsoft Commercial Business) |
Ghodsi's full quote: "For nearly a decade, Databricks and Microsoft have helped enterprises innovate with data and AI. Today, our partnership is stronger than ever. With Databricks Genie and Unity AI Gateway deeply integrated across Microsoft's products, we're helping enterprises unify their data and ground AI in business knowledge. This lets customers get the full benefits of agents and models while controlling costs and ensuring governance."
"Controlling costs" isn't filler there. Databricks published its own Azure Databricks update blog the same week, and the substance of Unity AI Gateway is a centralized runtime registry inside Unity Catalog that enforces real-time rate limits, content filtering, and hard spend caps. That's aimed squarely at the actual pain of enterprise AI in 2026, which is not model quality — it's the invoice.
That blog also pins down integration maturity, which the press release doesn't. Querying OneLake data directly through Unity Catalog without copying is GA. Storing managed Delta tables natively in OneLake, with zero-copy availability across Fabric engines, is public beta. Genie for Microsoft Teams and M365 Copilot is beta, delivering context-aware answers from the lakehouse and anchoring M365 Copilot Cowork tasks with Genie Ontology. The Excel add-in with Unity Catalog metric views and permissioned write-back is public preview; native Excel ingestion is GA; the SharePoint connector via Lakeflow Connect is public beta. Around that sit Lakebase (serverless Postgres with copy-on-write database branching), Lakehouse//RT (millisecond-latency serving on the vectorized Reyden engine), CustomerLake (an agentic customer data platform), and the Genie family: Genie Agents, Genie App Builder, Lakeflow Designer, Genie ZeroOps, Genie Code.
Summarize it this way: the plumbing between the two platforms is done, the faucets people actually touch are still under construction. Data exchange is GA. The human-facing surfaces in Teams, Copilot, and Excel are beta or preview.
Who wins what — and the question nobody in the release wants asked
Microsoft is buying context. The most consistent criticism of Copilot over the past two years is that it doesn't know anything about your company. It reads your email and your documents, but it has no idea which SKU slipped in last week's logistics data or by how much. That answer lives in the warehouse, and in large enterprises that warehouse is very often Databricks. Wire Genie into Teams and M365 Copilot and Copilot can, for the first time, cite your numbers. For a company that has to renew 30 million paid seats, that's not a feature — that's retention.
Databricks is buying distribution. Genie could be excellent and still lose, because using it meant logging into a Databricks workspace. The population willing to do that is the few dozen people on the data team. Teams is the window everyone leaves open all day. Excel is the tool finance will die holding. Push Genie into both and reach multiplies by an order of magnitude. Databricks monetizes on consumption, so reach equals queries equals revenue. This is the cleanest possible growth lever to show an IPO roadshow.
Both sides get something from the customer list, too. Banco Bradesco, Cincinnati Reds, Electrolux, SMBC, Unilever — a Brazilian bank, an MLB club, a European appliance maker, a Japanese megabank, a global consumer goods company. Five names spanning wildly different industries, geographies, and regulatory regimes. That's a deliberately constructed reference catalog whose message is "this combination works regardless of your vertical." In enterprise sales that's ammunition.
Now the uncomfortable part. Microsoft sells Fabric. Fabric overlaps substantially with what Databricks does — a figure of roughly 60% functional overlap gets thrown around in the analyst and practitioner community, though nobody can actually verify that number and it swings hard by workload. What's not in dispute is the direction of travel: Databricks bolted on SQL Warehouses and marched into BI territory, Fabric bolted on Spark runtimes and marched into data engineering territory. They are standing on each other's lawn.
And Fabric is not a hobby. Microsoft disclosed at FY26 Q2 that Fabric's annual revenue run rate had passed $2 billion with more than 31,000 customers, and around FY26 Q4 the figure cited was 35,000 paid customers, up 60% year over year (that one comes from earnings commentary rather than the press release body, so treat the citation path accordingly). So Microsoft is pulling Databricks all the way into Teams while simultaneously growing a product that competes for the same budget line at 60% a year.
Which brings us back to the missing dollar figure. Undisclosed terms are normal in partnership announcements. But Snowflake put $6 billion over five years in a headline two months ago. Naming a number buys investor credibility and costs you flexibility. Databricks declining to name one could mean it wants minimal single-cloud concentration on paper before an S-1, or it could mean the commercial terms simply aren't finished. We won't know until the S-1 lands. Anyone asserting which it is right now is guessing.
The same two companies already ran this play in 2017 — and two partnerships that came apart
The success precedent doesn't require looking elsewhere. It's these two companies. On November 15, 2017, at Connect(), Microsoft announced Azure Databricks — Apache Spark based, designed by Databricks in collaboration with Microsoft, shipped as a first-party Azure service. That last bit is the whole trick. Third-party software normally lands in a marketplace with separate contracts, separate billing, separate support. Azure Databricks went into the Azure portal proper: one bill, Microsoft taking first-line support.
General availability followed on March 22, 2018, and the following nine years validated the structure. Databricks stayed multicloud but its Azure presence was disproportionately large, and Microsoft protected big-enterprise data workloads by seating the best product in its own portal instead of forcing a homegrown one. This week's announcement is that formula's second run: in 2017 they put Spark in the Azure portal, in 2026 they're putting Genie in Teams.
Now the cautionary side, and the sharpest example is Microsoft's own most famous partnership. OpenAI. From 2019 it was effectively exclusive in both directions — Azure as OpenAI's compute, OpenAI as Microsoft's frontier model. Then came the October 2025 restructuring. Microsoft's stake was valued around $135 billion for roughly 27% on a diluted basis, but it gave up its first right of refusal on compute, and OpenAI became free to serve its products from any cloud.
Microsoft diversified just as fast. In September 2025 it said Anthropic's Claude models would come to Microsoft 365 Copilot, and from January 7, 2026 Claude was enabled by default for most commercial tenants worldwide, with Anthropic onboarded as a subprocessor under Microsoft's data protection terms. Its in-house MAI model family expanded with Foundry as the distribution point. The lesson is uncomfortable and unavoidable: the deepest AI partnership in the industry was disassembled and rebuilt in six years. Hold "into the 2030s" up against that and read it again.
The second cautionary case sits inside the data industry itself. In October 2013, Microsoft made Azure HDInsight generally available, built on Hortonworks' Hadoop distribution. At the time it looked like the answer for cloud big data. Hortonworks was absorbed into Cloudera via merger in January 2019, and Azure's data narrative moved on through Synapse and then Fabric. Part of that was just Hadoop losing to Spark. The more durable lesson is that a hyperscaler's data roadmap moves independently of any partner's survival. Databricks is not Hortonworks — 20,000 customers and 70% of the Fortune 500 is a completely different position of strength. But that precedent is one plausible explanation for why nobody wrote a number on this contract.
Snowflake, Google, Amazon — and Microsoft competing with itself
Snowflake already picked its side. Q1 FY27, reported May 27: product revenue $1.3343 billion, up 34%, total revenue $1.391 billion, up 33%, net revenue retention 126%, 779 customers above $1 million in trailing-twelve-month product revenue (up 29%, with 46 crossing the line in the quarter), 813 Forbes Global 2000 customers, and RPO of $9.21 billion, up 38%. Full-year FY27 product revenue guidance went up from $5.66 billion to $5.84 billion, or 31% growth. In the same release: the $6 billion multi-year AWS agreement, deepened OpenAI and SAP partnerships, and a planned acquisition of Natoma for AI agent capabilities. Microsoft appears nowhere in it. Two lakehouse giants, two different hyperscalers: Databricks–Azure versus Snowflake–AWS.
Google Cloud played a different card entirely. It struck a strategic AI partnership to make Gemini models available as native products inside the Databricks platform — callable straight from SQL queries and model endpoints, with Gemini usage billable through the Databricks contract. Databricks pushes the broader version of that promise too: committed dollars usable against OpenAI, Anthropic, or Gemini tokens across AWS, Azure, or Google Cloud, with model coverage now extending to Kimi and Grok. So Databricks is signing into the 2030s with Azure while selling cloud-and-model neutrality as a product. That's not a contradiction — it's an intentional double position. It's just not a comfortable one from Redmond's side of the table.
AWS is the quietest player with the most exposure. Historically a large share of Databricks workloads ran there. Databricks publicly stating that its own operations now run on Azure Databricks is an awkward slide for AWS account teams. AWS already matched the symmetry by locking in $6 billion with Snowflake, but the old default assumption — that Databricks runs best on AWS — has started to wobble.
The most interesting competitor is Microsoft itself. Fabric and Azure Databricks chase the same customers, the same budgets, frequently the same meeting. Microsoft's official position is complementarity, and technically the argument holds up: OneLake speaks Delta, Unity Catalog queries OneLake without copying (GA), and Fabric can consume Databricks data via mirroring or direct Delta access. Running both without duplicating data is genuinely achievable now.
But technical coexistence and commercial coexistence are separate problems. Out in the field, when a Microsoft rep carrying a Fabric quota gets asked "so should we be on Databricks instead?", that conversation does not follow the documentation. Whether this partnership changed anything real will be settled not by the press release but by how the two products get positioned in the field a year from now.
So what actually changes for you
If you're a data engineer or platform owner, two things deserve your attention this quarter. First, how genuinely zero-copy the OneLake ↔ Unity Catalog path is in practice — querying is GA, managed tables in OneLake are public beta, and those maturity levels behave very differently under production load. Second, whether Cobalt 200 actually lowers your bill. Don't take the 135% database figure on faith; it's a first-party measurement, and Arm migration burns human hours on container images, JIT behavior, and native library compatibility. Eight preview regions is also a real constraint if your data residency requirements don't intersect that list.
If you're a developer, the Genie lineup is the concrete part. Genie Agents let non-technical users build reusable personal agents. Genie App Builder gives a low-code governed environment. Lakeflow Designer builds pipelines from natural language. Genie ZeroOps provisions infrastructure autonomously, and Genie Code positions itself as a debugging and optimization partner. All of it governed through Unity Catalog — that's the pitch. The practical translation: a meaningful share of data-access requests will arrive as Teams messages instead of tickets, and what you end up maintaining is not queries but the semantic layer and the permission model. Genie Ontology claims to auto-extract table relationships, column metrics, and query popularity signals to eliminate manual curation. Fine. When the auto-extracted ontology is wrong, fixing it is still a human problem, and now it's wrong at conversational speed.
If you're a data leader or enterprise buyer, there is one genuinely usable clause here: Unity AI Gateway's hard spend caps. The most common way an AI pilot dies in 2026 is an unforecast token bill, and being able to enforce a ceiling at the gateway changes the finance conversation. But flip it around for procurement. This announcement means the two vendors are more tightly bound, which means your leverage went down. When Azure Databricks renewal comes up, "we could move to AWS" may not carry the weight it did last year.
On the investment angle, carefully. The announcement helps the Databricks IPO narrative, and how a hyperscaler relationship running into the 2030s shows up in the S-1 risk factors will be worth reading closely. On Microsoft's side, this flows through as Azure revenue — and the pattern of a partner spending on a partner's cloud, with that spend recognized as revenue, has been contested since the OpenAI and Nvidia deals of 2025. Here the scale isn't disclosed, so there's not even enough information to apply the critique. An announcement with no number can't go into a financial model. That's the honest conclusion.
If you just work at a company that uses this stuff, the effect is simple. Within a few months you may be able to type "what was churn in my region last month" into Teams and get an answer without opening a dashboard. That's good. It also introduces a new failure mode: a natural language answer hides its own definition. Nobody sees how churn was calculated. A dashboard with a wrong metric definition eventually gets caught by someone; a chatbot answer with a wrong metric definition goes straight into a deck. And most of these surfaces are still in beta — gate your internal rollout decisions on GA, not on this press release.
🥄 Three Things You're Probably Wondering
— So what does this mean for me? No direct impact. But if your company runs both Microsoft 365 and Databricks, there's a decent chance you'll be querying internal data in plain English from a Teams window within a few months. When that happens, get in the habit of questioning how the number in the answer was defined.
— Isn't this just Databricks getting locked into Azure? The signals are real — it moved its own operations onto Azure Databricks and committed to a CPU Microsoft designed. But the same company is embedding Google's Gemini natively in its platform and letting committed dollars be spent across AWS, Azure, or Google Cloud. And no contract value was disclosed. Calling it lock-in is premature.
— Does one of Fabric or Databricks eventually die? On the current numbers, both are growing — Fabric's paid customer count was cited as up 60% year over year, and Databricks holds 70% of the Fortune 500. Technically, OneLake and Unity Catalog now connect without copying data at GA, so running both is viable. The sales-floor collision over the same budget hasn't gone anywhere, though, so the real answer shows up in field positioning a year out.
Sources
- Databricks and Microsoft expand partnership to help enterprises bring business context to enterprise AI — Microsoft Source
- Databricks and Microsoft Expand Partnership to Help Enterprises Bring Business Context to Enterprise AI — Databricks newsroom
- Unifying Data and Governance in the Agentic Era: What's New with Azure Databricks — Databricks blog
- New Azure Cobalt 200 VMs deliver 50% performance improvement, fully optimized for modern agentic AI workloads — Microsoft Azure blog
- Announcing Cobalt 200: Azure's next cloud-native CPU — Microsoft Community Hub
- Microsoft Cloud and AI strength drives fourth quarter results (FY26 Q4) — Microsoft Investor Relations
- Snowflake Reports Financial Results for the First Quarter of Fiscal 2027 — Snowflake Investor Relations
- Microsoft announces Azure Databricks powered by Apache Spark (2017) — Microsoft Source
- Databricks Announces Strategic AI Partnership with Google Cloud to Bring Gemini Models Natively to the Data Intelligence Platform — Databricks newsroom
- Microsoft and Mistral expand strategic partnership to give enterprises and regulated industries frontier AI they can control — Microsoft Source
Numbers and criteria are as of announcement and may change. Investment calls are yours to make!



