An Oil Painter Built a Village of 25 Fake People. Now He Wants to Build 8 Billion.
Back in the spring of 2023, a Stanford PhD student built a little pixel town that looked like it had been ripped out of The Sims. He called it Smallville. Twenty-five residents lived there, except none of them were people — each one was an agent driven by a language model. They woke up, made breakfast, went to work, ran into each other at the cafe, and at the end of the day they reflected on what had happened and planned tomorrow. The researchers planted exactly one seed into this world: one agent wants to throw a Valentine's Day party. Over the next two simulated days, invitations spread through the town on their own. Agents asked each other out. They showed up at the right place at the right time. Nobody scripted any of that. The paper was called Generative Agents: Interactive Simulacra of Human Behavior, the first author was Joon Sung Park, and it took the Best Paper award at UIST 2023.
Three years and change later, on July 30, 2026, the company that grew out of that paper announced a Series B of more than $200 million at a $2 billion post-money valuation. The company is called Simile. TechCrunch reported the round as led by Greenoaks; Simile's own announcement says it was co-led by Greenoaks and Index Ventures, with participation from Hanabi, Bain Capital Ventures, A*, Factory, CVS Health Ventures, and Definition. That small discrepancy is worth noting, but the timing is the part that made people sit up. Simile only came out of stealth on February 12, 2026, with a $100 million Series A led by Index. Five months later the valuation had doubled, and total disclosed funding hit $300 million.
Here's the mission, in the company's own words: simulate all eight billion people on earth, accurately and honestly. That reads like a pitch-deck flourish, except Simile treats it as an engineering roadmap. Co-founder and Stanford professor Percy Liang published a piece the same day arguing that after the age of prediction and the age of reasoning, we're entering the age of simulation — and that simulation, not raw capability, is the honest path to robust superintelligence. What Simile sells isn't a survey replacement so much as a foundation model of human behavior. CVS Health, Deloitte, Gallup, and Wealthfront are already running on it, and the company says tens of millions of simulations have gone through the system for Fortune 100 customers.
But if we stop there, this is just another funding post. The interesting questions are underneath. Why is this kind of valuation showing up in market research, of all places? What exactly is the "85% accuracy" number that gets quoted everywhere actually measuring? Revenue grew 5x in five months — five times what, though? And the one that matters most: can AI actually replace asking people things? We pulled the original papers, the company's own posts, and a 2026 benchmark study that goes straight at this methodology, and read them side by side.
Who's Actually Standing on This Board
Start with Simile itself. It began in Palo Alto and is now a global team of 50-plus. The company describes itself, flatly, as "The Simulation Company." There are three co-founders and the lineup is unusual. CEO Joon Sung Park used to paint in oils. Index Ventures partner Shardul Shah, writing about the Series A, described him as someone with a genius for holding contradictions comfortably — deeply creative but operationally relentless, sky-high in ambition but grounded, ruthlessly competitive but humane. The Smallville paper became one of the most-cited AI papers of its moment, and it became the company's founding artifact.
The other two co-founders are Stanford faculty. Michael Bernstein is a leading human-computer interaction researcher and was senior author on the Smallville work. Percy Liang led Stanford's Center for Research on Foundation Models. Index flagged a detail that's easy to skip past but does a lot of work in the narrative: this founding group introduced generative agents, rich agentic simulation, and literally coined the term "foundation model" in the 2021 CRFM report. So the people who named the category that the entire AI industry now runs on are saying the next frontier is simulation. If you're an investor, that's a heavy piece of framing to argue with.
On the capital side, Greenoaks is known in the Valley for concentrated, quiet, long-hold positions rather than spray-and-pray. Index led the A and doubled down here. Shah's line from February is the one that stuck with people: he hadn't felt market pull like this since the early days of Wiz — a reference to one of the fastest enterprise ramps in cybersecurity history. Translated out of VC-speak, that means the sales cycle is short and large enterprises are knocking first rather than being chased.
The customer list isn't just a logo wall either. CVS Health is a customer and an investor via CVS Health Ventures. Gallup — the company that has been polling humans for ninety years — is on that list too, which is either an endorsement or an irony depending on your mood. Deloitte covers professional services, Wealthfront covers fintech. Healthcare, consulting, and finance don't have much in common on the surface, but they share one thing: these are industries where a single wrong decision is catastrophically expensive, which is exactly the market where "let's test it in simulation first" sells itself.
Finally, the backdrop. ESOMAR pegs the global insights industry at roughly $153 billion — about $56 billion in market research proper, $62 billion in research software, and $35 billion in reporting. It grew from around $130 billion in 2022 to $142 billion in 2023 and has kept climbing. Companies like Simile aren't trying to take a slice of that pie. They're trying to change its unit economics, turning something that took weeks and six figures per study into something that takes hours and an API call.
What Was Actually Announced, Line by Line
Simile's own post lists five accomplishments from five months of operating: grew revenue 5x, built a new foundation model for human behavior that has run tens of millions of simulations for Fortune 100 enterprises, trained a first-of-its-kind confidence model that predicts the accuracy of every simulation, shipped the first product that lets organizations verifiably predict the future, and scaled from a small house in Palo Alto to 50-plus people globally. The item that got the least press attention is the one that matters most commercially: the confidence model. Producing a simulation result and producing a calibrated estimate of how much you should trust that specific result are entirely different products. Enterprises pay for the second one.
The CVS Health case study is the most concrete thing the company has published. Simile says its agents are built on 2.9 million consented responses from more than 400,000 participants across 200-plus behavioral scenarios, modeled on real people's interview responses and past choices. CVS used that four ways: reproducing conclusions from prior research fast enough to pre-screen ideas and reserve fieldwork for the promising ones; testing end-to-end journeys — pharmacist access, wait times, communication clarity — to isolate which levers actually move NPS; testing reminder cadences, education content, and benefit designs against intent to refill before spending money on pilots; and benchmarking brand perception against grocery and mass-market competitors. The company specifically calls out value for hard-to-reach populations like chronic-condition patients who are slow, expensive, and unevenly represented in traditional panels.
Now the accuracy number. The "85%" you see quoted everywhere traces to arXiv 2411.10109, authored by Park, Liang, Bernstein and colleagues. The study recruited a diverse national sample of 1,052 Americans and ran two-hour semi-structured interviews using the American Voices Project schedule, plus structured surveys including General Social Survey items and the Big Five inventory. Agents were built from interviews only, surveys only, or both, and then tested on held-out GSS items. Results: interview-only agents 83%, survey-only 82%, combined 86%, and demographics-only agents 74%.
Here's the thing, and it's the single most misread number in this entire story. Those percentages are not hit rates. The paper's benchmark is participants' own two-week test-retest consistency — how well a real person reproduces their own answers when you ask them the same questions fourteen days later. Set that to 100, and the agents land at 83 to 86. So the honest sentence is not "the AI is 85% accurate." It's "the agent reproduces you about as well as you reproduce yourself, minus about 15%." That is still a genuinely striking result. It's also a completely different claim from the one that gets made in boardrooms. And the paper itself notes that combining interviews and surveys only produced modest gains over either alone, which suggests predictive benefit asymptotes once the model has seen enough evidence in a domain.
| Item | Detail | Basis |
|---|---|---|
| Round size | More than $200 million | Simile official post ("over $200 million") |
| Valuation | $2 billion post-money | Simile post and TechCrunch agree |
| Lead investor | Official post says Greenoaks and Index co-led; TechCrunch says Greenoaks led | Wording differs between the two sources |
| Other participants | Hanabi, Bain Capital Ventures, A*, Factory, CVS Health Ventures, Definition | Simile official post |
| Prior round | $100M Series A led by Index Ventures, Feb 12, 2026 | Index Ventures perspective post |
| Revenue | 5x growth in five months (starting base not disclosed) | Simile official post |
| Headcount | 50+, global | Simile official post |
| Customers | CVS Health, Deloitte, Gallup, Wealthfront | Simile and Index official posts |
| Data scale | 400,000+ participants, 2.9M consented responses, 200+ scenarios | Simile CVS case study |
| The "85%" | Actually 83–86% of participants' own two-week test-retest consistency on held-out GSS items | arXiv 2411.10109 abstract |
| Counter-evidence | No LLM beats the strongest non-LLM baseline at the individual level; segment gaps inflated 2–4x | arXiv 2607.26348 abstract |
That last row is the uncomfortable part. A 2026 benchmark paper titled When Synthetic Users Fail ran four models — spanning 8B to frontier scale, across two model families — against real human response data from the General Social Survey and the World Values Survey. Two failures replicated across every model, both families, and both domains. First, at the individual level, no LLM beat even the strongest non-LLM baseline fit on held-out human data; on cross-cultural values every model fell well below it. Second, models systematically over-determine demographics — they treat identity as far more predictive of attitudes than it actually is in real people. On a segment-targeting task, that distortion inflated between-segment gaps by two to fourfold, would have pointed a team at the wrong segment in half of the U.S. cases and most cross-cultural ones, and manufactured segment splits that don't exist in real humans. Scaling the model did not fix either failure.
The two research lines aren't actually contradicting each other, and that's the key to reading this story correctly. Simile's work is about agents grounded in two hours of real interview data from a specific named person. The critical benchmark is about LLMs prompted with demographics and told to role-play a population. The fault line in this whole industry isn't model quality — it's how honestly you sourced the human data underneath. That's almost certainly why Simile puts "2.9 million consented responses" and a public participant agreement front and center. It's a signal that they took the expensive road.
Who Wins, and Who Quietly Loses
The founders and early investors are the obvious winners. Index put $100 million in during February at an undisclosed but presumably sub-billion valuation and re-upped five months later at $2 billion. The paper markup is great, but the more consequential thing is how the company got priced: not as a research SaaS business, but as a foundation model company. Those two categories trade at multiples that aren't in the same neighborhood. The founding team's strategy of leading with papers, terminology, and research lineage converted directly into valuation, and other AI founders are absolutely taking notes.
CVS Health occupies a strange and interesting seat here: customer and shareholder simultaneously. That's good for both sides in the short run — it de-risks the anchor account and gives Simile a healthcare reference nobody else can easily get. But it carries a known risk in early enterprise AI, which is that the product quietly over-fits to one investor-customer's use cases and stops generalizing. On the other hand, healthcare is a genuine moat: it's regulated, the data is sensitive, and the CVS post explicitly highlights value in "sensitive health behaviors and privacy-constrained data that are difficult to study directly in the wild." That's a defensibility claim dressed as a case study, and it's a good one.
The quiet losers are traditional panel operators — the firms that maintain online respondent pools, pay per-completion incentives, and run fieldwork. Their cost structure scales linearly with the number of responses collected, and Simile is selling a product designed to break that linearity. This won't happen overnight, and there's an irony worth noting: Simile's approach actually requires buying better human data at higher cost per participant, not less human data. But once budgets get reallocated so that exploration happens in simulation and only validation happens in the field, total fieldwork volume goes down. Industry surveys already show the tension — researchers report near-universal use of AI tools in their workflow while reporting much lower trust in synthetic participants specifically.
And then there's a stakeholder almost nobody discusses: the participants themselves. Simile's models are built on real people's two-hour interviews and past choices. Once your digital twin exists, you never have to be recruited again — the twin just keeps answering, potentially thousands of times, for years. The company states responses are consented and publishes a participant agreement, but the compensation structure for ongoing use of a person's twin isn't something we could verify from public materials. This is the same fight that already blew up around music and image training data, except the training data here is somebody's personality, politics, and taste. It has more room to get uglier.
Inside enterprises, one group's job changes shape rather than disappearing. Consumer insights teams shift from "the org that goes and collects data" to "the org that validates simulations and rules on when to trust them." That's a role change, not a downgrade — but teams that fail to make the transition may find their budget line absorbed into marketing or data science. Simile's confidence model is genuinely good news for these people, because it hands them the instrument they need to make that ruling defensible.
We've Run This Movie Before — the Wins and the Wrecks
Start with a win. Automotive crash simulation actually did change an industry's unit economics. Developing a new vehicle used to mean driving dozens of physical prototypes into walls. As finite element analysis matured, most crash scenarios moved inside the computer. Physical crash tests never went away — regulatory certification and final validation still demand real steel — but the division of labor settled into explore in simulation, validate in reality, and development cycles and costs collapsed. Semiconductors tell the same story: without SPICE-class circuit simulators and modern EDA, designing a chip with billions of transistors wouldn't be conceivable. The trajectory Simile is drawing is exactly this one.
The second precedent is a little ironic given the customer list. Gallup's founding myth is a story about beating brute-force data collection with better sampling. In the 1936 U.S. presidential race, Literary Digest collected 2.4 million responses and confidently predicted a Landon victory. George Gallup, working with a far smaller sample, called Roosevelt correctly. The lesson was that sample size is not the thing — how well your sample represents the population is the whole ballgame. Simile's own paper makes a structurally similar argument: agents grounded in real interviews (83%) beat agents built from demographics alone (74%), and notably reduce accuracy disparities across racial and ideological groups.
Now the wrecks. Google Flu Trends is the textbook case, and it should be required reading for anyone selling behavioral prediction. Using search queries to nowcast flu prevalence worked beautifully at first. Then in February 2013, GFT predicted more than double the proportion of influenza-like-illness doctor visits that the CDC actually recorded — a model built specifically to predict CDC numbers missing CDC numbers by 2x. Lazer, Kennedy, King, and Vespignani's Science paper named the failure mode "big data hubris": assuming scale substitutes for validity checks, and ignoring that the environment generating the data (Google's own search algorithm, which kept changing) was drifting underneath the model.
The second wreck is the 2016 U.S. election polls, and the shape of that failure is the one that should worry Simile's customers most. The AAPOR evaluation found a striking asymmetry: national polls were actually accurate by historical standards. What broke was state-level polling, driven mainly by a late swing in vote preference and a widespread failure to adjust for over-representation of college graduates. In other words, being right in aggregate tells you nothing about being right at the segment level. And that's precisely the failure mode the When Synthetic Users Fail benchmark documented for LLM respondents. It isn't a coincidence — it's the same statistical trap. Enterprises don't buy simulation to learn about the average consumer; they buy it to learn about this specific segment, which is exactly where the method is weakest.
How the Competition Punches Back
The most direct rival is Aaru, a New York company founded in March 2024 by a group of teenage founders. It raised a Series A led by Redpoint in December 2025 at a reported $1 billion "headline" valuation, though multiple outlets reported the round used a multi-tier structure that put the blended valuation below $1 billion, on under $10 million of ARR — so comparing the two headline numbers directly would be sloppy. Aaru's likely counter-play is speed and price. Interview-grounded agents are slow and expensive to build; the alternative strategy is to ship "good enough" predictions at a fraction of the cost and win on volume.
The second front is the incumbent research infrastructure. Qualtrics has already bolted synthetic respondents onto the most widely deployed survey platform on earth, and industry surveys suggest a majority of market research professionals have touched synthetic responses in some form. Their weapon is distribution. A startup can build a demonstrably better model and still lose if the platform already embedded in every enterprise survey workflow ships a comparable feature as a checkbox. Simile leading with a confidence model and "verifiably predict the future" language is the natural response: if you can't win on distribution, force the fight onto provable quality where the checkbox feature can't follow.
The third threat is a different animal entirely — the frontier labs. If OpenAI, Anthropic, or Google DeepMind decide persona simulation is a feature rather than a market, the bottom of this category gets vaporized. But Simile's positioning here is genuinely clever. Frontier labs are optimizing for models that are smarter and more correct. Simile wants models that fail the way an ordinary person fails — carrying the same biases, making the same mistakes, holding the same mediocre taste in breakfast cereal. That objective is close to the inverse of RLHF. Heavily aligned models are documented to collapse toward safe, socially desirable, average answers, which is poison for simulation fidelity. Building a deliberately un-optimized model is not something a frontier lab can do as a side project without hurting its main product.
The fourth counterparty isn't a competitor so much as a referee, and referees can end games. Polling and social research bodies keep publishing the same warning: calling a synthetic sample "representative" without valid anchor data is methodologically indefensible, and for seldom-heard groups where no reliable ground truth exists, synthetic samples compound the equity problem rather than solving it. One high-profile corporate blowup traceable to a synthetic-data-driven decision could push this industry's adoption curve back several years, regardless of how good anyone's model is.
Finally, watch Gallup. Right now it's listed as a customer. But an organization sitting on ninety years of time-series data and a brand built entirely on measurement credibility has to eventually ask whether it wants to run on somebody else's model. When data owners start building their own, the board resets. Media companies licensing content to AI labs walked this exact path before discovering they'd rather own the leverage.
So What Actually Changes
For developers and ML engineers, the concept that shifts is evaluation. Product A/B testing has always required live user traffic, which means shipping first and learning second. A mature simulation layer lets you throw a change at a synthetic population before deployment. One caveat you should tattoo somewhere: simulation is pre-screening, not evidence. Even Simile's own CVS write-up describes the role as narrowing candidates so fieldwork can be reserved for the promising directions — not eliminating fieldwork. Cross that line and you get to rediscover Google Flu Trends personally.
For investors, this round says two things. First, in 2026 the market still pays a hefty premium for research lineage. A single paper became a company's narrative, and that narrative got the business priced on AI-infrastructure multiples rather than research-SaaS ones. Second, doubling a valuation in five months is probably more a function of competitive round dynamics than of operating results. 5x revenue growth is real and impressive, but the starting base wasn't disclosed, and 5x from a company that exited stealth five months ago can still be a small absolute number. That's not a knock on Simile — it's just how you read the number honestly.
If you run an enterprise function, the thing to prepare isn't whether to adopt but how you'll adjudicate. Decide in advance which questions get simulated and which never do, what you trust when simulation and fieldwork disagree, and what confidence threshold automatically escalates to real humans. Translating the benchmark paper's warning into operational language: simulation is comparatively safe for directional and aggregate questions, and dangerous for decisions that allocate budget based on the size of gaps between segments. If your model inflates those gaps two to fourfold, you'll pick the wrong segment roughly half the time — and you'll do it with a very confident-looking chart.
For ordinary people, nothing changes tomorrow. But the direction is worth knowing. Over the next few years, a rising share of the features in your apps, the wording in your insurance documents, and the cadence of your pharmacy's refill reminders will be decided based on simulations of people like you. That's not automatically bad. Done well, it means fewer useless features shipped and fewer occasions where you're the unwitting subject of a live experiment. The failure mode is when the simulation reduces you to a demographic label and confidently asserts that "people like you think X" — which is exactly the distortion the counter-benchmark measured, and it's a representation problem before it's a technical one.
For researchers and policymakers, the clock is already running. The capital is deployed, the product is inside Fortune 100 workflows, and the argument about whether synthetic respondents are permissible is basically over. The live question now is what disclosure standard applies. What has to appear in a report: the provenance of the sample, the nature of the grounding data, calibrated intervals, whether any human validation was run at all. Without a minimum standard, we end up a few years from now with a market full of decks where "AI-simulated" quietly does the work that a methodology section used to do.
🥄 Three Things You're Probably Wondering
— So what does this mean for me? Not much directly. But an increasing share of the product decisions and marketing copy you encounter will be based on how a digital twin of someone like you responded, rather than on an actual survey. The felt difference is subtle — mostly a change in how often you think "why did they build that?"
— If it's 85% accurate, do we still need surveys? That 85% (really 83–86%) isn't a hit rate. It's a ratio against how consistently a real person reproduces their own answers two weeks later. And a separate benchmark found LLMs losing to conventional statistical baselines at the individual level while inflating segment gaps two to fourfold. Explore in simulation, validate decisions with real humans — that's the safe line right now.
— Doubling the valuation in five months, isn't that a bubble? Too early to call. The 5x revenue growth is real, but the starting base is undisclosed, and $2 billion looks more like a price on the "foundation model company" framing and the founders' research lineage than on trailing performance. On the other side, having your anchor customer show up as an investor is a real signal that deployment isn't shallow. The next two quarters of renewals will tell you more than this round does.
Sources
- Announcing Simile's Series B — Simile official blog
- Simulation: The Next Frontier for AI — Percy Liang, Simile official blog
- CVS Health x Simile: Simulations for faster, safer decisions
- Simulating Society at Scale: Our Investment in Simile's $200M Series B — Index Ventures
- Life, the Universe, and Simile: Leading Simile's $100M Series A — Index Ventures
- LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals — arXiv 2411.10109
- Generative Agents: Interactive Simulacra of Human Behavior — arXiv 2304.03442
- When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses — arXiv 2607.26348
- An Evaluation of 2016 Election Polls in the United States — AAPOR full report
- The Parable of Google Flu: Traps in Big Data Analysis — Science, 2014
- Inside the $153bn Insights Industry — ESOMAR Research World
- Synthetic-user startup Simile raises $200M at $2B valuation — TechCrunch
Numbers and criteria are as of announcement and may change. Investment calls are yours to make!



