This week, 200 ordinary Koreans sit down and grade the national AI
On the morning of August 8, two hundred people recruited under gender and age quotas will sit down in front of screens. Their job for four days is deceptively simple: use four Korean-built AI models and score them on an absolute scale. Not a panel of experts. Not an automated benchmark harness. Just regular people, using the thing and rating it. Those scores get added to benchmark scores and expert scores, and the combined result lands around August 12. When it does, one of the four teams is out.
Exactly one year ago today, that tournament started. On August 4, 2025, Korea's Ministry of Science and ICT announced five elite teams for its Sovereign AI Foundation Model project — Naver Cloud, Upstage, SK Telecom, NC AI, and LG AI Research. Fifteen teams had entered by the July 21 deadline. A paper review cut that to ten. A presentation evaluation cut it to five. The government promised GPUs, data, and salary support, and attached one condition: fail a stage evaluation and you leave.
That condition turned out to have teeth. The first-stage results came out on January 15, 2026, and the advance guidance had been "five teams down to four." When the envelope opened, only three were left. Naver Cloud and NC AI went out together. The company that essentially opened Korea's LLM era with HyperCLOVA X had been cut from the national team. The government reopened recruitment to backfill, and on February 20 a thirty-person startup named Motif Technologies came in as the fourth team. The four now in the ring are LG AI Research, SK Telecom, Upstage, and Motif.
Here's what this piece is actually about. Why a 213.6 billion won national program was designed as an elimination bracket rather than a grant. How that design bent the behavior of everyone in it over twelve months. Why Naver was cut on "originality" rather than performance, and what that signal really said. Why four large Korean models suddenly dropped inside ten days in late July — not a coincidence. What a mid-size economy is actually buying with this money. And finally, the far more aggressive alternative reportedly under review inside the government right now: stop splitting the GPUs and hand roughly ten thousand of them to a single team.
Four in the ring, two on the outside, and the referee
Start with the referee. MSIT's framing is blunt: a country that runs entirely on models built in the US and China is dependent, and dependence is a policy problem. Per the government's own policy briefing, the program runs up to three years with a total budget of about 213.6 billion won, funding GPUs, data, and talent through December 31, 2027. There's a stated performance target too — at least 95% of the capability of the best global models. And when MSIT explained why it picked those original five, it cited three things: sovereign AI, open source, and ambitious scaling. It specifically credited the teams for proposing "a high level of open source policy allowing other companies to use the resulting models commercially." Hold onto that. It explains the Apache 2.0 flood that came later.
Team one is LG AI Research, which finished first in round one. It took the top mark in all three components — 33.6 out of 40 on benchmarks, 31.6 out of 35 from experts, and a perfect 25 out of 25 from users — for a total of 90.2. Its round-one submission was K-EXAONE at 236 billion parameters. For round two it brought K-EXAONE 2.0, the 750-billion-parameter model (37 billion active) it dropped on Hugging Face under Apache 2.0 on July 31, the one spoonai covered on August 2. It's the largest model any Korean lab has published, with a 262,144-token context window.
Team two is SK Telecom. In round one it fielded A.X K1, Korea's first model in the 519-billion-parameter class. For round two it brought A.X K2, released July 29 at 688 billion parameters with 33 billion active. SKT's play has been consortium width — pulling SK AX and Technomatrix in to thicken the industrial deployment side rather than chasing raw size alone. Team three is Upstage, the only startup among the original five, which tied LG at the top on some round-one benchmarks. Its Solar Open2, released July 23, is far smaller at 250 billion total and 15 billion active parameters, and it leans hard into agent capability, operating efficiency, and an actual deployment on the Daum portal.
Team four is the wildcard. Motif Technologies joined late, developing from February through July — the government explicitly equalized the clock, since the original three had January through June — and then detonated the standings on July 21 with a preview of Motif 3, a 314-billion-parameter model. It scored 44 on Artificial Analysis's composite intelligence index, tying DeepSeek's latest release, landing in the top three among open-weight models and eleventh overall including closed frontier models. That came from a company with roughly thirty employees, in five months. The stated reason it was selected matters too: pure proprietary architecture, built without borrowing foreign open-source designs. Its government package was 768 B200 GPUs plus 1.75 billion won of bespoke data development and 10 billion won of shared data procurement.
Outside the ring are two. NC AI's exit was ordinary — it simply fell short on the combined benchmark, expert, and user score. Naver Cloud's was not. Naver scored in the top tier on performance and still failed, on an "originality" gate. Its submitted HyperCLOVA X SEED 32B Think model used Alibaba Qwen's vision encoder and pretrained weights without reinitialization, and MSIT ruled that a model that reuses open-source encoders and weights that way doesn't satisfy the project's independence requirement. Naver argued that "the core areas responsible for reasoning and judgment were developed with our own technology" and that the vision encoder was chosen for efficiency and global compatibility. The argument was rejected. Neither Naver nor Kakao entered the supplementary call afterward — the reputational downside of being cut twice was judged worse than sitting out.
The tournament's blueprint, in numbers
The core of this program isn't the money. It's the shape. 213.6 billion won — call it roughly 150 million dollars — is not a frontier-lab number. It's a rounding error against what a single US lab spends on compute in a year. So MSIT didn't try to fund adequacy. It tried to manufacture intensity through elimination. Teams start at about 500 GPUs each, survivors scale past 1,000, and the compute freed by an eliminated team flows to those who remain. From a participant's seat this isn't a subsidy, it's a prize purse, and prize purses change behavior.
The scoring deserves a look. Round one was 40 points of benchmarks, 35 points of expert assessment, and 25 points of user evaluation, out of 100. The five-team benchmark average was 30.4, so LG's 33.6 sat more than three points clear of the field. But sitting on top of that numeric scale was a separate pass/fail gate labeled "originality," and as Naver's case showed, the gate operates independently of the score. You can win on points and still be disqualified. That's the moment it became unambiguous that the state isn't buying the best model available. It's buying a model Korea built.
Round two moved the goalposts on purpose. Round one tested from-scratch development and LLM benchmark performance. Round two is weighted toward agentic capability — models that plan and execute multi-step work on their own — and operating efficiency in real B2B settings inside companies and public institutions. In plain terms, the criterion shifted from "scores well" to "actually does the job." The 200-person citizen panel running August 8-11 is the same instinct expressed differently: wire lived usage directly into the score. And final selection of the last two teams, originally scheduled for February 2027, has been pulled forward into this year. A schedule that accelerates is a government in a hurry.
| Item | Detail |
|---|---|
| Program | Sovereign AI Foundation Model project (nicknamed "dokpamo" / "national team AI") |
| Total budget | About 213.6 billion won, running through December 31, 2027 |
| Original call | Closed July 21, 2025 with 15 teams; paper review to 10; presentation evaluation to 5 |
| The five (announced Aug 4, 2025) | Naver Cloud · Upstage · SK Telecom · NC AI · LG AI Research |
| GPU support | ~500 per team to start, scaling past 1,000 for teams that clear a stage gate |
| Talent support | Up to about 2 billion won per year matched for recruiting senior researchers from abroad |
| Performance target | 95%+ of the best global models |
| First-stage evaluation | Run January 1-15, 2026; announced January 15 — 5 teams → 3 |
| Round-one leader | LG AI Research, 90.2 points (benchmarks 33.6/40, experts 31.6/35, users 25/25) |
| Round-one cuts | Naver Cloud (failed the originality gate) · NC AI (composite score shortfall) |
| Backfill | Motif Technologies, February 20, 2026 — 768 B200 GPUs, 1.75B won bespoke data + 10B won shared procurement |
| Second-stage evaluation | 200-person citizen panel August 8-11 → results as early as August 12, 4 teams → 3 |
| Final selection | 2 teams, now targeted within 2026 (moved up from February 2027) |
The line that should bother you is the GPU row. Five hundred scaling to a thousand. Motif's 768 B200s. Those numbers are strange because the reference points sit an order of magnitude higher. The estimates circulating in Korean coverage put GPT-4-class training at roughly ten thousand A100-equivalents and the following generation north of twenty thousand. Add up everything handed to all four Korean teams and you still don't reach one frontier lab's single training cluster. That is the structural weakness of this whole program, and it's exactly what the domestic industry has been complaining about for a year: splitting scarce compute four ways may widen the gap rather than close it.
But the constraint produced a strange side effect. Starved of compute, the teams competed on efficiency rather than size — and as the evaluation approached, they started publishing. Motif 3 preview on July 21. Upstage Solar Open2 on July 23. SK Telecom A.X K2 on July 29. LG K-EXAONE 2.0 on July 31. Four large Korean models in eleven days, most of them under Apache 2.0. They published their exam answers to the entire world right before walking into the exam room. Motif filed its final model on August 4, which means all four teams have now finished development.
Who actually walks away with what
What the government buys is leverage, and June gave everyone a very concrete illustration of why that matters. On June 12, the US Commerce Department ordered Anthropic to immediately block foreign nationals from two of its top models. Unable to verify user nationality in real time, Anthropic pulled both models globally and didn't restore them until June 30. Korean users lost access along with everyone else. The message was unmistakable: frontier models are being handled as strategic materiel, and a country without its own models can have its services switched off by another government's administrative letter. A domestic model doesn't have to be frontier-class to be a bargaining chip. It has to exist.
What the companies buy is compute and time. A thirty-person shop does not produce a 314-billion-parameter model that ranks eleventh globally in five months without 768 B200s appearing on its doorstep. The same logic applies at the other end of the size spectrum. For LG or SK Telecom to burn a 700-billion-parameter training run on their own balance sheet, someone has to win a board argument. "It's a national project" cuts the cost of that argument dramatically, and the difference between being able to say "the state asked us to" after a failure and not being able to say it is larger inside a chaebol than outsiders assume.
What developers and Korean enterprises get is Apache 2.0 weights, which was close to a byproduct rather than the plan. MSIT put open source policy into the selection criteria; teams kept loosening their licenses to score against it; and the result is that Hugging Face now hosts four Korean-tuned models between 250B and 750B parameters that you can use commercially without asking anyone. For banks, hospitals, and public agencies that cannot send data outside their perimeter, that's a real option rather than a press release. The practical ceiling still applies — serving K-EXAONE 2.0 in BF16 takes sixteen H200s per its model card — but the legal friction is gone.
The losses are just as legible. Naver Cloud opened Korea's LLM market and then got cut from the national team, and chose not to reapply. That isn't one company's problem. The way this program defines originality effectively penalizes the strategy of assembling excellent open-source components quickly, which happens to be how the global open-weight ecosystem actually works right now. The world rewards borrowing the best available parts; the national evaluation deducts points for it. Watching which signal Korean companies follow when those two conflict is the real experiment running here.
Then there's the citizen panel itself. Feeding 25 of 100 points from two hundred laypeople using the models for four days is politically shrewd — it inoculates the program against "the experts just handed it to their friends." Whether two hundred people can reliably resolve fine-grained quality differences between four large models is a separate question, and if perfect user scores keep appearing (LG took 25 out of 25 in round one), the component stops discriminating between teams at all.
We have seen state-run technology tournaments before
Korea has a genuinely successful precedent, and it's worth being specific. The TDX electronic switching system in the 1980s and CDMA commercialization in the 1990s were both national projects where the government set a target and had multiple companies develop against it in parallel. Both worked. CDMA didn't just deliver a world-first commercialization headline; it became the foundation of Korea's entire telecom equipment and handset industry. The lesson people usually take from it is about engineering, and that's the wrong lesson. What the state actually manufactured was demand. There was a guaranteed buyer — nationwide switch replacement, a licensed commercial mobile service — locked in before the technology existed.
The canonical failure is Japan's Fifth Generation Computer Systems project. Starting in 1982, the government spent a decade and enormous sums pulling researchers from multiple companies together to build parallel inference machines. By the time it wrapped, the world had gone somewhere else entirely. While the program built specialized hardware for logic programming, general-purpose workstations and a different branch of AI research took the field. The lesson stings: the fatal risk in a national program isn't technical failure, it's a goal that ages badly. "95% of the best global model" is admirably concrete, and it quietly assumes that the frontier in 2027 will be shaped roughly like the frontier today.
There are open-weight failures to learn from too. The UAE's TII released Falcon 180B with heavy fanfare as the largest open model of its moment, and almost no ecosystem formed around it. It was too big for anyone to run and the follow-through was thin. BigScience's BLOOM 176B was academically admirable and barely adopted in production. Both were models funded by a state or an international consortium, and both stopped at "we built it." Size buys headlines; usability buys ecosystems.
The instructive contrast is France's Mistral. No tournament, no elimination bracket. The state shaped regulation, procurement, and the capital environment, and the company fought in the market. Mistral laid down a dense ladder of model sizes, won developer adoption first, and converted that into enterprise contracts. If Korea wants a Mistral-shaped outcome, what it needs isn't more judging — it's the guaranteed demand that CDMA had and this program currently lacks. The fact that round two pivoted toward "field applicability" reads like the government already knows that.
How the competition punches back
The most unsentimental opponent is China. On July 31, Huawei released openPangu-2.0-Pro's weights, inference code, and technical report all at once — 505 billion total parameters with 18 billion active per token, a 512K context window, 34 trillion training tokens, and a claim that all of it ran on Huawei's own Ascend NPUs. Alibaba's Qwen line covers the entire size spectrum from sub-1B to hundreds of billions. DeepSeek has weaponized training efficiency itself, repeating "same capability, far cheaper" until it became a market fact. Their counter-strategy is not clever. It's cadence: ship the next version faster. A national program moving on six-month stage evaluations is structurally poorly matched against a release treadmill.
The awkward part is that Korean teams have to benchmark against those models specifically. The comparison set on K-EXAONE 2.0's model card contains no American models at all — it's Qwen, GLM, and DeepSeek. That reflects the actual terrain: the top closed models belong to the US, while the open-weight frontier that anyone can download belongs to China. Korea is entering the second arena, which unfortunately is the fastest-moving neighborhood on the map.
The American counter works differently. Frontier labs keep their best capability closed and toss mid-size open-weight models into the ecosystem to prevent developer defection. And June's Anthropic episode revealed a second lever: access itself can simply be revoked. That's less a competitive move than a statement about who writes the rules, and paradoxically it strengthens the case for programs like this one. When the party holding the rulebook can flip the table, having your own table starts to look cheap.
The most interesting counter-punch is domestic. There's a proposal reportedly under review that would abandon the current split — roughly 700 to 800 GPUs per team across four teams — and concentrate about ten thousand Nvidia Blackwell chips on a single team. The funding would come from windfall tax revenue off the semiconductor boom, pushed through as a supplementary budget within 2026 rather than waiting for the next fiscal year. The figure discussed is around 5 trillion won, which is comparable to MSIT's entire annual AI budget of 5.1 trillion won and roughly half of the government's total AI spending of about 10 trillion. Selection would reportedly weigh foundation-model project scores, technical capability, and each company's willingness to co-invest. If that lands, the tournament being decided this week is suddenly playing for a completely different prize.
And then there are Naver and Kakao, outside the bracket. Not being judged means not being bound by the originality gate. They're free to assemble the best open-source parts, partner with foreign labs, or do both — and they already own the one thing the program lacks, which is guaranteed demand in the form of their own products and users. One camp has the national-team badge and is still building its customer. The other has the customers and no badge. Which set of models is more widely used in Korea in 2027 is genuinely not obvious today.
So what actually changes
If you're a regular user. Nothing in your apps changes on August 12, and no prices move. The effect reaches you along two paths instead. The first is public services: the government has consistently said this program's output goes into administration, healthcare, defense, and manufacturing, so it surfaces as civil-service chatbots and public helplines quietly getting better at Korean. The second is the backend of the Korean apps you already use — as domestically tuned models start displacing foreign APIs, both answer quality and the cost structure behind your subscription shift together. Also worth noticing: you could have been one of those two hundred graders. A procedure where citizens directly score a state-funded AI is actually running.
If you're a developer or practitioner. This is the most useful moment of the whole cycle. Evaluation pressure pushed all four teams to publish inside eleven days, and most of it is Apache 2.0, so commercial use carries no strings. The range is unusually wide, from Upstage's Solar Open2 at 250B total and 15B active — realistically servable — up to K-EXAONE 2.0 at 750B, which needs sixteen H200s. Choose on your workload, not on an average benchmark row. Long-form Korean and domestic norm compliance point toward the EXAONE line; tight on-prem budgets point toward Solar; curiosity about a from-scratch architecture recipe points toward Motif. One caveat that applies to all of them: nearly every published number here is vendor self-reported.
If you're an investor. Don't map the August 12 result onto share prices one-to-one. Clearing the gate produces no immediate revenue, and being cut doesn't make the model disappear. Watch three other things instead. First, GPU reallocation — an eliminated team's compute flows to survivors, and if the 10,000-chip concentration plan becomes real, compute access diverges sharply at the company level. Second, real deployment references: since round two is scored on field applicability, the meaningful signal is which team actually converts into public-sector and industrial contracts. Third, the pairing with Korean AI silicon; a Korean model running on Korean chips is a separate industrial story with its own economics. Valuing any of this off self-published benchmarks is the mistake to avoid.
If you're a policymaker. Separate what twelve months proved from what it didn't. Proved: speed. Fifteen teams at the starting line produced four real models between 250B and 750B parameters within a year, and by Artificial Analysis's reckoning Korea now sits in the leading group outside the US and China. Not proved: durability. If the machine only turns while the state is feeding it GPUs, December 2027 is a cliff, and the guaranteed demand that made CDMA work has no equivalent here yet. You also have to decide how long the originality gate stays this strict. As written, it excluded the pragmatic assembly strategy that Naver was running — and the bill for that choice arrives in a few years, not this month.
If you're evaluating vendors for a company. Resist the temptation to treat the round-two result as a procurement shortlist. That score is a blend of two hundred laypeople's absolute ratings, expert assessment, and benchmarks, and it has no relationship to your call transcripts, your documents, or your workflows. Three things actually matter: total cost of ownership on-premises (the weights are free, the sixteen H200s are not), tool-call reliability inside agent workflows, and the exact license text — Apache 2.0 removes almost all legal review, but terms vary by team, so read the model card yourself rather than trusting a summary.
🥄 Three Things You're Probably Wondering
— So what does this mean for me? Directly, nothing. Whichever team gets cut on August 12, your apps stay the same. What the program did produce is four large, Korean-tuned, commercially usable models sitting in public right now — and once Korean services start running on them in the backend, the quality and the pricing of the Korean-language products you use can quietly shift over the following months.
— Why is this happening now, specifically? Because one year ago today, five teams were picked out of fifteen, and the condition attached to that selection was that failing a stage evaluation means leaving. January took five down to three. February brought Motif in to make four. The citizen panel runs August 8-11 and the result lands around the 12th, taking it to three. The burst of Korean model releases in late July was entirely driven by this calendar, and Motif's final submission on August 4 closed out development for all four teams.
— Has Korea caught up with the US and China? Too early to say that. Motif 3 tying DeepSeek's latest at 44 on Artificial Analysis's composite index and landing in the open-weight top three is a genuine result, and Stanford's AI Index 2026 puts Korea fourth overall and third for notable models produced. But the release counts for 2025 were 50 for the US, 30 for China, and 5 for Korea. That's an order of magnitude, and so is the GPU gap. "Holding position in the third tier" is the honest description, not "caught up."
Sources
- MSIT — Five elite teams selected following presentation evaluation, Sovereign AI Foundation Model project
- MSIT — Results of the open call for elite teams
- Korea.kr policy briefing — Seeking elite teams for the sovereign foundation model project, up to three years of support
- Hankyung — Korea finalizes five elite teams for its sovereign AI model (Aug 4, 2025)
- ZDNet Korea — Second selection round begins: four teams become three this month
- Herald Business — The decisive month for the second-stage evaluation
- HelloDD — Two hundred citizens will personally evaluate Korea's sovereign AI models
- Financial Today — Agentic AI and field applicability are the deciding axes
- Dealsite — Four-way fight decided in August: originality versus applicability
- Byline Network — Naver and NC eliminated in round one; LG takes first place
- ZDNet Korea — Naver Cloud and NC AI both eliminated; extra call opened
- VentureSquare — The originality dispute and the score shortfall behind the cuts
- Byline Network — Motif Technologies selected in the supplementary call
- Hankyung — Motif 3 unveiled, scoring on par with DeepSeek's latest
- e-focus — From GPUs split across four teams to 10,000 for one: a 5 trillion won plan under review
- Boannews — US Commerce order blocked foreign access to Anthropic's top models
- ZDNet Korea — Artificial Analysis calls Korea a clear number-three AI nation
- TheAI — Questions about the first-stage evaluation: schedule, technology gaps, and criteria
- Hugging Face — LGAI-EXAONE/K-EXAONE-2.0-750B-A37B model card
- Herald Business — SK Telecom unveils A.X K2 with 688 billion parameters
- DigitalDaily — Upstage releases Solar Open2, a 250B-parameter agent model
Numbers and criteria are as of announcement and may change. Investment calls are yours to make!



