It Starts From the Premise That Pseudonymization Breaks Some Training

Here's the deal: an amendment to Korea's Personal Information Protection Act (PIPA) cleared the National Assembly's Political Affairs Committee. The core of it is an exception for personal data used in AI training.

Until now there were effectively two legal routes to train AI on personal data in Korea: get separate consent from the data subject, or pseudonymize/anonymize until individuals can't be identified. The amendment adds a third. Meet certain conditions and you can use raw personal data for training without consent.

Three conditions, all of which must hold:

  1. Anonymization or pseudonymization alone makes AI development infeasible
  2. Safe processing measures and protective safeguards are in place
  3. The purpose is public interest and the risk of rights infringement is markedly low

Plus a procedural requirement: prior review and approval by the Personal Information Protection Commission (PIPC) is mandatory. A company cannot self-certify. And the scope is limited to data already lawfully collected — this does not authorize new scraping.

Compressed to a sentence: consent was replaced with prior regulatory review. The step of asking individuals becomes a step where the regulator examines the request on their behalf. It's a design that emerges when the data volumes make consent practically impossible to collect — and it correspondingly weakens an individual's direct control over where their data goes.

One caveat before going further: the conditions and process described here reflect the amendment as passed by committee. Wording can change through the plenary vote and the enforcement decree. The specific review requirements and the data subject rights provisions in particular are likely to be settled at the decree stage, so none of this should be read as final.

Why This Law Exists

Start with the technical background. Pseudonymization means stripping names and replacing national ID numbers with random values. For statistical analysis that's usually fine. For AI training it can gut performance.

Medical AI is the canonical case. Training a diagnostic model requires a patient's age, sex, history, and test values to stay linked together. Pseudonymize and bucket each field separately and the relationships between variables blur. In rare-disease work, where data is scarce to begin with, pseudonymization can erase the learnable signal entirely.

Speech and face recognition have the same problem from a different direction. A voice and a face are identifying information. "Pseudonymized voice" isn't a coherent object — remove identifiability and there's nothing left to learn.

The government announced a full overhaul of the personal data protection framework on July 3 this year, and this amendment is the legislative product of that push. PIPC had already published a guide to personal data processing for generative AI development and a guide to processing publicly available personal data, establishing working standards. But guides carry weak legal force, and companies kept asking whether following them actually protected them. Writing the exception into statute is aimed at removing that uncertainty.

There's an institutional backdrop too. Korea has framed its AI Basic Act regime as balancing industrial promotion with trust-building. In actual legislation the two axes don't always move at the same speed, which is the substance of the criticism that this amendment leans toward promotion. PIPC has separately published a Third Basic Plan for Personal Data Protection covering 2027 to 2029 to set medium-term direction.

Where the Criticism Lands

The objection is less about permitting AI training as such and more about the specificity of the safeguards.

First, all three conditions leave wide interpretive room. Who determines that pseudonymization makes development "infeasible," and where is the line for "markedly low" risk of infringement? Neither is fixed numerically in the statute. In practice PIPC's review casework will define the standard, which means low predictability early on.

Second, how data subject rights operate after the fact is unclear. Can you learn that your data entered a particular model's training? If so, can you demand removal? Deleting a specific record's influence from a trained model is technically very hard. Machine unlearning research is active but not practically mature, and the reliable alternative is retraining, which is expensive.

Third, the scope of "public interest." Some domains are unambiguous, like healthcare or disaster response. Others — "strengthening industrial competitiveness" — can stretch. Widen that scope and the exception becomes the default rather than the exception.

Fourth, oversight after approval. Clearing prior review doesn't guarantee processing proceeds as described. How training data gets reused for other purposes, outsourced to third parties, or transferred abroad still needs verification. That's the standard weakness of approval-centric regimes: without supervisory headcount and post-hoc audit capacity, review becomes a formality.

How Other Jurisdictions Handle It

Jurisdiction Approach Characteristic
Korea (amendment) Conditional exception + prior review Regulator approves case by case
Japan Transparency rules in preparation Requires disclosure of training data and methods
EU GDPR + AI Act Requires a legal basis; separate high-risk regime
US State-by-state legislation No unified federal standard

Japan's direction is the contrast. It opens use broadly while requiring AI developers to disclose training data and processing methods — transparency-centric, weighting after-the-fact verification over prior approval.

Both approaches have real trade-offs. Prior review can filter problems in advance, but throughput becomes the bottleneck. Applications pile up, review backs up, and you end up with "permitted but unusable." Transparency is fast, but without a party with the capacity to verify what's disclosed, it degenerates into paperwork.

The EU stacked the AI Act on top of GDPR. Using personal data in AI training requires a GDPR legal basis, and whether legitimate interest can serve as that basis remains contested. The norms are dense and enforcement burden is heavy — which is precisely what European AI companies keep complaining about.

The US has no unified federal standard. California, Colorado and others legislate separately, so companies operating across states default to the strictest applicable rule. Regulatory predictability is low, but the environment also permits shipping first and adjusting later.

Who Gets What Out of This

Korean AI companies get legal certainty. The prior model was: read the guides, make your own call, absorb the consequences if challenged. Review settles that judgment in advance. For startups blocked from fundraising or enterprise partnerships by regulatory risk, that's a material change.

Healthcare and biotech get data access. Korean medical data is high quality by international standards, and the long-standing complaint is that access barriers left it unused. This is the domain where the exception should bite most clearly.

PIPC gets authority and burden. Holding review power expands influence, but review capacity has to keep up. Assessing how a model trains and what risks it carries requires technical staff, not just legal expertise.

Cloud and infrastructure providers get a side effect. Training projects that clear review will generally carry access control and audit logging requirements, which concentrates demand on environments that satisfy them. If data residency conditions attach, providers with domestic regions gain.

Data subjects — individuals — get something unclear. That's the heart of the criticism. The provisions widening use are specific; what an individual can know and demand is comparatively less so. Whether after-the-fact notification duties, training data disclosure, and objection procedures make it into the enforcement decree, and at what strength, is what determines the real level of protection.

For foreign companies it's an entry condition. Global firms wanting to train on Korean user data face the same process. Whether identical conditions apply and whether extraterritorial enforcement works in practice could become a fairness issue.

Previous Rounds of Data Deregulation

The 2020 "Data 3 Acts" amendment is the closest precedent. It introduced pseudonymized data and opened use for statistical, research, and public interest purposes. Industry welcomed it; civil society worried. In retrospect both were half right. Pseudonymized data use never grew explosively, and the feared mass-breach scenario didn't materialize from that provision either. The actual lesson was that opening a channel doesn't mean companies use it. Complicated procedures and unclear liability push the person responsible toward not doing it.

Medical MyData followed the same arc: the framework existed, but hospitals and companies took years to actually exchange data, because standardization, security certification, and liability allocation remained unsolved.

The early GDPR period teaches the opposite. Strong regulation arrived and companies over-complied, avoiding data use altogether. Where rules are unclear, organizations behave as conservatively as possible. Korea's exception could produce the same outcome if review standards stay vague.

Korea's cloud security certification (CSAP) is also instructive. Built to enable public sector cloud adoption, its early requirements were strict enough that foreign providers were effectively excluded — procedural requirements functioning as a substantive market barrier. Depending on how the AI training review process is designed, it can do the same thing.

One pattern runs through all four: data regulation succeeds or fails on the specificity of working guidance, whether it opens or restricts. The Data 3 Acts opened and went unused because guidance was ambiguous; GDPR restricted and produced over-avoidance because early guidance was ambiguous. The outcome here will likely be determined by the enforcement decree and review standards rather than the statutory text. Watch the documents that come next, not the passage.

So What Actually Changes

If you build AI services, nothing changes today. Using the exception requires PIPC review, and both the standards and procedure will be specified in the enforcement decree and working guidelines. What to do now is document how your training data was collected. Proving "lawfully collected" requires records of the consent scope and legal basis at collection time.

If you handle sensitive data in healthcare or finance, the operative question is how you'll substantiate the third condition — "markedly low risk of rights infringement." Re-identification risk assessment, access controls, and post-training data disposal procedures are likely to become the substantive basis for review decisions.

If you run an education service or a community platform, revisit the data use clauses in your terms of service. What counts as "lawfully collected" turns on the purposes disclosed at collection, and older terms were almost universally written without contemplating AI training. That's likely the first gate you hit when trying to use the exception.

If you're an individual, there isn't much immediate action available. You can submit comments during the enforcement decree's public notice period, and it's worth watching how data subject rights get specified — particularly whether after-the-fact notification becomes a duty.

If you work in policy or legal, compare against Japan's transparency rules. A Korean company operating in both markets has to satisfy both regimes, and prior review and after-the-fact disclosure demand different documentation.

If you're a startup, budget for the review process itself. Preparing the application, running re-identification risk assessment, and designing safeguards requires legal and security staff. A cost a large company absorbs easily can be a barrier to a small team. Watch for whether industry-standard templates or published review checklists emerge.

If you're a researcher, check how this exception interacts with academic research. Provisions for research use of pseudonymized data already existed, and the relationship between the two needs clarifying to avoid confusion in practice.

🥄 Three Things You're Probably Wondering

— Is my data already being used to train AI? This amendment governs future use. But because the scope is "data already lawfully collected," data gathered during past service use could go through review and into training. Whether it's disclosed which data cleared review for which purpose is the key question, and that hasn't been settled.

— Can I ask to be removed later? Technically difficult. There's no practical method for removing one record's influence from a trained model; the reliable approach is retraining without it, which is costly. That's likely why the design weights prior review — controlling what goes in — over a deletion right.

— Does Korea need this law to compete in AI? It's not that simple. In domains where data access is the bottleneck, healthcare above all, the difference should be real. In domains like general-purpose conversational models, where web data dominates the training mix, this exception barely applies. Expect the effect to vary sharply by field.

References

Numbers and criteria are as of announcement and may change.