
If a person requests deletion because their name and phone number appear verbatim in an AI chatbot's response, how far must the AI development company go? It was unclear whether simply removing data would suffice or if retraining the model would be required.
Song Kyung-hee, Chairperson of the Personal Information Protection Commission, stated that while the right to delete personal data will be guaranteed, companies should not be forced to adopt specific technologies. Instead, authorities must assess whether they have made sufficient efforts considering technical characteristics, constraints, and proportionality.
Song (Chairman) explained that "companies should guide users through privacy policies indicating whether personal data was used in training and establish channels for exercising rights. Upon receiving a request, they should first filter out the information so it does not appear in responses, make partial adjustments to the model, and consider retraining if necessary."
Starting September 11, companies that have made significant preventive investments in personal data protection will see their administrative fines reduced by up to 40% even if an incident occurs. Conversely, for major or repeated incidents, punitive fines of up to 10% of sales revenue may be imposed. Additionally, special provisions allowing the use of original personal data in AI development will be promoted where public interest and social necessity are recognized.
The following is a Q&A with Song (Chairman).
-If I request deletion of personal data used to train generative AI, must 'machine unlearning' also be performed?
▶Rights such as correction and deletion must be guaranteed even in the AI environment. However, AI training data and models differ from traditional databases. Data is not structured, making it difficult to identify and extract specific items, and accurately evaluating and removing the influence of specific data within a model remains challenging at the current level of technology. Rather than mandating specific methods, there is a need to develop practical countermeasures that consider technical characteristics, constraints, and proportionality.
-How should companies respond to requests for deletion of personal data used in AI training?
▶When a data subject makes a request, we guide them to promptly take safety measures such as output filtering and fine-tuning, followed by considering model retraining. Machine unlearning can be a useful tool, but it is still an evolving technology. France's CNIL and Canada's Office of the Privacy Commissioner hold similar positions. Retraining after removing datasets requires excessive costs for large models, and approximate unlearning cannot perfectly verify whether all data has been completely erased.
-Why do personal data breach incidents keep occurring?
▶Our IT infrastructure developed rapidly through concentrated investment in a short period and is now connected to AI. In contrast, investment in protection and security remains insufficient, as evidenced by the numbers. While the United States invests 13% of its total IT spending on protection and security, South Korea's share is less than half, at around 6%. Although laws and awareness are at a global level, companies and institutional investors on the ground have not kept pace, creating a gap where incidents can occur. Once an incident happens, costs include both penalties and infrastructure improvements. When adding the mental distress suffered by citizens, the social cost becomes too high. Ultimately, increasing preventive investment before incidents occur to establish a management system is crucial.
-The administrative fine system will change starting in September.
▶Starting September 11, companies that have made significant preventive investments will see their fines reduced by up to 40% even if an incident occurs. The judgment criteria are based on whether disclosed investment amounts exceed industry standards. While the law requires only minimum safety measures, we intend to clearly reflect cases where investments far exceed those requirements. At the same time, punitive fines of up to 10% of sales revenue will be imposed for major or repeated incidents. Strengthening sanctions cannot be the ultimate goal; the final objective is to prevent incidents by encouraging investment.
-What about cases where evidence is destroyed before an investigation?
▶If discovered after an investigation, criminal penalties or administrative fines can be applied, but there are indeed limitations under current law if data is erased before the investigation begins. If a structure exists where companies profit by deleting and hiding data in advance, honest reporting companies would suffer losses instead; that is why we also requested prosecution in the recent KT case. We are creating a system to impose administrative fines of up to 3% even before an investigation starts. Since concealment or destruction often occurs internally and limits investigations, we are also considering whistleblower reward programs.
-There are also criticisms that case processing is slow.
▶About 30 investigators must handle all these cases. For the Tving incident, which we consider important, five to six personnel were assigned alongside the Korea Internet & Security Agency (KISA). We are trying to deploy as many personnel as possible. Litigation is also a burden; government institutional investors typically incur very small litigation costs because cases usually go through three levels of trial. Unlike the Fair Trade Commission, where legal principles are established through precedents, our cases involve diverse types, requiring us to develop new arguments for each one.
-You are promoting special provisions allowing the use of original personal data in AI development.
▶Since anonymized or pseudonymized data makes AI development difficult, this is a separate procedure that allows the use of original data only when public interest and social necessity are recognized, subject to enhanced safety measures and approval by the Commission. It does not constitute an unconditional exemption from consent.
-Will private commercial AI services or medical and biometric information also be eligible for special provisions?
▶Private companies and commercial AI services can be included. Medical and biometric information will not be excluded from special provisions unless specifically regulated differently under other laws. In cases where the impact on data subjects' rights or interests is significant, we plan to conduct a separate risk factor assessment for thorough review. Ultimately, decisions will be made by comparing the public benefits created by AI against the risks that may arise for data subjects.
-What do you most want to achieve during your term?
▶Article 1 of the Personal Information Protection Act states that its purpose is to realize individual dignity and value. I have frequently recited this provision since taking office, almost like an oath. This law does not aim to prohibit all use of personal data; doing so would not necessarily help realize dignity and value. In the AI era, personal data must be used safely and actively, and companies must be able to innovate for benefits to return to individuals. Post-incident sanctions are merely a means to move toward a preventive system, not the ultimate goal.
