State audit finds duplicate, unusable data in $1.1 billion AI training project

BAI audit reveals duplicated, mislabeled and unusable datasets in Korea's government-funded AI training data program.

Published
The Board of Audit and Inspection sign in Jongno District, central Seoul

A government project that has spent more than 1.6 trillion won ($1.1 billion) since 2017 to build datasets for AI training also produced numerous cases of duplicate or unusable data, a state audit found.

The Board of Audit and Inspection (BAI) released the findings Wednesday in an audit of AI training data under its review of efforts to foster the AI industry. The audit found that inadequate procedures for checking whether existing datasets were similar to or could replace newly proposed data led to repeated investments in overlapping datasets across government ministries and public institutions.

Of the 1.6 trillion won, the Ministry of Science and ICT spent 1.6 billion won in 2020 to build a dataset containing 150,000 images of 128 types of household waste.

But Seo District in Daejeon separately spent 130 million won in 2025 to create 9,000 images covering 45 types of household waste. Of those, 5,200 images covering 26 types were found to be similar to existing data, accounting for 57.8 percent of the total.

A wildlife image dataset created by the Korea Expressway Corporation at a cost of 284 million won also contained substantial overlap with existing materials. Of its 60,000 images, 25,000, or 41.7 percent, were found to be similar to existing data.

The audit also found cases in which government agencies separately built similar datasets at around the same time.

The logo of the Ministry of Science and ICT is seen in this undated file photo.

The Science Ministry and the Personal Information Protection Commission spent 1.75 billion won and 308 million won, respectively, to build oral image datasets in 2023.

All 1,000 images of upper and lower teeth produced by the Personal Information Protection Commission were found to be similar to data already created by the Science Ministry.

Some datasets were found to be unusable for actual AI training.

A total of 2,680 road traffic CCTV images released by the Korea Expressway Corporation lacked coordinate information identifying the locations of vehicles, making them unusable for training AI models.

Errors were also found in a dataset of bulky waste images created by the Seoul Metropolitan Government. Of the 2,303 images, 538, or 23.4 percent, had item labels that did not match the objects shown in the photos.

The BAI concluded that the lack of a system for agencies to share or coordinate their data-building plans in advance had resulted in overlapping budget expenditures on similar datasets and shortcomings in quality control.

The BAI instructed the Science Ministry to assess whether existing data could be used as a substitute before launching new projects and to establish a system for sharing and coordinating data-building plans among government agencies.

With the government investing heavily in fostering the AI industry, the findings have prompted calls for stronger safeguards against duplicate investments and more rigorous quality checks when building AI training datasets.

BY JEONG JAE-HONG [kim.jiye@joongang.co.kr]

This article was originally written in Korean and translated by a bilingual reporter with the help of generative AI tools. It was then edited by a native English-speaking editor. All AI-assisted translations are reviewed and refined by our newsroom.