Dokpamo’s 5 Teams Analyze 35.44 Million AI Data Points… Available to Any Company or Researcher
1.56 trillion in tokens
Upstage Unveils 29 Types of AI Training Data, Including 1 Trillion Tokens
Free Download from AI Hub
Startups and Small and Medium-Sized Enterprises Can Also Utilize It for Model Development and Service Building
[Edaily Reporter Kim Hyun-ah ] The government will make 35.44 million AI training data sets—built by five elite teams participating in the Independent AI Foundation Model (Dokpamo) project—available to the private sector. Domestic AI startups, small and medium-sized enterprises (SMEs), and researchers, who have struggled to secure large-scale training data, will be able to download these data sets for free from the AI Hub and use them to develop AI models and build services.
The Ministry of Science and ICT and the National Information Society Agency (NIA) announced that they will make 29 types of AI training data—built by Naver Cloud, Upstage, SKTelecom(017670), NC AI, and LG Corp.(003550) AI Research Institute during the first-stage evaluation—available on the AI Hub.
Deputy Prime Minister and Minister of Science and ICT Bae Kyung-hoon delivers a welcome address at the presentation ceremony for the Independent AI Foundation Project held on the afternoon of December 30 at the COEX Auditorium in Gangnam-gu, Seoul. Photo: News1 The release comprises approximately 35.44 million records and 1.56 trillion tokens (estimated). In addition to large-scale pre-training data, the collection includes inference, Agent AI, multimodal, and AI safety data, making it useful for developing new AI models or enhancing the performance of existing ones.
Of particular note is the 1 trillion-token pre-training dataset built by Upstage. It can be used to train large-scale AI models from scratch, and post-training data—including knowledge and intelligence data, user preference data, and agent behavioral capability data—is also being released.
Naver Cloud provides video-based text and audio data, including 2.34 million publicly available videos and 520,000 broadcast videos. This data can be utilized for the development of multimodal AI, such as video understanding and image generation.
SKTelecom is releasing high-difficulty reasoning data from specialized fields such as mathematics, science, and law, along with voice and image data, as well as 10,000 cases of Korean-style red-team testing data. This data can be used to enhance the reasoning capabilities and safety of specialized AI models.
NC AI provides data based on real-world industrial settings, such as manufacturing technical documents and audio recordings of customer service consultations. This data can be used for long-context understanding, step-by-step reasoning, multi-turn dialogue, and multimodal AI development; it also includes image data related to medical AI.
LG Corp. AI Research is releasing a “Scene Understanding” dataset for humanoid robots. This image and text dataset, built based on domestic home environments, can be used for the development of physical AI, vision, and VLM models.
Free Access via AI Hub…Some Data Available in the “Safe Zone”
The
newly released
data
can be searched and viewed by team or data type under the “Independent AI Model Data” menu on AI Hub. Domestic companies, researchers, and students can download and use it free of charge. Some data is available in AI Hub’s “Safe Zone” after a separate application process to ensure privacy protection.
Prior to the release, the government verified data quality, personal information protection, and potential harm through the National Information Society Agency (NIA) and the KoreaInformation&Communication Technology Association (TTA), and excluded data subject to licensing restrictions.
This release is expected to provide AI startups and small and medium-sized enterprises (SMEs)—which often struggle to secure large-scale datasets on their own—with an opportunity to obtain the data necessary for model development. Since the data can be utilized for a wide range of purposes—from pre-training to inference, Agent AI, multimodal applications, and safety—it is also expected to help reduce research and development time and costs.
The Ministry of Science and ICT plans to conduct quality verification on the data being built during the second-stage evaluation process and make it available as well. Kim Kyung-man, Director General of the Artificial Intelligence Policy Bureau at the Ministry of Science and ICT, stated, “The Independent AI Foundation Model Project is significant not only for AI model development but also for contributing to the qualitative growth of the entire AI ecosystem through the experience and know-how accumulated during the development process.” He added, “The data released today is a valuable asset that embodies the AI training strategies of elite teams, and it will serve as a foundation for driving the growth and self-sustainability of the AI ecosystem.”
As NVIDIA achieved record-high revenue for the 13th consecutive quarter, the company once again emphasized its collaboration with major memory manufacturers such as SamsungElectronics and SK hynix. Th…
#Shinsegae Co.,Ltd is accelerating the final stages of the spin-off involving siblings Chung Yong-jin and Chung Yu-kyung. Since establishing an independent management structure for the siblings in 201…
The government has unexpectedly released the detailed scores from the second-stage evaluation of the “Independent AI Foundation Model (DOKPAMO)” project. This came after Motif Technologies requested t…