Ministry of Science and ICT Releases Detailed Scores for Second Round of “Dokpamo” Evaluation… SKT Takes First Place, Upstage Second, and LG Corp. Third
Ministry of Science and ICT Releases Detailed Findings in Response to Motif’s Concerns
SKT Takes First Place Overall with 70.6 Points
Upstage Scores 69.9 Points; LG Corp. AI Research Scores 69.0 Points
Motif Topped the AAII Benchmark, But...
Lags Behind in Expert and User Evaluations… Overall Score of 65.8 Points
Ministry of Science and ICT: “Evaluation Criteria to Be Determined in Consultation with Elite Teams”
[Edaily Reporter Kim Hyun-ah ] The government has unexpectedly released the detailed scores from the second-stage evaluation of the “Independent Artificial Intelligence (AI) Foundation Model (DOKPAMO)” project. This came after Motif Technologies requested the disclosure of detailed evaluation results, and other elite teams agreed to the release.
According to the released results, SKTelecom(017670)took first place overall with 70.6 points, followed by Upstage (69.9 points), LG Corp. (69.0 points), and Motif Technologies (65.8 points).
Notably, while Motif received the highest score among the four elite teams on the AAII benchmark by Artificial Analysis—a global AI evaluation agency—it ranked fourth in the expert evaluation, the evaluation by AI power users, and the public evaluation.
Detailed Scores
Released
at Motif’s Request… “Balancing Fairness and Transparency”
The Ministry of Science and ICT (Deputy Prime Minister and Minister Bae Kyung-hoon) released the detailed scores for the second-stage evaluation of the Dokpamo Project on the 27th.
The government stated that it had initially been cautious in determining the scope of disclosure to minimize direct and indirect harm to companies resulting from the release of evaluation results, while simultaneously ensuring the fairness and transparency of the evaluation process.
However, on the same day, Motif Technologies requested the disclosure of detailed evaluation results in a statement, and as other elite teams agreed, the ministry decided to release the full detailed scores.
The Ministry of Science and ICT stated, “Dokpamo is not a project aimed at simple ranking competitions or selecting teams to be eliminated,” adding, “Its purpose is to enable domestic AI companies to grow beyond competition and drive the qualitative growth and expansion of the AI ecosystem.”
Evaluation Across Five Areas… Benchmark, Expert, and User Evaluations Conducted in Parallel
The second-stage evaluation of ‘Dokpamo’ was broadly conducted through △benchmark evaluation, △expert evaluation, and △user evaluation.
The benchmark evaluation was further divided into △the AAII benchmark and △the National Information Society Agency (NIA) benchmark. The user evaluation consisted of △evaluations by a panel of AI experts and △evaluations by the general public.
The total score is 100 points, allocated as follows: 40 points for benchmark evaluation, 35 points for expert evaluation, and 25 points for user evaluation.
The Ministry of Science and ICT explained that, considering the rapidly evolving nature of AI technology, it is applying a “moving target” approach that adjusts development goals every six months. Accordingly, the ministry stated that the criteria and methods for each phase of the evaluation are established through consultation with elite teams each time.
The government explained that the criteria for this second evaluation were also finalized after numerous consultations with the elite teams, and that the evaluation was conducted only after the final draft was officially announced and no objections were raised.
Motif Takes First Place in AAII… Raw Score of 47.4 Points
The most notable result comes from the AAII evaluation.
In the AAII benchmark—conducted using the criteria announced by Artificial Analysis in late June—Motif ranked first with a raw score of 47.4 points. When converted to a 25-point scale, this equated to 11.9 points.
Upstage ranked second with a raw score of 37.4 and a converted score of 9.4, while SKTelecom ranked third with 35.0 and 8.8, respectively. LG Corp. AI Research ranked fourth with 31.0 and 7.8.
Motif recorded relatively high scores in several categories, including GDPval-AA v2, τ3-Banking, Terminal-Bench v2.1, AA-LCR, AA-Omniscience Accuracy, and HLE. In particular, on Terminal-Bench v2.1, it scored 74.9 points, significantly outperforming the other three teams.
SKTelecom Takes First Place in NIA Benchmark… Scores Among the Four Teams Are Close
In the NIA benchmark
, SKTelecom scored the highest at 13.4 points. Upstage followed with 13.3 points, while LG Corp. AI Research and Motif scored 12.8 and 12.7 points, respectively.
The NIA evaluation covered eight areas: mathematics, general knowledge, long-passage comprehension, task execution, Korean language, reliability, safety, and social safety.
In terms of raw scores, Upstage scored the highest at 732.4 points, followed by SKTelecom with 728.7 points, LG Corp. AI Research with 705.8 points, and Motif with 702.8 points. These scores were converted to a 15-point scale.
Looking at the details, Upstage showed strengths in mathematics, following instructions, and reliability, while SKTelecom scored highly in knowledge and Korean. LG Corp. AI Research received relatively high scores in long-passage comprehension and safety.
LG Corp.
Takes First Place
in Expert Evaluation… Motif
Ranks
Fourth with 27.1 Points
In the expert evaluation, LG Corp.
took first place
with
29.5
points
. SKTelecom followed
with
29.3 points, Upstage
with
29.1 points, and Motif
with
27.1 points.
The expert evaluation was conducted by an evaluation committee composed of 10 external experts: 3 from industry, 5 from academia, and 2 from the research sector. Based on the evaluation documents submitted by the elite teams, the committee conducted a parallel process of written evaluation and question-and-answer (Q&A) sessions over approximately five days.
The evaluation criteria were: △ Development Strategy and Technology (10 points), △ Development Achievements and Plans (10 points), and △ Ripple Effects and Contribution Plans (15 points). An assessment of the minimum standard for originality was also conducted.
Rather than simply adding up all the scores given by individual committee members, the final score was calculated as the arithmetic mean of the remaining scores after excluding the highest and lowest scores.
LG Corp. AI Research received consistently high evaluations across all categories: development strategy and technology, development achievements and plans, and impact and contribution plans.
In contrast, Motif received high scores from some evaluators but scored only 27.1 points overall.
Evaluation of
49 AI Startup CEOs… SKTelecom
Ranked
1st, Motif 4th
In the evaluation by AI experts, SKTelecom took first place with 11.6 points.
LG Corp. AI Research came in second with 11.3 points, Upstage third with 10.8 points, and Motif fourth with 8.4 points.
The Ministry of Science and ICT explained that the panel of AI experts was selected from AI startup CEOs who possess professional insight and experience in AI and have extensively used AI models in various actual AI development processes.
Although 50 members were initially selected, 49 participated in the actual evaluation. Potential conflicts of interest were also verified.
The evaluation was conducted by assessing usability and other factors based on the AI-powered websites provided by each elite team. The evaluation panel assigned scores ranging from 1 to 5, which were then converted to a 15-point scale.
SKTelecom had the highest raw score total at 189 points, followed by LG Corp. AI Research with 184 points, Upstage with 176 points, and Motif with 137 points.
LG Corp. AI Research Also
Tops
Public Evaluation…Motif
Scores
5.7 Points
In the public evaluation, LG Corp. AI Research
To ensure greater representativeness of the general public, the government recruited evaluation panel members over a sufficient period, set quotas by gender and age based on domestic resident registration statistics, and randomly selected 200 individuals. Of these, 185 actually participated in the evaluation.
The panel members were also finalized after verifying the absence of any conflicts of interest. Each model was evaluated on an absolute scale of 1 to 5, which was then converted to a 10-point scale.
The total raw scores were as follows: LG AI Research (701 points), SKTelecom (698 points), Upstage (675 points), and Motif (528 points).
Motif is contesting the fairness of the user evaluation. They argue that when general users evaluate models developed by well-known large corporations and relatively lesser-known startups, it is difficult to rule out the possibility that brand recognition or preconceptions will influence the results if users are aware of the developers’ identities beforehand.
Motif Technologies argued, “If the performance differences observed in the AAII benchmark changed significantly in the general user evaluation, a convincing explanation is needed regarding which evaluation methods and conditions influenced these results.”
They also emphasized that the detailed results of user evaluations should not merely serve as data to determine which models pass or fail, but should be provided as practical feedback to help participating companies improve the usability of their models.
SKTelecom
Takes
Overall First Place… Motif Ranks 4th Overall Despite Topping the AAII Category
After combining the scores from all five evaluation categories, SKTelecom secured first place with a total score of 70.6 points.
Upstage came in second with 69.9 points, LG Corp. AI Research third with 69.0 points, and Motif Technologies fourth with 65.8 points.
Consequently, Motif took first place in the benchmark evaluation with 24.6 points out of a possible 100. However, it received 27.1 points in the expert evaluation and 14.1 points in the user evaluation, causing it to drop to fourth place in the final rankings.
In contrast, SKTelecom ranked third in the AAII evaluation but took first place in the NIA evaluation, second in the expert evaluation, first in the evaluation by AI expert users, and second in the general public evaluation, securing the top spot overall.
Upstage also secured second place overall with a score of 69.9 points, having ranked second in the benchmark evaluation, third in the expert evaluation, and third in the user evaluation.
LG Corp.(003550) The AI Research Institute ranked 4th in the AAII evaluation but took 1st place in both the expert evaluation and the public evaluation, ultimately placing 3rd overall.
The Ministry of Science and ICT emphasized that, with the release of these detailed results, companies should contribute to the domestic and international AI ecosystems based on their own technological capabilities and strengths, rather than becoming fixated solely on scores and rankings.
An official from the Ministry of Science and ICT stated, “We expect our companies to make broad contributions to the domestic and international AI ecosystems based on their inherent potential and capabilities, and for the ecosystem to evolve into a dynamic one where new and competitive AI companies continuously take on new challenges.”
SamsungElectronics has ranked first in the general home appliance category for the fourth consecutive year in the “Digital Customer Experience Index (DCXI)” evaluation organized by the Korean Standard…
Global cosmetics Original Design Manufacturer (ODM) company COSMAX, INC.(192820)is participating in the launch of Indonesia’s first department specializing in cosmetics. The company plans to take the …
Update on the Phase 1b clinical trial results for “CS5001 (LCB71)” released by LigaChem Biosciences’ partner, Siston Pharmaceuticals (Source: Siston Pharmaceuticals’ first-half earnings report)
Li…