[Edaily Han Kwangbeom Reporter Kim Hyun-ah] In the second-stage evaluation of the independent AI foundation model project, Upstage, SKTelecom(017670), and LG Corp.(003550) AI Research Institute advanced to the third stage. In contrast, Motif Technologies, which had ranked first with a score of 47 on the Artificial Intelligence Comprehensive Index (AAII) compiled by the global AI model evaluation agency Artificial Analysis, was eliminated.
According to detailed evaluation results obtained by Edaily, Motif demonstrated competitiveness in benchmark and expert evaluations but received relatively low scores in user evaluations, resulting in an overall score of just 65.8 points. The fact that Motif was ultimately eliminated despite receiving the highest score on the AAII—which measures technical capability—can be attributed to an evaluation structure that placed relatively greater emphasis on usability and applicability than on technical capability.
Taking this evaluation as an opportunity, the government has decided to review a plan to comprehensively restructure the program, shifting the focus from “Dukpamo” to concentrating resources on the development of frontier-level AI models. However, it is necessary to separately examine whether the scoring structure for technical capabilities versus usability and applicability was appropriate during this elimination process, and whether the evaluation formula was sufficiently shared with the participating companies.
Ryu Je-myeong, Second Vice Minister of Science and ICT, stated at a briefing at the Seoul Government Complex on the 18th, “Based on a comprehensive evaluation of benchmarks, expert reviews, and user assessments of the four elite teams that participated in the second phase, the elite teams from Upstage, SKTelecom, and LG AI Research have advanced to the next stage.”
The government notified each company of the selection results that day and plans to begin the third-stage evaluation after completing the objection and appeal procedures. The policy is to proceed with the third-stage evaluation as originally planned to select the final two companies. Teams advancing to the third stage will be provided with approximately 1,000 B200 GPUs per team. The leasing budget is approximately 40 billion won per team, totaling 120 billion won.
Upstage, SKT, and LG Corp. to the Third RoundIn this second evaluation, Upstage received high marks for unveiling “SOLAR-pro-2,” a model with 250 billion (250B) parameters, and presenting an ecosystem that integrates with the “Daum” and “Timely” platforms and collaborates with the domestic AI semiconductor company FuriosaAI’s NPU.
SKTelecom showcased “A.X-K2,” a model with 688 billion (688B) parameters. Its strengths were recognized as top-tier performance on the International Mathematical Olympiad (IMO) and Korean language benchmarks, as well as its applicability in industrial sectors such as defense, manufacturing, legal affairs, and taxation.
LG Corp. AI Research, based on its “K-ExaOne 2.0” model with 750 billion (750B) parameters, ranked 9th globally in hallucination suppression metrics and received high marks for its collaboration strategy with global international organizations and its reliability assurance system.
Meanwhile, Motif unveiled a MoE (Mix of Experts) model with 314 billion (314B) parameters. Notably, it scored 47 points on the AAII, outperforming Upstage (37 points), SKTelecom (35 points), and LG Corp. (31 points). It was the only Korean model to rank 10th on a global scale.
However, the final results were different.
Motif, which ranked first in the AAII evaluation, scored only 14.1 points in user evaluationsAccording to detailed results from the second-stage evaluation obtained by Edaily, Motif’s overall score was 65.8 points, resulting in its elimination from the competition.
Motif received 24.6 points in the benchmark evaluation (11.9 points from AAII and 12.7 points from NIA) and 27.1 points in the expert evaluation. However, in the user evaluation, it scored only 14.1 points (8.4 points from expert users and 5.7 points from the general public).
The government explained that while Motif’s technical capabilities were outstanding, its relatively low ratings in terms of usability and practical applicability contributed to its elimination.
The question is whether this result can be fully explained simply by saying that it “fell short in usability.”
AAII: 47 points; Final evaluation: 11.75 pointsThe Dokpamo evaluation consisted of a total of 100 points, broken down into a benchmark score of 40 points, an expert evaluation of 35 points, and a user evaluation of 25 points. Of these, the AAII accounted for 25 of the 40 benchmark points.
The raw AAII scores were 47 points for Motif, 37 for Upstage, 35 for SKTelecom, and 31 for LG Corp. AI Research. When converted to a 25-point scale, these become 11.75 points for Motif, 9.25 for Upstage, 8.75 for SKTelecom, and 7.75 for LG Corp. AI Research.
While the gap between Motif and LG Corp. in the AAII raw scores is 16 points, the point difference reflected in the final evaluation is reduced to 4 points. Adding the 15 points from the NIA evaluation brings the total benchmark score to 40 points.
In essence, even if there is a large point difference in the global technology evaluation, the gap becomes relatively smaller in the final evaluation.
Of course, this does not mean that the usability and applicability evaluations themselves are flawed. For a national AI project, it is important to verify whether technological capabilities translate into actual services and real-world industrial applications.
However, it is necessary to examine how much weight should be assigned to technical capability and applicability, respectively, and whether the scoring criteria and calculation formulas are designed to align with the project’s objectives.
In particular, since 60 points out of the total 100—including 35 points for expert evaluation and 25 points for user evaluation—are heavily influenced by qualitative judgments, there are calls for the evaluation criteria, detailed items, and conversion methods to be fully disclosed so that the results can be objectively verified.
Government to Restructure Initiatives Around “Frontier AI”Meanwhile
,these results are expected to lead to changes in the government’s AI strategy.
Vice Minister Ryu explained that as the model performance of global frontier companies is advancing rapidly, a new competitive landscape—different from the existing “Dokpamo”—is necessary.
The government believes that relying solely on the current decentralized competition model makes it difficult to compete with global frontier companies, and is therefore reviewing plans to link and integrate related projects, such as the “Dokpamo” and “Mitos”-level frontier AI development projects. Specific project directions, including whether to continue the “Dokpamo” initiative, are expected to be announced shortly.
Kim Kyung-man, Director General of the Artificial Intelligence Policy Bureau at the Ministry of Science and ICT, stated regarding the development of frontier-level models, “In addition to forming a special-purpose company (SPC) involving multiple companies, it is also possible for a specific company to lead the project.”