Motif 47 points, Upstage 37 points, SK 35 points, LG Corp. 31 points… AAII Also Highlights ‘Limitations of Benchmarks’ (Comprehensive)
AAII Results Released as the 3rd Round of the "Dokpamo" Evaluation Begins
Motif Ranks No. 1 in Korea and No. 10 Globally Among LLMs
It Is Difficult to Fully Reflect Practical Capabilities in the Era of Agent AI
All Eyes on Final Results, Including NIA’s Confidential Evaluation
[Edaily Reporter Kim Hyun-ah ] Global AI evaluation results for the four companies participating in the “Domestic AI National Team” selection process—the Independent AI Foundation Model Project (Dokpamo)—have been released. Motif Technologies’ “Motif 3” scored the highest at 47 points, followed by Upstage with 37 points, SKTelecom(017670) with 35 points, and LG Corp.(003550) AI Research Institute with 31 points.
While it is significant that a global evaluation agency compared domestic AI models using uniform criteria, some point out that relying solely on AAII scores has limitations when assessing a model’s overall competitiveness or practical applicability.
In particular, as AI evolves beyond simple question-and-answering to become Agent AI capable of independently performing complex tasks, the extent to which existing benchmarks can reflect actual task performance is emerging as a new evaluation challenge.
Motif Takes First Place with 47 Points… Upstage, SKT, and LG Corp. Follow
According to the “Artificial Analysis Intelligence Index (AAII) v4.1.1” released on the 13th by the global AI analytics firm Artificial Analysis, Motif’s “Motif 3” scored 47 points.
Upstage’s “Solar Open 2 250B” took second place with 37 points, SKTelecom’s “A.X K2” came in third with 35 points, and LG Corp. AI Research’s “K-EXAONE 2.0 0803” followed with 31 points.
AAII is a metric that synthesizes evaluation results across various fields—including mathematics, science, coding, and reasoning—to present an AI model’s performance as a single score. This evaluation utilized nine categories, including financial tasks, coding, science and math problems, long-passage reasoning, and information accuracy.
Motif 3 ranked 10th among global large language models (LLMs) and 4th among OpenWeight models. It demonstrated particular strengths in certain categories, such as AI agents performing financial tasks and coding capabilities.
Motif also announced its plans to expand the model ecosystem by releasing not only the model weights for Motif 3 but also its training code and libraries.
AAII 25-point score factored in… Separate from the final “Dokpamo” rankings
These results are drawing attention because they are directly linked to the Ministry of Science and ICT’s third “Dokpamo” evaluation.
In the third evaluation, one of the four companies—LG Corp. AI Research, Upstage, SKTelecom, and Motif Technologies—will be eliminated.
In the Dokpamo evaluation, the benchmark section accounts for a total of 40 points, of which the AAII evaluation accounts for 25 points and the National Information Society Agency (NIA)’s proprietary benchmark accounts for 15 points. The remaining 60 points consist of 35 points from expert evaluations and 25 points from user evaluations.
Since AAII accounts for 25% of the overall evaluation, this point difference could have a significant impact on the final results. However, AAII does not determine the final Dokpamo rankings.
In particular, since the NIA’s confidential evaluation and the expert and user evaluations are still pending, it cannot be assumed that the AAII rankings will necessarily match Dokpamo’s final rankings.
“AAII Has
Limitations
in Evaluating Agent AI’s Actual Work Capabilities”
Regarding these results
,
experts acknowledged the usefulness of AAII while pointing out the limitations of its evaluation method.
A member of the Technology Innovation and Infrastructure Subcommittee of the National AI Strategy Committee stated, “Evaluations like Artificial Analysis are meaningful in that they allow for a quick comparison of the performance of various models,” but added, “There are limitations to fully assessing Agent AI’s actual task-performance capabilities using only the current evaluation method.”
This is because AI is evolving beyond simply answering questions to engaging in multiple rounds of conversation, verifying information by searching websites or reference materials, and then performing actual tasks.
He explained, “Agent AI requires multiple stages of reasoning and execution even when performing a single task,” adding, “It is difficult to fully evaluate such complex capabilities using existing benchmarks alone.”
Since evaluating actual task-performance capabilities requires examining the entire problem-solving process, the time and cost involved in evaluation also increase. For this reason, while existing benchmarks that allow for rapid comparison of multiple models’ performance are widely used today, they may differ from real-world AI service environments.
However, he noted, “This does not mean that the AAII is a meaningless metric.” He added, “While it is useful for comparing the performance of various models under the same criteria, it should not be viewed as the sole metric for evaluating model performance in the era of Agent AI.”
Upstage and SKTelecom: “It’s Difficult to Judge Model Maturity Based on a Single Metric”
Upstage and SKTelecom, which trailed Motif in the AAII rankings, also maintain that model
maturity
should not be judged
based
solely
on
AAII scores. LG Corp. did not issue a separate statement.
SKTelecom stated, “AAII is only one part of the Dokpa-mo evaluation criteria; expert and user evaluations are still pending,” adding, “Please take into account that it is difficult to assess a model’s maturity based on a single metric.”
The company added, “A model’s competitiveness should be evaluated not only by metrics but also by its performance, cost, and stability in real-world usage environments,” and noted, “We plan to improve areas that require refinement based on the metrics while simultaneously enhancing performance in actual usage environments.”
Upstage also commented, “While the AAII benchmarking results are not meaningless, they should not be viewed as an absolute standard for model performance.”
Will the “AAII No. 1” ranking lead to a first-place finish in the Dokpamo evaluation?
These AAII results are significant in that they compare domestic AI models against the same global standards. In particular, Motif 3’s high ranking (10th) compared to global models is regarded as a noteworthy achievement within the domestic AI industry.
However, there are clear limitations to judging a model’s overall competitiveness based solely on AAII scores. Since AAII is an indicator that aggregates multiple evaluation results into a single score, the outcome can vary depending on which evaluation criteria are included. Factors such as Korean language proficiency, the ability to perform tasks in specific industries, and aspects like cost, speed, and stability in actual service deployment are also difficult to fully capture with a single score.
In fact, Artificial Analysis provides various metrics separately—such as response speed, usage costs, and coding agent performance—in addition to the Comprehensive Intelligence Index. This means it is difficult to fully capture all aspects of an AI model’s competitiveness with a single composite score.
The current global evaluation rankings are: Motif in 1st place, Upstage in 2nd, SKTelecom in 3rd, and LG AI Research in 4th. However, this ranking does not necessarily reflect the final Dokpamo ranking.
Since the final results incorporate 25 points from the AAII, 15 points from the NIA evaluation, and assessments by experts and users, the rankings may change in the final outcome.
These AAII results serve as a significant reference point for gauging the relative performance of domestic AI models. However, for AI to be effectively utilized in actual industries and service environments, factors such as task performance, cost, speed, and stability must be considered alongside benchmark scores.
The key focus of the third Dokpamo evaluation is also expected to extend beyond the question of “who ranks first in the AAII” to how global benchmark scores and actual AI application capabilities will be comprehensively assessed.
Tension is mounting as domestic credit rating agencies adopt an extremely cautious stance ahead of the introduction of the new accounting standard, International Financial Reporting Standard (IFRS) 18…
F&F has secured the global trademark rights and related intellectual property (IP) for “Discovery Expedition,” marking the start of full-scale overseas expansion for the lifestyle outdoor brand Discov…
Vaccine specialist EuBiologics Co., Ltd.(206650)announced on the 13th that it recorded sales of 48.9 billion won and operating profit of 10.8 billion won in the first half of this year. Compared to th…