National AI Team Fails at ‘Point Allocation Design’… The Paradox of AAII’s No. 1 Motif Being Eliminated [Kim Hyun-ah’s “Reading the IT World”]
AAII at 47 points, despite the benchmark at 24.6 points and experts at 27.1 points
Failed to make the cut with a user rating of 14.1 points… Overall score: 65.8 points
Expert and User Ratings: 60-point scale
Emphasis on ‘Usability’ Over Technological Capabilities
“What Is the Purpose of the Dokpamo Project?”
[Edaily Reporter Kim Hyun-ah ] My first thought upon seeing the results of the second evaluation of the “Dokpa-mo” (Proprietary AI Foundation Model) was, “Is this really right?”
This is because Motif Technologies—which had ranked first among domestic companies and 10th globally with a score of 47 on the Artificial Intelligence Index (AAII), conducted by the global AI model evaluation agency Artificial Analysis (AA) at the government’s request—was ultimately eliminated.
In contrast, the AI Research Institute at LG Corp.(003550), which scored 31 points on the AAII, advanced to the third round of evaluation. Upstage scored 37 points, while SKTelecom(017670)scored 35 points.
However, a closer look at the detailed evaluation results obtained by Edaily reveals the reasons behind Motif’s elimination more clearly.
Motif received 24.6 points in the benchmark evaluation and 27.1 points in the expert evaluation. However, it scored only 14.1 points in the user evaluation—8.4 points from expert users and 5.7 points from the general public.
The overall score was 65.8 points. While Motif outperformed all domestic competitors in the AAII, which measures global technological capabilities, its weakness in user evaluations ultimately led to its elimination.
The government also explained that while Motif’s technological capabilities were outstanding, its relatively low ratings in usability and practical applicability contributed to its elimination.
This raises an unavoidable question:
What is the essence of the Dokpamo project? Is it about finding a company to develop a world-class foundation model, or is it about selecting the AI that is best utilized right now?
Source:Ministry of Science and ICT
AAII score of 47 points translates to 11.75 points in the final evaluation
The Dokpamo evaluation consisted of a total of 100 points, broken down into 40 points for the benchmark, 35 points for expert evaluation, and 25 points for user evaluation. Of these, the AAII accounted for 25 of the 40 benchmark points.
The raw AAII scores were 47 points for Motif, 37 for Upstage, 35 for SKTelecom, and 31 for LG Corp. AI Research.
Converted to a 25-point scale, Motif scored 11.75 points, Upstage 9.25 points, SKTelecom 8.75 points, and LG Corp. 7.75 points.
The gap between Motif and LG Corp. in the AAII raw scores is 16 points. However, the point difference reflected in the final evaluation is reduced to 4 points. When the 15-point NIA evaluation is added to this, the total benchmark score is 40 points.
The structure is such that even if there is a significant gap in the global technology assessment, that difference is greatly reduced in the final evaluation.
In contrast, the 60 points—comprising 35 points from expert evaluations and 25 points from user evaluations—are heavily influenced not only by the model’s performance but also by its practical applicability and impact.
The recent Motif case demonstrates just how significantly this scoring structure can influence the results.
Commenting on the evaluation results, one AI expert questioned the purpose of the Dokpamo project, stating, “It all came down to the actual usability score—I can’t help but wonder if this is how it’s supposed to be.”
In particular, he pointed out that “if the goal is to improve usability, you can simply use an open-weight model,” suggesting that usability is a factor to be considered only after the model’s performance has reached a certain threshold.
Is “Usability” the Essence of Model Development
?
Of course, this does not mean that evaluating usability and applicability is inherently wrong.
Verifying whether technological capabilities translate into actual services and industrial applications is important even in national AI projects. The issue lies in the weight given to this factor.
Good usability does not necessarily mean that the model itself is outstanding. This is because it can vary depending on what kind of service environment, tools, or so-called “harness” is built on top of the model.
Conversely, if a foundation model’s performance has reached a world-class level, determining how to turn it into a service could be considered a task for later.
In particular, it is worth considering whether it is appropriate to evaluate the practical applicability of startups—such as Motif, which achieved world-class model performance with a small team of about 30 people—using the same yardstick as competitors with large organizations and customer bases.
In this second round of evaluation, the weighting for “ripple effect and contribution plan” was also increased from 10 points to 15 points.
The point here is not to conclude that Motif’s elimination was unfair.
The question is whether the scoring structure—which allowed a company that demonstrated global competitiveness with a score of 47 in the AAII evaluation (where even top-tier global models struggle to exceed 50 points) to be eliminated solely due to a user evaluation score of 14.1—truly aligned with the objectives of the Dopa Moe project.
Furthermore, 60 out of the total 100 points—including 35 points for expert evaluation and 25 points for user evaluation—are areas heavily influenced by qualitative judgments.
Given this level of weighting, the evaluation criteria, detailed subcategories, scores by each evaluator, and the final conversion method should all be made public. Only then can we verify in which categories Motif’s score diverged from that of its competitors and to what extent that difference contributed to its ultimate elimination.
What Is Needed Is Not “Corporate Rescue” but Evaluation Verification
The Ministry of Science and ICT has notified the companies of the second-round evaluation results and plans to proceed with the objection and explanation process.
Motif, too, needs to use this procedure to verify the raw scores and detailed calculation formulas for the benchmark, NIA, expert, and user evaluations.
If there were no issues with the evaluation, the government should simply disclose the information. Conversely, if there were scoring criteria or evaluation standards that companies could not have predicted in advance, they must be improved.
There is no need to view this controversy solely as a matter of Motif and LG Corp. winning or losing.
In national AI projects—which will involve investments of hundreds of billions, or even trillions, of won in taxpayer money—it is a policy choice to determine how much weight to place on technical capabilities and to what extent usability and practical application should be reflected.
However, the scoring criteria directly dictate the direction of the project.
If, while claiming to build a world-class foundation model, the evaluation structure actually allows usability assessments to have a greater impact on the results than model performance, we must ask again what the purpose of this national AI project truly is.
Companies must be able to trust the evaluation criteria and formulate their development strategies accordingly, and once the results are in, anyone should be able to verify them.
What the national AI initiative needs is an evaluation system that companies striving to create world-class models can trust and compete under.
What we really need to reflect on in this Dokpamo controversy is not Motif’s elimination itself. Rather, it is whether the “scoring design” that led to that elimination was truly aligned with Dokpamo’s objectives.
Keystone Private Equity (Keystone PE) has secured the contract to operate public sports facilities in the Sinbanpo 4 District of Seocho-gu, Seoul. This is not a typical buyout—where a company is acqui…
Daesang(001680), which began with the MSG brand “Miwon,” is now expanding its fermentation technology into the field of pharmaceutical raw materials. Moving beyond the mass production of general-purpo…
My first thought upon seeing the results of the second evaluation of the “Dokpa-mo” (Proprietary AI Foundation Model) was, “Is this really right?”This is because Motif Technologies—which had ranked fi…