Internet

National AI Team Fails at ‘Point Allocation Design’… The Paradox of Motif’s Elimination Despite Topping the AAII Rankings [IT World Insights byKim Hyun-ah]

AAII at 47 points, despite the benchmark at 24.6 points and experts at 27.1 points Failed to make the cut with a user rating of 14.1… Overall score: 65.8 Expert and User Ratings: 60-point scale Emphasis on ‘Usability’ Over Technological Capabilities “What Is the Purpose of the ‘Dokpamo’ Project?”

Kim Hyun-ah
2026-08-18 16:44:45
[Edaily Reporter Kim Hyun-ah ] My first thought upon seeing the results of the second evaluation of the “Dokpamo” (Proprietary AI Foundation Model) was, “Is this correct?”

This is because Motif Technologies—which ranked first among Korean companies and 10th globally with a score of 47 on the Artificial Intelligence Comprehensive Index (AAII), conducted by the global AI model evaluation agency Artificial Analysis (AA) at the government’s request—was ultimately eliminated from the competition.

In contrast, the AI Research Institute at LG Corp.(003550), which scored 31 points on the AAII, advanced to the third round of evaluation. Upstage scored 37 points, and SKTelecom(017670)scored 35 points.


However, a closer look at the detailed evaluation results obtained by Edaily reveals the reasons behind Motif’s elimination more clearly.

Motif received 24.6 points in the benchmark evaluation and 27.1 points in the expert evaluation. However, it scored only 14.1 points in the user evaluation—8.4 points from expert users and 5.7 points from the general public.

The overall score was 65.8 points. While Motif outperformed all domestic competitors in the AAII, which measures global technological capabilities, its weakness in user evaluations ultimately led to its elimination.

The government also explained that while Motif’s technological capabilities were outstanding, its relatively low ratings in usability and practical applicability contributed to its elimination.

This raises an unavoidable question.

What is the essence of the Dokpamo project? Is it about finding a company to develop a world-class foundation model, or is it about selecting the AI that is best utilized right now?

Source:
Ministry of Science and ICT

AAII score of 47 points translated to 11.75 points in the final evaluation
The Dokpamo evaluation consisted of a total of 100 points, broken down into 40 points for the benchmark, 35 points for expert evaluation, and 25 points for user evaluation. Of these, the AAII accounted for 25 of the 40 benchmark points.

The raw AAII scores were 47 points for Motif, 37 for Upstage, 35 for SKTelecom, and 31 for LG Corp. AI Research.

Converted to a 25-point scale, Motif scored 11.75 points, Upstage 9.25 points, SKTelecom 8.75 points, and LG Corp. 7.75 points.

The gap between Motif and LG Corp. in the AAII raw scores is 16 points. However, the point difference reflected in the final evaluation is reduced to 4 points. Added to this are 15 points from the NIA evaluation, bringing the total benchmark score to 40 points.

The structure is such that even if a company shows a large gap in the global technology evaluation, that difference is significantly reduced in the final evaluation.

In contrast, the 60 points—comprising 35 points from expert evaluations and 25 points from user evaluations—are heavily influenced by assessments of not only the model’s performance but also its practical utility and impact.

The Motif case illustrates just how significantly this scoring structure can influence the results.

Commenting on the evaluation results, one AI expert questioned the purpose of the Dokpamo project, stating, “They made the outcome hinge entirely on the ‘real-world usability’ score—I can’t help but wonder if this is really how it’s supposed to be.”

In particular, he pointed out, “If the goal is to improve usability, you might as well just use an open-weight model,” suggesting that usability is a secondary consideration once a model’s performance exceeds a certain threshold.

[Edaily Reporter Moon Seung-yong]

Is “Usability” the Essence of Model Development
? Of course, this does not mean that usability and applicability evaluations are inherently flawed.

Verifying whether technological capabilities translate into actual services and industrial applications is crucial, even in national AI projects. The issue lies in the emphasis placed on it.

Good usability does not necessarily mean that the model itself is outstanding. This is because it can vary depending on the service environment and tools—known as a “harness”—built on top of the model.

Conversely, if the performance of a foundation model has reached world-class levels, determining how to turn it into a service could be considered a task for later.

In particular, it is worth considering whether it is appropriate to evaluate the practical applicability of startups—such as Motif, which achieved world-class model performance with a small team of about 30 people—using the same yardstick as competitors with large organizations and customer bases.

In this second round of evaluation, the weighting for “impact and contribution plans” was also increased from 10 points to 15 points.

The point here is not to conclude that Motif’s elimination was unfair.

The question is whether the scoring structure—which allowed a company that proved its global competitiveness with a score of 47 in the AAII evaluation (where even top-tier global models struggle to exceed 50 points) to be eliminated solely because of a user evaluation score of 14.1—truly aligned with the objectives of the Dobamo project.

Furthermore, 60 points out of the total 100—including 35 points for expert evaluation and 25 points for user evaluation—are areas heavily influenced by qualitative judgments.

Given this level of weighting, the evaluation criteria, detailed sub-categories, scores by each evaluator, and the final conversion method must be disclosed. Only then can we verify in which categories Motif’s score diverged from that of its competitors and to what extent that difference influenced its ultimate elimination.

What Is Needed Is Evaluation Verification, Not “Corporate Bailouts”
The Ministry of Science and ICT has notified the companies of the second-round evaluation results and plans to proceed with the objection and clarification process.

Motif, too, needs to use this process to verify the raw scores and detailed calculation formulas for the benchmark, NIA, expert, and user evaluations.

If there were no issues with the evaluation, the government should simply make the information public. Conversely, if there were scoring criteria or evaluation standards that companies could not have predicted in advance, they must be improved.

There is no need to view this controversy solely as a matter of Motif and LG Corp. winning or losing.

In national AI projects—which will involve investments of hundreds of billions, or even trillions, of won in taxpayer money—the extent to which technical capabilities are prioritized and how much usability and practical application are factored in are policy decisions.

However, the scoring criteria directly dictate the direction of the project.

If, while claiming to build a world-class foundation model, the structure actually allows usability evaluations to have a greater impact on the results than model performance, we must ask again what the purpose of this national AI project truly is.

Companies must be able to trust the evaluation criteria and formulate their development strategies accordingly, and once the results are in, anyone must be able to verify them.

What the national AI initiative needs is an evaluation system that allows companies striving to create world-class models to compete with confidence.

What we really need to reflect on in the recent Dokpamo controversy is not Motif’s elimination itself. Rather, it is whether the “scoring design” that led to that elimination was truly aligned with Dokpamo’s objectives.

Economy

Corporation

IT·Science

Economy

[VC's Pick] From Virtual IP to Space... A Flurry of Interest in Early-Stage Companies

This week (August 31–September 4), startups across various sectors—including virtual intellectual property (IP), freight transportation, social apps, and space technology—secured investments from vent…
2026-09-05 08:01:05

Corporation

SK BIOPHARMACEUTICALS’s $400 Million Gamble… A Bigger Picture Than ‘A Second Excofry’ [Invest Bio]

SK BIOPHARMACEUTICALS(326030)is attributing significance beyond that of a mere follow-up pipeline to Opakalim (BHV-7000), a new epilepsy drug candidate it acquired from the U.S.-based Biohaven. It is …
2026-09-05 06:01:02

IT·Science

GC Wellbeing and JW Pharmaceutical Rally [K-bio Pulse]

South Korea’s pharmaceutical and biotech sector experienced significant volatility on August 13, driven by growth themes in obesity/medical aesthetics, hair loss, and immuno-oncology. #GCWellbeing sur…
2026-09-05 08:02:02