[Park Yong-hoo / Perspective Designer] The results of the second evaluation for the “Independent Artificial Intelligence (AI) Foundation Model” (DOKPAMO) development project, announced by the Ministry of Science and ICT on the 18th, are clear on the surface. LG Corp.(003550) AI Research Institute, SKTelecom(017670), and Upstage have advanced to the next stage, while Motif Technologies has been eliminated. However, upon closer examination of the announcement, the only thing that is clear is “who remained.” “Why this happened” remains difficult to verify with hard data.
This evaluation was based on a composite score from three categories: benchmark (40 points), expert evaluation (35 points), and user evaluation (25 points). In the benchmark evaluation, the average score for the four teams was 22.5 points, with a 4.0-point gap between first and fourth place. The expert evaluation averaged 28.8 points with a gap of 2.4 points, while the user evaluation averaged 17.6 points with a gap of 5.0 points. Of the three categories, the user evaluation showed the widest gap.
[Park Yong-hoo / Perspective Designer]
The benchmark leader was eliminated
.
Looking more closely at the numbers, the most striking aspect of this announcement is not the “gap” but the “reversal.” Motif Technologies’ “Motif 3” scored 47 points on the AAII (Artificial Analysis Intelligence Index), a global AI performance index, ranking first among the four teams and tenth among models worldwide. This score overwhelmingly outperformed Upstage (37 points), SKTelecom (35 points), and LG Corp. (31 points).
Nevertheless, Motif was eliminated. This is believed to be because it failed to outperform the three teams that had advanced from the first round in domestic benchmark tests and expert and user evaluations. The government explained that the top team differed across the three evaluation categories, and no single company consistently ranked first. The team with the best global benchmark performance was eliminated after hitting a wall in “usability” and “applicability.” This is the real story behind this second round of evaluation.
How Did a 16-Point Gap in AAII Narrow to a 4-Point Gap in Benchmark Scores
?
However, there is one figure here that is difficult to comprehend. In the AAII raw scores, the gap between first-place Motif (47 points) and fourth-place LG Corp. (31 points) was a whopping 16 points, yet the gap between first and fourth place in the benchmark evaluation (40 points) was only 4.0 points. The 16-point performance gap was drastically compressed to just 4 points.
The secret lies in the internal composition of the 40-point benchmark. The benchmark is the sum of the AAII score (25 points) and the National Information Society Agency (NIA)’s own evaluation (15 points). While the AAII measures global performance metrics such as agent capabilities, coding, general knowledge, and scientific reasoning, the NIA assesses mathematics, knowledge, long-text comprehension, Korean language proficiency, safety, and reliability. It’s as if two completely different sets of criteria have been mixed into a single basket. If Motif ranked first in the AAII but the gap in the overall benchmark narrowed to 4 points, it is highly likely that the lead it established in global performance was largely eroded in the NIA’s domestic language and safety categories. The government’s own statement that “the top performers differed across the three evaluation categories” also suggests that Motif, which ranked first in the AAII, may not have been first in the overall benchmark.
The problem is that it is impossible to verify exactly which items or which calculation formula caused this “compression from 16 points to 4 points.” The government did not disclose the conversion formula used to convert each team’s detailed scores and the AAII raw scores to a 25-point scale. While it revealed how narrow the benchmark gap was, it blocked any way to quantitatively determine whether that gap was due to the NIA scores or the nature of the conversion formula.
Scores Are Missing, Only the Gap Remains
What stands out is that the Ministry of Science and ICT did not disclose the actual composite scores for individual teams. Since only the average and the gap between 1st and 4th place were presented, it is impossible to tell whether Motif was edged out by the 3rd-place team by a narrow margin or lagged far behind. While some metrics—such as the AAII raw scores—were disclosed, there is no way to verify the numerical weight of the overall score that determined the elimination. The phrase “lowest in the comprehensive evaluation” merely indicates the ranking; it does not reveal how narrow the gap actually was.
While this non-disclosure approach may be an attempt to avoid objections or controversy, the fact that it blocks any means of verifying the reasons for disqualification with concrete numbers—in a project funded by the national budget—is in itself a matter worthy of journalistic investigation. However, the government has offered an explanation that prevents this issue from being framed solely as a matter of fairness. The Ministry of Science and ICT stated that, regarding allegations of “benchmaxing” by some participating companies, it requested official verification from Artificial Analysis, the operator of AAII, and received a response stating, “We found no evidence to substantiate claims of systematic memorization or overfitting.” Benchmaxing refers to the practice of excessively optimizing a model to fit a specific test set or evaluation method.
The government further explained that, given the minimal differences between each category, it is difficult to conclude that any single category was decisive. While the criticisms surrounding the non-disclosure of scores are valid, a balanced perspective requires considering the broader context as well.
“Please verify reproducibility and accountability” — Formal media requests for disclosure
Calls for the
disclosure
of detailed evaluation data have also emerged from within the industry. According to an E-Daily report, a tip was received requesting the disclosure of the following nine items, arguing that the government’s
explanation of
“marginal differences” alone is insufficient to verify the process by which the results were derived. ① Detailed scores by category and final rankings for the four teams; ② The conversion formula for the raw scores from AAII and NIA; ③ the procedures for selecting expert evaluators and excluding conflicts of interest, ④ the distribution of anonymized expert evaluation scores, ⑤ whether the expert user and public evaluations were conducted blindly, ⑥ the order in which evaluation questions and models were presented, ⑦ the method for preventing duplicate participation, ⑧ the actual total score difference between 3rd and 4th place, and ⑨ the appeal process.
The whistleblower stated, “The intention is not to overturn the results, but to verify the reproducibility and accountability of the national AI selection process, which was funded by taxpayer money.” The nature of the requested items underscores the significance of this statement. The nine items mentioned here are not intended to overturn specific results, but rather constitute the minimum conditions required to verify “reproducibility” in a scientific experiment.
Disclosing the raw scores and conversion formula would allow external verification of the “AAII 16 points → Benchmark 4 points compression” structure pointed out in this article; the procedures for excluding conflicts of interest among evaluation committee members and whether the process was blind would verify the reliability of expert and user evaluations; the order in which models are revealed and the methods for preventing duplicate participation would help verify the control of bias in user evaluations, and the actual total score difference between third and fourth place would help verify the weight of elimination, respectively. If even just one of these were disclosed, the vague descriptions of “averages and gaps” used so far would be replaced by much more robust numbers.
In other words, these nine elements—which are currently undisclosed—constitute the black box of this evaluation. While the government may have grounds to claim the process is “fair,” the public will only have the basis to verify for themselves whether it is “fair” once this list is made public.
Although it is called a “model competition,” the expert evaluation is, in fact, a business evaluation
.
Another key point of contention regarding quality is the expert evaluation (35 points). According to the evaluation committee’s opinions released by the government in a press release, Upstage was commended for “linking its efforts with the ‘Daum’ portal and ‘Timely’ platform so that the public can tangibly experience the development outcomes” and for “reducing dependence on foreign hardware through collaboration with FuriosaAI’s NPU.” SKTelecom was recognized for “providing the A.X K1 model for the defense sector, demonstrating agents specialized for the manufacturing industry, demonstrations in the legal and tax sectors,” while LG Corp. AI Research was highly evaluated for “establishing a collaboration strategy with global international organizations” and “the distinctiveness of its agentic AI.”
Here, the only part that actually directly evaluated the “performance of the model itself” was SKTelecom’s claim of “top-tier performance among peers in mathematical reasoning and the Korean language domain.” The rest of the content—such as platform integration, domestic hardware production, industrial pilot projects, global collaboration strategies, and safety and trust frameworks—is closer to “business expansion strategies.” In fact, the evaluation criteria for the expert assessment are defined as “development strategy and technology, development achievements and plans, and ripple effects and contribution plans.” This means the criteria are designed to assess development strategies and business plans, not to measure performance through benchmarks.
This prompts us to reconsider the true nature of the “model competition” known as Dokpamo. Out of a total score of 100 points, only 40 points—the benchmark score—are based purely on model performance, while the remaining 60 points (35 from experts + 25 from users) are allocated to strategy, expansion, and user experience. The fact that Motif, which ranked first in AAII performance, was eliminated is not unrelated to this structure. Although the competition is called “AI Model,” more than half of the evaluation focuses not on the model’s capabilities but on “how to disseminate it, with whom to partner, and in which industries to apply it.”
25 points for user evaluation—yet the door to customer touchpoints was closed
To understand Motif’s elimination, one must also examine the interplay between the evaluation design and the consortium’s composition. To the best of my knowledge, Motif attempted to join forces with Kakao during the formation of the “AI for All” consortium to secure customer touchpoints. The strategy was to leverage the overwhelming subscriber base of KakaoTalk—South Korea’s largest messaging app—as the actual user base for its AI service. However, the government reportedly restricted Motif to forming a consortium with only one partner. Motif was forced to partner with KTCorporation alone, instead of the two partners it had originally sought, and Kakao—the largest potential customer touchpoint—was pushed out of the consortium.··
It is precisely at this point that the scoring structure reveals a fundamental contradiction. The “user evaluation” (25 points) in this second round of evaluation consisted of scores assigned by 49 AI experts and 185 members of the general public who directly tested the service. If the government truly placed such high importance on user feedback and actual usability, the logic would hold only if it had allowed collaboration with Kakao—which has the most extensive customer reach—and expanded the consortium to a two-company structure, similar to how KTCorporation partnered with Motif and Upstage. Requiring Motif to undergo evaluation without a partner possessing a large user base, while then judging that very usability and perceived quality with a weighting of 25 points, constitutes a set of conflicting criteria.
The fact that Motif, which entered the user evaluation without securing a major customer touchpoint, fell behind—and that the gap (5.0 points) was the largest among the three categories, exceeding both the benchmark (4.0 points) and the expert evaluation (2.4 points)—is consistent with this structural disadvantage. In an evaluation design where passing is not guaranteed by benchmark performance alone, if elimination is determined by “user evaluation” while institutionally blocking opportunities to expand user touchpoints from the outset, it is difficult to attribute the full weight of that elimination solely to Motif’s technology. This point, too, cannot be resolved by the government’s explanation alone unless the previously mentioned “request for disclosure of detailed data” is fulfilled.
What Does “Advancement” Actually Mean
?
Finally, we must examine what the statement that LG Corp., SKT Corporation, and Upstage will “continue the final competition in the next stage” actually means. The government announced plans to select the final two teams early next year following the third evaluation. The number of teams to be selected has been disclosed. However, the detailed structure of the funding flow—such as how the budget is allocated per team and how resources originally assigned to the eliminated Motif will be redistributed to the remaining teams—remains unexplained. Behind the spotlight of the performance competition, there is no answer as to how taxpayer money will flow to the remaining teams. Only when news reports on this three-way race delve into this point will the term “National AI” truly hold meaning.
With the detailed scoring criteria, conversion formula, selection process for evaluators, appeal process, and funding redistribution method all still undisclosed, only one thing remains at this stage. The government’s statement that “it was a very close call” stands in contrast to the demands from the industry and the public to “demonstrate that this narrow margin is reproducible.” Whether this announcement will remain merely the “result” of the national AI selection process or serve as the “starting point for questions” about how the national budget is evaluated will be determined by the government’s next response to these demands.
#NEXTBIOMEDICAL CO., LTD., a company specializing in innovative medical devices, has been named to the Korea Exchange’s “KOSDAQ Rising Star” list for the second consecutive year, once again demonstrat…
Portable X-ray manufacturer Remedi saw its stock price surge after reporting explosive earnings growth. The company is expanding rapidly not only in South Korea but also overseas, raising expectations…
The results of the second evaluation for the “Independent Artificial Intelligence (AI) Foundation Model” (DOKPAMO) development project, announced by the Ministry of Science and ICT on the 18th, are cl…