Internet

Reason for AAII No. 1 Motif’s Elimination: ‘Usability Evaluation’… National AI Initiative Linked to Frontier Project [Q&A]

Deputy Minister Ryu Je-myeong: “Motif’s technological capabilities are world-class” "Rated lower than other companies in terms of usability and practicality” “Benchmarking Controversy Not Reflected in Evaluation” "AAII Review Finds No Evidence of Memorization or Overfitting" The 3rd Round of the "Dokpamo" Evaluation Proceeds as Scheduled Reviewing a Restructuring Plan to Compete in the Frontier-Class Model Market

Kim Hyun-ah
2026-08-18 12:36:44
[Edaily Reporter Kim Hyun-ah ] The Ministry of Science and ICT cited the “usability and applicability evaluation” as the reason for Motif Technologies’ elimination from the second round of evaluations for the Independent AI Foundation Model (DOKPAMO) project, despite the company having ranked first in the AAII (Artificial Analysis Intelligence Index).

In the second round of the Dokpamo evaluation, Upstage, SKTelecom(017670), and the LG Corp.(003550) AI Research Institute passed.

Ryu Je-myeong, Second Vice Minister of Science and ICT, stated during a briefing on the results of the second round of the Dokpamo evaluation on the 18th “It is clear that Motif demonstrated outstanding performance with an exceptionally high AAII score,” but added, “While the AAII accounts for 25 points out of a possible 100, and considering the other evaluation criteria, Motif’s technical capabilities were excellent, it received a lower score than other companies in the usability and applicability category, which carried significant weight with a 75-point allocation.”

Vice Minister Ryu also stated, “Issues related to benchmarking were not reflected in the evaluation,” adding, “Based on AAII’s official response and expert evaluations, we concluded that there was no unfair technical support or similar practices.” Benchmarking refers to the practice of excessively optimizing a model to fit a specific test set or evaluation method.

Ryu Je-myeong, Second Vice Minister of Science and ICT. Photo by Reporter Kim Hyun-ah, Edaily


The following is a Q&A session between Second Vice Minister Ryu Je-myeong and other officials from the Ministry of Science and ICT and reporters.
Q. Why was Motif, which ranked first in the AAII evaluation, eliminated in the second round of the Dokpamo evaluation?

Motif achieved an exceptionally high AAII score. It is an indisputable fact that it demonstrated truly outstanding performance in that evaluation.

However, out of a possible 100 points, the AAII score accounts for only 25 points. Looking at the other evaluation criteria as a whole, while its technical capabilities were outstanding, it received lower scores than other companies in the areas of usability and applicability—which account for the remaining 75 points and carry significant weight in the evaluation. This is how the results can be interpreted.

I hope you understand it that way.

Q. Does Motif’s low rating for usability apply to both the expert evaluation and the AI power user evaluation?

We would be happy to disclose company rankings and specific scores if they are meaningful information, but we are concerned that score differences could lead to differing interpretations or cause unintended disadvantages for the companies involved.

As was the case in the first round, it is difficult to disclose in detail exactly which criteria resulted in which evaluations.

To reiterate, Motif has proven itself to be world-class in terms of technical capability, and this was fully reflected in the evaluations.

However, I would appreciate it if you could explain that Motif received a lower evaluation compared to other models in terms of usability and practicality.

Q. Are there any separate government-level support measures or follow-up actions for Motif?

Motif Technologies has gained global recognition through this result on the AAII Global Benchmark Leaderboard. We view this as a tremendous achievement, as it demonstrates our global competitiveness in AI technology, starting with the design of a highly original algorithmic model.

However, we have been a technology company that has built our R&D-focused capabilities with a small research team. We had previously stated that we would expand our focus on usability and applicability in the second evaluation following the first round, and these efforts were reflected in the results this time as well.

As our AI models are set to be utilized across various industrial sectors and the public sector in the future, we believe Motif has already laid a solid foundation in this regard.

As for the specific question of how to support Motif, other than Motif applying for and receiving support through future government programs, I do not believe it would be appropriate to discuss any special measures or initiatives specifically for Motif at this time.

Source: Artificial Analysis


Q. Was Motif excluded because it lacked partnerships, collaboration, or contributions to the ecosystem compared to large corporations? This is listed as an evaluation criterion.

I would ask you to understand this as a result of a comprehensive evaluation—the gaps between the top-ranked company and the second- and third-ranked companies in each category were so narrow that it is difficult to pinpoint any single criterion as decisive.

The gaps were very narrow across all three categories: the benchmark evaluation, expert evaluation, and user evaluation.

Overall, the gap between first and fourth place was 4 points in the benchmark evaluation and 5 points in the user evaluation, so you can think of it as fluctuating within that range.

Q. Did the benchmarking controversy affect Motif’s elimination?

In conclusion, issues related to benchmarking were not reflected in the evaluation.

We based our decision on AAII’s official response, and after verifying these various matters through expert evaluations, we concluded that there was no unfair technical support or similar issues that would have compromised the fairness of the competition.

I understand there was some controversy surrounding the benchmarking issue. We officially requested an analysis from Artificial Analysis regarding the benchmark evaluation and received an official response.

Artificial Analysis explained the measures it takes to prevent models that receive high scores from being optimized solely to perform well on tests. They responded that, after reviewing the specific matter we had requested, they found no evidence to substantiate claims of systematic memorization or overfitting.

Therefore, we were informed that there are no major issues in that regard. Regarding the benchmarking issue, I will consider the fact that we received an official response from the relevant evaluation agency to be our answer.

Q. Is there a possibility that the 40-point weighting for benchmark scores will be revised in the future?

The allocation of points was determined by combining the major benchmarks established during the first round—which were created through consensus among the participating elite teams—with the individual benchmarks that the participating elite teams proposed to compare their performance against.

Since that approach had both pros and cons, we discussed standardizing on AAII—a globally recognized benchmark—for the second stage, which is why we implemented this change.

While there are differences of opinion regarding scoring and related matters from various perspectives, overall, the scoring evaluation method and point allocation were finalized after considering the opinions of participating companies and discussions among experts.

Please understand that this was achieved by thoroughly addressing the objections raised by the participating companies and working to build consensus until those objections were resolved.

DOKPAMO Phase 3 Evaluation to Proceed as Scheduled… Reviewing Restructuring of Frontier-Level Model Competition System
Q. Could the existing plan to ultimately select one or two companies in the DOKPAMO Phase 3 evaluation be changed?

The stated goal of the Independent AI Foundation Model Development Project was to select two companies as the final winners based on the results of the third round, with these two companies then working for one year in 2027 to develop AI models possessing approximately 95% of global-level capabilities.

However, as we mentioned during the first briefing and are reiterating now, the purpose of this project is not to select one or two winning companies, but rather to provide direct and indirect incentives for our AI ecosystem to advance to a global level through this competition.

The goal is to enhance the competitiveness of the AI ecosystem itself.

Regarding the third evaluation, the government budget will be finalized shortly, and we will provide a detailed explanation at that time.

Q. Are you creating a new “Frontier Model” project to complement the existing “Dokpamo,” or are you restructuring the existing “Dokpamo” into a “Frontier” competition format?

We are not yet at the stage where we can provide more specific details, so we ask for your understanding.

I would like to reiterate our overall perspective on this issue.

The model performance of global frontier companies is advancing exponentially, and the expansion of the global service ecosystem driven by this progress is unfolding at an extremely rapid pace.

The companies participating in the Dokpamo project also agree that a new competitive landscape is needed.

We need a new approach to develop models within our domestic ecosystem that can rival those of global frontier companies and to ensure that both the domestic and international AI ecosystems remain sustainable and continue to maintain their capabilities.

We are holding extensive discussions with the companies themselves—the direct stakeholders—on whether to distribute resources as we have done so far or to concentrate them, and on what kind of consortium would be most desirable.

Within the government as well, discussions are underway regarding the feasibility and effectiveness of financial support. I will have the opportunity to provide further details as the plans take shape.

Q. What is the relationship between the third evaluation of the Dokpamo project and the Frontier AI Project?

The third phase of the current Dokpamo project is underway in the second half of this year, and if we proceed as planned in 2027, a new competitive landscape will emerge. As for how we will address this issue in relation to that, I believe I will be able to provide an update on the discussions as soon as the government budget is finalized.

There is a consensus that we should move away from the current competitive model and restructure the competition in a way that allows us to develop models that are truly on par with those of global frontier companies.

Q. What is the schedule for the third evaluation, and what are the specific evaluation methods?

We notified the companies of the results of the third-stage evaluation today. Since the objection period is still open, if any objections are raised, we will address them, and the process will proceed immediately thereafter.

Regarding the specific evaluation plan for the third phase, we plan to review the first and second evaluations to identify any areas for improvement and, after consulting with the three participating companies, finalize and announce the plan as soon as possible.

Q. What is the scale of GPU support for the three teams that passed the third evaluation?

(Director Choi Dong-won) Regarding GPU costs, we estimate that leasing GPUs for six months would cost approximately 40 billion won.

Since there are three teams, you can estimate the total cost of the GPUs to be provided to them during this third phase at approximately 120 billion won.

Public Evaluation Has No Impact on Selection Results… Collaboration with AAII for the Third Evaluation
Q. What was the public’s reaction to the evaluation, and did it influence the selection results?

We initially recruited participants for the public evaluation with the goal of forming a panel of 200 members from the general public. At the time, approximately 1,400 people applied, and we randomly selected 200 of them to serve on the evaluation panel.

We assess that public enthusiasm for participation was extremely high. Ultimately, 185 members of the public participated.

Upon final review, the public evaluation did not influence the selection results.

Q. Why were the user evaluation scores generally low?

A total of 185 people ultimately participated.

We randomly selected participants from among the 1,400 applicants and verified each individual by requiring them to submit a declaration regarding potential conflicts of interest with their employer and provide information on their current organization to rule out such conflicts.

There were some cases of conflicts of interest, and even among those who actually participated, there were some who did not submit their responses.

I’m not sure if the evaluation results can be described as “harsh.” The variance across each indicator is not significant.

Since there are only four companies, the internal variations within each indicator are not significant, and the differences in evaluation scores are also relatively minor.

In particular, for the user evaluations, we filtered out cases where votes were overwhelmingly concentrated on a specific company, or where participants awarded either very high scores, perfect scores, or the lowest possible scores.

Taking these factors into account, the resulting variance could not be considered significant compared to other indicators.

Q. What kind of collaboration are you pursuing with Artificial Analysis?

AAII continued to show interest in Korea’s independent AI foundation model project during this evaluation as well.

For this evaluation, they created a dedicated platform and provided concrete cooperation to ensure the evaluation could be conducted accurately and fairly within a short timeframe.

We plan to continue this collaborative relationship regarding AAII’s evaluations, including the third round, to facilitate the exchange of expert opinions and consultations.

The emergence of
new startups like Motif… a process that enhances the competitiveness of the domestic AI ecosystem”
Q. Isn’t the ultimate goal of the Dokpamo Project to select one or two national AI representatives?

The purpose of this project is not to select one or two winning companies, but rather to provide direct and indirect incentives for our AI ecosystem to develop to a global level through this competition.

There are about 58 companies producing AI models registered with the AAII Global Benchmark. As you may know, the majority of these 58 companies are based in the United States and China.

However, looking at the countries producing these 58 competing companies, it appears that 12 nations—including the United States, China, and South Korea—are developing AI models. Seven of these companies are from South Korea.

In a global AI ecosystem increasingly dominated by the United States and China, we are closely watching as South Korea grows into a key player, building a robust AI ecosystem unlike that seen in other countries.

Just as Motif—which was not selected in the first round—participated in the second round and demonstrated outstanding performance, a small AI startup with about 30 employees achieved remarkable results through a government project despite a lack of government support.

We view this process itself as evidence that South Korea’s AI ecosystem is growing into a major player in the global AI landscape, second only to the United States and China, and that it is steadily moving toward that goal.

Q. Was there a team that took first place in all three areas in the second round of evaluation?

Yes, they were all different. They were all different, and no single company consistently ranked first.

Economy

Corporation

IT·Science

Economy

Shouting “Incredible” and giving a thumbs-up… Madison Hwang Holds a 2-Hour, 30-Minute ‘Robot Meeting’ with LG Corp.

“Incredible.” Madison Huang, Senior Director of Product Marketing for NVIDIA’s Omniverse and Robotics, waves to greet reporters through her car window after a meeting at LGELECTRONICS’ Yangjae Rese…
2026-08-18 11:37:40

Corporation

"The Japanese Market: Close Yet Challenging"... K Retail Launches a 'Speed Campaign' to Conquer the Archipelago

Domestic retailers are increasingly expanding into the Japanese market. This trend spans across various sectors, from e-commerce to department stores and cosmetics. While past expansions were primaril…
2026-08-18 11:41:03