Internet

LG Corp. Ranks Last in AAII Survey, but Tops Dokpamo Experts’ Evaluation… “Evaluation Criteria Should Be Disclosed”

LG Corp. AI Research Receives Top Marks in Expert Evaluation Ranked Last Among Four Teams at AAII AI Industry: “We Need to Verify Who Conducted the Evaluation and What Criteria Were Used”

Kim Hyun-ah
2026-08-20 20:44:59
[Edaily Reporter Kim Hyun-ah ] Following the announcement that the “ LG Corp.(003550) ” AI Research Institute received the highest score in the expert evaluation during the second-stage assessment of the “Independent AI Foundation Model (DOKPAMO)” project, critics in the AI industry have called for the disclosure of more specific criteria used in the evaluation.

Given the significant discrepancy between the AAII benchmark results and the expert evaluation, critics argue that it should be possible to verify which technical elements led to high scores in the expert evaluation and what criteria the evaluators used to make their judgments.

The Ministry of Science and ICT explained that this evaluation did not determine rankings based solely on specific benchmark scores, but rather comprehensively assessed factors such as expertise, usability, and industrial impact.



On the 20th, following questions raised by Edaily, the Ministry of Science and ICT released additional details regarding the criteria and results of the second-stage evaluation of the Dokpamo competition. According to this information, in the expert evaluation—scored out of 35 points—LG Corp. AI Research received 29.5 points, the highest score among the four elite teams. The average score for the four teams was 28.75 points.

In contrast, in the Artificial Intelligence Index (AAII)—a global AI benchmark—Motif Technologies took first place with 11.9 points out of a possible 25, while LG Corp. AI Research ranked last among the four teams. The AAII rankings were Motif Technologies, Upstage, SKTelecom, and LG Corp. AI Research, in that order.

While Motif Technologies, which ranked first in the AAII, was eliminated, LG Corp. AI Research—which had ranked last in the AAII—took first place in the expert evaluation, resulting in significantly divergent outcomes across the different evaluation categories.

Within the AI industry, there are calls to ensure transparency regarding the evaluation process and the basis for judgments, rather than focusing solely on the results themselves.

One expert stated, “It is crucial to determine whether the evaluation panel included members with the expertise to properly assess AI model development methodologies and technical achievements,” adding, “We need to verify who evaluated the models’ technical performance and development strategies, and based on what criteria.”

He continued, “Due to potential conflicts of interest with the elite teams, it is possible that the experts with the deepest understanding of actual AI model development and technology were excluded from the review,” noting, “If that is the case, we need to examine the specific expertise on which the evaluators based their assessment of technical performance.”

Another expert argued that, based on the technical materials released by LG Corp.'s AI Research division, there is a need to explain in greater detail why the model received the highest score in the expert evaluation.

This expert pointed out, “Looking at the LG Corp. Tech Report, while the model size was increased, SFT was not properly implemented. So, what exactly is innovative about that?” They added, “It is difficult to understand how the team could have ranked first in the expert evaluation, which is inevitably linked to benchmark results.” The point is that it is difficult to explain why the model received a high technical score in the expert evaluation simply by increasing its size. SFT (Supervised Fine-Tuning) is a supervised fine-tuning method that uses data with human-provided correct answers to further train an AI model’s ability to generate responses.



Expert Evaluation: LG Corp. Scores 29.5 Points… What Were the Evaluation Criteria
? According to the Ministry of Science and ICT
,
the expert evaluation was conducted by 10 external experts—including 3 from industry, 5 from academia, and 2 from the research sector—via written evaluations and a question-and-answer session.

The evaluation criteria totaled 35 points, with 10 points for development strategy and technology, 10 points for development achievements and plans, and 15 points for ripple effects and contribution plans. This is higher than the 25-point weighting assigned to the global benchmark, AAII.

In the “Development Strategy and Technology” category, the evaluation assessed the excellence and innovation of the developed technology, its technical originality, the level of performance improvement, resource utilization and processing efficiency, and technologies for developing safe and reliable AI models. “Development Achievements and Plans” as well as “Ecosystem and Global Impact” were also included in the evaluation.

Each team’s score was calculated as the arithmetic mean of the scores assigned by the evaluators, excluding the highest and lowest scores.

However, the specific basis for the evaluation—such as the exact questions asked by the evaluators, how the companies responded, and which technical factors led to high scores—was not disclosed.

Industry insiders argue that while there is no need to disclose the personal identities of individual evaluators, it is necessary to make public the detailed scores for each team, scores for each evaluation category, and the key criteria used in the assessment.

One expert stated, “If development strategies and technical achievements were evaluated, it should be possible to verify what questions were asked, how the companies responded, and the basis on which the evaluators assigned their scores,” adding, “Otherwise, it is merely a label of ‘expert evaluation,’ and it is difficult to verify what actual expertise the evaluation was based on.”



Government: “The Evaluation Was Not Determined by AAII Alone”
The Ministry of Science and ICT explained that this evaluation did not determine the final rankings based solely on a specific global benchmark.

The second-round evaluation was out of a total of 100 points, consisting of 40 points for the benchmark, 35 points for the expert evaluation, and 25 points for the user evaluation.

In the global AAII benchmark, Motif Technologies took first place with 11.9 points, while SKTelecom scored the highest in the NIA benchmark with 13.4 points. SKTelecom also ranked first in the AI expert user evaluation with 11.6 points, and LG Corp. received the highest score of 7.6 points in the general public evaluation.

The Ministry of Science and ICT explained that different teams demonstrated strengths in each evaluation category.

Regarding Motif Technologies’ elimination, the Ministry explained that it was not simply due to low practicality, but rather that the team received low scores overall in both expertise and usability, adding, “It is true that their performance fell short.”

The ministry explained that when considering the NIA benchmark—which reflects the domestic environment—along with expert evaluations and actual user evaluations, in addition to the global public benchmarks where Motif Technologies had shown strength, the final results were bound to differ.

The Ministry also explained that the user evaluations focused not merely on website design or UI/UX, but on the content and quality of the output generated by the AI models. A total of 49 AI experts and 185 members of the general public evaluated the services by actually using them.
“Why Did These Results Emerge?”
The focus of this controversy lies not so much on the elimination of Motif Technologies—which ranked first in the AAII—or the fact that LG Corp. AI Research topped the expert evaluation, but rather on how transparently the ministry can explain the criteria and weighting used to synthesize the results of models that demonstrated strengths in different evaluation categories.

From the Ministry of Science and ICT’s perspective, a national representative AI cannot be selected based on a single global benchmark alone; performance in the domestic environment, technological development capabilities, actual usability, and industrial impact must all be considered together.

On the other hand, some in the industry, while agreeing with this rationale, argue that since the expert evaluation accounts for 35% of the total score, the public should be able to verify the specific reasons why LG AI Research received 29.5 points and the exact margin by which it scored compared to other teams.

While the Ministry of Science and ICT disclosed the first-place finishers in each category, the highest overall scores, and the average scores of the four teams, it did not reveal the detailed scores for each team or the specific score differences between them. Critics point out that, given this is a selection process for a national AI initiative funded by public money, it is necessary to disclose evaluation information at a level that allows for verification of why such results occurred—beyond simply identifying which team took first place.

Economy

Corporation

IT·Science

Economy

[Market Insight] The Corporate Bond Market Is Heating Up… But Will a Cold Snap Return by Year-End?

Credit spreads—an indicator of corporate bond investment sentiment—continue to narrow (signaling a bull market in corporate bonds), despite concerns about some investors pulling out. This trend is dri…
2026-08-20 20:07:04

Corporation

SILICON 2 Co.,Ltd., a 'K-Beauty Distributor,' Sets Sights on the Wellness Market… Global Launch of 'Gamtan Bra'

SILICON 2 Co.,Ltd., a K-Beauty distribution platform company, is expanding its business into the wellness sector by introducing the innerwear brand “Gamtan Bra” to the global market. SILICON 2 Co.,Ltd…
2026-08-20 18:39:16

IT·Science

LG Corp. Ranks Last in AAII Survey, but Tops Dokpamo Experts’ Evaluation… “Evaluation Criteria Should Be Disclosed”

Following the announcement that the “ LG Corp.(003550) ” AI Research Institute received the highest score in the expert evaluation during the second-stage assessment of the “Independent AI Foundation …
2026-08-20 20:44:59