LG Corp. Ranks First in Expert Evaluations by Dokpamo… “Evaluation Criteria Should Be Disclosed”
LG Corp. AI Research Receives Top Marks in Expert Evaluation
Ranked Last Among Four Teams at AAII
AI Industry: “We Need to Verify Who Conducted the Evaluation and What Criteria Were Used”
[Edaily Reporter Kim Hyun-ah ] Following the second-stage evaluation of the “Independent AI Foundation Model (DOKPAMO)” project, in which the “ LG Corp.(003550) ” AI Research Institute received the highest score in the expert evaluation, the AI industry has called for the disclosure of more specific criteria used in the assessment.
Given the significant discrepancy between the AAII benchmark results and the expert evaluation, critics argue that it should be possible to verify which technical elements led to high scores in the expert evaluation and what criteria the evaluators used to make their judgments.
The Ministry of Science and ICT explained that this evaluation did not determine rankings based solely on specific benchmark scores, but rather comprehensively assessed factors such as expertise, usability, and industrial impact.
On the 20th, following questions raised by Edaily, the Ministry of Science and ICT released additional details regarding the criteria and results of the second-stage evaluation of the Dokpamo competition. According to this information, in the expert evaluation—scored out of 35 points—LG Corp. AI Research received 29.5 points, the highest score among the four elite teams. The average score for the four teams was 28.75 points.
In contrast, in the Artificial Intelligence Index (AAII)—a global AI benchmark—Motif Technologies took first place with 11.9 points out of a possible 25, while LG Corp. AI Research ranked last among the four teams. The AAII rankings were as follows: Motif Technologies, Upstage, SKTelecom, and LG Corp. AI Research.
While Motif Technologies, which ranked first in the AAII, was eliminated, LG Corp. AI Research—which had ranked last in the AAII—took first place in the expert evaluation, resulting in significantly divergent outcomes across the different evaluation categories.
Within the AI industry, there are calls to verify the evaluation process and the basis for judgment rather than focusing solely on the results themselves.
One expert stated, “It is crucial to determine whether the evaluation panel included members with the expertise to properly assess AI model development methodologies and technical achievements,” adding, “We need to verify who evaluated the models’ technical achievements and development strategies, and based on what criteria.”
He continued, “Due to potential conflicts of interest with the top-tier teams, it is possible that the experts with the deepest understanding of actual AI model development and technology were excluded from the review,” and pointed out, “If that is the case, we need to examine the basis of the evaluators’ expertise when judging technical performance.”
Another expert argued that, based on the technical materials released by LG Corp.'s AI Research division, there is a need to explain in greater detail why the model received the highest score in the expert evaluation.
This expert pointed out, “Looking at the LG Corp. Tech Report, they increased the model’s size but failed to properly implement SFT. So what exactly is innovative about that?” He added, “It’s hard to understand how they ranked first in the expert evaluation, which is inevitably linked to the benchmark results.” The point is that it is difficult to explain why the model received a high technical score in the expert evaluation simply by increasing its size. SFT (Supervised Fine-Tuning) is a supervised fine-tuning method that uses data with human-provided correct answers to further train an AI model’s ability to generate responses.
Expert Evaluation: LG Corp. Scores 29.5 Points… What Were the Evaluation Criteria
?
According to the Ministry of Science and ICT
,
the expert evaluation was conducted by 10 external experts—including three from industry, five from academia, and two from research institutions—through written evaluations and a question-and-answer session.
The evaluation criteria totaled 35 points, broken down as follows: 10 points for development strategy and technology, 10 points for development achievements and plans, and 15 points for ripple effects and contribution plans. This is higher than the 25 points allocated to the global benchmark, AAII.
In the “Development Strategy and Technology” category, the evaluation focused on the excellence and innovation of the developed technology, its technical originality, the level of performance improvement, resource utilization and processing efficiency, and technologies for developing safe and reliable AI models. Development achievements and plans, as well as ecosystem and global impact, were also included in the evaluation.
Each team’s score was calculated as the arithmetic mean of the scores given by the evaluators, excluding the highest and lowest scores.
However, the specific basis for the evaluations—such as the exact questions posed by the evaluators, how the companies responded, and which technical elements led to high scores—was not disclosed.
Industry experts argue that while there is no need to disclose the personal identities of individual evaluators, it is necessary to make public the detailed scores for each team, the scores for each evaluation category, and the key criteria used in the assessment.
One expert stated, “If development strategies and technical achievements were evaluated, it should be possible to verify what questions were asked, how the companies responded, and the basis on which the evaluators assigned their scores,” adding, “Otherwise, it is merely called an ‘expert evaluation’ in name only, and it is difficult to verify what actual expertise the evaluation was based on.”
Government: “The Evaluation Was Not Determined by AAII Alone”
The Ministry of Science and ICT explained that this evaluation did not determine the final rankings based solely on a specific global benchmark.
The second-round evaluation was scored out of a total of 100 points, consisting of 40 points for the benchmark, 35 points for the expert evaluation, and 25 points for the user evaluation.
In the global AAII benchmark, Motif Technologies took first place with 11.9 points, while SKTelecom scored the highest in the NIA benchmark with 13.4 points. SKTelecom also ranked first in the AI expert user evaluation with 11.6 points, and LG Corp. received the highest score of 7.6 points in the general public evaluation.
The Ministry of Science and ICT explained that different teams demonstrated strengths in each evaluation category.
Regarding Motif Technologies’ elimination, the Ministry explained that it was not simply a matter of low practicality, but rather that the company received low scores overall in both expertise and usability, stating, “It is true that their performance fell short.”
The ministry explained that when considering not only the global public benchmarks—where Motif Technologies demonstrated strength—but also the NIA benchmarks reflecting the domestic environment, expert evaluations, and actual user evaluations, the final results were bound to differ.
The Ministry also explained that the user evaluations focused not merely on website design or UI/UX, but on the content and quality of the content generated by the AI models. A total of 49 AI experts and 185 members of the general public evaluated the services by actually using them.
“Why Did These Results Emerge?”
The focus of this controversy lies not so much on the elimination of Motif Technologies—which ranked first in the AAII—or the fact that LG Corp. AI Research topped the expert evaluation, but rather on how transparently the ministry can explain the criteria and scoring system used to synthesize the results of models that demonstrated strengths in different evaluation categories.
From the Ministry of Science and ICT’s perspective, a national representative AI cannot be selected based on a single global benchmark alone; performance in the domestic environment, technological development capabilities, actual usability, and industrial impact must all be considered together.
On the other hand, some in the industry, while agreeing with this rationale, argue that since the expert evaluation accounts for 35% of the overall score, the public should be able to verify the specific reasons why LG AI Research received 29.5 points and the exact point differential compared to other teams.
While the Ministry of Science and ICT disclosed the first-place finisher and the highest score for each category, as well as the average score of the four teams, it did not release the detailed scores for each team or the score differences between them. Critics point out that, given this is a national AI selection process funded by public money, it is necessary to disclose evaluation information at a level that allows for verification of why such results occurred—beyond simply identifying which team came in first.
MSEED, the largest shareholder of artificial intelligence (AI) diagnostics company Noul Co., Ltd.(376930), is facing repercussions from a repurchase agreement (RP) signed during its participation in a…
The retail and e-commerce sectors are stepping up efforts to attract customers ahead of the Chuseok holiday. Department stores and outlet malls are hosting character-themed and fashion pop-up shops as…
BLACKPINK’s Rosé holding an iPhone 18 Pro model. (Photo: Apple, Rosé’s Instagram)
Pre-orders for Apple’s next-generation premium smartphones, the “iPhone 18 Pro” and “iPhone 18 Pro Max,” begin in …