
Just ahead of the second evaluation for "Dokpa-MO" (Domestic AI Foundation Model), which selects Korea's national representative AI, an unexpected result has emerged. Motif Technology's 'Motif3' surpassed models from LG AI Research, SK Telecom, and Upstage on global comprehensive AI performance indicators that include agent task capabilities.
According to Artificial Analysis's "AI Model Comprehensive Performance Index (AAII)" released on the 13th, Motif3 scored 47 points—the highest among the four teams participating in the second DOKPA-MO evaluation.
Upstage's 'Solar Open2' followed with 37 points, SK Telecom's 'A.X K2' recorded 35 points, and LG AI Research's 'K-Exaone' scored 31 points. Motif3 led the second-place model by 10 points and the fourth-place model by 16 points.
Motif3 also achieved a high score compared to all models in the AAII overall rankings. According to Artificial Analysis, Motif3's 47-point score significantly exceeds the median of 16 points for comparable models. The latest AAII comprehensively measures multiple capabilities beyond simple knowledge, including coding, scientific problem-solving, and agent-based task execution.

However, this result does not automatically translate into the final ranking of the second DOKPA-MO evaluation. In the second DOKPA-MO evaluation, benchmark assessments account for 40 points out of a total of 100: AAII contributes 25 points, while the National Intelligence Service (NIA) benchmark accounts for 15 points. The remaining 60 points consist of expert evaluation (35 points) and user evaluation (25 points). Thus, the score gap observed in AAII could be reversed during the final evaluation process.
AAII itself has limitations. It is an index that combines multiple benchmarks covering knowledge, reasoning, coding, etc., to represent a model's overall performance as a single score. While useful for comparing models at a glance, results can vary depending on which benchmarks are included. When the index is restructured and evaluation items are replaced, scores of existing models may shift significantly. Notably, AAII lacks a benchmark that separately evaluates Korean language performance. This makes it difficult to judge the competitiveness of domestic AI models—including their Korean understanding and reasoning abilities—based solely on AAII scores.
In the AI industry, concerns about "benchmarking," where models are optimized for specific benchmarks to artificially inflate evaluation scores beyond actual practical performance, have been consistently raised. The government's decision to incorporate not only a single benchmark but also NIA benchmarks along with expert and user evaluations in the DOKPA-MO assessment aims to address these limitations.
With the public evaluation concluding on the 12th, the government is set to finalize the second DOKPA-MO evaluation process and announce results soon. Only three of the four participating teams will advance to the next stage. The key question remains whether Motif3, which took an early lead in global comprehensive performance indicators, can maintain its advantage through expert and user evaluations, or if the other three models can close the gap in the remaining assessments.