Motif Technologies, which was eliminated in the second-stage review of the government’s independent AI foundation model project, was 3.2 points behind third-ranked LG AI Institute in overall score. Motif received the highest score in the benchmark evaluation among the four teams, but lagged in expert and user evaluations and finished fourth overall.
The Ministry of Science and ICT on Wednesday disclosed detailed evaluation results for four teams that took part in the second-stage review: Upstage, SK Telecom, LG AI Institute and Motif.
In overall scores, SKT ranked first with 70.6 points, followed by Upstage with 69.9, LG AI Institute with 69.0 and Motif with 65.8. The gap between LG AI Institute in third and Motif was 3.2 points.
The review comprised 100 points in total: 40 for the benchmark evaluation, 35 for the expert evaluation and 25 for the user evaluation. The user evaluation was split into 15 points from AI expert users and 10 points from the general public.
By category, Motif scored 24.6 points in the benchmark evaluation to rank first, ahead of Upstage with 22.7, SKT with 22.2 and LG AI Institute with 20.6.
The benchmark evaluation comprised 25 points for the Artificial Analysis Intelligence Index (AAII) and 15 points for a benchmark by the National Information Society Agency (NIA). In AAII, Motif scored the highest at 11.9, followed by Upstage with 9.4, SKT with 8.8 and LG AI Institute with 7.8.
In the NIA benchmark, SKT ranked first with 13.4 points, followed by Upstage with 13.3, LG AI Institute with 12.8 and Motif with 12.7.
Rankings reversed starting with the expert evaluation. LG AI Institute scored the highest with 29.5 points, followed by SKT with 29.3, Upstage with 29.1 and Motif with 27.1.
The expert evaluation was conducted by 10 external experts: 3 from industry, 5 from academia and 2 from research institutions. They assessed development strategy and technology, development results and plans, and expected impact and contribution plans. The ministry said each team’s score was calculated by excluding the highest and lowest scores from each evaluator and averaging the remaining scores.
The gap widened further in the user evaluation. SKT scored 19.1 points, LG AI Institute 18.9 and Upstage 18.1, while Motif had the lowest score at 14.1. The user-evaluation score gap between LG AI Institute and Motif was 4.8 points.
In a breakdown of the user evaluation, SKT ranked first with 11.6 points in the AI expert user group assessment, in which 49 expert users including AI startup founders took part. LG AI Institute scored 11.3, Upstage 10.8 and Motif 8.4.
In the general user assessment involving 185 members of the public, LG AI Institute ranked first with 7.6 points, followed by SKT with 7.5, Upstage with 7.3 and Motif with 5.7.
As a result, Motif led LG AI Institute by 4.0 points in the benchmark evaluation but trailed by 2.4 points in the expert evaluation and by 4.8 points in the user evaluation. Its overall score was 3.2 points lower.
The disclosure of detailed scores came shortly after Motif formally objected on Wednesday to the second-stage review result. In a statement, Motif said its model, "Motif 3", received the highest AAII score but was eliminated, and it requested disclosure of detailed evaluation scores and a re-review.
Motif in particular said the benchmark score gap was reduced in the process of converting scores into final weighted points. It demanded disclosure of the detailed criteria for the expert evaluation and the methodology for the user evaluation. It also raised the need for blind evaluations, saying brand awareness or preconceptions could influence results if general users evaluated the models while knowing the developers.
Motif said its objection was not intended to demand that it be reselected. It said it would not participate in the third stage of the project regardless of the outcome of the objection.
A ministry official said, "Dokpamo is not intended to be a simple ranking contest or to select eliminated teams, but aims to drive the growth of our AI companies and the qualitative growth and expansion of the domestic AI ecosystem." The official added, "We have decided the scope of disclosure by considering both potential harm to companies from disclosing evaluation results and the fairness and transparency of the evaluation process and outcomes."