Dokpamo. [Photo: ChatGPT]

As competition among global large language models moves beyond simple question-and-answer to “agentic AI” that carries out real work using external tools and data, attention is also focusing on how the government’s “independent AI foundation model” project evaluates such capabilities.

According to relevant ministries and other officials on Aug. 14, the second-phase Dokpamo evaluation consists of the same elements as the first: benchmark evaluation, expert evaluation and user evaluation. Unlike the first phase, where 49 professional users including AI company CEOs took part in the user evaluation, the second phase splits user evaluation into a professional-user assessment and a general-user assessment involving 200 members of the public.

The Ministry of Science and ICT says it differentiates between core model performance and agentic capability by varying whether external tools can be used depending on the purpose of the evaluation.

A ministry official said the professional-user assessment restricted tool calls and similar functions to focus on the LLM’s own performance. The official said tool calls can be made freely in the expert evaluation. The official added that there is sufficient structure in place to objectively verify the agentic AI performance of models participating in Dokpamo.

Agentic AI refers to a model that does more than answer questions. It breaks down tasks needed to achieve a goal, then selects and calls external tools to carry out multi-step work. It uses functions such as search, database queries and running external applications as needed.

Benchmarks are also used to evaluate agent capabilities. In the first-phase evaluation, the National Information Society Agency used a domestic AI model performance evaluation set it prepared, along with global common and individual benchmarks. The global common benchmarks included 13 types that evaluate agents, mathematics, and knowledge and reasoning.

The ministry is not disclosing in detail the list of global benchmarks used in the actual evaluation, citing considerations such as fairness.

In the global market, various benchmarks are used to measure agent capabilities in more detail. Global AI model evaluator Artificial Analysis classifies “tau³-Banking,” “Terminal-Bench v2.1” and “GDPval-AA v2” as indicators measuring, respectively, agentic tool use, agentic coding and terminal use, and real task performance capability.

All 4 models participating in the second-phase Dokpamo evaluation have publicly available results for these three benchmarks.

In GDPval-AA, which looks at performance on real job tasks, Motif Technologies’ “Motif 3” scored 39 percent, Upstage’s “Solar Open2” 31 percent, SK Telecom’s “A.X K2” 30 percent, and LG AI Research’s “K-EXAONE 2.0” 24 percent.

In tau³-Banking, which evaluates the ability to use external tools, Motif 3 scored 35 percent, Solar Open2 22 percent, A.X K2 16 percent, and K-EXAONE 2.0 12 percent.

In Terminal-Bench v2.1, which measures agentic coding and terminal use capability, Motif 3 scored 75 percent, Solar Open2 44 percent, K-EXAONE 2.0 40 percent, and A.X K2 39 percent.

These benchmarks each measure some agent capabilities of a model in specific environments. There are limits to interpreting individual benchmark scores as equivalent to overall agentic performance or practical usability in real industrial settings.

Experts see a process to verify usability in real work environments as important beyond benchmark scores. Choi Jae-sik (최재식), a distinguished chair professor at the Korea Advanced Institute of Science and Technology, said agent AI is the key in global AI competition for now, making it important to raise the performance of independent models to a level that can be used in the field. He said usability must be verified through demonstrations in specific areas such as coding and work automation, rather than stopping at benchmark scores.

Keyword

#Dokpamo #Ministry of Science and ICT #Artificial Analysis #KAIST #Terminal-Bench
Copyright © DigitalToday. All rights reserved. Unauthorized reproduction and redistribution are prohibited.