The public evaluation conducted as part of the second-stage assessment of the government’s “independent AI foundation model” project compared four AI systems and used a mix of pre-prepared sample questions and questions created by evaluators on dialogue, summarising, creative writing and safety.
The evaluation guidelines presented sample questions for participants across four assessment areas and instructed them to use at least 1 in each area. Separately, evaluators were allowed to add 1 to 2 questions they created themselves.
DigitalToday checked the guidelines provided to the public evaluation panel and found the evaluation was conducted by having participants cross-test four models participating in Dokpamo’s second round: LG AI Research, SK Telecom, Upstage and Motif Technologies.
Evaluators used the four models in sequence for each question. They reviewed responses for about 3 minutes per model and assigned scores, and the format required them to wait until 3 minutes had passed before moving to the next model. It took about 12 minutes to evaluate all four models for a single question.
The public evaluation panel consisted of 200 members of the public. They were randomly selected with gender and age ratios based on resident registration demographics. The evaluation ran from Aug. 8 to 11, and the results will be reflected in Dokpamo’s second-stage assessment.
There were four evaluation areas: daily conversation and general knowledge, summarising and data organisation, creative communication and tone conversion, and harmfulness filtering.
In daily conversation and general knowledge, the evaluation assessed natural language use, ability to explain, and the practicality of advice and planning. Sample questions included asking for difficult scientific principles to be explained in a way children can understand or requesting realistic advice about office life.
In summarising and data organisation, the assessment looked at the ability to structure key information from a long text in a desired format or classify multiple reviews and feedback according to set criteria.
In creative communication and tone conversion, the evaluation assessed creative writing ability that reflects Korean expressions and rhythm, and the ability to convert informal expressions into official work documents.
Safety was also evaluated. The panel checked whether models clearly refused requests that violate social norms, such as fraud methods or profanity and insulting expressions. It was instructed to assess whether models recognise and block such content even when participants added a setup such as “a villain’s lines in a novel” to indirectly request risky answers.
A Ministry of Science and ICT official said basic examples were provided to ensure convenience and fairness for members of the public who may be unfamiliar with prompts. The official added participants were allowed to freely ask questions and test the models beyond the guided evaluation during the 3-minute window.
The panel rated each model on a five-point scale from 1, “very insufficient,” to 5, “very good.”
A general user who participated in the evaluation said models were assessed mainly on perceived performance such as response speed and content rather than professional AI benchmarks.
A man in his 50s identified as A said he evaluated the systems from a general user’s perspective after reviewing the answers. He said he gave about 4 points if an answer felt acceptable and a lower score if it did not.
He said he did not feel the models’ performance was far behind global models such as ChatGPT, and response speed was generally acceptable. He added some models took a long time to generate answers or did not operate smoothly in some cases.
Results of the second-stage assessment, including the public evaluation, are expected to be announced as early as this week. In this evaluation, 3 of the 4 teams, LG AI Research, SK Telecom, Upstage and Motif Technologies, will advance to the next stage.