"AI safety should be approached through a separate evaluation framework, not as an extension of existing cyber security. Because generative AI produces probabilistic answers, it is necessary to assess risks that differ from traditional security, such as hallucinations, hate, bias and harmfulness."
Kim Min-woo (김민우), lead of Selectstar's AI Safety Team, said, "Hallucinations, factuality, and hateful or biased answers fall under safety, not security."
A leading way to verify such risks is AI red-teaming. It tests in advance whether a model produces unintended or harmful answers through various attacks and identifies vulnerabilities. Kim described it as "a series of processes or techniques that deceive AI to induce unintended answers."
AI safety evaluation is becoming more important as generative AI evolves beyond simple question-and-answer into agent forms that use external tools. Global AI companies such as OpenAI and Anthropic are strengthening in-house red-teaming and risk assessments before releasing models. In South Korea, the Korea Internet & Security Agency (KISA) published an 'AI Security Red-Teaming Guide' last month.
Selectstar is developing safety evaluation data, automated attack models and technology to determine whether AI answers are safe. It is also researching blue-teaming technology, including guardrails that block answers when problems are found.
AI red-teaming does not stop at simply inputting prohibited questions. Automated attack models rewrite them into prompts that can bypass restrictions and check whether the AI ultimately produces harmful answers.
Kim said, "Simply attacking in many ways is low-level red-teaming," adding, "There are areas AI can find and areas of attacks people are likely to actually attempt, and both are important." Selectstar currently operates both a model that automates human attack methods and an AI-specific attack model.
What counts as risk differs by industry. In general-purpose AI, information related to chemical, biological, radiological and nuclear (CBRN) may be a key risk, but in enterprise AI, leaks of corporate secrets or personal information may matter more. Domain experts who understand the risk characteristics of the industry also participate in the evaluation process.
Kim also saw limits in applying overseas AI safety benchmarks in South Korea as they are. He said, "If you simply translate or literally translate overseas benchmarks, the culture or evaluation purpose may not match our situation," adding, "We need to keep creating safety benchmarks that fit domestic conditions."
The spread of AI agents is making safety evaluation more complex. As AI connects and uses internal databases (DB), retrieval-augmented generation (RAG), and external APIs and plugins, problems can arise in external components or permission design even if the model itself is safe.
Kim said, "Because the final answer is ultimately given by AI, even if risk arises in components, it is seen as a risk of the AI service," adding, "With prompt-based red-teaming alone, it is difficult to know where the problem occurred, so there are cases where you have to check DBs, plugins and permissions one by one."
He said these characteristics create opportunities for South Korean AI companies in the safety field. While global big tech leads the competition in model performance, safety tailored to each country's language, culture, laws and systems requires local expertise.
Kim said, "In AI safety, there is global safety, but there is also clearly localized safety that fits a country," adding, "Because risk standards and domestic laws also keep changing, global big tech cannot keep up with all of that."
He added, "I think AI safety is a blue ocean in that we can develop unique technology that reflects this in the domestic environment and industry."