(From left) KAIST professor Junmo Kim (김준모) and doctoral student Minchan Kwon (권민찬). [Photo: KAIST]

A safety verification technology has been developed that can uncover hidden vulnerabilities in generative AI about 7 times more diversely than existing techniques.

The Korea Advanced Institute of Science and Technology (KAIST) said on Wednesday that a research team led by Junmo Kim (김준모), a professor in the School of Electrical Engineering, has developed a new red-teaming framework, Stable-GFlowNet, to verify the safety of large language models (LLMs).

AI red-teaming is the process of checking vulnerabilities by creating attack prompts that induce harmful or dangerous responses before applying a model to a service. Beyond attack success rates, the ability to identify a wide range of different types of vulnerabilities is important.

Existing reinforcement learning-based methods had a limitation in that they can suffer from "mode collapse," repeatedly generating only a few high-success attacks, making it difficult to discover diverse vulnerabilities. Generative Flow Networks (GFlowNet), designed to produce diverse outcomes, were also presented as an alternative, but had issues including a complex and unstable training process and high rewards being assigned to meaningless sentences.

To address this, the team introduced a "contrastive trajectory balance" technique that stabilises learning by comparing the relative quality of outputs. It also applied "noise gradient pruning," which excludes data with small reward differences from training, and a "min-K fluency stabiliser" that filters out unnatural sentences.

In experiments, Stable-GFlowNet identified 134 unique attack types, about 7 times more than the 17 found by existing techniques. It maintained a 92 percent attack success rate. A defense model trained with Stable-GFlowNet was also found to effectively defend against multiple types of attacks in cross-attack tests that applied other attack methods.

The team explained that the technology could be used not only for AI safety verification but also in areas that need to find diverse and strong candidates, such as generating drug candidates and exploring materials.

The research was selected as a spotlight paper, representing the top 2.2 percent, at the AI international conference International Conference on Machine Learning (ICML) 2026. Minchan Kwon (권민찬), a doctoral student in KAIST's School of Electrical Engineering, participated as the first author.

Kim said it is significant in that it can reliably identify diverse AI vulnerabilities even with limited data and in realistic environments with substantial noise. He said he expects it to become a foundational technology for finding and defending against a broad range of risks before applying generative AI to real services.

Keyword

#KAIST #Stable-GFlowNet #GFlowNet #LLM #ICML 2026
Copyright © DigitalToday. All rights reserved. Unauthorized reproduction and redistribution are prohibited.