A South Korean research team has developed technology to reduce hallucinations in which AI confuses multiple sensory inputs or generates information that does not exist.
The Korea Advanced Institute of Science and Technology (KAIST) said on Thursday that a team led by professor Yong-man Ro (노용만) of the School of Electrical Engineering developed two technologies to improve multimodal large language models' understanding of sensors and reduce confusion between visual and audio information.
Multimodal AI processes not only text but also images, audio and various sensor data. Some existing models have struggled to distinguish the characteristics of each sensor or answered as if information obtained through one sense was also confirmed through another.
The team developed an optimisation method called Sensor Understanding AI (DNA) to help AI accurately understand the physical meaning carried by sensors such as thermal imaging, depth images and X-rays.
For example, existing AI sometimes misread bright areas in thermal images as light in ordinary photos rather than high temperatures. The team trained the AI by focusing on cases it often gets wrong, so it can distinguish physical information measured by sensors such as heat and distance. This enabled AI to recognise objects more accurately by using various sensor data even in environments where ordinary cameras struggle to identify objects, such as darkness or smoke.
The team also developed a sensory confusion prevention (MAD) technology to reduce hallucinations that occur as visual and audio information influence each other.
The team enabled AI to distinguish on its own, before answering a question, which sensory information to use as the basis for judgment between video and sound, and to focus on the most relevant information. KAIST said this effectively reduced sensory confusion and hallucinations without retraining the AI from scratch, in a training-free approach.
The team said DNA improves performance so AI can understand various sensors with a small amount of data, while MAD was designed to be applied immediately, like adding a program to existing AI. The team expects the technology can be used in industrial settings such as autonomous vehicles, rescue robots and thermal-imaging drones. It can also be applied to medical AI that analyses multiple medical images together.
Yong-man Ro said it is important for multimodal AI to accurately understand the characteristics of various sensors and avoid confusing multiple sensory inputs in order to be used in real-world environments. He said the work reduced sensory bias and hallucinations without large-scale retraining and laid the groundwork for multimodal AI that can be trusted in industrial settings.
Sang-yoon Jung (정상윤), a doctoral student in KAIST's School of Electrical Engineering, participated as first author on both studies. Young-joon Yoo (유영준) was listed as co-first author on the DNA study.
The MAD study was presented in June at the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). The DNA study was published in the journal IEEE Transactions on Image Processing.
The research was supported by projects including the Institute of Information & Communications Technology Planning & Evaluation (IITP)'s core source technology development programme for human-centered AI and an assignment from the Agency for Defense Development's AI specialized research center.