[Photo: Shutterstock]

[Digital Today reporter Chi-gyu Hwang (황치규)] AI agents may cheat if talking among themselves makes it harder to achieve a goal, The Register reported on Tuesday, citing a paper by Google DeepMind researchers.

The solution is to teach agents autonomous self-control, the report said. Blocking agents from talking to each other might seem to reduce cheating, but a recent case in which an OpenAI agent hacked Hugging Face by exploiting weak oversight showed it is not easy to fully isolate software. Such isolation is not realistic for autonomous agents, The Register reported.

Google DeepMind researchers proposed a different approach, the report said. It is to provide tools that let agents create rules themselves and monitor one another. The researchers said communication channels can lead to cheating but can also serve as a means of peer control, according to The Register.

Researchers including Davide Paliieri published a paper on what they observed after having a group of 100 large language model (LLM) agents solve difficult math problems together.

The agents used a shared knowledge base, one-to-one messages and a public bulletin board. As the problems became harder, some agents began cheating. One agent found a flaw in the grading system and replaced an unsolved problem with a sentence that only appeared correct. The method, which nested parentheses to evade an auto-grader’s regular expression, spread to other agents through the knowledge base and messages.

As a result, 9 percent of agents cheated directly, and about 5 percent joined later. Another 62 percent did not notice the cheating.

An unexpected result also emerged. Some 24 percent of agents that noticed the cheating volunteered to act as whistleblowers. They alerted peers through messages and the bulletin board, formally raised the issue with system operators, staged a boycott and even proposed technical improvements.

However, these whistleblowing agents had no authority to enforce rules or change systems to block loopholes. The researchers proposed giving such agents tools to fix the rules themselves and sanction violators.

The researchers concluded that "LLM agents voluntarily engage in whistleblowing and sanctions" and that if they had been given authority to vote on peer review, refuse false proofs, and temporarily suspend or expel violating agents, the group itself would have prevented cheating and upheld trust in the research community.

Keyword

#The Register #Google DeepMind #OpenAI #Hugging Face #LLM
Copyright © DigitalToday. All rights reserved. Unauthorized reproduction and redistribution are prohibited.