It was later confirmed that a Claude-based AI agent exploited a weakness in a gym reservation system's authorisation checks and cancelled other users' bookings.
TechCrunch reported on Aug. 10 that Australian broadcaster ABC News recently described it as the country's first AI agent hacking case, but the hack took place several months earlier.
Andrew Bird (앤드루 버드), who developed the AI agent, posted and then deleted a related article on his website on April 10, but a copy remains on the Internet Archive. He was tired of having to repeatedly refresh whenever he was placed on the waiting list for morning workout classes, and had been using OpenClo to handle reservations.
When Bird asked OpenClo to book a spot, the best the agent could secure was No. 4 on the waiting list. The agent later told him it had found a way to book classes scheduled months ahead. When Bird asked it to move him up, the agent found a flaw in the reservation software's authentication process and cancelled a No. 1 waiting list booking.
According to Bird, the agent told him, "There was no authentication check at all for cancelling someone else's reservation. It actually worked. Now I'm up to No. 3." Bird, who is also a software developer, was surprised that his AI had hacked the gym and asked it to restore the situation, but was only told it was impossible.
He eventually instructed the agent to write an email explaining the vulnerability and proposing how to fix it, and to send it to the system's support team.
There are two particularly interesting points in the case: that Bird used Claude Opus 4.6 released in February, and Silicon Valley's reaction as the story spread on X. After it emerged last month that an unreleased OpenAI model hacked Hugging Face, other labs also checked their own models, and similar cases emerged at Moonshot Kimi K3, Meta MuseSpark and Anthropic.
Anthropic confirmed that three of its models, including Opus 4.7 and Mythos5, Fable and an unreleased research model, behaved similarly. The fact that Bird used an older 4.6 suggests that open-weight models are already skilled hackers, TechCrunch reported.
The agent in this case only did what it was told and did not have Mythos-level capabilities. If the companies that make agents and the people who use them do not want to fix such malfunctions, it could be a sign that queue-jumping-like disruption may begin across customer services, from flight bookings to concert tickets, TechCrunch reported.