A benchmark has emerged that tracks whether AI model performance declines after release. [Photo: ChatGPT]

Japanese media outlet Gigazine introduced an open-source benchmark called Livenerf on Oct. 1 to track over the long term whether an AI model’s performance actually drops after its release.

Livenerf was developed from the idea that even for an AI model with the same name, performance and output characteristics can change during service as inference settings or routing are altered. Past research on GPT-4 also found that some evaluation results, including math and coding, differed sharply between versions from different points in time.

More recently, a case of actual quality degradation was confirmed at Anthropic. Anthropic said an investigation into quality issues seen in Claude Code, the Claude Agent SDK and Claude Co-Work found that changes to inference intensity made on March 4 and separate software issues, among other factors, had an impact. It said the API and the underlying inference system itself were not affected.

Livenerf aims to confirm such changes with repeated measurement data rather than user perception. It was developed by Nathan Langley (네이선 랭글리), a student at the University of North Carolina at Greensboro, and was built on the U.K. AI Security Institute’s open-source evaluation framework, Inspect.

The current measurement target is Claude Opus 5.5, released on Sept. 22. After repeatedly testing 2,336 questions in a preliminary evaluation, it selected 78 questions whose results varied after excluding those that were always answered correctly or incorrectly. It uses the average of the first 10 days as a baseline and then compares changes across two subsequent 10-day periods.

The experiment, however, measures the subscription-based Opus 5.5 provided through Claude Code, so it does not represent the performance of the entire set of API models. The Livenerf developer also explains that it is not possible to judge performance degradation based on the current data alone, and that the core of the experiment is distinguishing actual change from perceived differences stemming from rising user expectations.

Keyword

#Livenerf #Gigazine #Anthropic #Claude Opus 5.5 #Inspect
Copyright © DigitalToday. All rights reserved. Unauthorized reproduction and redistribution are prohibited.