| Mobile Web

Open-source benchmark Livenerf tracks whether AI model performance changes after release

An open-source benchmark called Livenerf has emerged to track over time whether an AI model’s performance declines after release. Built to detect changes caused by shifts in inference settings or routing, it aims to verify differences through repeated measurements rather than user perception. The current target is Claude Opus 5.5, with 78 questions selected for ongoing monitoring. The test does not represent overall API performance, and its developer says current data are insufficient to judge degradation.