Research

Open-Source livenerf Tracks Claude Opus 5.5 Drift

Developer ninjahawk has launched livenerf, an open-source benchmark that tracks daily performance drift in Claude Opus 5.5 to help developers detect post-launch capability changes.

AlphaSignal4 days agoResearch
Image: AlphaSignal

Developer ninjahawk has released livenerf, an open-source benchmark designed to detect post-launch capability drift in frontier artificial intelligence models. The initial v0 release targets Anthropic's Claude Opus 5.5, running through headless Claude Code on a Max subscription without requiring an API key. By recording a fixed benchmark over time, the tool aims to provide concrete evidence of post-launch changes in model accuracy and output length, addressing a common frustration among developers who suspect their hosted models are getting dumber.

To track these changes, livenerf runs a frozen 78-question panel daily for 30 days. The benchmark is built on Inspect, the UK AI Security Institute's evaluation framework, and incorporates statistical methods from Anthropic's error-bar analysis, including paired comparisons and uncertainty intervals. The system is highly controlled: it pins the Claude Code command-line interface version, disables tools, Model Context Protocol servers, and memory files, and uses exact-match grading to avoid relying on an LLM judge that could also drift.

The 78-question panel was selected from a pool of 2,336 candidate questions sourced from GPQA Diamond, MMLU-Pro, competition mathematics, and AIME 2025-26. During calibration, Claude Opus 5.5 recorded approximately 93% accuracy on the first sample, with 97% of the questions yielding consistent results across all runs. The resulting daily test is highly efficient, utilizing only about 3.6% of a user's weekly subscription plan while remaining sensitive enough to detect accuracy swings of roughly 7.5 points within a 10-day window.

For AI practitioners, livenerf shifts the conversation around model degradation from subjective complaints to verifiable time-series data. Because the decision rules, panel selection, and metrics are pre-registered in a public git commit, developers can trust the integrity of the drift measurements. This allows engineering teams to make data-driven decisions about whether to update their prompts, adjust their workflows, or switch models when a provider silently updates a hosted service.

This is our own summary of reporting by AlphaSignal

More in Research