Is Claude getting dumber after launch? livenerf runs and ranks the same task over and over again to find out.
- livenerf is an open-source benchmark that started tracking Claude Opus 5.5 on its launch day (September 22, 2026) to test whether the model gets worse over time, a recurring complaint with no prior clean baseline to verify against.
- The project uses frozen prompts, a pinned CLI version, pre-registered statistical methods, and a control arm running an older model to separate genuine degradation from platform noise.
- Output token count is tracked as a leading indicator: in validation tests, reducing effort level cut tokens by 26-62% before accuracy dropped meaningfully, suggesting token volume may be a more sensitive early warning than benchmark scores.