AI Weekly Malaysia

Back to items Summaries

Livenerf: Has Opus 5.5 been nerfed yet?

ID
30158
Status
summarized
Published
30 Sep 2026, 6:36 AM
Fetched
30 Sep 2026, 12:35 PM
Provider
Hacker News
Category
dev-community
Original URL
https://github.com/ninjahawk/livenerf
Source URL
https://hnrss.org/best

Summary

Score
7.5
Created
30 Sep 2026, 12:35 PM
Tags
Audience
developersai_ml_learnersai_agent_usersvibe_coders

What happened

livenerf is an append-only, pre-registered benchmark built to test whether a frontier model quietly degrades after launch, and it started the clock on Claude Opus 5.5, released 2026-09-22. Day 1 ran 2026-09-24 22:10 UTC, roughly 2.5 days after launch, and it now samples once a day for 30 days: days 1-10 form the baseline, then two 10-day windows, so the first possible drift call lands around 2026-10-24 and the first Results row after day 20. It runs through headless Claude Code (claude -p) on a Claude Max subscription with no API key, using frozen prompts, a pinned CLI version, exact graders and raw logs, built on the UK AI Security Institute's Inspect framework with error bars per Anthropic's 'Adding Error Bars to Evals'; the repo has 366 stars and the Hacker News thread has 343 points and 147 comments.

Why it matters

If you ship anything on Claude models, this is the closest thing to a day-0 baseline anyone has published, and it says plainly that sampling parameters are gone and thinking can't be turned off, so reproducibility has to come from pinning the CLI version, freezing prompts and keeping raw logs. The concrete decision: pin your model and CLI version in a file the way this repo does, log raw outputs now, and treat any post-launch quality claim as unproven until there are thousands of samples with error bars - not vibes. Note the timeline: no drift verdict exists before roughly 2026-10-24, so anything claiming Opus 5.5 was 'nerfed' before then is speculation.

Discussion angle

Would a stripped-down version of this - pinned version file, a dozen frozen prompts, raw logs, error bars - be worth building for the one model your product actually depends on, and what would you do differently if the benchmark did show a 10% drop after day 20?

Top