TL;DR
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
The Livenerf project has collected six of 30 planned daily benchmark runs for Claude Opus 5.5, which Anthropic released on September 22, 2026. It has not yet reported whether the model’s performance has changed: the baseline period is still underway, and the first comparison is expected around October 24.
Livenerf has collected six of 30 planned daily runs in a benchmark tracking Claude Opus 5.5 after its September 22 release, but it has not yet found or reported a post-launch performance change. The project’s first comparison is due after its baseline and follow-up periods, with an initial results row expected around October 24, 2026.
The GitHub project describes itself as a small, append-only benchmark for checking whether a model’s measured performance changes after launch. Its progress update, dated September 29, says all six collected days completed the full set of 90 samples, using the same harness hash and pinned Claude Code CLI version, 2.1.280. No collection days had been missed. One run, on day five, used an override to the project’s budget guard; the repository records that as a deviation.
Livenerf plans to run the benchmark daily for 30 days. Days one through 10 establish the baseline; the next two 10-day windows provide comparisons. The first results row is scheduled after day 20, and the project says its first possible call about a change is around October 24. Until those periods are complete, the available status is progress on data collection, not evidence that Opus 5.5 has been weakened or has stayed unchanged.
The benchmark uses a panel of 78 questions selected from 2,336 questions across GPQA Diamond, MMLU-Pro, competition math and AIME 2025–26. The project reports that the model answered about 93% of the screened questions correctly on its first try; questions that were sometimes answered correctly formed the panel. It says fresh samples raised the panel’s pass rate from 54.7% to 62.0%, and that this adjustment is used in its power calculation.
How the Benchmark Could Detect Drift
The project addresses a recurring dispute about whether model performance changes after release. Users can report that a system seems weaker, but without a consistent launch-period measurement, those impressions can be difficult to separate from differences in prompts, sampling, tools or task difficulty. Livenerf’s stated purpose is to create a reference point and retain the raw logs, so later runs can be compared with launch-week performance under a fixed procedure.
That comparison could matter to people relying on Opus 5.5 for technical or academic tasks. A measured drop might prompt closer scrutiny of the model’s behavior or serving conditions. A stable score would also be useful evidence, within the benchmark’s limits. The project says its primary metric is the paired per-item score difference against the baseline, with clustered standard errors; it also tracks output tokens as a secondary signal, since reduced output could indicate a change in how much the model produces before accuracy shifts.
The benchmark is designed to detect a specific scale of change, not every possible alteration. Its validation says a daily run can detect an accuracy change of about 7.5 percentage points per 10-day window, at about 3.6% of the weekly plan. Smaller shifts, changes on tasks outside the panel, or behavioral differences that do not affect the measured outcomes may not appear in the result.
Top picks for "livenerf opus nerf"
As an affiliate, we earn on qualifying purchases.
A Launch Baseline for Opus 5.5
Anthropic released Claude Opus 5.5 on September 22, 2026, according to the project’s report. Livenerf says its first run took place on September 24 at 22:10 UTC, about two and a half days after launch. The project presents this timing as an opportunity to begin a prospective record rather than infer a change from scattered comparisons made later.
Its method uses Inspect, an open-source evaluation framework from the UK AI Security Institute. The project says prompts are frozen, the command-line tool is pinned, graders are fixed and raw logs are retained. It also says the statistical approach follows Anthropic’s published guidance, “Adding Error Bars to Evals.” These choices aim to limit sources of measurement variation; they do not make a hosted model fully deterministic. Livenerf notes that sampling controls are unavailable in its current Claude Max and headless Claude Code setup, and that thinking cannot be turned off.
The project’s validation describes limits on identifying model changes. In its tests, replacing Opus 5.5 with Opus 5 was not distinguishable at the 99% level in a validation-sized sample: the reported accuracy difference was −3.8 ± 6.3 points, alongside a 23% reduction in output tokens. Livenerf says a 10-day window has about 2.5 times as many samples, but has not shown that this would be enough to detect a swap of that size.
““The series is running.””
— Livenerf project description
What Six Runs Cannot Show
No Opus 5.5 drift result is available yet. Six collection days are still part of the 10-day baseline period, so they cannot establish whether later performance differs from launch-week performance. The project’s schedule leaves the first comparison for a later window.
The benchmark also has known question-quality limits. A report-only audit identified eight answer keys that appear wrong and 30 ambiguous questions among the selected and later excluded items. Livenerf says it has retained the items and pre-registered a sensitivity analysis that will rerun the result without them. It also says a safety classifier sometimes responds using Opus 5 or refuses some biology and math questions; affected samples are rejected and counted, and questions touched by that behavior are excluded.
Any eventual result will be limited to this panel, setup and measurement period. The project’s own validation says its test may miss a same-family model swap of a certain size. The data also cannot, by itself, establish why a measured change occurred. Quantization, routing, effort settings or other serving changes are possibilities raised by the project, not findings established by the current runs.
The October Comparison Window
Livenerf plans to continue one full run each day through the 30-day series. After completing the 10-day baseline and the next 10-day window, it expects to publish the first results row around October 24. That table is intended to compare the paired scores with the launch-week baseline and report standard errors, median output tokens, a control comparison and a decision.
Readers should look for the completed sample count, any additional deviations, and the sensitivity analysis alongside the headline score. The project says it will report improvements as well as regressions. Until the first comparison is published, whether Opus 5.5 has changed remains unanswered by this benchmark.
Key Questions
Has Livenerf found that Opus 5.5 was nerfed?
No finding has been reported. As of September 29, the project had collected six days of a 30-day series and was still building its baseline.
When could the first result appear?
The project expects its first results row after day 20, around October 24, 2026, if its planned collection schedule proceeds.
What does Livenerf measure?
It compares performance on a panel of 78 questions with a launch-week baseline. It also tracks output tokens as a secondary signal.
Can the benchmark detect every model or serving change?
No. Livenerf reports that its validation could not distinguish Opus 5 from Opus 5.5 at the 99% level in a validation-sized sample. The project has not shown that its longer comparison window would detect a change of that size.
Source: hn
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
