AI Weekly Malaysia

Summaries

Short AI and tech summaries with source links, signal scores, and why each update matters for builders, founders, and Malaysian tech workers.

Reset

Showing 1-2 of 2 results

DateProviderScoreSummary
30 Sep 2026, 6:36 AMHacker News7.5 Livenerf: Has Opus 5.5 been nerfed yet?

livenerf is an append-only, pre-registered benchmark built to test whether a frontier model quietly degrades after launch, and it started the clock on Claude Opus 5.5, released 2026-09-22. Day 1 ran 2026-09-24 22:10 UTC, roughly 2.5 days after launch, and it now samples once a day for 30 days: days 1-10 form the baseline, then two 10-day windows, so the first possible drift call lands around 2026-10-24 and the first Results row after day 20. It runs through headless Claude Code (claude -p) on a Claude Max subscription with no API key, using frozen prompts, a pinned CLI version, exact graders and raw logs, built on the UK AI Security Institute's Inspect framework with error bars per Anthropic's 'Adding Error Bars to Evals'; the repo has 366 stars and the Hacker News thread has 343 points and 147 comments.

Why: If you ship anything on Claude models, this is the closest thing to a day-0 baseline anyone has published, and it says plainly that sampling parameters are gone and thinking can't be turned off, so reproducibility has to come from pinning the CLI version, freezing prompts and keeping raw logs. The concrete decision: pin your model and CLI version in a file the way this repo does, log raw outputs now, and treat any post-launch quality claim as unproven until there are thousands of samples with error bars - not vibes. Note the timeline: no drift verdict exists before roughly 2026-10-24, so anything claiming Opus 5.5 was 'nerfed' before then is speculation.

28 Sep 2026, 8:00 AMOpenAI News3.0 Basis completes a tax workbook 2x faster with GPT-6 Astra

OpenAI published a customer case study saying Basis, which builds AI agents that automate accountants' manual work, completed a 50-tab tax workbook in half the time with GPT-6 Astra compared with GPT-5.6 Sol, and saw roughly a 20% improvement in its internal evaluation scores. Basis co-founder Mitch Troyanovsky is quoted saying Astra better understands user intent, makes better decisions at the start of a task, and can dial reasoning effort up or down mid-task while keeping its cache intact, which he says lowers cost and response time on long-running tasks. Every number is self-reported by the vendor and its customer: there is no independent benchmark, no pricing, no context-window or token figures, and no availability date.

Why: The only transferable detail here is cache-preserving adaptive reasoning on long agent runs, which is a cost lever if it is real - but this post gives you no price, no rate limits, and no way to verify the 2x claim, so do not plan a migration on it. Instead, treat the 50-tab workbook as a template for your own eval: pick your longest multi-step task, count the tabs-equivalent steps, and measure wall-clock time and token spend per run before believing any vendor speed claim.

Top