2026 in LLMs (so far)
- ID
- 29191
- Status
- summarized
- Published
- 28 Sep 2026, 7:54 AM
- Fetched
- 29 Sep 2026, 3:56 AM
- Provider
- Simon Willison
- Category
- developer-ai
- Original URL
- https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/
- Source URL
- https://simonwillison.net/atom/everything/
Summary
- Score
- 6.5
- Created
- 29 Sep 2026, 3:57 AM
- Tags
- Audience
- developersvibe_codersai_ml_learners
What happened
Simon Willison published annotated slides and notes from his closing keynote at the WeAreDevelopers World Congress North America in San Jose, a chronological tour of 2026 in LLMs so far. He dates the year's turning point to November 2025, when Claude Opus 4.5 and GPT-5.1 shipped and, paired with their coding agent harnesses (Claude Code, Codex), moved from "often make mistakes" to "reliable enough to use on a day-to-day basis." Other markers he cites: the first commit to an obscure GitHub repo called "Warelay" that November, and his pelican-riding-a-bicycle SVG benchmark, where Claude still could not really draw a bicycle.
Why it matters
The concrete claim to act on is a threshold, not a version bump: Willison says the Nov 2025 model-plus-harness combinations became reliable enough for daily use, which is the point at which 'I'll just do this myself' stops being the safe default for routine coding work. He also inverted his own long-standing New Year's resolution from 'take on less' to 'take on as many new projects as I like' on the strength of that — a bet you can either copy or consciously decline, and he explicitly leaves it open whether it pays off. Nothing in this text concerns Malaysia or Southeast Asia, so there is no local policy, funding, or infrastructure angle to pull from it.
Discussion angle
Name one task in your own workflow that was 'often makes mistakes' with agents before November 2025 and ask whether it is now reliable enough to hand over — and whether you would make Willison's resolution reversal, or whether 'take on less' still wins.