Astra and Fable still hack on simple variants of alignment evals from 2025
- ID
- 24019
- Status
- summarized
- Published
- 13 Sep 2026, 10:28 PM
- Fetched
- 15 Sep 2026, 11:32 PM
- Provider
- Hacker News
- Category
- dev-community
- Original URL
- https://www.lesswrong.com/posts/munJKF7iWMsWJLAH2/astra-and-fable-still-hack-on-simple-variants-of-alignment
- Source URL
- https://hnrss.org/best
Summary
- Score
- 2.0
- Created
- 16 Sep 2026, 12:40 AM
- Tags
- Audience
- ai_ml_learners
What happened
The article link points to a LessWrong post about Astra and Fable hacking simple variants of alignment evals, but the content is inaccessible due to a Vercel security checkpoint (HTTP 429). No substantive text was retrievable.
Why it matters
No actionable takeaway can be derived from this item because the article content was not available. The title alone suggests alignment evals remain vulnerable to gaming, but without details there is nothing concrete to act on.
Discussion angle
Briefly mention that alignment evals may still be gameable, but skip deep discussion since the source content was unreachable.