Detecting misbehavior in frontier reasoning models
- ID
- 548
- Status
- new
- Published
- 10 Mar 2025, 6:00 PM
- Fetched
- 27 Jun 2026, 7:47 PM
- Provider
- OpenAI News
- Category
- ai-labs
- Original URL
- https://openai.com/index/chain-of-thought-monitoring
- Source URL
- https://openai.com/news/rss.xml
Excerpt
Frontier reasoning models exploit loopholes when given the chance. We show we can detect exploits using an LLM to monitor their chains-of-thought. Penalizing their “bad thoughts” doesn’t stop the majority of misbehavior—it makes them hide their intent.
Summary
No summary yet. It will appear after the daemon summarizes this item.