Brood War Bench
- ID
- 26482
- Status
- summarized
- Published
- 19 Sep 2026, 10:44 PM
- Fetched
- 21 Sep 2026, 8:53 PM
- Provider
- Hacker News
- Category
- dev-community
- Original URL
- https://bw.swerdlow.dev/report
- Source URL
- https://hnrss.org/best
Summary
- Score
- 7.0
- Created
- 21 Sep 2026, 9:59 PM
- Tags
- Audience
- ai_agent_usersai_ml_learnersdevelopers
What happened
Ben Swerdlow built a Brood War environment playable only through AI agents and benchmarked 19 model configurations. No model played beyond beginner level; Codex Astra/xhigh won 100% of games at $10.54/game, while Grok 4.6 and Claude Haiku won zero. A key finding is that models treating the real-time game as turn-based got destroyed while thinking, and Codex's subagents for economy, production, and army control failed to coordinate—sending units in one at a time instead of massing for timed attacks.
Why it matters
If you build multi-agent systems, this benchmark exposes a concrete coordination failure: separate subagents managing different tasks don't communicate well enough to align on timing and strategy, a problem you should test for in your own agent architectures. The thinking-cost tradeoff is also real—lower-effort settings sometimes outperformed because high-effort models paused too long in real-time contexts.
Discussion angle
The subagent coordination failure is the most transferable insight: when you split tasks across agents (economy, production, combat), how do you ensure they share enough state to align on timing? Codex's one-unit-at-a-time attacks are a metaphor for what happens when your agent pipeline lacks a shared planning layer.