Why Dwarkesh is Wrong about Computer Use + How OpenAI shipped its Jev competitor in 1 Week
- ID
- 30542
- Status
- summarized
- Published
- 01 Oct 2026, 6:23 AM
- Fetched
- 01 Oct 2026, 6:35 AM
- Provider
- Latent Space
- Category
- developer-ai
- Original URL
- https://www.latent.space/p/devday-2026
- Source URL
- https://www.latent.space/feed
Summary
- Score
- 6.5
- Created
- 01 Oct 2026, 6:36 AM
- Tags
- Audience
- developersai_ml_learnersai_agent_usersvibe_coders
What happened
Latent Space's DevDay 2026 episode (its first DevDay pod) features OpenAI's Computer Use (CUA) team and API platform leads, pushing back on Dwarkesh Patel's June 27, 2026 argument that computer use progress has been slow because the domain is 'clearly verifiable.' The counter-frame offered is that 'grindability is just as important as verifiability,' with guest Ari Weinstein (Sky cofounder, now working on CUA) describing computer use as '180 degrees different' from months ago as agents learn to debug and recover from failures. Concrete shipped artifact referenced: Computer History in ChatGPT, released Aug 14, 2026, which lets ChatGPT learn from everything you do on your computer, with a timeline view for reviewing that history.
Why it matters
Two decisions here. First, the episode's stated architecture claim is that combining screenshots with accessibility data, the DOM, Playwright, and generated code is what changed the speed of computer-use agents - if you're building or evaluating an agent that drives a browser, that's a direct input into how you wire it up, versus screenshot-only loops. Second, Computer History (Aug 14, 2026) makes reviewable screen-activity capture a shipped consumer default in ChatGPT, so if you ship anything that records user screen or workflow data, users will now compare your privacy controls against a timeline view they can inspect. No Malaysia or SEA angle appears in the text.
Discussion angle
Is 'grindability' - whether an agent can practice a task repeatedly with usable feedback - a better predictor of progress than verifiability? Test it against your own agent work: which tasks did you abandon not because you couldn't check the answer, but because you couldn't get enough repetitions to learn from?