🎙️ How I AI: Grok Bot + Grok 4.6—what’s great (and what’s still hype) & Lessons from spending $20,000 on Devin in one month
- ID
- 17307
- Status
- summarized
- Published
- 24 Aug 2026, 11:02 PM
- Fetched
- 25 Aug 2026, 1:01 AM
- Provider
- Lenny's Newsletter
- Category
- product-startup
- Original URL
- https://www.lennysnewsletter.com/p/how-i-ai-grok-bot-grok-46whats-great
- Source URL
- https://www.lennysnewsletter.com/feed
Summary
- Score
- 7.5
- Created
- 25 Aug 2026, 1:01 AM
- Tags
- Audience
- developersvibe_codersai_agent_usersai_ml_learners
What happened
Claire runs blind evaluations of Grok 4.6 against GPT-5.6 Sol, Claude Sonnet 5, and Opus 5, finding Grok 4.6 ties GPT-5.6 Sol at the top of her personal index. She also tests Grok Bot, whose multi-account connectors (handling her 4 email addresses and 7 Slack workspaces) solve a problem Codex and Claude still ignore, and Cursor Origin, an agent-native GitHub alternative that looks attractive but lacks the features teams relying on GitHub Actions and code owners need to migrate.
Why it matters
If you're choosing a frontier model for agent workflows, Claire's blind evals suggest Grok 4.6 and GPT-5.6 Sol are now ahead of Sonnet 5 and Opus 5 for general tasks, though Sonnet 5 still wins for conversational agent interactions. If you manage multiple Slack/Gmail accounts for agent access, Grok Bot's multi-account connectors are a concrete reason to try it over Codex or Claude. Don't migrate to Cursor Origin yet if your team depends on GitHub Actions, code owners, or existing automations.
Discussion angle
Claire's blind eval methodology—grading outputs herself with 70% weight on her own judgment—is something any team could replicate cheaply to stop relying on public leaderboards for model selection decisions.