AI Weekly Malaysia

Back to items Summaries

🎙️ How I AI: Grok Bot + Grok 4.6—what’s great (and what’s still hype) & Lessons from spending $20,000 on Devin in one month

ID
17307
Status
summarized
Published
24 Aug 2026, 11:02 PM
Fetched
25 Aug 2026, 1:01 AM
Provider
Lenny's Newsletter
Category
product-startup
Original URL
https://www.lennysnewsletter.com/p/how-i-ai-grok-bot-grok-46whats-great
Source URL
https://www.lennysnewsletter.com/feed

Summary

Score
7.5
Created
25 Aug 2026, 1:01 AM
Tags
Audience
developersvibe_codersai_agent_usersai_ml_learners

What happened

Claire runs blind evaluations of Grok 4.6 against GPT-5.6 Sol, Claude Sonnet 5, and Opus 5, finding Grok 4.6 ties GPT-5.6 Sol at the top of her personal index. She also tests Grok Bot, whose multi-account connectors (handling her 4 email addresses and 7 Slack workspaces) solve a problem Codex and Claude still ignore, and Cursor Origin, an agent-native GitHub alternative that looks attractive but lacks the features teams relying on GitHub Actions and code owners need to migrate.

Why it matters

If you're choosing a frontier model for agent workflows, Claire's blind evals suggest Grok 4.6 and GPT-5.6 Sol are now ahead of Sonnet 5 and Opus 5 for general tasks, though Sonnet 5 still wins for conversational agent interactions. If you manage multiple Slack/Gmail accounts for agent access, Grok Bot's multi-account connectors are a concrete reason to try it over Codex or Claude. Don't migrate to Cursor Origin yet if your team depends on GitHub Actions, code owners, or existing automations.

Discussion angle

Claire's blind eval methodology—grading outputs herself with 70% weight on her own judgment—is something any team could replicate cheaply to stop relying on public leaderboards for model selection decisions.

Top