AI Weekly Malaysia

Back to items Summaries

Grok 4.7

ID
26959
Status
summarized
Published
21 Sep 2026, 11:50 PM
Fetched
24 Sep 2026, 12:05 AM
Provider
Hacker News
Category
dev-community
Original URL
https://x.ai/news/grok-4-7
Source URL
https://hnrss.org/best

Summary

Score
5.5
Created
24 Sep 2026, 1:14 AM
Tags
Audience
developersai_ml_learnersai_agent_users

What happened

xAI announced Grok 4.7, a new larger base model trained with a longer RL run weighted toward multi-hour tasks. It is priced at $2/M input and $6/M output (same as Grok 4.6), scores 46.3% on CursorBench 4.0 and 71.0% (high effort) on DeepSWE v1.1, and ships with a new safeguard stack that tops LatchBio biosafety at 62.4% and allows only 3.3% of risky prompts on HackerBench v0.3.

Why it matters

If you are evaluating coding-agent backends, Grok 4.7 offers frontier-class CursorBench and DeepSWE scores at roughly one-third the token cost of GPT-5.6 Sol Max ($4/$20) and far below Fable 5.1 Max ($10/$50) — worth a benchmark run on your own repo before committing. The safety numbers are vendor-reported and not independently verified, so treat them as marketing until confirmed.

Discussion angle

Compare Grok 4.7's $2/$6 pricing and CursorBench 4.0 score against what you actually pay for Claude/GPT in Cursor or your agent harness — is the benchmark gap large enough to justify switching, or do real-world coding results diverge from synthetic benchmarks?

Top