Grok 4.7
- ID
- 26959
- Status
- summarized
- Published
- 21 Sep 2026, 11:50 PM
- Fetched
- 24 Sep 2026, 12:05 AM
- Provider
- Hacker News
- Category
- dev-community
- Original URL
- https://x.ai/news/grok-4-7
- Source URL
- https://hnrss.org/best
Summary
- Score
- 5.5
- Created
- 24 Sep 2026, 1:14 AM
- Tags
- Audience
- developersai_ml_learnersai_agent_users
What happened
xAI announced Grok 4.7, a new larger base model trained with a longer RL run weighted toward multi-hour tasks. It is priced at $2/M input and $6/M output (same as Grok 4.6), scores 46.3% on CursorBench 4.0 and 71.0% (high effort) on DeepSWE v1.1, and ships with a new safeguard stack that tops LatchBio biosafety at 62.4% and allows only 3.3% of risky prompts on HackerBench v0.3.
Why it matters
If you are evaluating coding-agent backends, Grok 4.7 offers frontier-class CursorBench and DeepSWE scores at roughly one-third the token cost of GPT-5.6 Sol Max ($4/$20) and far below Fable 5.1 Max ($10/$50) — worth a benchmark run on your own repo before committing. The safety numbers are vendor-reported and not independently verified, so treat them as marketing until confirmed.
Discussion angle
Compare Grok 4.7's $2/$6 pricing and CursorBench 4.0 score against what you actually pay for Claude/GPT in Cursor or your agent harness — is the benchmark gap large enough to justify switching, or do real-world coding results diverge from synthetic benchmarks?