GPT 5.6 Sol is the best "vision" model OpenAI ever released
- ID
- 14942
- Status
- summarized
- Published
- 17 Aug 2026, 8:09 PM
- Fetched
- 19 Aug 2026, 6:08 AM
- Provider
- Hacker News
- Category
- dev-community
- Original URL
- https://blog.roboflow.com/openai-gpt-5-6/
- Source URL
- https://hnrss.org/best
Summary
- Score
- 7.0
- Created
- 19 Aug 2026, 6:09 AM
- Tags
- Audience
- developersai_ml_learnersai_agent_users
What happened
Roboflow benchmarked OpenAI's new GPT-5.6 lineup (Sol, Terra, Luna) on vision tasks and found Sol is a major leap for object detection, scoring 46.2 mAP@50 versus GPT-5.5's 13.8. The models perform best when prompted to return absolute XYXY pixel coordinates; using the wrong format (e.g., normalized YXYX like Gemini 3.5 Flash) drops performance by ~15 mAP points. Sol still occasionally hallucinates bounding boxes in random layouts unrelated to actual objects.
Why it matters
If you're building document-layout or object-detection pipelines with VLMs, GPT-5.6 Sol is now a practical option where GPT-5.5 was unusable—but you must prompt for absolute XYXY pixel coordinates or lose ~15 mAP points. Watch for hallucinated boxes in dense scenes; consider post-processing validation before trusting outputs in production.
Discussion angle
Coordinate format as a prompt-engineering lever: the same model loses 15 mAP points just from asking for the wrong coordinate scheme—how much of 'model quality' in production is actually prompt discipline?