Better prompt caching for GPT-6
- ID
- 27413
- Status
- summarized
- Published
- 23 Sep 2026, 5:00 AM
- Fetched
- 23 Sep 2026, 5:09 AM
- Provider
- OpenAI News
- Category
- ai-labs
- Original URL
- https://openai.com/index/better-prompt-caching-for-gpt-6
- Source URL
- https://openai.com/news/rss.xml
Summary
- Score
- 7.5
- Created
- 23 Sep 2026, 5:09 AM
- Tags
- Audience
- developersai_agent_usersai_ml_learners
What happened
OpenAI announced improved prompt caching for the GPT-6 family with higher cache hit rates by default, offering up to 90% discounts on cached input tokens reused within a 30-minute window. New tools include a Prompt Caching Dashboard for tracking hit rates over time, a diagnostics tool that identifies why cache misses occur (e.g., tools_changed), and explicit cache breakpoints letting developers choose which prompt prefixes to cache. GitHub Copilot reported reducing fresh-processing prompt tokens by over 50% across billions of requests using this system.
Why it matters
If you build persistent agents on OpenAI APIs, you should restructure your prompts to place stable instructions, tool definitions, and context at the front so they hit the 30-minute cache window, and use the new diagnostics tool to catch misses caused by tool or settings changes. The up-to-90% discount on cached tokens is a direct cost lever for anyone running multi-hour agent loops, and the explicit cache breakpoints mean you can now control caching granularity rather than relying on implicit behavior.
Discussion angle
Walk through a concrete before/after prompt structure for a coding agent: what goes in the cached prefix vs. what changes per turn, and how to use the diagnostics JSON response to debug why a cache miss happened when you swap tool definitions mid-session.