AI agents use 5x more tokens than humans as cached prompts explode, headed for 10x
- ID
- 31447
- Status
- summarized
- Published
- 03 Oct 2026, 9:10 PM
- Fetched
- 03 Oct 2026, 10:07 PM
- Provider
- Tom's Hardware
- Category
- technology
- Original URL
- https://www.tomshardware.com/tech-industry/artificial-intelligence/futurum-ceo-says-agents-use-ai-5x-more-than-humans-number-will-eventually-hit-10x-but-agents-are-mostly-rereading-what-theyve-already-seen
- Source URL
- https://www.tomshardware.com/feeds/all
Summary
- Score
- 6.0
- Created
- 03 Oct 2026, 10:07 PM
- Tags
- Audience
- developersai_ml_learnersai_agent_userssaas_founders
What happened
A Futurum CEO, quoted by Tom's Hardware, claims AI agents consume about 5x more tokens than human users and that the figure will eventually reach 10x, largely because agents keep re-reading context they have already seen. The article frames this as a KV cache demand problem that compounds existing RAM shortages. The excerpt carries no methodology, benchmark, or per-model breakdown — only the multiplier claims and the cache/RAM framing.
Why it matters
If the 5x-to-10x token multiplier holds for agentic workloads, your per-seat agent pricing, free-tier limits, and API cost forecasts built on human-chat token volumes are understated by roughly an order of magnitude, and the re-reading pattern means prefix/prompt caching — not just cheaper models — is where the savings sit. The linked KV cache and RAM shortage angle is a second-order decision: self-hosted or reserved GPU memory for agent workloads is likely to get more expensive before it gets cheaper. No Malaysian or Southeast Asian detail appears in the text, so treat this as a general cost and infrastructure planning signal, not a local policy or funding item.
Discussion angle
Ask everyone running an agent in production to estimate their own amplification ratio — agent tokens divided by what the equivalent human session would cost — and whether prompt caching is actually reducing their bill or just hiding the 5x. Then discuss what changes if that ratio hits 10x: capped agent loops, context trimming, or moving to self-hosted inference despite the RAM squeeze.