AI & Compute

AI Agents Burn 5x the Tokens Humans Do, and 85% Is Cache Rereads

OpenRouter data shows agents consumed 7.3T tokens vs 1.4T for humans, with 85%+ coming from cached prompts — pressure that is straining HBM supply and worsening RAM shortages.

By
Grace Kim
Filed
Channel
AI & Compute
Read
3 min read

AI agents consumed 7.3 trillion tokens on OpenRouter in August, against 1.4 trillion for humans — a 5x gap that Futurum Group CEO Daniel Newman expects to widen to 10x and beyond. The figure comes from an Andreessen Horowitz (a16z) chart of OpenRouter data that Newman posted on X, capturing a crossover that happened six months earlier: February marked the first month agent usage surpassed human usage on the platform.

The composition of that traffic matters more than the headline. More than 85% of agent tokens come from cached prompts, a16z wrote, citing OpenRouter. Agents are mostly rereading what they have already seen rather than processing fresh input. A call center consultancy that tested DeepSeek on rented Nvidia H200 GPUs found the same pattern in its own logs: in September usage on Claude Code, 96% of all input was rereading old conversation.

OpenRouter, which describes itself as a leading AI model gateway and routing platform, sorts each API key into one of three categories — agentic, mixed, or human — using what Peter Walker, the company's head of insights, calls a "7-signal weighted composite score that includes inputs such as tool call rate, turn count, gap timing, and others." Walker's chart tracks 7-day average token usage split by type. Since the February crossover, agents have increased their token consumption 14x while human usage rose 2.8x. The mixed category, possibly covering traffic that is part agent and part human, grew 4.7x over the same period — and depending on how that traffic splits, the agents' lead over humans could shift in either direction.

There are caveats. The data covers one platform only. Token volume is not spending; cached tokens cost far less than prompts processed from scratch. The trend also shows dips in April and July rather than a clean exponential curve. But agent adoption is broadening beyond OpenRouter: McKinsey's 2026 State of AI survey found 40% of respondents from large organizations are scaling AI agents, up from 27% a year earlier.

The hardware implications are concrete even if the billing math is soft. Cached tokens may be cheap to reuse, but models must keep stored context resident in memory in what's called the KV cache — and, according to reporting cited by a16z, "the KV cache is outgrowing GPU HBM capacity." a16z, which is an OpenRouter investor, ties the surge in cached tokens directly to rising demand for high-bandwidth memory. Cached tokens also account for nearly all of the relative growth in token usage, a16z wrote, which means the marginal load lands disproportionately on memory rather than compute.

That demand lands on a supply base already under strain. Memory makers are prioritizing HBM for AI data centers over conventional DRAM, and Micron expects RAM and storage shortages to worsen in 2027 and 2028, with customers paying more than they do this year. The consultancy's log data reinforces the point: the token count overstates the bill, but the hardware cost of holding context in memory is real and recurring.

Newman put the trajectory bluntly: "AI is currently used by AI 5x more than it is used by humans. That number will accelerate to 10x and then higher and higher." He has since suggested 10x could become "10X, 20X, 30X." If agent token consumption climbs anywhere near that path while HBM stays allocated to data centers first, PC buyers will find themselves bidding for memory against an ever-larger fleet of agents — a competitive dynamic that could keep DRAM pricing elevated through the back half of the decade.

Source: Tom's Hardware

Share this article:

More from Grace Kim

Grace Kim

Show full bio

Market editor covering industry trends and analytics at Chip Dispatch.

174 articles

Related articles

« Previous article