TLDR:
- OpenRouter logged 7.3T agentic tokens by Aug. 10, more than 5X the human-driven total of 1.4T tokens.
- Frontier firms generated 8.3X more output tokens per user than typical enterprises, up from 2.6X in January.
- Codex produced 64% of combined Codex and ChatGPT enterprise output tokens by June as workflow use expanded.
- More than 85% of agentic token usage came from cached prompts, helping lower repeated input costs and latency.
Enterpriseartificial intelligence is moving from conversation toward execution as automated systems consume far more model capacity than ordinary human users. OpenRouter recorded about 7.3 trillion agentic tokens on a seven-day average by August 10, around 14 times the level seen in early February.
Human-driven usage reached about 1.4 trillion tokens over the same period, rising 2.8 times from early February. That left AI agents consuming more than five times as many tokens as humans across traffic routed through the platform.
OpenRouter separates agentic activity using signals including tool calls, conversation turns, and timing patterns. The figures cover OpenRouter traffic rather than the entire generative AI market, but they show how quickly automated workloads are scaling.
Unlike a chatbot answering one prompt, AI agents can inspect information, call tools, test outputs, revise steps, and repeat actions before completing work.
Enterprise AI Shifts From Chat to Automated Workflows
OpenAI’s Enterprise Signals report points to the same change inside companies, especially among the heaviest users. Frontier firms, defined as the top 10% of enterprise AI users, generated 8.3 times more output tokens per active user than typical firms.
That gap stood at 2.6 times in January, showing a widening difference between ordinary enterprise adoption and the most intensive users. By June, Codex produced 64% of combined Codex and ChatGPT output tokens among enterprise customers.
Weekly active Codex users had increased 108 times in legal since February, compared with 41 times in sales and recruiting. The mix shows enterprise use expanding beyond chat into coding, document creation, research, and other multi-step workflows.
OpenAI also reported that 21% of active users at frontier firms used Plugins weekly, compared with 9% at typical firms. That difference reinforces the shift toward systems that can use tools, create files, and complete tasks instead of only answering questions.
Codex, Plugins and Caching Drive the Shift Beyond Chat
Andreessen Horowitz, citing OpenRouter data, further acknowledged that more than 85% of agentic token usage now comes from cached prompts. Basically, prompt caching lets systems reuse previously processed context instead of recomputing identical input during repeated calls.
OpenAI says caching can reduce latency and input costs, making repeated-context workloads cheaper to operate. However, those workloads still require infrastructure able to store and rapidly reuse large contexts alongside the GPUs handling inference.
That increases the importance of memory capacity and bandwidth as enterprises delegate more complex assignments to automated systems. The shift also creates a test for traditional workflow software.
A16z argued that rising agent adoption could pressure products built around fixed automation steps as users adopt systems that reason through tasks. Nevertheless, web-traffic changes alone do not prove displacement, while Zapier, Make, and n8n are also adding AI features.
Overall, AI agents are generating more token demand as enterprise software is performing longer chains of work on behalf of users. As that pattern expands, enterprise AI is becoming less defined by chat volume and more by the amount of work delegated to software.



