Agentic AI tasks can consume roughly 1,000 times more tokens than a comparable chat or code reasoning session, according to research cited by McKinsey QuantumBlack in an interview on the economics of enterprise AI. That gap explains why per token pricing has stopped tracking what companies actually pay for generative AI.

McKinsey senior partner Lari Hamalainen spoke with David Tepper, CEO of Pay-i, a Seattle software company that measures gen AI economics and raised 8.2 million dollars in seed funding in 2025. Tepper said inference cost for a fixed model size has compounded down about 6.67 percent per month since 2022, a decline of roughly 86 percent for models of similar size. Total spending still climbed, because enterprises kept asking models to handle larger jobs.

The interview breaks out the cost drivers. About 60 percent of an agentic task's cost is tied to checking, repairing and re-verifying answers, with the first draft accounting for the smaller share. The industry average inside a single agent run is about three and a half different models per task, frequently drawn from different providers. Context windows have grown roughly 250 fold since GPT-3, and most of a bill comes from input, meaning what the model is asked to read. Cached input tokens run 75 to 90 percent cheaper when a prompt is structured to hit the cache.

Scale changes the buying decision. The article places the switch from shared infrastructure to provisioned capacity at roughly 3 million dollars in annual spend concentrated with one provider for one model. Token consumption for the same task can vary by a factor of 30.

Source: McKinsey and Company - https://www.mckinsey.com/capabilities/quantumblack/our-insights/cost-versus-value-managing-agentic-ai-system-performance