A July 2026 report from QuantumBlack, AI by McKinsey, published in the McKinsey Quarterly by the New York headquartered firm, finds that 60 percent of total agentic AI spend goes to the iterative cycles of checking, correcting and improving that agents run before delivering a usable output. The remaining 40 percent covers other operational costs, including the first inference call.

Budgets are already stretched. The McKinsey Enterprise AI FinOps Survey, conducted in May 2026 with 75 qualified respondents across five major industries, found that 93 percent of those organizations had exceeded their AI budgets. In the firm 2026 State of AI survey, which gathered 1,719 responses between May and June, one in five participants said operating costs had constrained their use of the technology.

Model prices fell sharply over the same period. Inference on GPT-3.5 level capability cost 20 dollars per million tokens in early 2024 and dropped to 0.07 dollars by the end of that year, according to Stanford AI Index figures cited in the report. Enterprise spending on large language models tripled across the twelve months ending in late 2025, based on Menlo Ventures data also cited by McKinsey.

The report identifies two structural causes: consumption pricing that ties the bill to answer length, and the routing of simple tasks through frontier models priced for complex work. It argues that the measure worth tracking is whether the output an agent produces is worth more than the full cost of producing it, counting refinement cycles, human supervision time and downstream error correction.

Source: McKinsey QuantumBlack - https://www.mckinsey.com/capabilities/quantumblack/our-insights/is-that-ai-agent-worth-it-agentic-economics-and-the-modern-operating-model