AI agents raised their accuracy on OSWorld, a benchmark of real computer tasks across operating systems, from roughly 12 percent to 66.3 percent in a single year, landing within 6 percentage points of human performance, according to the 2026 AI Index from Stanford HAI. The same chapter records agents still failing about one in three attempts on structured benchmarks.

Model quality has converged at the top of the field. As of March 2026, six developers sit within a narrow band on the Arena leaderboard: Anthropic at 1,503 Elo, xAI at 1,495, Google at 1,494, OpenAI at 1,481, Alibaba at 1,449 and DeepSeek at 1,424. The report says that clustering pushes competition toward cost, reliability and domain specific performance. The gap between the top closed model and the top open model widened to 3.3 percent as of March 2026, up from 0.5 percent in August 2024, and six of the top ten leaderboard models are closed.

Frontier models gained 30 percentage points in one year on Humanity's Last Exam, an evaluation built to resist AI systems and favor human experts. In professional domains including tax, mortgage processing, corporate finance and legal reasoning, top models scored between 60 and 90 percent, with the leading 15 models separated by as little as 3 percentage points on each benchmark.

Benchmark reliability drew scrutiny in the same chapter. A review of widely used evaluations found invalid question rates ranging from 2 percent on MMLU Math to 42 percent on GSM8K, and separate research suggests leaderboard standing may partly reflect adaptation to the platform.

Robotics results lag the software gains. Robots succeed at 12 percent of real household tasks while reaching 89.4 percent success on RLBench simulations.

Source: Stanford HAI AI Index - https://hai.stanford.edu/ai-index/2026-ai-index-report/technical-performance