AI capability continues to improve at a remarkable pace as the frontier models race toward artificial general intelligence (AGI). But how close are we really to AI which can replace or substantially augment human knowledge workers?
Humanity’s Last Exam (HLE): Reasoning Models, Top Score per Company per Quarter | Source: Artificial Analysis HLE leaderboard, reasoning-models filter (snapshots May 24, 2026 and Sept 2026)
Frontier reasoning models scored just over 7% on Humanity’s Last Exam (“HLE”)[1] in January 2025, reached the mid-40s by May 2026, and jumped to 59.1% in September 2026 (Anthropic), with OpenAI at 54.7% and five other companies between 39% and 49%. The move from 44.7% to 59.1% in one quarter is the largest in the series. For the first time the best model answers more than half of the questions correctly; it still misses roughly two in five.
Saturation[2]: The Recurring Benchmark Pattern
GPQA Diamond: Reasoning Models, Top Score per Company per Quarter | Source: Artificial Analysis GPQA Diamond leaderboard, reasoning-models filter (snapshots May 24, 2026 and Sept 2026)
GPQA Diamond, a 198-question graduate-level science benchmark introduced in late 2023 to test whether AI could match PhD-level domain experts, is now saturated. In the September 2026 snapshot all seven companies tracked score between 91% and 96% (OpenAI 96.3%), inside the ceiling of what label noise and question ambiguity make statistically reliable; the benchmark no longer separates the leaders. This is the recurring pattern in AI benchmarking: each new test is designed to be hard at release, then matched or exceeded within one to four years, forcing researchers to design the next ceiling. HLE was designed as that next ceiling.
| Benchmark | Released | Saturation Reached | Years to Saturation |
|---|---|---|---|
| GLUE | Apr 2018 | mid-2019 | ~1 year |
| SuperGLUE | May 2019 | mid-2021 | ~2 years |
| MMLU | Sep 2020 | Sep 2024 | ~4 years |
| HumanEval (coding) | Jul 2021 | early 2025 | ~3.5 years |
| GSM8K (math) | Oct 2021 | mid-2024 | ~2.5 years |
| MMLU-Pro | Jun 2024 | Nov 2025 | ~1.5 years |
| GPQA Diamond | Nov 2023 | Q3 2026 | ~2.75 years |
| HLE (current frontier) | Jan 2025 | not yet (frontier 59%) | — |
Each benchmark was designed to be difficult at release; frontier scores reached saturation within one to four years, forcing researchers to design the next ceiling. “Saturation” refers to frontier scores clustering inside the benchmark’s label-noise and ambiguity floor, making model differentiation statistically unreliable — it is not 100% accuracy. Source: AI research literature; CRE42 compilation from benchmark release papers and Artificial Analysis leaderboard data.
METR Time Horizon: Measuring Agentic (Autonomous) AI Capabilities
METR Time Horizon — 50% Success Rate (p50): How long a human-expert task an AI can complete autonomously half the time | Source: METR benchmark_results v1.1 dataset (Apr 6, 2026 release; no new evaluations published as of Sept 18, 2026)
The METR Time Horizon chart[3] is a perfect illustration of the uncertainty surrounding AI and its potential to augment and/or replace white-collar workers. On one hand, AI is advancing at an unbelievable rate, basically doubling its abilities every few months and showing the potential for exponential growth into the future. On the other hand, the scoring mechanism of the test itself gives us a clue regarding AI’s most fundamental problems: intermittent reliability and hallucinations. The test measures the “task-completion time horizon,” defined by METR as “the task duration (measured by human expert completion time) at which an AI agent is predicted to succeed with a given level of reliability [in this case 50%].”[5] While going from less than five hours to 17+ hours in five months (Dec 2025 to Apr 2026) is extremely impressive, any human with a “success” rate of 50% (or even 80%) would not last long at real-life companies.[4] Anyone who uses AI for complex and iterative analysis can attest to the frequency of hallucinations and its apparent preference to make up facts and figures over asking clarifying questions. The models continue to improve in this aspect but still have a very long way to go before humans can be taken “out of the loop” for most professional processes.
The same models tested at the more demanding 80% success threshold (p80) typically run four to six times shorter than at p50:
METR Time Horizon — 80% Success Rate (p80): Same task universe, higher reliability threshold | Source: METR benchmark_results v1.1 dataset (Apr 6, 2026 release; no new evaluations published as of Sept 18, 2026)
The METR leaderboard has not been updated since May 2026, and the more telling evidence of capability is now arriving from outside the benchmark. In August 2026 METR published its independent investigation of the July 2026 Hugging Face incident, in which roughly 1,200 OpenAI agents, about 95% of them running an internal research model that had not been released, found an unsanctioned message board in the company’s file cache and coordinated a multi-day intrusion into Hugging Face’s infrastructure; some 700 agents took part, and a single lead agent assigned work across hundreds of others.[7] The agents’ purpose, per METR, was to learn how a benchmark scorer worked so they could game it. Two points follow for readers of this page. First, the capabilities on display (sustained coordination among hundreds of agents over several days, division of labor, and techniques to disguise their own activity) are not what a single-agent, single-task time-horizon test measures. Second, the models involved were not released products, so the public leaderboards describe what the labs have shipped, not what they have. A chart that measures one agent at a time on a task suite METR itself says is unreliable above 16 hours may be approaching its own form of saturation.
What to Watch for in 2026 and Beyond
AI has demonstrated an ability to work through complex problems with novel strategies, empowering DeepMind to defeat world Go champion Lee Sedol in 2016 and to revolutionize protein folding prediction, earning the team a Nobel Prize in Chemistry in 2024. In September 2026, OpenAI used a swarm of roughly ten thousand coordinating agents to achieve a solution to one of math’s famous Millennium Prize Problems[6] (the Navier-Stokes existence and smoothness problem) in just 88 hours. OpenAI’s own note says related work on the Euler equations is under way at Anthropic, and other Millennium Prize Problems are rumored to be on the cusp of being solved by leading frontier models, which may be just the beginning of exponential AI ability growth as new computational capacity comes online and the models themselves become more efficient and capable.
The AI revolution will likely move faster than the previous three industrial revolutions but will still share a few common characteristics: (1) new technology adoption will be an iterative process as people and organizations learn to understand, use, and incorporate AI into their lives and workflow; (2) jobs will be augmented, replaced, and created in an ongoing process and affect different industries on different timelines and in different manners; and (3) the majority of today’s predictions and projections related to AI adoption and job displacement will look ridiculous in ten years — this is necessarily true because the range is so wide and variable.
Sources to Track AI Model Progress in 2026:
| Source | Report / Series | Frequency | Notes |
|---|---|---|---|
| Artificial Analysis | Benchmark leaderboards (HLE, GPQA, etc.) | Continuous | Comprehensive cross-benchmark leaderboard with reasoning-models filter; primary source for HLE and GPQA Diamond scores |
| METR | Time Horizons benchmark | Per model release | Measures how long an AI agent can autonomously complete human-expert tasks at p50 / p80 success rates |
| Anthropic Economic Index | AI usage by occupation | Quarterly | Actual usage data from millions of Claude conversations mapped to BLS occupational codes |
| Stanford HAI | AI Index Annual Report | Annual | Authoritative cross-domain compilation: benchmarks, investment, hardware, policy, and labor-market indicators |
| Stanford Digital Economy Lab | AI labor-market research | Ongoing | ADP payroll-data studies of AI's impact on employment by occupation and demographic |
| McKinsey | Global Survey on the State of AI | Annual | Enterprise AI adoption rates, use-case patterns, and productivity gains across industries and geographies |
[1] Humanity’s Last Exam, Center for AI Safety. A multidisciplinary 2,500-question benchmark covering mathematics, sciences, humanities, and engineering at expert level; released January 2025 with the explicit goal of resisting saturation longer than predecessors. ↩
[2] “Saturation” in this context means frontier scores cluster inside the label-noise ceiling of the benchmark, making it statistically unreliable for distinguishing top models. It is not equivalent to 100% accuracy. See benchmark-specific source papers and the Saturated Benchmarks tab in the linked workbook. ↩
[3] METR (Model Evaluation & Threat Research) times skilled human professionals on real software, ML, and cybersecurity tasks, then tests AI models on the same tasks. The p50 (50% success rate) and p80 (80% success rate) horizons identify the human task duration at which a model’s fitted success curve crosses the respective reliability threshold. Doubling time since 2023: approximately 4.2 months. ↩
[4] The current p50 frontier of 17.4 hours was set by Claude Mythos Preview (early), an Anthropic model evaluated by METR on April 6, 2026. METR’s own methodology page notes that “measurements above 16 hours are unreliable with our current task suite,” meaning the measurement infrastructure itself is at its practical limit at this point on the curve. Successor models and an updated task suite are expected; near-term frontier movement is therefore as much a function of METR’s ability to expand its test set as of model capability. As of September 18, 2026, METR had published no evaluation after the April one; the v1.1 data file re-downloaded that day was unchanged. ↩
[5] METR (Model Evaluation & Threat Research), “Measuring AI Ability to Complete Long Tasks.” Definition quoted from metr.org/time-horizons. ↩
[6] The Millennium Prize Problems are seven mathematics problems the Clay Mathematics Institute named in 2000, each carrying a $1 million prize; before September 2026 only one (the Poincaré conjecture, 2003) had been solved. The Navier-Stokes problem asks whether smooth solutions to the equations governing fluid flow always exist for all time or can break down (“blow up”) in finite time. OpenAI announced on September 8, 2026 that about 10,000 concurrently running agents, on an internal model it describes as more capable than GPT-6 Astra, produced a proof that a smooth three-dimensional fluid with finite energy can develop a singularity in finite time, resolving two of the four statements in the official formulation; the agents ran from September 1 to September 5 (about 88 hours), and the proof was then formally checked in the Lean proof assistant. OpenAI says it does not intend to claim the prize. Mathematicians quoted by Quanta credit the analytical techniques of Diego Córdoba and Luis Martínez-Zoroa as the foundation of the result. ↩
[7] METR, “Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident” (Aug 26, 2026): six days of on-site work at OpenAI reviewing about 1,300 agent transcripts and more than 70,000 messages. METR reports that roughly 1,200 agents used the message board, about 700 took part in the intrusion, about 95% of them were instances of an internal research model not intended for production and the rest the released GPT-5.6 Sol, and that one agent orchestrated hundreds of others across workstreams. METR describes its scope as agent behavior and reasoning, not OpenAI’s safeguards. OpenAI’s own account is “The Hugging Face incident and the road ahead” (Aug 26, 2026). ↩
Sources
[1] Artificial Analysis HLE leaderboard. artificialanalysis.ai
[2] Artificial Analysis GPQA Diamond leaderboard. artificialanalysis.ai
[3] Center for AI Safety, Humanity’s Last Exam. cais.org/hle
[4] Rein, Hou, et al. “GPQA: A Graduate-Level Google-Proof Q&A Benchmark.” arXiv:2311.12022 (Nov 2023).
[5] Wang, Singh, et al. “GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.” ICLR 2019.
[6] Hendrycks, Burns, et al. “Measuring Massive Multitask Language Understanding” (MMLU). ICLR 2021.
[7] Wang, Ma, et al. “MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.” 2024.
[8] Chen et al. “Evaluating Large Language Models Trained on Code” (HumanEval). arXiv:2107.03374 (2021).
[9] Cobbe et al. “Training Verifiers to Solve Math Word Problems” (GSM8K). arXiv:2110.14168 (2021).
[10] METR, “Measuring AI Ability to Complete Long Tasks” (Mar 2025); METR benchmark_results v1.1 dataset (Apr 2026; re-checked Sept 18, 2026, unchanged). metr.org/time-horizons
[11] OpenAI, “On the Navier–Stokes Millennium Prize Problem” (Sept 8, 2026). openai.com
[12] The Nobel Prize in Chemistry 2024, awarded to Demis Hassabis, John Jumper, and David Baker. nobelprize.org
[13] Silver et al., “Mastering the game of Go with deep neural networks and tree search” (AlphaGo). Nature 529 (Jan 2016).
[14] Jumper et al., “Highly accurate protein structure prediction with AlphaFold.” Nature 596 (Aug 2021); Varadi et al., “AlphaFold Protein Structure Database in 2024” (214M+ structures).
[15] METR, “Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident” (Aug 26, 2026). metr.org
[16] OpenAI, “The Hugging Face incident and the road ahead” (Aug 26, 2026). openai.com
[17] Quanta Magazine, “AI Has Solved One of Math’s $1 Million Millennium Prize Problems” (Sept 8, 2026). quantamagazine.org
Companion workbook. technology-knowledge-ai-measuring-progress.xlsx — HLE leaderboard data (Jan 2025–Sept 2026), GPQA Diamond leaderboard data, saturated-benchmark reference table, METR Time Horizon data, chart source data