2 min read

AI explains discoveries but fails to predict them

A CUSP benchmark across 4,760 scientific events found top models could explain plausible mechanisms, but were near chance at forecasting real breakthroughs.

Image: iXBT

A new CUSP benchmark suggests today’s top models can talk convincingly about science without reliably predicting where science will actually go next.

Researchers used Cutoff-conditioned Unseen Scientific Progress to test 6 leading systems across 4,760 verifiable scientific events in 8 disciplines that occurred after each model’s training cutoff. The lineup included GPT-5.4, Claude S4.5, and DeepSeek R1. Models were evaluated on four tasks: whether they could predict if a discovery would happen by a deadline, identify the mechanism that would lead to the result, devise a strategy for solving the scientific problem, and forecast when the achievement would arrive.

The central finding was a clear split between understanding scientific mechanisms and making real forecasts. On mechanism-selection tasks, models performed well above chance: GPT-5.4 reached 79.2% accuracy and Claude S4.5 scored 69.9%. But when asked whether a specific achievement would actually happen in the real world, results were close to coin-flip territory. GPT-5.4 scored 49.1%, while Claude S4.5 reached 52.6%.

The models also struggled with timing. They tended to predict that scientific advances would become public later than they actually did. Date-forecasting errors ranged from 4.9 months for LLaMA 3.3 to 15.7 months for DeepSeek R1, while some previous-generation models missed by more than 2 years.

Expert comparison and limits of web access

Giving models more information did not solve the problem. When researchers added access to time-limited web search, performance improved in cases where missing data had caused the error. But even with all publicly available information, the gap remained between knowing facts and forecasting outcomes.

Recommended reading

Google readies Frozen v2 chip for Gemini efficiency leap

Human experts still did better. In judging the feasibility of scientific events, experts achieved 73% accuracy, versus 52.5% for GPT-4o.

The study’s authors say current models show strong retrospective scientific competence: they can analyze existing knowledge and generate scientifically plausible ideas. But predicting future discoveries appears to be a separate challenge altogether, with obvious implications for research planning, funding decisions, and broader scientific decision-making.

Ava Chen

AI Editor

Ava covers the rapidly evolving world of artificial intelligence, from foundational models and research labs to the real-world economics of intelligence. With a background in computational linguistics, she cuts through the hype to find out what actually works. She firmly believes that benchmarks are just marketing until reproduced in the wild.

via iXBT

// Keep reading