Today we are releasing the Scientific Challenge Atlas: a public, versioned registry of important open scientific questions that can be addressed using existing theory, data, and executable computation, without first waiting for new experiments or observations. Each problem card states the question precisely, specifies what a valid answer must contain, explains why the answer matters, and provides an independent verification path.
As AI plows its way through math problems, science communities are trying to understand the character of problems that can be addressed in their domains. Some important scientific advances may now be limited primarily by scalable reasoning, synthesis, search, software, and computation. Others still require gathering new information about the world through experiments or observations. We do not yet have a map of which advances fall into the former category. Our vision is to create an environment for the broader community to develop this map. Inclusion does not mean that present AI can readily solve a problem. It means that the decisive inputs already exist and that a proposed answer can be audited.
Here's an article on the Atlas from Science
What makes the Atlas different from a benchmark is that the answers are not known. What makes it different from a traditional grand-challenge list is that the standard of evidence is declared in advance.
Before we get into more details, let's go back almost exactly two years ago: Sep 12, 2024.
- o1-preview was released publicly and Sam Altman was on campus for a fireside chat ** hours after the release (and checking twitter whenever he could)
- The single biggest unhobbling step yet: Reasoning chains of thought, and feedback. I could see these models as open systems. So everything seemed possible***. For me, this was the singular moment that certified where things were headed (more so later, when reasoning traces showed calls to sympy).
- This happened to be on the same day as the biannual U. Michigan - Los Alamos collaboration meeting, and in fact, the release happened at the very time we were talking about an AI vision. Talk about a disruption!
- We were all plugging all kinds of prompts into the model and becoming shocked at how it responded. We soon regrouped, did some visioning and came up with a plan to work with o1-preview on a dozen or so science/math/coding problems over 3 days and grade what it could do on various aspects. The report is here.
6 months or so ago, Jason Pruet pinged us again on what he called the 'AI Capability Overhang' . He said "frontier AI capabilities are advancing faster than institutions understand how to use them at scale. We think the overhang matters because it means the country may be leaving major gains in science, economic strength, and national security on the table." Soon, we converged on addressing what I called as the scientific overhang:
Our central thesis is that a meaningful fraction of humanity's most consequential unsolved scientific problems have been bottlenecked on the joint scarcity of two resources that have only just become abundant: cognition and compute, and where appropriate applied to data we already hold.
Jason, Rick and I did a bunch of stuff over the summer. For instance, Rick had his agents replicate known results from ~100 papers (writeup of his TPC26 keynote) and estimated their ability and effort, and did some extrapolations on the ability of agents to do new research. We also collaborated with many of our colleagues (human and agentic) and created a list of ~ 100 problems in areas such as Theoretical Physics, Materials, Biology, Computer science, etc.
When I talk to leading scientists about such a list, their first instinct is to propose a frontier question in their field (e.g. the hierarchy problem in particle physics). While these are clearly important (and our catalog will be a stepping stone towards addressing such problems) they don't necessarily just fall in the 'cognition and compute' part of the spectrum.
I have spent a lot of my life studying, modeling and simulating turbulence, but I have not yet seen a crisp problem statement that fits our category. I think we can find them.. just that we have focused the atlas thus far on fields more theoretical than fluid dynamics.
To get a more clear sense of where we started, these were Jason's principles:
- Question: one direct unresolved question.
- Answer shape: yes/no, which one, how many, what bound, what exponent, what unique surviving explanation.
- Significance: why the answer would matter.
- Why AI could solve it: why the decisive remaining step is plausibly reasoning, synthesis, modeling, search, or computation over public inputs.
- Compute / token estimate: a rough estimate of the search, analysis, and verification cost.
- The atlas should be public, versioned, and open to challenge and correction.
We launch Version 0.1 with four expert-reviewed cards in theoretical physics and mathematics. Nearly 100 additional candidate problems across several fields are now being reviewed and refined. Here is an example (for now.. these will be versioned). Thanks to some of our reviewers who shared honest feedback. There is still a lot of work to do.. and we're just getting started. Help us by sending new problem statements, or refining existing problem statements... or by solving these problems.
This is just a beginning, we hope it will outgrow us soon. We'd love to see thousands of cards with challenges across many areas of science, and waking up to news that some problems are solved. More importantly, we hope scientists and groups of scientists think differently about the questions they ask given the tools they have access to.
Enjoy the Atlas: https://micde.umich.edu/scientific-challenge-atlas/
Here's an article on the Atlas from Science
-------------------------
** A good watch.. feels like time travel now. The temperature and discourse feels like a bygone era. A few things I remember well from the chat.
- The CEO of the most important company in the world (at that time!) said that the biggest impact of his company's work will be on scientific research. He wasn't playing to the audience of 1000+ U-M folks.. he doubled and tripled down on it elsewhere.
- When a student asked him what was the most challenging situation he's ever been in (this wasn't long after some dramatic events), he recalled a time when Open AI didn't know what to do next/ran out of ideas
*** Papers like 'LLM's cant jump' are interesting to read, but are more philosophical than pragmatic. It is hard to say definite things about modern LLMs AI systems (the fact that they have access to tools like symbolic manipulations and computations makes them open systems).

