DeepMind CEO Demis: frontier models need a 30-day checkup from a dedicated body before release — or they don't get to market
- AGI is probably just a few years out. He compares its scale to electricity and fire — roughly 10x the impact of the Industrial Revolution at 10x the speed
- Core proposal: the US builds a "Frontier AI Standards Body" first, using dynamic benchmarks to define frontier models and frontier labs
- Process: voluntary submission up to 30 days before release at first; once the protocol proves solid, it can become a hard gate for entering the US market
- Only frontier systems are covered; open and closed models, and labs of any nationality, are treated the same. Most startups and academic work fall below the threshold and are exempt. Coordinated slowdowns remain an option if needed
What he's saying, and why the urgency
Demis Hassabis, CEO of Google DeepMind, believes AGI — a system with the full cognitive range of a human brain — is probably just a few years away. Decades from now, he argues, people will realize we were standing at the foot of the singularity, at the dawn of a new era for humanity.
He separates AGI from ordinary technological breakthroughs: it's not even comparable to the internet or the smartphone — it's closer to discovering electricity or learning to use fire. Chips are mostly made of silicon, which is why he wrote: "We've essentially figured out how to make sand think." His own position is that, built and deployed responsibly, AGI could be one of the most beneficial, most transformative technologies in history.
The upside he lists includes accelerating drug discovery, developing new clean energy, and creating new advanced materials; further out, he sees the possibility of reaching a point where resources are no longer a hard constraint on human progress — an age of abundance. At the same time, he stresses that AI is already delivering real-world benefits, but fulfilling its larger promise means taking this critical development period seriously and addressing the risks that could emerge as we approach AGI, as quickly as possible.
He believes humanity's collective intelligence can handle the technical risks that come with AI — but only if we give ourselves the time and space to get the next steps right. His read is that, as a field and as a broader society, we haven't done that yet. Right now everyone is caught in an intense, multi-layered commercial and geopolitical race. That race accelerates progress and expands the upside, but it's also pushed the frontier ahead of our understanding of the technology itself. "Nobody in the world can say for certain what happens next, and experts disagree with each other." Given that uncertainty and the stakes involved, he argues for advancing with cautious optimism: public policy should both spur innovation and reward responsibility, push international cooperation on key safety issues, and seriously consider how AI should be deployed for society's benefit.
The core proposal: a Frontier AI Standards Body
Given how fast things are moving, he's calling for a dynamic, adjustable, sufficiently rigorous way to test frontier model capabilities. The US is positioned to go first, given its economic and technological standing.
A model that clears the benchmark set by the standards body — and updated as capabilities evolve — is designated "frontier-class"; the organization that holds such a model is designated a "Frontier Lab." That designation carries prestige, and it's open to any organization: what matters is building a model that hits the threshold, not where you came from.
At first, frontier labs would voluntarily submit models to the standards body for review up to 30 days before release. Once the evaluation protocol proves effective and robust, it can quickly become formalized: frontier models would have to pass the test to enter the US market. If a critical vulnerability surfaces after release, the lab would still work with the body to address it.
Submit for review up to 30 days before release
Frontier labs voluntarily hand models to the standards body for review, to get the evaluation protocol working first. The goal isn't full mandatory enforcement from day one — it's proving the "exam" is effective and robust.
Must pass the test to enter the US market
Once the protocol is proven effective, it can quickly escalate: frontier models must pass evaluation to be deployed in the US market. If a critical vulnerability surfaces after release, the lab still has to work with the body to address it.
Coordinate frontier labs to slow development
If things get serious enough, the framework can escalate further — including coordinating frontier labs to slow their pace of development. This isn't a default setting; it's the highest-level response option, built into the design.
What gets tested, what labs have to do, and who's exempt
Model evaluations need to cover high-risk capability areas like cybersecurity and biological threats, using rigorous scientific testing. For agentic systems (ones that can take multi-step actions and call tools), the checks look for attempts to circumvent safety guardrails or signs of deception — and push best practices such as digitally watermarking AI-generated images and producing human-readable intermediate tokens so the model's reasoning process can be understood.
It could start with quarterly updates; benchmarks that go stale or get saturated by score-gaming get retired and replaced. Early on, questions could be negotiated with frontier labs, but eventually the body needs to build its own independent question bank — one labs don't see in advance — to prevent overfitting, and to grow a third-party audit ecosystem.
The approach needs to be technically grounded while still supporting innovation and rewarding responsible behavior; it needs to keep pace with an accelerating field and adjust as the biggest identified risks shift. The US going first is meant to give international shared standards a starting point.
Once designated a frontier lab, an organization would be encouraged to adopt a set of best practices:
- Frontier-class models that clear the benchmark threshold
- Open-source or closed-source, either counts
- No restriction on country of origin
- Holders = Frontier Labs
- Models that don't reach the frontier threshold
- Most startups
- Typical academic projects
- Don't need to go through this review process
Because this technology will affect the whole planet, he hopes that this framework — built first by the US — can help push the international community toward consensus on managing the most serious risks, while making sure everyone gets access to the opportunities AI creates.
Even if the technology gets solved, the societal questions remain
He frames AGI as the ultimate tool for advancing science and medicine, and for lifting productivity and economic growth — provided the technical foundations are done right, coordination happens around a shared global framework, the most rigorous scientific methods are used, and the best minds are brought to the same table.
Even if the hard technical challenges get solved, there are still economic and philosophical questions waiting: what new economic model does a post-scarcity world need for everyone to actually thrive? What values do we want to live by? What becomes of meaning and purpose? How might the human condition itself change? These are clearly not questions that can, or should, be left to technologists alone — society as a whole needs to help define this next chapter.
He writes: there is both enormous excitement and enormous uncertainty around AI, and both are valid. But the future isn't written yet — we have to use this window before AGI arrives to shape the technology so it benefits all of humanity.
What we collectively do now will determine how the next stage of civilization unfolds. If we can safely bring AGI into the world, we can enter a new golden age of scientific discovery and progress, and human flourishing.
Demis Hassabis · paraphrased from the closing of the source postSource: X · 2026-07-14 · Demis Hassabis. This explainer only organizes his proposal and the mechanism design — it does not predict how legislation will actually unfold. FINRA = Financial Industry Regulatory Authority, the US self-regulatory body for securities firms; referenced here only as an institutional analogy.
Before the most powerful models ship, they get a checkup first
DeepMind CEO Demis: the US should build a standards body first — frontier models submit for review up to 30 days before release, which can then harden into a hard market gate
↓ Read the whole page · one figure animates
Demis thinks AGI is probably just a few years out — but right now the whole field is being driven by commercial and geopolitical competition, and frontier progress has already outpaced our understanding of the technology.
Huge upside ahead — he compares the scale to the Industrial Revolution's impact times its speed, both roughly 10x
✘ There's still no dynamic, upgradable checkup system for frontier models that can keep pace
Cybersecurity challenges are already here; nuclear/bio risk could surface soon; agentic and self-improving systems are even harder to manage
He proposes the US build a "Frontier AI Standards Body" first, using dynamic benchmarks to define frontier models and frontier labs — starting with voluntary submission, later escalating to a market-entry requirement.
Labs self-test and self-report; benchmarks are easy to saturate through score-gaming
Under competitive pressure, safety evaluation can't keep up with release schedules
Startups/academia and big labs get lumped into the same imagined framework, which makes it hard to implement
An independent standards body sets and updates the exam
Submit for checkup up to 30 days before release (voluntary first, mandatory later)
Only frontier systems are covered; non-frontier models are explicitly exempt
XiaoHu breaks Demis's proposed system down into a checkup pipeline: who writes the exam, who submits, and how the three tiers escalate.
This isn't a leaderboard — it's an institutional time window. Here's what "up to 30 days" and "can be tightened" feel like in practice:
XiaoHu shrugs: but with everyone racing this hard, who's got time for a checkup?
Cybersecurity's already a mess
Nuclear/bio risk is next in line, and agentic systems might start improving themselves
Demis: the US builds a FINRA-style standards body first. Industry funds it, experts plus open-source reps sit on the board, and it tests national-security-relevant capabilities alongside national labs.
Voluntary
Submit up to 30 days before release, to prove the exam works
Hard gate
Pass the test or don't enter the US market
Tighten
Coordinate frontier labs to slow down, if needed
Open or closed source, domestic or foreign — same treatment. The US goes first mainly to give international standards a starting point.
Checkup first
then ship