The most consequential technology of this century is being built by a few private labs, behind closed doors, on infrastructure that costs more than most countries spend on science. The people deciding how capable these systems become, and how safe they are, answer mostly to shareholders and to each other's release schedules.
Physicists faced a version of this problem in the 1950s. No single lab could afford the machine they needed, so they pooled resources and built one together, with shared data and shared credit. This piece sketches what the same move could look like for advanced AI: a public institution at the frontier, built to make powerful systems trustworthy and to let others check that they are.
The short version
- Trust has to be built into how systems are made. Hospitals, grid operators, banks and public agencies cannot responsibly deploy what they cannot inspect. That requires public capacity at the frontier.
- Combine ARPA and CERN. Empowered programme directors and open collaboration for most work. Secure in-house teams for the dangerous work.
- Protect the foundations. Reserve a fixed share of compute and staff for the unglamorous science of interpretability, guarantees and verification.
- Decide the guardrails in advance. Release stages, safety thresholds and governance triggers are written down in calm conditions, before anyone is under pressure on release day.
- Tier the membership and fund for the long term. Start with a small trusted group, grow carefully, and commit money for longer than one hardware cycle.
The problem is who holds the capability
The organisations that understand frontier AI best also have the strongest incentive to move fast and keep findings private. That is how competition works, and it says nothing about anyone's intentions. A lab that pauses to verify safety loses ground to one that doesn't.
Everyone else is mostly on the sidelines. Governments can regulate what they can see but cannot easily build what they want to see. Universities have talent but little compute. Startups have speed but not capital. Public institutions have legitimacy but rarely the agility.

The third step in that loop is the one we care about most. Sectors that run critical services want the benefits of advanced AI and cannot get them responsibly without a way to inspect and verify what they are deploying. Trust is the bottleneck, and it cannot be added after the fact. It has to be engineered into how the technology is made and checked.
A public frontier institution would break the loop. It would not replace private labs. It would change who gets to build trustworthy systems, and who gets to verify them.
"The money gap is too large"
Private labs are announcing data-centre projects measured in hundreds of billions of dollars. A public body cannot match that, and it does not need to.
First, progress no longer depends only on scale. Recent gains have come heavily from post-training methods such as reinforcement learning and from spending more compute at inference time. Technique matters as much as raw size, and smaller teams have shown they can get surprisingly far.
Second, a public institution should compete on the most trustworthy science, not the biggest model. Making advanced systems safe, interpretable and verifiable is an open research problem. The labs building them say openly that today's safety techniques may not hold for much more capable systems. Open research frontiers are exactly where CERN-style institutions do well. The useful question is which problems the private market will consistently under-invest in, and to build there.
Why CERN, and why ARPA
Two institutional models come up again and again in this conversation, and they solve different problems.
CERN is shared infrastructure for fundamental science: an expensive facility no single member could build, with results that benefit everyone. It was also deliberately a peace project and a magnet for talent.
ARPA, the agency behind the early internet, GPS and much of modern computing, is a model of empowered individuals with ambitious missions and little red tape. Its main ingredients:
- programme directors with the authority to pick bold, unusual bets;
- time-boxed missions, usually five to seven years;
- autonomy over hiring, procurement and funding;
- tolerance for failure, because a few successes on very large bets pay for the rest;
- close ties to academia and industry, so ideas reach the world.
Neither model fits on its own. A pure ARPA works through distributed grants and avoids large in-house teams, which is awkward when you need to train a frontier model or protect dangerous capabilities. A pure accelerator-style lab is too slow for a field that reinvents itself every eighteen months. The institution we have in mind is a hybrid: CERN's serious shared infrastructure and long horizon, with ARPA's speed and talent-first culture.
What it has to get right
Some requirements matter from the first day:
- Parallel moonshots. Many ambitious programmes at once. Most will fail and a few will matter enormously.
- Startup speed with public accountability. Fast hiring, procurement and partnerships, without losing democratic legitimacy.
- Deep links to industry. Research that never reaches products changes nothing.
- A world-class resource base. Competitive pay and serious compute, or the best people will not come.
- Strong founding leadership. The first few years set the culture for decades.
Others matter in year ten:
- Adaptive oversight that tightens when capabilities demand it and stays light when they don't.
- Serious security for model weights and dual-use findings.
- Defences against mission drift, built into the structure.
- International engagement, because a technology this general cannot be shaped by one bloc.
- An ecosystem role. The institution should help a whole field grow until it can stand on its own.
Many of these pull against each other: speed against oversight, openness against security, public control against private partnership. The design work is managing those tensions on purpose.
The mission
A mission statement is the one thing that survives leadership turnover. CERN's fits in a sentence about understanding what the universe is made of. A version for this era could read:
To pioneer the science of trustworthy, controllable and increasingly general AI, up to and beyond human-level capability, and to turn that science into applications that benefit society.
Two choices in that sentence need defending.
General-purpose systems. A single general model now matches or beats specialists across coding, translation, analysis and diagnosis. That is the opportunity and the risk at once: one technology touching everything. If it cannot be trusted, the problem is everywhere at the same time. The science of making general systems trustworthy is therefore the highest-leverage research available. Specialised models still matter, and there should be room for targeted work in climate, materials, medicine and robotics.
Naming superintelligence. The trajectory matters more than today's products. If systems keep improving, the hard questions concern systems that exceed people at most cognitive work. An institution designed only for current models will be under-equipped when those questions arrive. It is better to build the governance and security capacity before it is needed. Reasonable people disagree about timelines, so the institution should be useful whether progress plateaus soon or arrives quickly.
Four focus areas, one flywheel
Programme selection can rest on three tests. Does it advance the mission? Is it a big bet with transformative upside? Would it happen without this institution? If a programme would be funded anyway, it does not belong.

Foundations: the science of trust. Interpretability tools that show what is happening inside a model. Architectures with mathematical guarantees about behaviour. Reasoning systems that are robust and predictable.
Scaling: making it real. Trustworthy frontier models, domain models for areas such as materials science, and carefully built datasets that are high quality and preserve privacy.
Applied: problems people care about. Cyber-defence systems, laboratory automation with very high precision, and AI that manages energy grids running on variable renewable power.
Hardware: the physical layer. Chips that do more per watt, sensors that process data on the device, and hardware mechanisms that can verify how AI chips are being used. That last one gives oversight a physical anchor instead of relying on trust alone.
Protect the boring part
One rule deserves emphasis: reserve at least a quarter of compute and staff for foundations. Without a floor, resources drift toward work with a near-term payoff. Applications have customers and press releases. Foundational research has neither until the day it matters more than anything else. The floor is how the institution keeps its own promise.
Two research tracks
ARPA-style organisations work through networks. A programme director funds and coordinates outside teams. That suits ideas that benefit from openness and a diverse talent pool. Some work cannot be done that way:
- findings with dual-use potential, such as a defensive cyber tool that could be turned to attack;
- work that needs sustained, tight collaboration, such as training a frontier model;
- anything involving model weights that must never leak.
For these, a second track is needed: a dedicated in-house team in a secure environment, working on one defined challenge over the long term.

The two tracks provide three things together:
- Access control by design. Think of a building with a public lobby and a locked vault. Partners work on lower-tier programmes according to their trust level. Only the most trusted get near the vault.
- Geographic spread. Sensitive work needs one carefully chosen facility. Distributed programmes can run wherever there is a strong partner, so many members can host real work.
- The right tool for each problem. Open ideas stay open. Dangerous ones stay contained.
Scrutiny first, then freedom
Giving directors freedom is only responsible if the scrutiny comes first. A good pattern is a short incubation of around three months, in which a proposed director defends their thesis, resources, success metrics and checkpoints before leadership and independent experts. Only then does the programme get real operational autonomy, followed by light reviews at agreed milestones.
Systems engineers hold it together
Research organisations often fail quietly because each team optimises its own corner. The algorithm team doesn't talk to the chip team, and the applied team builds something the foundations team has already made obsolete. Industrial labs in their best years solved this with systems engineers: technically strong people whose job is the wide view. They track new capabilities, find real needs, judge feasibility and break coordination deadlocks, such as researchers refusing to build an algorithm for hardware that doesn't exist while chip designers refuse to build hardware without proof the algorithm matters.

The time limit is deliberate. It keeps programmes honest, forces a decision about what happens to the results, and frees resources for the next bet.
How results leave the building
Most work should simply be published, open-sourced or licensed to spin-offs. The most capable models change the calculation. At some point, publishing the weights of a system that could help synthesise a pathogen or automate a cyberattack is not responsible. A release ladder sets out in advance how access narrows as capability and risk rise.

- Open weights. Early, lower-risk models are published so small firms and researchers can build without frontier-scale resources.
- Licensed weights. When risks in areas such as cyber or biology rise, access moves to vetted partners under licence. Their fees help fund programmes.
- Hosted inference. At the highest thresholds the weights never leave. A dedicated team runs the model as a service, and surplus revenue can go into a benefit-sharing pool for members.
What matters most is that the thresholds are written down beforehand and not improvised under commercial or political pressure.
Governance: autonomy inside, accountability outside
Researchers need freedom and the public needs assurance. Two principles reconcile them: transparency by default and accountability aimed at the right level.
Transparency. Inside, the default is sharing: regular updates, independent internal evaluations and a culture where questions are welcome. Withholding information without a clear security reason should be treated as serious enough to end a programme. Outside, concrete tools matter more than slogans: whistleblower protections, public safety assessments, honest disclosure of capabilities, and regular explanation of why high-risk bets are being taken.
Where transparency is applied matters. Many public bodies audit small expenses in detail while major strategic decisions escape scrutiny. The aim here is the reverse: light at the operational level and rigorous at the strategic one. ARPA itself lost effectiveness in later years as bureaucracy built up.
Accountability. Final authority should sit with political representatives of member states. AI is already geopolitical, and putting politics openly in the boardroom is better than leaving it to back channels. Representatives shape strategy. Leadership runs operations. Space agencies offer a working template, with a council for strategy and a director-general for execution. Private companies get an advisory voice, because their expertise is valuable and their incentives differ from the public's.

The two expert boards matter most. The Mission Alignment Board asks whether the institution is still doing what it said it would do. The Scaling and Deployment Board reviews whether and how powerful models are trained and released. Together they give the political board competent, independent technical advice and real visibility.
Two frameworks written in advance
AI moves faster than governance documents can be rewritten, so two if-then frameworks belong in the founding design:
- A safety and security framework with capability thresholds that trigger stronger measures automatically. For example, once a system passes a set level of biological capability, extra monitoring and tighter security apply before anything else happens.
- A governance plan for the institution itself, specifying how oversight changes as stakes rise: more sign-offs for major releases, wider security vetting, more central control if capabilities become extreme.
The moment new checks are most needed is the moment there is most pressure to skip them. A plan written in calm conditions is worth far more than improvisation in a crisis.
Membership
The right starting point is a small trusted group, designed to grow carefully. Several founding members speed adoption, because scientists take techniques home with them, and bring more funding, more talent and more legitimacy than one country could. Growth also brings risks: sensitive technology must not leak to hostile actors, and a large membership can paralyse decisions.

- Full members vote, access top-tier programmes, contribute the most and meet the strictest criteria.
- Associate members have a limited voice, access to lower-security programmes and hosted products, and a lower contribution.
- Observers and partners have no vote but join open research and shared evaluation work, admitted programme by programme.
The same tiering applies to companies and universities. A firm with troubling ties can be refused even if its home country is a member, while a university from a non-member country might join one open programme under strict protocols.
Three mechanisms keep growth from becoming gridlock. An existing member sponsors each applicant and handles the negotiation, while the vote stays collective. Strategic changes need a supermajority that includes founding members. And misconduct, such as unauthorised technology transfer, carries real consequences up to expulsion.
Money and legal form
This will be expensive. The CFG study estimates tens of billions of dollars over the first few years, mostly for data centres. That is a small share of what governments already say they must invest in digital technology, and it buys more than AI: climate tools, cyber resilience, scientific capacity, and the ability to build instead of only buying.
Core funding has to be committed upfront from founding governments, public research budgets and strategic partners such as data-centre and chip companies, because nobody can recruit top researchers or sign compute contracts on maybes. Variable income from industry programmes, model licences and spare compute is welcome upside but should not be budgeted on. Members should commit for a minimum multi-year term, as they do at CERN. AI hardware ages quickly and the competition for talent never pauses.
The legal vehicle needs to hire competitively, accept public and private money, commercialise and reinvest, share benefits, stay under public control, admit outside members and be set up in months. A treaty organisation takes years to negotiate. A research council cannot partner commercially. A university-style non-profit lacks public control. The CFG study found the best fit in a joint public-private research entity created under an existing legal framework, and pointed to a comparable compute body that became operational about ten months after it was proposed. Its warning is worth repeating: legal flexibility is wasted if founders then add their own restrictions on hiring and leave unclear who decides what.
What could go wrong
- It could speed up the race it is meant to calm. Building frontier systems in public might push others to move faster. The safety framework needs real restraint, including the willingness not to train or release something. An institution that cannot say no is just another lab.
- Capture is a permanent risk. Industry, a dominant member state or the institution's own leadership could bend it over time. The independent boards and the public layer need real powers.
- Security cuts both ways. A concentrated facility is a target, and strong security can slide into unnecessary secrecy. Openness has to stay the default.
- "Superintelligence" is a loaded word. It can read as hype. We use it because being honest about the trajectory is better than a euphemism.
- Execution decides everything. Diagrams don't build institutions. Leadership, culture and the first three years do.
Why this matters to us
Cygnux Labs works on making AI systems verifiable: records of what agents did that cannot be quietly changed, and methods that test why they acted. The institution described here puts the same idea at a much larger scale. Its foundations programmes, release ladder, capability thresholds and chip-level verification all depend on being able to check what a system is and what it did. That is the layer we think deserves far more public investment than it gets.
We are about to make some of the largest decisions in history about a technology nobody fully understands. Those decisions should not be made in private, under commercial pressure, with no shared infrastructure for verification. The physicists built a machine to look inside the atom, together. We can build one to look inside the minds we are creating.
This essay is our own reinterpretation of ideas from "Building CERN for AI: An institutional blueprint" by Alex Petropoulos, Bálint Pataki, Daan Juijn, David Janků and Max Reddel at the Centre for Future Generations (January 2025). Cost estimates, the legal-form analysis and the ten-month precedent come from that study. The risk section and the framing for the superintelligence era are our own.
