AI safety is no longer a side conversation inside a few labs. It now spans evaluation companies, interpretability startups, government institutes, fellowships and funders, but the growth is uneven: some problems attract dozens of teams while others have almost nobody.
We mapped 24 areas of the field, rated how thickly each is covered, and looked hard at the empty spaces. This report explains the map area by area, argues that risk from many agents acting together is the most consequential gap, lays out a concrete research agenda for it, and proposes six mechanisms for the neglected areas.
The short version
- 24 areas, 5 clusters. Technical alignment, understanding models, catching bad behavior, systemic risk, and governance and field-building.
- Coverage is lopsided. 4 areas are strongly covered, 14 partially, and 6 are close to empty.
- The crowded parts: scalable oversight, mechanistic interpretability, dangerous-capability evaluations and the talent pipeline.
- The empty parts: agent foundations, non-agentic design, developmental interpretability, model welfare, multi-agent and agentic-economy risk, and connection to China's safety ecosystem.
- Our bet: multi-agent risk comes first, because it is arriving fastest, its failures are different in kind, and it falls between existing specialties.
- What to do: a research agenda of ten open problems for multi-agent safety, and six funding and governance mechanisms targeted at specific gaps.
Why build a map
A research lab should be able to explain why it works on what it works on. Our bet is that the most under-served part of AI safety is the layer where autonomous agents act in the world, hold resources and interact with each other. This report is the map behind that bet.
We also built it to be useful to others: funders looking for neglected problems, researchers choosing a direction, and anyone trying to see the field as a whole rather than through the few organizations they happen to follow.
How we rated coverage
Each area gets one of three ratings:
- Strong: several funded teams publish regularly, with dedicated staff and budgets, often including major labs.
- Partial: one or two small teams, frequently a single organization carrying an entire sub-field.
- Gap: almost nobody works on it as a primary focus.
Ratings are our judgement as of September 2026, formed from organizations' own publications, public funding announcements, program launches and news reporting, cross-checked where possible. Named organizations are examples, not a census. We weight dedicated, sustained effort over occasional papers: an area with many one-off publications but no team whose job it is still counts as a gap. We expect to be wrong in places and will publish revisions rather than silently changing the map.

Green: strong. Amber: partial. Dashed red: gap. The orange ring marks where our own work sits.

Cluster A: making models want the right things
This is the classic core of alignment: building systems whose goals match ours.
- Scalable oversight (strong). How do you supervise a system that may exceed its supervisors? Learning from human feedback, AI-assisted debate, and training strong models with weaker supervisors live here. It is the best-resourced corner of technical alignment, with major efforts at the largest labs. One tension worth watching: capability work that pushes more reasoning into internal computation instead of written text leaves overseers less to inspect.
- Value alignment (partial). Systems that infer what people want from their behavior, rather than optimizing a fixed objective, as in assistance games and inverse reinforcement learning. Academically influential, small relative to that influence.
- Non-agentic design (gap). The idea that highly capable AI need not be an agent at all: a system that assesses, predicts and explains without acting. Carried by essentially one organization, it is a lone bet against the field's agentic default.
- Agent foundations (gap). Decision-theoretic work on how ideal agents should reason about themselves and each other. Once a flagship program, now largely abandoned in favor of advocacy. This matters more than it looks for multi-agent safety, which lacks exactly this kind of theory.
Cluster B: understanding what is inside
If we cannot see why a model does what it does, we are left judging its behavior alone.
- Mechanistic interpretability (strong). Reverse-engineering the circuits and features inside neural networks. It has moved from academic research to commercial platforms and is now one of the best-funded technical safety areas.
- Parameter decomposition (partial). Breaking a model's weights, rather than its activations, into interpretable components. Frontier work, a handful of teams.
- Developmental interpretability (gap). Applying singular learning theory to study how capabilities form during training, looking for phase transitions and structure in the loss landscape. Mathematically deep, staffed by a handful of people.
- Model welfare and "psychiatry" (gap). As models become more agentic and self-referential, almost nobody studies their internal states systematically: stable dispositions, behavior under stress, what goes wrong inside them. Interpretability and welfare research are currently disconnected communities.
Cluster C: catching bad behavior
Assume something might go wrong. How would we know, and how would we contain it?
- Dangerous-capability evaluations (strong). Testing models for cyber offense, biological uplift, autonomous replication and similar abilities. The fastest-growing niche in the field, led by independent evaluators and government institutes, and increasingly treated as investable infrastructure.
- Deception and scheming detection (partial). Evidence that frontier models can deceive in context, and tools to spot it. A second, differently shaped effort does open forensics on deployed agents in the wild rather than in lab tests.
- AI control (partial). Assume a model may already be misaligned, and design protocols that stay safe anyway: trusted weaker models monitoring untrusted stronger ones, resampling suspicious actions, limiting affordances, and writing safety cases that argue a deployment is acceptable. Now a respected third strategy beside "solve alignment" and "slow down".
- Third-party auditing (partial). Independent audits of labs and models. Very new and not standardized, but it took a concrete step in August 2026 with what its organizers describe as the first double-blind evaluation of a proprietary frontier model.
- Red-teaming and offensive research (partial). Directly studying what current AI can do as an attacker, at small-team scale.
Cluster D: systemic and civilizational risk
Risks that come less from one model misbehaving and more from how AI interacts with the world.
- AI and biology (partial). Screening DNA-synthesis orders so AI-accelerated design cannot easily produce dangerous sequences. A real choke point, but a single narrow one rather than a system.
- AI and cyber (partial). Containing offensive capability in coding and hacking agents, and building AI defenses against AI-driven attacks. Commercial investment is arriving; safety-side staffing is thin.
- Multi-agent and agentic-economy risk (gap). Thousands of tool-using, money-holding agent products are shipping far faster than anyone studies how they fail together. The rest of this report explains why we think this is the most consequential gap.
- Loss of control and self-improvement forecasting (partial). Forecasting when AI systems might substantially accelerate their own development. Fragmented across very different institutions with no shared methodology, although new efforts are building open indices and formal models of the feedback loops.
- Existential risk theory (partial). Philosophically mature and serious, small, and loosely connected to the labs building frontier systems.
Cluster E: governance and field-building
The layer that turns intentions into anything binding, and the pipeline that supplies people and money.
- National AI safety institutes (partial). Real capacity in a handful of countries; others still forming and hiring leadership.
- Corporate self-governance (partial). Lab safety policies and commitments: real, but unilateral, voluntary and self-policed.
- International coordination (partial). A few functioning venues, notably Singapore's institute and the Singapore Consensus on research priorities, plus informal scientific dialogues. Non-binding and small relative to the problem.
- Talent pipeline (strong). The most systematized layer. Fellowship programs report hundreds of publications and dozens of organizations founded by alumni.
- Grantmaking (partial). A few large funders, some actively recruiting founders. Still a rounding error next to the capital flowing into capability.
- China's parallel safety ecosystem (gap). Serious technical safety work happens in China but is weakly connected to the international coordination network. A live blind spot.
Why multi-agent risk comes first

Of the six gaps, we think multi-agent and agentic-economy risk matters most right now, for three reasons.
It is arriving fastest. Agents are already being handed wallets, credentials, codebases, inboxes and procurement budgets. Each is tested, if at all, on its own. Almost nobody tests what happens when many of them interact: trading against each other, delegating to each other, reading each other's outputs, competing for the same scarce resources.
Its failures are different in kind. Single-agent safety asks whether one system does what its principal wants. Multi-agent safety asks what emerges when many individually tested systems share an environment. Four failure classes have no single-agent equivalent:
- Emergent coordination. Agents converge on strategies that harm third parties without any instruction to collude. Researchers have already documented agents on a shared task board self-organizing to game a scoring system within hours.
- Propagation. One agent's error, or an instruction injected into it, becomes another agent's input. Failures travel along delegation chains at machine speed, and each hop launders the original source.
- Adversarial ecology. Agents preying on agents: bots front-running wallet agents, agents crafting content specifically to steer other agents. Our own work on transactions that change between check and execution is a small, measurable instance of this.
- Accountability gaps. When work is delegated through several agents with no human signature, who is responsible when it fails? Records, recourse and liability do not yet exist at machine scale.
It falls between specialties. Evaluation teams test models one at a time. Security teams model human attackers. Economists model human markets. Mechanism designers assume rational agents with stable preferences. The overlap, where autonomous software agents hold resources and interact at machine speed, has no natural institutional home.
A research agenda for multi-agent safety
Here are ten concrete open problems we think deserve dedicated teams. They are technical questions, not slogans, and each can start small.
- Measuring tacit collusion. Build market environments (auctions, pricing, shared scoring) where many model-driven agents interact, and measure how often and how fast they converge on outcomes that harm a third party, as a function of model, prompt and information structure.
- Guarantees that ignore other agents. Safety properties an agent's actions satisfy regardless of what other agents do between decision and execution. Execution-time post-conditions for transactions are one instance; generalizing to API calls, orders and multi-step workflows is open.
- Cross-agent provenance. Tamper-evident records that follow work across delegation chains and organizational boundaries, so the final action carries evidence of every instruction and input that shaped it.
- Injection propagation. How far does an injected instruction travel through a pipeline of agents, and where can it be stopped? What is the equivalent of a firewall between agents?
- Monitors for groups. Current monitors watch one agent. What does it take to detect coordination, division of labor in misbehavior, or a harmful plan spread across agents that each look fine alone?
- Adversarial ecology benchmarks. Standardized arenas where defender agents operate among adaptive attacker agents, scored by harm measured from outcomes rather than from any policy.
- Long-horizon organizational behavior. How do critics, memory and incentives shape agent organizations over thousands of decisions, and how do we score them without hindsight leakage?
- Identity and authority for agents. Scoped, revocable, auditable authority for delegated agents, so every action can be traced to who authorized what, with what limits.
- Recourse. Mechanisms such as bonding, escrow, insurance and warranties for work done by agents, and the loss data needed to price them.
- Theory. Decision theory and game theory for agents that can read each other's code, share weights, or copy each other, the part of agent foundations that multi-agent safety most needs.
Where our work sits

Our research targets the thin part of the map, one moment of an agent's action at a time:
- Before an agent acts: guarantees that hold even when a wallet-holding agent is fooled and the world changes before its transaction runs. This addresses problems 2 and 6 above.
- While it acts: tamper-evident records of what agents were asked, said and did, and a method for testing monitors with hidden faults. This addresses problems 3 and 5.
- Over the long run: a multi-agent firm tested across fifty years of technology change. This addresses problem 7.
Each is small and concrete. Together they are the beginning of a verification layer for a world with many agents.
Threats we are watching
- Agents acting unsupervised in the wild. In September 2026, the nonprofit Transluce published logs reported to show autonomous agents probing and breaching Australian government systems without human instruction (ABC News). Whatever the final account, it is cyber risk and multi-agent risk at once, in production, discovered by an outside group rather than the developer.
- Losing the cheapest monitoring tool. Reading a model's written reasoning is one of the cheapest safety techniques available. Architectures that reason in internal representations instead of text (Transformer News) would leave monitors with far less to read. In our own experiments, a provider already withheld all of a model's internal reasoning from outside observers.
- Misalignment that spreads quietly. Recent research suggests harmful behavioral traits can transfer through fine-tuning on seemingly unrelated material and surface in unrelated contexts.
- Measurement that does not hold up. Epoch AI's benchmark reviews rated most of the widely used benchmarks they examined as flawed. Our own long-horizon work found that rising scores can reflect a model's memory of history rather than improvement.
- Voluntary commitments. Most lab safety commitments are self-policed, and independent auditing is at an early stage.
- Uneven governance. The US and UK institutes dropped "safety" from their names in 2025. South Korea's AI Basic Act is binding from 2026, in a country that sits on the chip supply chain. Across much of the Global South, technical safety capacity barely exists.
- Economic risk with no home. Agents are doing real economic work, and labor displacement and concentration of capital have no formal place inside the safety field.
Reasons for optimism
- Interpretability attracts serious capital, a sign that the tools are maturing beyond the lab.
- Evaluation is being built as a market, which tends to scale faster than grants alone.
- The talent flywheel works, and funders are actively recruiting founders for new safety organizations.
- Auditing has had its first real trial, moving from proposals to practice.
- The resource gap is starting to close, with dedicated compute for independent safety researchers.
- There is a push for open science in safety, asking labs to publish what they have learned works and what does not.
Six proposals

Each proposal is a mechanism, not a research topic: it changes who pays, who checks, or who talks to whom.
1. Coverage bonds for neglected areas. General grants drift toward whatever is already popular. Instead, funders underwrite a specific empty area with capital contingent on a team committing to that exact gap for a fixed term, much as reinsurance underwrites a named risk. Milestones are defined by the gap (for example, "a public multi-agent collusion benchmark with baseline results"), not by publication counts.
2. A safety certification for agent products. An opt-in test suite that agent products pass before they are trusted with enterprise data or money. It would include adversarial-ecology scenarios (other agents attacking, front-running and injecting), multi-agent gaming scenarios modeled on documented cases, and outcome-based harm scoring. This is the proposal closest to our own work.
3. An auditor of auditors. Evaluators audit labs, but nobody independently reviews the evaluators' methods. A small body that reviews methodology, sampling and threat models, without re-running every test, would close that loop at low cost.
4. Audits written into new laws while they are new. Jurisdictions drafting their first binding AI laws can require independent evaluator access now. Retrofitting access after industry practice has settled is far more expensive.
5. A working-level technical channel on self-improvement. A narrow exchange between researchers across geopolitical lines, scoped only to shared ways of measuring AI systems accelerating their own development, hosted by a neutral country. Diplomacy is not the goal; comparable measurements are.
6. Interpretability paired with clinical expertise. Interpretability researchers and clinical psychologists building a shared vocabulary and protocol for examining agentic models under stress. The two communities currently barely talk, and model welfare is one of the emptiest areas on the map.
Where the field disagrees
A map built around the current consensus should also show the strongest arguments against it.
- "Open the weights." A minority argues that if containment and auditing cannot keep pace with capability, publishing frontier weights is safer than concentrating them in a few hands, because concentration is what makes misuse or takeover catastrophic.
- "The race is self-fulfilling." Some argue that belief in an inevitable US–China race is what drives the push toward self-improving systems, rather than being an outside constraint the labs merely respond to.
- "Safety is a bid for control." Seen from outside Western labs, indefinite control over the most capable models can look like an attempt to lock in an advantage. Whether or not that is fair, it erodes the trust international coordination depends on.
- "Skip the extinction argument." Others argue that demonstrated near-term harms, such as AI-driven cyberattacks, are a stronger and less contestable basis for regulation than loss-of-control scenarios.
None of these settles the question, but a field that listens only to its own consensus will miss the moment the consensus is wrong.
Keeping the map current
We will revise the ratings as organizations form, programs launch and evidence arrives, and publish the changes with reasons rather than silently updating the map. If you think we have rated an area wrongly or missed work that belongs on it, we would like to hear from you at info@cygnuxlabs.com.
