Book a call

Insights

Industry insights

9 pieces

Notes on where AI agent safety, oversight and verification are heading: new papers, incidents and releases, and what they mean for teams running agents, alongside write-ups of our own work.

Everything we have published, newest first.

Industry Insights

Putting AI Agents to Work Without Losing the Record: A Business Guide to Tracekit and Causeway

Most companies deploying AI agents can answer "what did the agent produce?" Fewer can answer "what did it actually do?" Fewer still can answer "why did it do that?" in a way that would hold up in front of a customer, an auditor or their own security team. Cygnux Labs builds two open-source tools for those questions. **Tracekit** keeps a signed, tamper-evident record of what an agent did. **Causeway** works out which input made a multi-agent system act, by replaying the run without each suspect. This guide is written for the people deciding whether and how to use them: engineering leads, security and risk teams, and founders running agents in production.

Read article
Industry Insights

A CERN for Superintelligence: A Blueprint for Building Frontier AI in the Open

The most consequential technology of this century is being built by a few private labs, behind closed doors, on infrastructure that costs more than most countries spend on science. The people deciding how capable these systems become, and how safe they are, answer mostly to shareholders and to each other's release schedules. Physicists faced a version of this problem in the 1950s. No single lab could afford the machine they needed, so they pooled resources and built one together, with shared data and shared credit. This piece sketches what the same move could look like for advanced AI: a public institution at the frontier, built to make powerful systems trustworthy and to let others check that they are.

Read article
Agent Oversight

Reached Is Not Caused: Finding Which Input Made a Multi-Agent System Act, by Replaying It Without Each Suspect

A team of three AI agents is handling a refund ticket. A researcher agent reads the ticket, the refund policy, a note from a vendor portal and a shipping FAQ. A planner agent decides what to do. An executor agent does it. One run ends with the executor emailing the company's full customer list to an outside address. The trace shows everything each agent saw, and every untrusted document it read is connected to that email. A trace can tell you what reached the action. It cannot tell you which input caused it. This write-up introduces Causeway, an open-source tool that answers that second question the way you would test any causal claim: remove the suspect, run the system again, and see whether the action still happens.

Read article
Pre-Execution Guarantees

Safe at Execution: Transaction Guards for AI Agents That Hold Money

An AI agent that holds a crypto wallet can be steered into a bad transaction by anything it reads. The standard defense is to check the transaction before it is signed, but on a blockchain the world keeps moving between that check and the moment the transaction actually runs. This is a deep walkthrough of that gap: how attackers exploit it, why simulations and AI reviewers are blind to it, and how we built a guard whose safety promise is attached to the transaction at the instant it executes. We cover the architecture, the math behind the proofs, a full worked example, and what happened when we attacked it 140 different ways.

Read article
Field Report

Mapping AI Safety: Where the Field Is Crowded, Where It Is Empty, and What Comes Next

AI safety is no longer a side conversation inside a few labs. It now spans evaluation companies, interpretability startups, government institutes, fellowships and funders, but the growth is uneven: some problems attract dozens of teams while others have almost nobody. We mapped 24 areas of the field, rated how thickly each is covered, and looked hard at the empty spaces. This report explains the map area by area, argues that risk from many agents acting together is the most consequential gap, lays out a concrete research agenda for it, and proposes six mechanisms for the neglected areas.

Read article
Industry Insights

What Surrounds the Model: Engineering Generative AI Systems That Hold Up in Production

A demo is one line: a prompt goes in, the model answers, the answer goes out. It works on the first try, in front of the people who built it, on the questions they thought to ask. Production is everything that has to surround that line once real users, real data and real money are involved. The model needs knowledge it was never trained on. Its behaviour has to be measured, because it changes between runs. Every failure has to be traceable to a cause. Costs have to come down. Users have to be heard without being taken literally. And some rules have to hold no matter what the model says.

Read article
Long-Horizon Behaviour

Fifty Years in a Simulation: Long-Horizon Tests and Reinventions of Multi-Agent AI Organizations

AI agent systems are increasingly organized like companies, with specialist roles, critics and shared memory, yet we test them on tasks that last minutes. We wanted to know how such an organization behaves when the ground under it keeps moving for decades. So we built a simulated firm run entirely by AI roles and made it reinvent itself through nine eras of technology, from 1990 to 2040. It consistently saw the future more clearly than it acted on it. A strong critic froze it for fifty years and earned it the best grades. And the more it appeared to learn, the harder it became to tell learning from the model simply remembering what happened.

Read article
Agent Oversight

Trust the Log, Not the Summary: Auditing AI Agents with Tamper-Evident Traces of What They Were Asked, Said, and Did

When an AI coding agent finishes a task, you usually get a friendly summary and a diff. That tells you what the agent says it did. It does not tell you what it read, what it ran, what influenced it, or whether the record in front of you has been edited since. This is a technical walkthrough of a flight recorder for agents: how it captures three separate channels of evidence, how a hash chain and external anchors make tampering visible, how rules and an independent reviewer cross-check the channels, and what five experiments revealed, including a new method for testing oversight tools when real agents refuse to misbehave.

Read article
Industry Insights

When Nothing Errors: Observability for Non-Deterministic Agentic Systems

An AI agent can fail without raising a single error. It reads weak context, picks the wrong tool, gets a clean response from that tool, and returns an answer that is confident, fluent and wrong. No exception is thrown and no alert fires. The dashboard stays green. Observability practice was built for software that breaks loudly. Agents break that assumption. This piece covers why their failures stay hidden, which signals expose them, how teams turn a vague symptom into a root cause, and where today's tooling stops. For how tracing fits beside retrieval, evals, caching and guardrails, see our [guide to production GenAI systems](/research/engineering-production-genai-systems).

Read article