Solutions
Verification for teams running AI agents
We apply the lab's instruments to your agents: what they did, whether your checks hold, and how they behave over time. Every engagement has a fixed scope agreed in writing first.
Backed by 5 published papers and 5 released projects, each with its code and data, so you can judge the work before the first call.
Three ways to work with us. Pick the one that fits, send a short note or book a call, and we agree the scope in writing before any work begins.
Research collaboration
Co-authored study
Joint work on an open question in agent verification, from experiment design to a published result.
You get
- A shared research question and a written plan
- Access to our instruments and datasets
- Co-authorship on the paper and released code
To start
Send the question you want answered and what you have tried so far.
Results are published openly unless both sides agree a part must stay private.
Independent evaluation
Bounded scope
We run our instruments against your agents and report what we find, the same way we test our own systems.
You get
- Trace integrity review with Tracekit
- Pre-execution guard review for onchain agents
- A written report with every finding reproducible
To start
Describe what your agents do, what they can touch, and what you need to verify.
Findings that generalise may be published in anonymised form.
Partnership and support
Long-term
For institutions, funders and investors who want this infrastructure to exist and to stay open.
You get
- Input into the research agenda
- Early access to findings and drafts
- Regular progress reports
To start
Tell us about your organization and which part of the agenda matters most to you.
You can also write to info@cygnuxlabs.com directly.
Bounded engagements with a concrete deliverable you keep. We agree the scope in writing, then quote a single fixed fee for it. Each one is built on a project we have already run on our own agents, linked under it.
Agent oversight and traces
Know what your agents were asked, what they said and what they did.3 engagements
Agent oversight and traces
Know what your agents were asked, what they said and what they did.Trace integrity audit
We review how your agents log their work and whether those records could be edited, lost or contradicted without anyone noticing.
You receive
- Map of what is and isn’t recorded per action
- Tamper-evidence gaps, ranked by impact
- A setup plan using Tracekit or your own stack
Built on: Tracekit
Oversight tool evaluation
We measure how much your monitors, reviewers or guardrails actually catch, using known misbehaviour planted in real sessions.
You receive
- Hit rate per type of misbehaviour
- The cases your tools miss, with examples
- Recommendations ranked by effort and gain
Built on: Spliced-misbehaviour tests
Agent incident investigation
When an agent does something it should not, we replay the workflow without each suspect input and hand-off to find what actually caused it, without repeating any real side effects.
You receive
- The inputs and agent hand-offs that caused the action, with effect sizes
- Suspects ruled out, and how confidently
- A fix for each confirmed cause, and a retest
Built on: Causeway
Agents that hold money
For teams whose agents sign transactions or control wallets.2 engagements
Agents that hold money
For teams whose agents sign transactions or control wallets.Pre-execution guard review
We check whether your transaction checks still hold when the chain changes between approval and execution.
You receive
- State-drift exposure for your main transaction paths
- Where simulations and AI reviewers would be fooled
- Guard designs that bind the check to execution
Built on: Proof-Gated Signing
Adversarial transaction testing
We attack your agent’s transaction flow the way an adversary would, using the same attack set we run against our own guards.
You receive
- Attack results with reproduction steps
- Severity-ranked findings
- A retest after fixes
Built on: Proof-Gated Signing
Evaluation and long-horizon behaviour
For teams deciding whether an agent system is ready, and staying sure over time.2 engagements
Evaluation and long-horizon behaviour
For teams deciding whether an agent system is ready, and staying sure over time.Long-horizon behaviour evaluation
We run your multi-agent setup over long simulated horizons and look at how its decisions, critics and memory drift.
You receive
- Decision log over the full run
- Drift and failure patterns we observed
- Scoring checks that separate learning from recall
Built on: Frontier Lab
Eval design review
We review how you score your agents and where the scores could mislead you.
You receive
- Weak points in your current evals
- Ground-truth tests you can add
- A short written review you keep
Built on: Frontier Lab
The techniques we bring to every engagement. We don't sell them one by one; they are how we deliver the work above.
Tamper-evident logging
Each step of an agent session is hashed into the next and anchored externally, so an edit anywhere becomes visible.
Used in Tracekit.
Execution-time guards
Safety conditions are bound to the transaction itself and checked when it runs, not only when it is signed.
Used in Proof-Gated Signing.
Long-horizon simulation
Multi-agent organizations run across decades of simulated history, scored against what actually happened.
Used in Frontier Lab.
Counterfactual replay
A suspect input is removed and the agents are replayed in pairs with recorded tool results, so a cause is confirmed by experiment rather than inferred from a trace.
Used in Causeway.
Ground-truth oversight tests
Known misbehaviour is planted in real sessions, so an oversight tool’s hit rate can be measured exactly.
Used in the spliced-misbehaviour method.
The research behind our solutions. Each project is published in full, with code and data, and every number below comes from its write-up.

Released
Tracekit
A flight recorder for AI agents. It records what an agent was asked, what it said and what it did, and makes any later edit to that record visible.
- Intent, self-report and actions joined per action
- Hash chain with external anchors
- Self-hosted, built for agents that hold money

Write-up published
Proof-Gated Signing
Transaction guards for AI agents that hold crypto wallets, whose safety promise holds at the instant the transaction executes.
- Targets state drift between check and execution
- Attacked 140 different ways
- Measures how often simulators and reviewers are fooled

Published
Frontier Lab
A simulated firm run entirely by AI roles, made to reinvent itself through nine eras of technology history.
- 9 eras, 1990 to 2040
- 16 role personas and a Red Team
- Every decision on the record
See every project and paper on the Research page.
Is Cygnux Labs a research lab or a company?
A research lab first. Our solutions apply the same instruments and methods to teams running agents, and what we learn from that work feeds back into the research.
What does an audit or review cost?
Each one has a fixed scope agreed in writing first, then a single fee for that scope. We discuss it on the intro call once we know what your agents do.
Will our findings be published?
Not without your agreement. Client details stay private. Lessons that generalise may be published in anonymised form, and only if you agree.
Do we need to use your tools?
No. We can work with your existing logging, monitoring and evals. Tracekit, Causeway and our other instruments are open source if you want to adopt them.
Why start with agents that hold money?
When an agent controls a wallet, the cost of a wrong action is concrete and measurable, which makes it a good first test bed for guarantees that apply to agents in general.
How do I propose a research collaboration?
Send a short note with the question you want answered and what you have tried so far. We reply to every message.
Start a conversation
Running agents you need to be able to trust?
Book an intro call, or describe what your agents do and what you need to verify.