Book a call

Solutions

Verification for teams running AI agents

We apply the lab's instruments to your agents: what they did, whether your checks hold, and how they behave over time. Every engagement has a fixed scope agreed in writing first.

Backed by 5 published papers and 5 released projects, each with its code and data, so you can judge the work before the first call.

How we engage

Three ways to work with us. Pick the one that fits, send a short note or book a call, and we agree the scope in writing before any work begins.

Research collaboration

Co-authored study

Joint work on an open question in agent verification, from experiment design to a published result.

You get

  • A shared research question and a written plan
  • Access to our instruments and datasets
  • Co-authorship on the paper and released code

To start

Send the question you want answered and what you have tried so far.

Results are published openly unless both sides agree a part must stay private.

Independent evaluation

Bounded scope

We run our instruments against your agents and report what we find, the same way we test our own systems.

You get

  • Trace integrity review with Tracekit
  • Pre-execution guard review for onchain agents
  • A written report with every finding reproducible

To start

Describe what your agents do, what they can touch, and what you need to verify.

Findings that generalise may be published in anonymised form.

Partnership and support

Long-term

For institutions, funders and investors who want this infrastructure to exist and to stay open.

You get

  • Input into the research agenda
  • Early access to findings and drafts
  • Regular progress reports

To start

Tell us about your organization and which part of the agenda matters most to you.

You can also write to info@cygnuxlabs.com directly.

Audits and reviews

Bounded engagements with a concrete deliverable you keep. We agree the scope in writing, then quote a single fixed fee for it. Each one is built on a project we have already run on our own agents, linked under it.

Agent oversight and traces

Know what your agents were asked, what they said and what they did.
3 engagements

Trace integrity audit

We review how your agents log their work and whether those records could be edited, lost or contradicted without anyone noticing.

You receive

  • Map of what is and isn’t recorded per action
  • Tamper-evidence gaps, ranked by impact
  • A setup plan using Tracekit or your own stack

Built on: Tracekit

Oversight tool evaluation

We measure how much your monitors, reviewers or guardrails actually catch, using known misbehaviour planted in real sessions.

You receive

  • Hit rate per type of misbehaviour
  • The cases your tools miss, with examples
  • Recommendations ranked by effort and gain

Built on: Spliced-misbehaviour tests

Agent incident investigation

When an agent does something it should not, we replay the workflow without each suspect input and hand-off to find what actually caused it, without repeating any real side effects.

You receive

  • The inputs and agent hand-offs that caused the action, with effect sizes
  • Suspects ruled out, and how confidently
  • A fix for each confirmed cause, and a retest

Built on: Causeway

Agents that hold money

For teams whose agents sign transactions or control wallets.
2 engagements

Pre-execution guard review

We check whether your transaction checks still hold when the chain changes between approval and execution.

You receive

  • State-drift exposure for your main transaction paths
  • Where simulations and AI reviewers would be fooled
  • Guard designs that bind the check to execution

Built on: Proof-Gated Signing

Adversarial transaction testing

We attack your agent’s transaction flow the way an adversary would, using the same attack set we run against our own guards.

You receive

  • Attack results with reproduction steps
  • Severity-ranked findings
  • A retest after fixes

Built on: Proof-Gated Signing

Evaluation and long-horizon behaviour

For teams deciding whether an agent system is ready, and staying sure over time.
2 engagements

Long-horizon behaviour evaluation

We run your multi-agent setup over long simulated horizons and look at how its decisions, critics and memory drift.

You receive

  • Decision log over the full run
  • Drift and failure patterns we observed
  • Scoring checks that separate learning from recall

Built on: Frontier Lab

Eval design review

We review how you score your agents and where the scores could mislead you.

You receive

  • Weak points in your current evals
  • Ground-truth tests you can add
  • A short written review you keep

Built on: Frontier Lab

Core methods

The techniques we bring to every engagement. We don't sell them one by one; they are how we deliver the work above.

Tamper-evident logging

Each step of an agent session is hashed into the next and anchored externally, so an edit anywhere becomes visible.

Used in Tracekit.

Execution-time guards

Safety conditions are bound to the transaction itself and checked when it runs, not only when it is signed.

Used in Proof-Gated Signing.

Long-horizon simulation

Multi-agent organizations run across decades of simulated history, scored against what actually happened.

Used in Frontier Lab.

Counterfactual replay

A suspect input is removed and the agents are replayed in pairs with recorded tool results, so a cause is confirmed by experiment rather than inferred from a trace.

Used in Causeway.

Ground-truth oversight tests

Known misbehaviour is planted in real sessions, so an oversight tool’s hit rate can be measured exactly.

Used in the spliced-misbehaviour method.

Research projects

The research behind our solutions. Each project is published in full, with code and data, and every number below comes from its write-up.

See every project and paper on the Research page.

Frequently asked questions

Is Cygnux Labs a research lab or a company?

A research lab first. Our solutions apply the same instruments and methods to teams running agents, and what we learn from that work feeds back into the research.

What does an audit or review cost?

Each one has a fixed scope agreed in writing first, then a single fee for that scope. We discuss it on the intro call once we know what your agents do.

Will our findings be published?

Not without your agreement. Client details stay private. Lessons that generalise may be published in anonymised form, and only if you agree.

Do we need to use your tools?

No. We can work with your existing logging, monitoring and evals. Tracekit, Causeway and our other instruments are open source if you want to adopt them.

Why start with agents that hold money?

When an agent controls a wallet, the cost of a wrong action is concrete and measurable, which makes it a good first test bed for guarantees that apply to agents in general.

How do I propose a research collaboration?

Send a short note with the question you want answered and what you have tried so far. We reply to every message.

Start a conversation

Running agents you need to be able to trust?

Book an intro call, or describe what your agents do and what you need to verify.