Most companies deploying AI agents can answer "what did the agent produce?" Fewer can answer "what did it actually do?" Fewer still can answer "why did it do that?" in a way that would hold up in front of a customer, an auditor or their own security team.
Cygnux Labs builds two open-source tools for those questions. Tracekit keeps a signed, tamper-evident record of what an agent did. Causeway works out which input made a multi-agent system act, by replaying the run without each suspect. This guide is written for the people deciding whether and how to use them: engineering leads, security and risk teams, and founders running agents in production.
The short version
- Agents act, and their own logs are weak evidence. A log the agent's process can rewrite, or a summary the model wrote about itself, will not settle a dispute.
- Tracekit proves what happened. Every tool call is signed by a separate process, hash-chained, checkpointed to an outside witness and exported as a bundle anyone can verify offline. A policy gate blocks known-dangerous calls before they run.
- Causeway explains why it happened. It records exactly what each model call saw and how agents passed information to each other. When something goes wrong, it replays the run with each suspect input removed and measures which one changed the outcome.
- They answer different questions and work as layers. Tracekit is the evidence layer and Causeway is the analysis layer. Writing Causeway's records through Tracekit's signer is on the roadmap.
- Both are open source and self-hosted. Your traces stay on your infrastructure. Neither tool needs a hosted service to work.
- Both are early. Tracekit is a v0.2 release candidate and Causeway is a v0.3 alpha. The sections below say plainly what is ready and what is not.
The problem businesses are running into
A year ago most AI in a business answered questions. Now agents fix code, triage tickets, issue refunds, update records and call other agents. Each of those steps is a tool call with real effects.
When everything works, nobody looks at the record. When something goes wrong, three questions come up quickly:
- What exactly happened? Which agent did it, with which arguments, after which steps?
- Can we prove it? Is the record in front of us complete and unaltered, or could the agent, a bug or a person have changed it?
- Why did it happen? Which document, message or other agent led to the action, and was that input really the cause or just present?
Standard application logs and most agent tracing tools answer the first question reasonably well. They struggle with the other two. The log is usually writable by the same process that is being audited. And when an agent has read ten documents before it acts, every one of them is "connected" to the action. Pointing at the suspicious-looking one is a guess.
These gaps cost real time and money. An incident review that cannot establish the cause ends with a vague fix. A customer who asks "how do you know your agent didn't touch our data?" gets an answer built on a log nobody can verify. A security team that cannot see what the agent read cannot tell a prompt injection from an ordinary bug.
Two questions, two tools

Tracekit: a record you can show to someone else
Tracekit sits between an agent and the actions it takes. It records each tool call, applies policy before the call runs, and stores the result in a ledger the agent cannot quietly change.

What it does
- Separate signer. On Linux in system mode, the signing service runs as its own operating-system user. The agent never holds the signing key, so it cannot forge or rewrite records with it.
- Signed hash chain. Each record carries the hash of the previous record, a sequence number and an Ed25519 signature. Editing, deleting, reordering or forging a record breaks verification and names the first bad record.
- External witness. The head of the chain is checkpointed to a git repository or file outside the machine. That catches the case a chain alone cannot: someone rebuilding the entire log.
- Policy gate and approvals. Versioned rules with stable IDs block known-dangerous calls before they run, such as uploading a
.envfile, piping a download into a shell or force-pushing. "Ask" rules hold a call until a different user approves it. - Transcript and model cross-checks. Tracekit hashes the agent's session transcript at every step, so deleting or editing earlier parts of the conversation is reported. An optional local model proxy compares what the model requested with what the hooks recorded, which shows if hooks were disabled or bypassed.
- Offline evidence bundles. A run exports as a single file containing the events, policy snapshots, checkpoints and a replay page. Anyone with the public key can verify it without trusting you or Tracekit. It can also export OpenTelemetry data for tools like Jaeger.
- Redaction by default. Secrets are masked before anything is written, and content is hashed by default. The model's self-reported reasoning is optional, off by default and never treated as evidence.
Claude Code works out of the box through hooks. Custom agents can use the Python SDK, and there are adapters for LangChain and LangGraph tools. Agents on other machines can send events to an authenticated ingestion gateway.
Where it helps a business
Engineering teams using coding agents on real repositories. Agents work in repositories that hold credentials, deployment scripts and customer code. Tracekit blocks the obvious dangerous calls, holds risky ones for a human, and keeps a record of everything else. When someone asks "did the agent run anything against production?", the answer is a verifiable file instead of a recollection.
Security teams. Agents read text they did not write: READMEs, code comments, web pages, test output. Any of it can carry instructions aimed at the agent. Tracekit's alerts surface blocked calls, held calls, capture gaps and transcript tampering in one place, and the live observer shows every agent and subagent as it works.
Companies that have to show evidence to others. Agencies and contractors that use agents on client code, vendors whose customers ask about AI controls, and teams preparing for an audit all need something better than "trust our logs." A Tracekit bundle can be handed to a third party and checked offline.
Incident response. After an incident the first job is to establish the facts. A signed, witnessed record removes the argument about whether the log is complete.
What Tracekit does not do
Tracekit is careful to state its limits, and businesses should plan around them:
- It records only what passes through its integrations. It does not discover agents or monitor the whole system automatically.
- The policy rules are tripwires. They stop honest mistakes and obvious cases. A determined agent can evade pattern matching, so hard enforcement belongs in sandboxes and operating-system controls, with Tracekit as the independent witness.
- It cannot see effects below the tool-call boundary, such as a script that opens its own network connection.
- Full isolation (system mode) is supported on Linux and experimental on macOS. Dev mode runs anywhere but only checks integrity, and the verifier says so.
Causeway: finding out why
Causeway is built for the question that comes after "what happened": which input actually caused it.
Consider a support system with three agents. A researcher reads a ticket, the refund policy, a note from a vendor portal and a shipping FAQ. A planner decides what to do. An executor does it. One run ends with the executor emailing the full customer list to an outside address. Every document the researcher read is upstream of that email, so a trace shows all of them "reaching" it. A trace cannot tell you which one caused it.
Causeway answers the way you would test any causal claim: remove the suspect, run the system again, and see whether the action still happens.

What it does
- Causal log. An SDK records structured events: inputs with trust labels (trusted policy, untrusted web page), each model call with the exact content it saw, each tool call linked to the decision that asked for it, and messages between agents. Every value is stored under its content hash and the log is hash-chained.
- Investigation app. A browser app answers questions in order: what happened (timeline), what needs attention (alerts), why a given action happened (lineage, where each argument came from, candidate causes) and how the agents depend on each other.
- Safe replay. A run can be re-executed with tool results served from the recording. Calls the original run never made are stubbed, so a replay cannot send a real email or move real money.
- Causality tests. Causeway removes one input or one message channel, replays the run many times with paired seeds and reports how often the action happened with and without it, with a 95% confidence interval. Each candidate is marked causal, ruled out (with a bound on how large a hidden effect could be) or untested.
- Multi-agent evaluation. It checks whether each message is actually used by the agent that receives it, which channels change outcomes when cut, and where the critical path runs.
- Influence graph across runs. Over many runs it builds an index of which content gets used by decisions and which of those links were confirmed by tests.
- Alerts for injection-shaped flows. A sensitive tool using an argument value that appears only in untrusted content is flagged as high severity. That is the classic signature of prompt injection.
In the demo above, replaying without the vendor note drops the outside email from 80% of runs to 0%. Replaying without the FAQ or the ticket changes nothing detectable. Cutting the researcher-to-planner channel also stops it, which tells the team which hop to guard. The demo uses seeded mock models, so those numbers show the method works, not how any particular real model behaves.
Where it helps a business
Root cause analysis that ends in the right fix. "The agent was prompt-injected" is not an actionable finding. "This vendor note, passed from the researcher to the planner, caused the email in 80% of replays, and the FAQ had no detectable effect" is. Teams can fix the actual input path instead of adding broad restrictions that slow everything down.
Customer-facing and operations agents. Support, refunds, account changes and back-office workflows are where multi-agent systems touch money and customer data. Causeway shows which inputs drive sensitive actions, and it can confirm that a refund depends on the refund policy, which is what you want.
Third-party content risk. Businesses increasingly feed agents content from vendors, partners, customers and the web. Trust labels and alerts show where untrusted content reaches sensitive tools, and the cross-run influence graph shows which sources keep driving decisions.
Designing and evaluating multi-agent systems. Before scaling a system, teams want to know whether each agent is pulling its weight. Causeway shows which messages are used, which channels matter and where a single compromised hop could cause harm. That makes it useful during development as well as after incidents.
Bringing evidence to a conversation with customers or leadership. A causal claim with a measured effect and an interval is easier to defend than an engineer's best guess.
What Causeway does not do yet
Causeway v0.3 is an alpha. It works and is tested, and the gaps matter for production use:
- The app and read API have no authentication yet and bind to localhost by default. There are no users, roles or multi-tenancy.
- Storage is plain files. Scale will need a database and object store.
- The Anthropic adapter is built. OpenAI SDK, LangGraph and OpenTelemetry import are on the roadmap.
- The benchmark currently uses a simulated model. Real-model results are an open issue.
- Replay costs model calls, roughly two per trial per downstream model call, and there are no budget controls yet.
- Its own hash chain detects edits but is unsigned. Anyone who can rewrite the whole log can rebuild it. Signing through Tracekit is planned.
- A causal effect holds for that program on that task. It does not explain the model's internal reasoning and may not transfer to other tasks.
Using them together
The two tools fit as layers. Tracekit establishes that the record is real. Causeway establishes what the record means.
A realistic incident looks like this:

- Detection. Causeway raises a high-severity alert: a
send_emailcall used an address that appears only in untrusted vendor content. - Facts. The team pulls the Tracekit bundle for the run and verifies it against the witness. The record is intact, the policy that applied is bound to every decision, and nothing in the transcript was removed.
- Cause. In Causeway, the team replays the run without each untrusted input. The vendor note is confirmed as the cause and the other inputs are ruled out within a stated bound.
- Fix. The team strips instructions from vendor content before it reaches the planner, adds an
askrule in Tracekit for outbound email to new domains, and re-runs the test to confirm the effect is gone. - Evidence. The bundle and the test results go into the incident report. Anyone reviewing it can check both.
Today the two tools run side by side, each with its own chain. The planned integration writes Causeway events through Tracekit's signer, so every causal claim points at a record nobody could quietly alter.
A practical way to start

Neither tool requires a large rollout. A sensible pilot takes a few weeks.
- Pick one agent workflow with real stakes. A coding agent with repository access, or a support workflow that touches refunds or customer data.
- Run the demos first.
tracekit demoandcauseway demoneed no API key and show the full loop on a sample scenario in a few minutes. - Install Tracekit on the coding agent. Start in dev mode to see the data, then move to system mode on Linux with a git witness for real use. Begin with the default policy and add your own rules with
extends: default. - Instrument the multi-agent workflow with Causeway. Label inputs as trusted or untrusted and declare which tools are sensitive. That declaration drives the alerts.
- Decide who approves held calls. Approvals in Tracekit must come from a different user than the agent's, so the approval path needs an owner.
- Anchor every session. A chain without an outside anchor only proves internal consistency. Witnessing at the end of every session is the minimum.
- Rehearse one incident. Take a recorded run, verify the bundle, run a Causeway test on a suspect input and write the report. A team that has done this once will do it much faster under pressure.
Open source, self-hosted, offline-first
Both tools are open source (Tracekit under MIT, Causeway under Apache-2.0) and run on your own infrastructure. Traces never have to leave your environment. Evidence bundles verify offline. Causeway has no runtime dependencies, and Tracekit needs only a standard cryptography library.
For businesses that handle sensitive data, that matters. An oversight tool that ships every agent action to someone else's cloud creates the kind of exposure it is supposed to reduce. The same design principle runs through all Cygnux Labs work: systems should keep working, and keep their guarantees, without depending on an outside service.
Where this is going
The next steps for both tools come from what businesses need in production: a gateway that records any agent by changing one base URL, adapters for more frameworks, authentication and multi-tenant storage for Causeway, real-model benchmark results, and the signed integration between the two tools. Each item is tracked publicly on GitHub.
Both repositories are public: Cygnux-Labs/Tracekit and Cygnux-Labs/Causeway. If you are running agents in production and want help piloting either tool, or want to tell us what your team needs from agent oversight, book a call.
