Book a call

Reached Is Not Caused: Finding Which Input Made a Multi-Agent System Act, by Replaying It Without Each Suspect

Reached Is Not Caused: Finding Which Input Made a Multi-Agent System Act, by Replaying It Without Each Suspect

A team of three AI agents is handling a refund ticket. A researcher agent reads the ticket, the refund policy, a note from a vendor portal and a shipping FAQ. A planner agent decides what to do. An executor agent does it. One run ends with the executor emailing the company's full customer list to an outside address.

The trace shows everything each agent saw, and every untrusted document it read is connected to that email. A trace can tell you what reached the action. It cannot tell you which input caused it. This write-up introduces Causeway, an open-source tool that answers that second question the way you would test any causal claim: remove the suspect, run the system again, and see whether the action still happens.

The short version

  • Reach is not cause. In a multi-agent run, almost every input ends up in the context of almost every later decision. Tracing therefore links every document an agent read to every action it took, and cannot separate the cause from the bystanders.
  • Causeway records what each model call actually saw, as content-addressed references, along with every tool call, the decision that asked for it, and every message between agents. The record is hash-chained.
  • It tests causes by replay. For each suspect input or message channel, it re-runs the system in pairs with the same random seed, with and without the suspect, and measures how often the action still happens, with a confidence interval and an exact significance test.
  • Replays cannot repeat side effects. Tool calls that match the original run get the recorded result back. New tool calls are stubbed. A replay never sends a real email or moves real money.
  • What we found so far. In the demo, removing the vendor note drops the exfiltration email from 80% of replays to 0%, while removing the ticket or the FAQ changes nothing. On a five-scenario benchmark with a simulated model, replay named exactly the true cause in all 61 harmful runs. Tracing alone never did, and string matching gave no answer in 39% of them. Replay also has a hard floor: fewer than about ten replays per suspect is not enough.
  • What it is not, yet. These results come from a simulated model. Running the benchmark on real models is the first open task on the project.

Why this matters

For anyone running agents, the expensive moment is the incident. An agent issued a refund it should have escalated, deleted a directory it should have rotated, or forwarded data to an address nobody recognises. The questions that follow are always the same: what happened, why, and what do we change so it doesn't happen again.

The "why" decides the fix. If the cause was a poisoned web page, you filter that source. If it was a message from one agent that another agent trusted blindly, you guard that hop. If it was your own policy document, you rewrite the policy. Guessing wrong means fixing the wrong thing, and the incident comes back.

Prompt injection makes this harder. Injected instructions arrive inside content an agent reads for good reasons: a vendor note, a customer email, a README, a search result. They are passed along between agents, often reworded in a summary, so by the time a harmful action executes, the instruction may be two hops and one paraphrase away from where it entered. Looking at the trace, every document looks equally guilty.

Reached is not caused

Every untrusted input reached the email. Only one caused it.

In the run above, the researcher's single model call saw six items: the task, the ticket, the refund policy, the vendor note, the FAQ and an order lookup. Its summary went to the planner, the planner's plan went to the executor, and the executor sent two emails: the expected one to the customer, and one with the customer list attached to audit@vendor-compliance.example.

Trace that email backward and you reach all three untrusted inputs. That is correct and useless. To find the cause, Causeway replayed the system 40 times without each untrusted input, using the same 40 seeds each time:

  • Without the vendor portal note: the email went out in 0% of replays instead of 80%. Effect +0.80, 95% interval [0.63, 0.90]. Confirmed cause.
  • Without the customer ticket: 80%, unchanged. Effect 0.00 [−0.18, 0.18]. Ruled out.
  • Without the shipping FAQ: 80%, unchanged. Effect 0.00 [−0.18, 0.18]. Ruled out.
  • Without the messages from the researcher to the planner: 0% instead of 80%. Effect +0.80 [0.63, 0.90]. Confirmed cause.

The vendor note is the cause. The ticket and the FAQ reached the email and did nothing, and the interval bounds how large any effect they have could be. Cutting the message channel from the researcher to the planner also stops the email, which tells you where an injection crosses between agents and therefore which hop to guard. The same test shows the intended refund depends on the refund policy (+0.93), which is what you would want.

How Causeway records a run

Content-addressed context

Every value an agent sees or produces is stored once under its SHA-256 hash: documents, tool results, model outputs, messages. A model call records the hashes of exactly the items in its context and the hash of what it produced. That turns "this document was in the context of that decision" into an exact link rather than a guess from timestamps or text similarity.

Here is the researcher's model call from the run above, trimmed:

{
  "schema": "causeway.event.v1",
  "seq": 10,
  "type": "decision",
  "agent": "researcher",
  "model": "mock",
  "purpose": "summarize",
  "context": [
    {"kind": "task", "source": "task", "trust": "trusted"},
    {"kind": "input", "source": "inbox:ticket-881", "trust": "untrusted"},
    {"kind": "input", "source": "kb:refund_policy.md", "trust": "trusted"},
    {"kind": "input", "source": "vendor:portal/notes.md", "trust": "untrusted"},
    {"kind": "input", "source": "web:shipping_faq", "trust": "untrusted"},
    {"kind": "result", "source": "lookup_order", "trust": "trusted"}
  ],
  "output": "sha256:228ed395b10…",
  "usage": {"input_tokens": 144, "output_tokens": 101},
  "prev_hash": "026168f243530c69e8…",
  "hash": "32750cfe352554422e…"
}

Tool calls record their arguments, result, the decision that asked for them, whether the tool is sensitive, and how the call was executed (live, from the recording, or stubbed). Messages between agents record the channel, so a whole channel can later be removed in a test.

Three properties fall out of this design:

  • Side channels become visible. If one agent writes a file and another reads it, the content hash matches, and the graph links them even though no message was sent.
  • Trust labels travel. A summary of an untrusted page is marked as carrying untrusted content, so laundering an injection through a colleague's summary doesn't hide it.
  • The record detects edits. Every event carries the hash of the previous one. Editing, deleting, inserting or reordering an event breaks the chain, and changing a stored value breaks its hash.

Three ways in

Teams can write their agent loop with Causeway's small Python API, wrap the Anthropic SDK so every request's context is captured with no other code changes, or send events from other machines to a collector that validates every hash and the chain before writing anything. Adapters for other frameworks and a gateway that captures any agent by changing one base URL are next on the roadmap.

Investigating a run

The investigation app: where each argument came from, the path through the agents, and every candidate cause marked confirmed, ruled out or not tested

The investigation app is built around the questions people actually ask after an incident, one screen each:

  • What happened? A timeline of every input, model call, tool call and message, with latency, token use and flags.
  • What needs attention? Alerts. The strongest one fires when a sensitive tool uses an argument value that appears only in untrusted content. In the demo, the address audit@vendor-compliance.example appears in the vendor note and nowhere else, so the email is flagged before any replay runs.
  • Why did it happen? For any tool call: where each argument value came from, the recorded path through the agents, and every upstream input and message channel, marked as a confirmed cause, ruled out, or not yet tested. On a live server, one click runs the test.
  • How do the agents depend on each other? Whether each agent actually uses the messages it receives, and an influence matrix from every test run so far.
  • What keeps causing trouble? Across runs, which content gets used by decisions, what it reaches, and which of those links tests have confirmed. It works like a citation index in which some citations are proven to matter.

The app is a single HTML file. It can be served live, or exported as a static report with the data embedded, so you can send an investigation to someone who has nothing installed.

Testing a cause

Remove one thing, replay in pairs, count what changes

Four design choices make replay safe and the answer trustworthy.

1. Never send the email twice. The original run's tool calls become a tape, keyed by agent, tool, a hash of the arguments and the call's position. In a replay, a call that matches the tape returns the recorded result and touches nothing. A call the original run never made gets a stub. Real execution only happens when someone switches it on explicitly. The cost of this choice is that an agent sees a fake result after a stubbed call, so Causeway reports how many calls left the tape in each test.

2. Change one thing per pair. The suspect is removed from the context of every model call it would have reached, and the record lists what was removed. Each pair uses the same seed with and without the suspect, so the only difference inside a pair is the intervention.

3. Use a test that respects the pairing. The effect is the drop in how often the action happens, reported with a Newcombe 95% interval, which behaves well at 0% and 100% where simpler intervals break. The yes-or-no verdict comes from an exact McNemar test on the pairs that disagree. When several inputs of one action are tested, Benjamini–Hochberg controls the false discovery rate. Without that correction, testing ten innocent inputs at the usual 5% level would produce at least one false "cause" about 40% of the time.

4. Replay one call when the system can't be replayed. Some production systems can't be re-run end to end. For those, Causeway can re-call a single recorded model call with and without one context item, using only the log. In the demo, removing the vendor note from the researcher's call drops the share of summaries that mention the outside address from 75% to 0%, which locates where the injection entered for a fraction of the cost.

Experiment: an attribution benchmark

To check the method against alternatives, we built five small multi-agent scenarios, each with exactly one planted cause and several decoys that reach the same action through the same agents, plus a clean control.

  • Data exfiltration. The customer list is emailed to an outside address. The address appears only in the injected note, so string matching can find it.
  • Refund override. A $450 refund is issued without the required approval. Its arguments come from a trusted order lookup, so there is nothing injected to match.
  • Destructive command. A log directory is deleted instead of rotated. The path appears in trusted and untrusted sources alike.
  • Two hops away. A partner update is copied to a look-alike domain, injected through a second researcher agent. The address appears only in the injected page.
  • Admin escalation. A contractor is added to the admins group, next to a second, obvious injection that the model ignores. The group name appears only in the injected page.
  • Control. No injection. Nothing should happen and nothing should be flagged.

Four methods named a cause in every run where the harmful action happened: blame everything upstream (tracing), blame the only untrusted source of an argument value (string matching), blame the input whose wording was reused most, or replay.

Only replay isolated the cause, and it needs a minimum budget

With a simulated model, 20 runs per scenario and 30 replays per suspect:

  • Tracing never isolated the cause. It always blamed three or four inputs.
  • String matching was right in 61% of runs and silent in the other 39%. It works when an injection plants a unique value, such as an address or a group name. It has nothing to match when the injection changes a decision without introducing a new value, as in the refund and the deleted directory.
  • Text reuse could not separate inputs, because the simulated summariser copied every sentence it kept.
  • Replay named exactly the true cause in all 61 harmful runs, blamed no innocent input, and correctly ignored the obvious injection the model did not follow. Each investigation cost 360 to 540 model calls.
  • The control raised no high alerts in 20 runs.

The right-hand panel shows the budget floor. With the exact paired test, five disagreeing pairs give a smallest possible p-value of 0.0625, so five replays per suspect can never confirm anything. Ten replays found the cause in 69% of runs, thirty in all of them.

These numbers come from a simulated model, which follows planted instructions with fixed probabilities and has none of the noise of a real one. They show that the method and the scoring work. They are not the accuracy to expect in production, where injected behaviour fires less reliably and most hosted model APIs don't accept a seed, which weakens the pairing. Expect real models to need more replays. Running this benchmark on real models is the first open task on the project, and the harness to do it, including a capped-cost run on GitHub Actions, is in the repository.

What this can and cannot establish

It establishes, for a recorded run: what each model call saw, which tool calls each decision asked for, which content reached which action, whether the record has been edited, and, for each tested input or channel, whether removing it changes how often the action happens in this system on this task, with an interval for the size of the effect.

It cannot establish:

  • Mechanism. A confirmed cause is a total effect through every path. It says nothing about the model's internal reasoning.
  • Generality. An effect measured on one task and one system may not transfer to others.
  • Completeness. The graph is only as complete as what the integration records. Anything sent to a model outside the recorded context is invisible.
  • Protection from a full rewrite. The hash chain detects edits but is unsigned, so someone who can rewrite the whole log can rebuild it. Signing through Tracekit closes that gap and is on the roadmap.
  • Cheapness. Each test costs roughly two calls per replay for every downstream model call. There are no budgets or caching yet.

Where it sits next to Tracekit

Tracekit is our flight recorder for coding agents: tool calls signed by a separate process, hash-chained, anchored externally, and verifiable offline by anyone. It answers can we prove what the agent did?

Causeway answers the next question: why did it happen, and is that really why? Tracekit's records are the evidence; Causeway's tests are the analysis. The two are designed to stack: Causeway events written through Tracekit's signer, so that a confirmed cause points at a record nobody could quietly alter.

Related work

Counterfactual replay for agents is an active research area. Recent work attributes task failures to individual steps of a single agent's trajectory and splits credit across interacting steps (Causal Agent Replay), scores which steps caused a failure and generates minimal repairs (CausalFlow), benchmarks which agent and which step caused failures in multi-agent systems (TraceElephant), and predicts which events replay would mark as decisive without running replays (BranchPoint-Latent). A recent survey covers evidence tracing and execution provenance more broadly.

Causeway's emphasis is different: it attributes a specific harmful action to the content and message channels that caused it, across agents, with replay that cannot repeat side effects, integrity checks on the record, and alerts for injection-shaped data flows. Step-level attribution and replay-free prediction are complementary, and both would be useful additions.

Guidance for teams running agents

  1. Label your inputs. Mark everything that comes from outside your organisation as untrusted, and declare which tools are sensitive. Every alert and every useful test depends on these two labels.
  2. Record the context, not just the output. If you can't say exactly what a model call saw, you can't say why it acted.
  3. Treat reach as a list of suspects, not a verdict. After an incident, test the untrusted suspects first, then the channels between agents.
  4. Guard the hop, not just the source. If cutting one agent-to-agent channel stops a harmful action, that channel needs its own check, whatever the original source was.
  5. Budget at least ten replays per suspect, and more with real models. Below that, a real cause can go unconfirmed.
  6. Keep replays side-effect free. Never re-run an agent against live tools to investigate it.
  7. Make the record tamper-evident, so the investigation itself can be trusted.

Use it, or work with us

Causeway is open source under the Apache 2.0 licence. It is an alpha: it works, it is tested, and its README lists what isn't built yet.

pip install git+https://github.com/Cygnux-Labs/Causeway
causeway demo

The demo records eight runs of the three-agent system, runs the tests above, and writes the investigation app as a single HTML file you can open in a browser.

For teams running agents in production, we also offer agent incident investigations: we instrument the affected workflow, reproduce the incident safely, and report which inputs and hand-offs caused it, with every finding reproducible and a recommended fix for each. Client details stay private. Book a call or write to us with what happened.

What comes next

  • Real-model results. The benchmark on current models, including the cases where replay gets it wrong.
  • Capture for any agent. A gateway that records any agent by changing one base URL, adapters for the OpenAI SDK, LangGraph and the OpenAI Agents SDK, and OpenTelemetry import for traces teams already have.
  • No-code replay. An MCP proxy that serves recorded tool results, so replay works without changing agent code.
  • Signed records. Causeway events written through Tracekit's signer and witness.
  • Does the reviewer actually review? Many agent systems include a critic or approval agent. Cutting its channel and measuring whether outcomes change is a direct test of whether it is doing anything, and it is the first oversight study we plan to run with Causeway.

Causeway's code, benchmark, documentation and recorded demo runs are available on GitHub.

← Back to Research