Agent Autopsy

Dead-box forensics for the agentic workforce

Your agent can't testify.
Your evidence can.

Prompt injection. A poisoned memory. A stolen key. However it starts, it ends in the same room: counsel asking what the agent actually did. Agent Autopsy answers with real forensics — one fused timeline across every log source, the attack chain correlated end to end, and root cause walked back to patient zero. Sealed so it holds up.

Agentless · Vendor-neutral · Evidence never leaves your environment

One incident, fully reconstructed from real engine output. Decode the evidence yourself — no signup, no form.

One actor, stitched across four log sources — with every link labeled attested or inferred.

Reads the logs you already have 21 connectors
Microsoft 365 AWS CloudTrail Azure Activity GCP Audit Entra ID Okta Ping Identity GitHub Jira ServiceNow Slack Anthropic API OpenAI Audit LLM Gateways LangSmith Traces MCP Traces OpenInference / OTLP Agent Memory Database CDC Network Telemetry Falco auditd Sysmon
The gap

Agent monitoring only helps if you were already recording. Agent Autopsy reconstructs the incident from the logs you already have.

Every tracing gateway and replay tool earns its keep — but only if you bought it, deployed it, and aimed it at the right agent before anything broke. The breach doesn't wait for that. Most teams are left with what they already have: ordinary logs, scattered across providers, and counsel already asking how it happened.

If you saw it coming

Agent frameworks with built-in tracing. Gateways that replay every call. Real coverage — of the agents you set out to watch. But the one that goes rogue is rarely the one you were watching: it's the integration you forgot, the agent that shipped without anyone wiring it up, the path no one thought to cover. And even where the trace exists, it watches the agent — not the blast radius. What got touched in Microsoft 365, in the cloud, at the identity layer lives outside the replay. Telemetry is built for operations, not for evidence.

When you didn't

An agent with production access was compromised — and nothing was watching. What's left is the ordinary trail, scattered across four providers: Microsoft 365 audit logs, CloudTrail, identity-provider logs, model-provider exports. Agent Autopsy fuses them into a single account of what the agent read, called, changed, and decided — then walks it back to the one event that set everything in motion. Not a log search. A reconstruction that holds up when counsel asks how it happened.

What it does

From scattered logs to a defensible account.

Fused timeline reconstruction

Every agent action — model calls, tool use, sign-ins, file access, sends — normalized and fused into one time-ordered timeline across all your sources, filterable down to just the reads, the writes, or the egress.

Cross-source identity resolution

One agent is many identities — an API key here, a service principal there, a role ARN somewhere else. Agent Autopsy stitches them into a single actor across model-provider, cloud, identity, and M365 logs, and renders the investigation graph. At the host boundary it asks instead of guessing: correlating into host telemetry takes your asset inventory — without one, host findings stand alone, labeled as such.

Attack-chain correlation

Findings correlate into kill-chain phases — privilege escalation, collection, lateral movement, exfiltration, impact — with multi-agent cascade detection, incident scoring, and time-separated memory-poisoning chains: write → retrieve → act.

Patient-zero root cause

A causal walk-back along attested links — parent lineage, memory provenance — to the earliest enabling event. Confidence-rated, and honest: when the evidence is only inferential, it says so instead of guessing.

The adversary's agent

When the attacker brings their own agent, its reasoning never touches your logs — but its actions do: processes spawned, files written, connections opened. Falco, auditd, and Sysmon captures are ingested as sealed evidence, and every OS principal is typed from the record — NHI, human, or unknown. Never guessed.

Evidence-grade chain of custody

Every event lands in an append-only, SHA-256 hash-chained ledger. Seal a case into a portable bundle, sign it with Ed25519, and anyone — counsel, an insurer, a court — can re-verify it independently. Verification is free, forever.

Ask the case

Natural-language questions — "how did this start," "what did it touch" — answered by deterministic queries over the timeline, findings, and chains. Every answer carries the exact events that support it. No model in the loop.

Anti-forensics & readiness

An agent tampering with its own trail is a finding, not a gap. And what the evidence can't answer is reported precisely — with the minimal instrumentation that would make the next incident fully reconstructible.

Reports built for handoff

Examiner reports for counsel and the board. EU AI Act Art. 73 and Art. 12 regulatory artifacts. STIX 2.1 and raw JSON for your SIEM. The product of a forensic tool is the artifact you hand off — so it exports all of them.

How it works

Three steps. No scripting. Nothing leaves your environment.

  1. Collect from logs you already have

    A guided wizard scopes your incident and generates a tailored acquisition plan — exact exports, access roles, time budgets — for each platform. One-shot after an incident, or continuous standing custody before one. 21 connectors · cloud · identity · SaaS · dev tools · model providers · agent traces · host telemetry — JSONL or CSV exports

  2. Reconstruct and correlate

    Drop the exports in. The engine normalizes, fuses the timeline, resolves identities across sources, builds attack chains, and walks back to patient zero — labeling every link attested or inferred. 36 detection rules · OWASP Agentic · MITRE ATT&CK mapped

  3. Hand off evidence that holds up

    Export the examiner report, seal and sign the bundle, deliver the regulatory artifacts. Anyone can verify the chain of custody without trusting you — or us. Sealed .case · Ed25519 signed · independently verifiable

Validation

Validated against a 3,240-case known-truth benchmark.

A forensic tool makes consequential claims, so we measure ours — against generated incidents with planted ground truth the engine never sees, across graded levels of evidence completeness.

False alarms

0 / 1,500. Benign cases seeded with suspicious-looking-but-clean activity that produced a critical or high finding: zero. A tool that cries breach on a quiet week is finished.

Root-cause recall

100% on attested evidence. The walk-back landed on the planted root — exact event and actor — in all 720 attested-cohort cases. Where attestation doesn't exist, it misses honestly instead of guessing.

Confidence calibration

720 / 720 dropped. Remove the attested link and every single case's confidence fell below high. Zero held a confident answer on weakened evidence.

Not circular

We proved the benchmark isn't grading its own homework: rename everything the engine's name-matching keys on, and core detection holds — because it's behavioral, not fingerprint recognition.

The layer above

Every gateway makes us better.

Agent Autopsy is not a gateway, a guardrail, or a firewall. It never sits in your request path, never blocks a call, and never asks you to deploy anything before the incident.

That position is deliberate. An enforcement tool can only testify about the traffic it carried. An investigation is about everything else — the agent that bypassed the proxy, the credential used before the control existed, the adversary who brought their own AI.

It also means every enforcement tool you deploy makes Agent Autopsy stronger. Gateway decision logs, guardrail verdicts, proxy records — all of it is evidence, and better records mean sharper reconstruction. When your AI gateway improves its logging, our timeline improves with it. The reverse is not true.

Nothing in your request path · Zero added latency · Gateway & proxy logs ingested as evidence

Agents recommend.
Humans authorize.

Agent Autopsy never asserts what the evidence can't support. Every conclusion carries its confidence, every link is labeled attested or inferred, and ambiguity is stated — not papered over. The examiner draws the conclusions; the tool's job is to put defensible evidence under them, fast. In the benchmark, that honesty is enforced: zero confident-wrong root causes across all 3,240 cases.

Why this was built

The name is deliberate. In digital forensics, when an incident is already over and you need to know exactly what happened on a machine, you don't watch it live — you examine it after the fact. Dead-box forensics. The tool a generation of examiners learned that discipline on is Autopsy. Agent Autopsy carries the same idea into the agentic era: a post-incident examination of what an autonomous system actually did, held to the same evidentiary standard.

Most of the market answers that with more monitoring you'd need running before anything broke, or more AI that sounds certain and is sometimes wrong. Confidently wrong. Agent Autopsy takes the opposite bet.

The result is an account an examiner can take to counsel and defend, line by line. Evidence-grade ground to stand on — for the team the incident actually hit.

Built by Todd E. Nelson — incident-response executive · author of The First 72 Hours

Get Agent Autopsy

Run it on your own logs.

Agent Autopsy ships as a standalone binary — no agents to deploy, no data to send us, nothing to install in production. Licensing is straightforward and conversation-first: tell us about your environment and we'll get you running.