What to Demand in an Agent Vendor's SOC 2 and Audit Trail
The specific SOC 2 scope, audit log fields, and evidence artifacts to require from an agentic AI vendor before you sign, and the clauses that are vendor theatre.

Most agent vendor security reviews end the same way: a SOC 2 Type 2 report lands in your inbox, someone in procurement notes that it exists, and the deal moves forward. That check is close to meaningless for autonomous action execution. A chatbot's worst failure is a wrong sentence. An agent's worst failure is a wrong wire transfer, a mispriced contract, or a deleted production table, and none of those risks are addressed by a report whose scope stops at the vendor's marketing site.
The question at final vendor evaluation is narrower than "do they have SOC 2." It is whether the report's system boundary actually contains the agent runtime, whether the audit trail records the fields you would need to reconstruct an incident, and whether the evidence the vendor can hand you tomorrow morning would survive contact with your own internal audit team. What follows is the checklist to run before signing.
Read the System Description, Not the Logo
A SOC 2 Type 2 opinion is only as useful as the system description it covers. The system description outlines the scope, boundaries, infrastructure, and relevant policies that were reviewed, and everything outside that boundary carries no assurance. For an agentic AI vendor, the boundary matters more than the logo on the cover because the report often covers a corporate SaaS product while the agent runtime, model gateway, or tool execution layer sits in a separate environment.
Ask the vendor to point, in the report itself, to the components that execute agent actions on your behalf. Then check three specific things:
- Cloud environments in scope. If a vendor runs on both AWS and GCP but only AWS is described, the GCP environment is unexamined. Agent vendors frequently split inference and orchestration across providers.
- Subservice organizations. Model providers (Anthropic, OpenAI, Bedrock), vector stores, and orchestration frameworks are typically carved out. Confirm which controls are the vendor's and which are inherited, and get the subservice SOC reports.
- Trust Services Criteria selected. Security is the only category required in every SOC 2 audit, and Availability, Processing Integrity, Confidentiality, and Privacy are optional. For an agent that executes financial or clinical actions, Processing Integrity is not optional in any real sense — the absence of it in scope is a finding of its own.
Type matters as much as scope. SOC 2 Type 1 evaluates whether controls are designed properly at a point in time, whereas SOC 2 Type 2 evaluates whether controls are designed and functioning effectively over a period of time. A Type 1 for an agent vendor is a design document, not evidence. Do not accept it as a substitute in a final evaluation.
Ask Which Common Criteria Actually Apply to Agents
The Common Criteria are not equally relevant to an agent product. The SOC 2 Common Criteria (CC1–CC9) are derived from the 17 principles of the COSO Internal Control - Integrated Framework (2013), with supplemental criteria for logical access, system operations, and change management. For an autonomous action vendor, three sections carry disproportionate weight and deserve line-by-line reading before signing.
CC6 (logical and physical access) is where you check that the agent's identity is not a shared service account with production write privileges. CC7 (system operations) is where log monitoring lives — the section that determines whether the vendor can actually see an anomalous agent action within a useful window. CC8 (change management) is where model updates, prompt updates, and tool registry changes are governed. If the vendor cannot show controls that treat a new prompt as a change subject to review, their agent is a moving target and the audit report is describing a snapshot that no longer exists.
One useful trick from experienced SOC 2 readers: a SOC 2 report with no exceptions, no qualifications, and no noted deviations is often more concerning than one with a few documented issues. Real programs surface access review misses and late change tickets. A pristine report over a twelve-month period, on a product that ships agent updates weekly, is worth a follow-up question.

Specify the Audit Log Fields You Will Not Ship Without
A generic access log is not an agent audit trail. Application logs were built for deterministic software where recording the input, state change, and result reconstructs the run. Agents break that model in two places: they usually run under a shared service account with no link to the human who set the task, and the action itself was picked at runtime by a model rather than fixed in code.
Anchor the requirement to a standard so the conversation is not a negotiation. An IETF draft defines the Agent Audit Trail (AAT), a JSON-based record structure with mandatory fields for agent identity, action classification, outcome tracking, and trust level reporting, using tamper-evident hash chaining via SHA-256. Whether or not the vendor implements AAT specifically, the field set is a reasonable floor. The AI vendor security questionnaire should demand, at minimum, the following per action record:
- Agent identity and version. A stable identifier for the agent client, the model version, the prompt version, and the tool registry version at the moment of execution. Model version alone is not enough — the same model with a new system prompt is a new agent.
- Delegation chain. The human principal, the agent principal, the session ID, the scope granted, and whether the specific action was auto-approved or gated by a human. This is what makes accountability recoverable after the fact rather than speculative.
- Tool call detail. The tool invoked, the arguments (with PII hashed or redacted), the target system, the idempotency key, and the returned status. "API call succeeded" is not sufficient.
- Policy evaluation. Which guardrail or policy fired, what it decided, and why. If an approval gate was bypassed, the record must state that it was and on what basis.
- Outcome and downstream effect. The observable state change in the target system, linked to the tool call by a correlation ID that can be traced from the agent log to your own system of record.
- Tamper evidence. Hash chaining or equivalent so that a modified record is detectable, not just discouraged by policy.
Retention has a floor as well. Article 12 of the EU AI Act requires that high-risk AI systems technically allow for the automatic recording of events over the lifetime of the system, and Article 26 requires deployers to keep automatically generated logs under their control for at least six months unless a longer period applies. Six months is the regulatory minimum for European high-risk use cases, not a target. For financial or clinical action execution, most internal audit teams will want years, aligned to the record retention rules of the underlying process. Confirm who holds the logs, in what format, and how you get them out if the relationship ends. Our own approach to action execution keeps the primary log inside the client's environment for exactly this reason.
Demand Evidence Artifacts, Not Attestations
An attestation letter is not evidence. The artifacts a serious agent vendor can produce in a due diligence pack, without a two-week delay, tell you more about their posture than the SOC 2 cover page.
- The full Type 2 report under NDA, with the system description, the management assertion, the auditor's opinion, the detailed control testing, and any exceptions. A summary or SOC 3 is not a substitute.
- A subservice organization matrix naming every carved-out provider, the controls inherited, and the corresponding SOC 2 or ISO 27001 reports for each.
- A sample audit log export for a representative agent run, redacted, in the format you will receive in production. This is the single most predictive artifact. If they cannot produce one, the logging is not real.
- The change management record for a recent model or prompt change, showing the ticket, the reviewer, the test evidence, and the rollout path. This maps to CC8 and to your own agent versioning discipline.
- The approval gate configuration for the workflows you plan to run, including which action classes are auto-executed and which require a human. Our post on which agent actions actually need human approval is one starting point for what the tiers should look like.
- Log management practice. NIST SP 800-92, the Guide to Computer Security Log Management, remains the authoritative NIST reference and maps to SOC 2 CC7.2 for monitoring of security events. Ask which of its practices the vendor implements, specifically for the agent runtime rather than the corporate SaaS.
For anything running on sensitive data, extend the pack to include deployment topology. A vendor claiming air-gapped or VPC support should be able to produce an architecture diagram showing where model calls terminate, where logs are written, and which components ever egress. If the answer involves a hosted control plane phoning home with action metadata, the compliance story is different from what the sales deck implied. Our guide to air-gapped vs VPC vs SaaS deployment covers the trade-offs in more detail.
What Real Agent Vendor Due Diligence Looks Like at Signing
The pattern that separates a real agent vendor security posture from a performative one is boringly consistent. The system description names the agent runtime. The Common Criteria selected include Processing Integrity. The audit log fields match the AAT-style floor above. The retention policy exceeds six months and can be enforced inside your environment. The evidence pack arrives in days, not weeks, and includes a sample log export you can actually parse.
None of this eliminates risk. It does move the residual risk from "we hope the vendor is careful" to "we can reconstruct any action the agent took, attribute it to a principal, and prove or disprove that it followed policy." That is the standard autonomous agent compliance evidence has to meet, because the alternative is a Type 2 report that describes a different system than the one taking actions in your production estate.
Run the checklist before signing. The vendors who pass it will thank you for the rigor, because it is the same rigor their own engineering teams are already applying internally. The ones who cannot pass it are the reason final evaluations exist.
Eric Lamanna is a Digital Sales Manager with a strong passion for software and website development, AI, automation, and cybersecurity. With a background in multimedia design and years of hands-on experience in tech-driven sales, Eric thrives at the intersection of innovation and strategy—helping businesses grow through smart, scalable solutions. He specializes in streamlining workflows, improving digital security, and guiding clients through the fast-changing landscape of technology. Known for building strong, lasting relationships, Eric is committed to delivering results that make a meaningful difference. He holds a degree in multimedia design from Olympic College and lives in Denver, Colorado, with his wife and children.
Put an agent to work, the right way.
Start on Automatic and put the workflow you want to automate in front of engineers who have shipped agents in regulated environments.
Agentic AI, in your inbox.
Occasional, high-signal notes on building and operating AI agents — automation patterns, architecture, and governance. No spam.


