How to Audit What AI Agents Did (Logs, Evidence, SOC 2)

Last updated ·Read as Markdown
Answer

To audit what an AI agent did, open the session in your gateway logs and read its timeline: which provider handled the request, which tool was called, with which arguments, and what came back. In Metorial, that is Integrations, then Connection Logs, then Sessions or Tool Calls. For a SOC 2 review, treat the logs as supporting evidence for access and monitoring controls. They do not make a company compliant, and an auditor still needs your policies and proof that someone reviewed the records.

When someone asks what an agent did last Tuesday, the answer has to come from a record, not from the agent. This page covers how to pull that record, and what it can and cannot prove to an auditor.

If the question is
Look at
What happened in one conversation
The session and its timeline
Every use of one tool, across sessions
The Tool Calls table, filtered
What failed
Tool Errors
What people, agents, and service accounts did across the platform
Audit logs
How long records are kept
Data retention settings

What do you need before you start?

  • Roughly when the activity happened, and ideally the tool or the person involved.
  • Agent traffic that goes through a gateway. Calls made directly from a laptop to a local server leave no central record.
  • A decision on who reviews the logs and how often. That decision matters more than any filter.

How do you trace what an agent did?

1. Open the connection logs. In the Metorial dashboard, select Integrations. Under Connection Logs, select Sessions.

2. Find the session. Each row shows the session name, connection status, provider count, and creation time. Match the creation time to the activity you are investigating.

3. Open the session. Select its name. If it has more than one connection, select the right one under Recent Connections.

4. Confirm the context. Expand Connection for the connection ID, transport, message counts, and last active time. Expand Providers to see which provider, deployment, configuration, and authentication configuration handled the request.

5. Read the tool call. In the timeline, find a Client called tool event and open its Tool Call card. Record the tool name, the arguments, and the result. Read the events before and after it, since the sequence often explains the call.

How do you answer a question across many sessions?

Under Connection Logs, select Tool Calls. The table lists calls from all captured sessions with the tool name, status, source, agent, session, and creation time. Select Filter to narrow it, then select a row to open its session. To look at failures only, select Tool Errors.

For attribution beyond tool calls, audit logs cover people in the dashboard, agents, and service accounts in one history, filterable by user, agent, service account, and integration. Session and tool call records are part of Tracing. To compare tools in this category, see Best AI agent observability tools.

What can logs show in a SOC 2 review?

SOC 2 stands for System and Organization Controls 2. It is an attestation report, issued by a certified public accountant (CPA) firm, on a company's controls against the American Institute of CPAs (AICPA) Trust Services Criteria for security, availability, processing integrity, confidentiality, and privacy.

Agent logs can support evidence for a few of those criteria:

  • Logical access (CC6.3). The criterion asks that access be authorized, changed, and removed based on roles and least privilege. Logs show what an agent used, which you can compare against what it was granted. They do not show the grant itself. That evidence is your access configuration.
  • Monitoring (CC7.2). It asks that the entity monitor system components for anomalies and analyze them. Its points of focus list logging of unusual system activities as one example of a detection procedure.
  • Event evaluation (CC7.3). It asks that security events be evaluated to decide whether they are incidents. A tool call record is the raw material for that evaluation.

What can logs not do?

They do not make anyone compliant. The auditor decides whether evidence meets a criterion, and the criteria are about controls operating, not only data existing. A log that nobody reads supports little under CC7.2, which describes analysis of anomalies. Logs also cover only traffic that passed through the gateway, and they record what happened, not whether it was appropriate.

How do you keep the evidence usable?

1. Set retention deliberately. Metorial retention is configurable per data type, so audit logs can be held longer than high-volume session and tool call data. Enterprise customers can retain data for a year or more. See data retention.

2. Review on a schedule. For example, check Tool Errors and unusual tools weekly, and write down who looked and what they found. That written trail is what an auditor can test.

3. Tie it to access. Pair reviews with the groups and tool limits described in How to control which AI tools each team can use and How to restrict AI agents to read-only tools.

The MCP specification itself says clients should log tool usage for audit purposes, which is a baseline, not a compliance claim.

Frequently asked questions

What should an AI agent audit trail contain?

At minimum: the session, the tool called, the arguments, the result or error, the time, and the identity the call ran as. Without the arguments and result you can see that something happened but not what it did.

Does logging agent activity make us SOC 2 compliant?

No. SOC 2 is an attestation by an independent CPA firm about a company's controls, and logs are one kind of evidence for some of those controls. Policies, access reviews, and incident handling are assessed separately.

Which SOC 2 criteria do agent logs relate to?

Mostly the common criteria for logical access (CC6) and system monitoring (CC7). The monitoring criterion CC7.2 names logging of unusual system activities as one example of detection. Whether your logs satisfy a criterion is the auditor's call.

How long should agent logs be kept?

As long as your own policies and your auditor require. Metorial lets you set retention per data type, so audit logs can be kept longer than high-volume tool call data.

Can logs tell us whether an agent should have taken an action?

No. A log records what happened, not whether it was appropriate. That judgment comes from the person reviewing it, which is why a review step matters as much as the log itself.

Sources

  1. Metorial documentation: Review connection logs
  2. AICPA: 2017 Trust Services Criteria (with revised points of focus, 2022)
  3. Model Context Protocol specification: Tools

Ready to build with Metorial?

Connect any AI agent to any tool or data source. Govern every action.