AI agent audit trail: what to log under the EU AI Act
AI agents in software teams no longer only answer questions. They create Jira issues, comment on pull requests, query tables in Databricks and post messages in Slack. Each of these actions is a call against a real system, made with real credentials. When something goes wrong, or when an auditor or a customer asks what an agent did, the team needs a record that answers the question precisely. In many organisations that record does not exist, or it exists only in the logs of the AI vendor.
This post describes what an audit trail for AI agents should contain, why the location of the log matters as much as its content, and how the EU AI Act and the GDPR frame the topic. It is part of our series on AI agents in the software development lifecycle.
Why the question comes up now
Two developments meet here. First, agents act on more systems than before. With the Model Context Protocol (MCP), a client such as Claude Code, Cursor or VS Code can call tools in Jira, GitHub or Confluence directly. Second, regulation has started to ask for evidence. The EU AI Act (Regulation (EU) 2024/1689) contains explicit record-keeping duties, and the GDPR has required accountability for any processing of personal data since 2018.
The timeline is often misread, so it is worth stating it carefully. Prohibited practices apply since 2 February 2025, and obligations for general-purpose AI models since 2 August 2025. The transparency obligations of Article 50 apply from 2 August 2026. The obligations for high-risk systems were postponed by the Digital Omnibus (Regulation (EU) 2026/1744, in force since 27 July 2026): stand-alone high-risk systems listed in Annex III follow from 2 December 2027, and AI embedded in regulated products under Annex I from 2 August 2028.
It is also important to be honest about scope. Most agents that plan sprints, review code or update documentation are not high-risk systems in the sense of the AI Act. The strict logging duties of Article 12 therefore do not apply to them directly. However, this does not make the question irrelevant. As soon as an agent reads or passes on personal data, the GDPR applies, and the controller must be able to demonstrate what happened (Art. 5(2) GDPR). In addition, security teams and customers increasingly ask the same question an auditor would ask: which agent did what, with whose authority, and was it allowed?
What an audit trail for AI agents should contain
A useful audit trail answers five questions for every single call. The following fields are a reasonable minimum.
Who acted. The identity behind the call: which agent or person, with which API key, in which project or workspace. "The AI did it" is not an answer an auditor accepts. The record has to name the accountable identity.
What was requested. The target system, the operation (for example jira_create_issue or
github_merge_pr) and the parameters, such as the project key or the repository. Parameters can
contain personal or secret data, so they should be redacted before they are written. The log must
not become a second copy of the sensitive data it is meant to protect.
What was decided, and why. This is the part most logs miss. It is not enough to record successful calls. A complete trail also records calls that were denied, together with the reason (for example "operation not enabled for this workspace" or "repository not in allowlist"), calls that were held for human approval, calls that hit a rate limit and calls that failed with an error. Denied calls are often the most informative entries, because they show what an agent tried to do outside its scope.
When. A precise timestamp, in a consistent time zone, so that entries can be correlated with events in other systems.
Whether the record is intact. An audit trail is only evidence if nobody can quietly change it afterwards. This requires technical protection, not only a policy that says "do not edit".
A good test is to take one real incident and try to reconstruct it from the log alone. If the reconstruction needs screenshots, chat histories or somebody's memory, the log is incomplete.
Why the vendor's log is not independent evidence
Several AI vendors now offer audit logs for their own tools and connectors. These logs are useful for operating the tool, and there is no reason to ignore them. However, as evidence they have a structural weakness: the party that is being evaluated also keeps the record.
This matters for three reasons. First, independence: an auditor treats a log kept by the operator of the system under review differently from a log kept by a separate control. The situation is comparable to financial accounting, where the bookkeeping and the audit are separated on purpose. Second, coverage: a vendor log only covers that vendor's tools. A team that uses Claude Code for code, Cursor in the IDE and an internal chatbot for support has three partial logs with three formats, and none of them shows the complete picture. Third, control over retention and access: the retention period, the export format and the access rights are set by the vendor, not by the organisation that has to answer for the data.
An independent control point avoids these problems by design. If every agent call passes through one gateway operated by the organisation itself, the log is complete across vendors, and it belongs to the organisation that is accountable. The article What is an MCP gateway? explains this architecture in more detail.
How Vordix records agent activity
Vordix is a self-hosted gateway between AI agents and tools such as Jira, GitHub, Confluence, Slack and Databricks. Every request passes through it, so every request is recorded in one place. The audit log is designed along the questions above:
- Every outcome is recorded: allowed, denied with the reason, held for approval, rate-limited and error. Each entry contains the decision trace, which shows which rule led to the decision.
- Parameters are redacted before they are written, so personal data from a request does not end up in the log in clear text.
- The log is append-only and hash-chained with an HMAC. No role, including the administrator, can edit entries, and the integrity of the chain can be verified on demand. A changed or removed entry breaks the chain.
- Retention has a floor of 180 days, and a legal hold can keep records beyond that.
- Export and forwarding: entries can be exported as CSV or JSON, followed as a live stream or sent to a SIEM (security information and event management system) through a signed webhook.
- Regulatory tags: audit entries carry tags for Articles 10 and 12 of the AI Act, and the compliance page of the documentation maps the controls to Articles 10, 12 and 14 of the AI Act and to ISO 27001 Annex A.
Because Vordix runs on the organisation's own infrastructure, the log is stored where the organisation decides, and the Vordix company is not in the request path.
For write operations that need a second pair of eyes, for example merging a pull request (see AI code review permissions on GitHub and Azure DevOps), an administrator can configure an approval policy. Held calls then appear in the log as well, with the approval or rejection that followed. For data access, the same trail shows which tables and columns an agent read, as described in AI agents and Databricks data access.
Limitations and trade-offs
An audit trail is a necessary part of governance, but it is not sufficient on its own, and a few limitations should be stated clearly.
First, a log documents; it does not prevent. The value comes from combining it with enforcement: scoped permissions, allowlists and approvals decide what an agent may do, and the log proves what it actually did. A log without enforcement mainly documents the damage.
Second, a gateway can only record what passes through it. If an agent still holds a direct API token for GitHub, calls made with that token bypass the gateway and are not in its log. Removing direct credentials is therefore part of the setup, not an optional step.
Third, Vordix holds no certification, and using it does not make an organisation compliant with the AI Act, the GDPR or ISO 27001. The controls and the mapping help to produce evidence; the assessment of whether a specific use case is high-risk, and whether the overall setup is sufficient, remains with the organisation and its advisers.
Finally, redaction is a trade-off. The more parameters are masked before writing, the less personal data the log contains, but the harder some investigations become. The right balance depends on the data the agents handle and should be decided together with the data protection officer.
Conclusion
An audit trail for AI agents should record who acted, what was requested, what was decided and why, and when, for denied calls as well as successful ones, and it should be protected against later changes. For most development-tool agents the AI Act's high-risk logging duties do not apply directly, but GDPR accountability and customer expectations lead to the same requirement. Keeping that record in an independent, self-hosted control point, instead of in each vendor's tool, gives one complete log that belongs to the organisation that has to answer for it.
The Vordix documentation describes the audit log in detail, and a demo shows it with real agent calls.