At a Glance
AI agent governance is provable only through the audit trail. For every agent action, record who directed it, which identity it ran as, the model and version, the inputs and data it read, the policy applied, its tool calls, its output and the downstream action. On a lakehouse, audit, lineage, trace and gateway tables already produce most of those fields.
- The gap: agents act under a service principal, so the directing human is the field most often missing.
- Effort: most fields come from tables the platform already writes; the work is joining them on a request ID.
- Risk: default retention on preview system tables is shorter than many record-keeping rules.
- Fit: start with agents that write data, call external tools or touch a consequential decision.
An agent that reads a customer table, calls a pricing tool and updates a CRM record has made three decisions on your behalf. When someone asks why, the answer has to come from a record, not from the engineer who built it. That is the whole of agent governance in practice: a trail complete enough that a reviewer can replay what happened.
Frameworks already point at this. In NIST’s crosswalk from the AI RMF to ISO/IEC 42001, the RMF subcategory for monitoring a system in production, MEASURE 2.4, maps to the 42001 Annex control for recording event logs. Colorado’s SB 26-189, effective 1 January 2027, requires developers and deployers of covered automated decision-making technology to keep records demonstrating compliance for at least three years.
What neither tells you is which fields to log. Generic checklists list a dozen. Most mid-market teams log the model call and miss the two fields an auditor asks for first: which human set the agent in motion, and what the agent changed downstream. This article lists the fields and the lakehouse source of each, scoped tightly to the agent audit trail.
Which Fields Does an Agent Audit Trail Need?
The table is our field list and our mapping. The column names in the sources are from Databricks and MLflow documentation; the design of joining them is ours.
| Field | Question it answers | Lakehouse source | Note |
|---|---|---|---|
| Event time | When did it happen? | event_time in system.access.audit; trace timestamp | Audit timestamps are UTC |
| Agent identity | Which identity was authorized? | identity_metadata.run_as in system.access.audit | One service principal per agent makes this meaningful |
| Directing human | Who set the agent in motion? | identity_metadata.run_by in audit; mlflow.trace.user on the trace | The field most often missing |
| Session and request ID | Which conversation or job run is this? | mlflow.trace.session; gateway request ID | The join key for every other row |
| Model and version | Which model, which version, which prompt release? | Registered model version in Unity Catalog; app version tag on the trace | Without it, a regression cannot be dated |
| Inputs | What did the agent see? | Gateway inference table request payload; trace span inputs | Mask personal data before it is logged |
| Data read | Which tables and columns fed the answer? | system.access.table_lineage and column_lineage | Lineage system tables hold a rolling one year |
| Policy applied | Which access rule or guardrail fired? | ABAC policy and governed tags; gateway guardrail outcome | Record blocks, not just passes |
| Tool calls | What did it invoke, with what arguments? | Trace spans; gateway logging of tool calls | MCP governance maturity: check current release notes |
| Output | What did it return? | Gateway inference table response; trace span output | Payload size limits apply |
| Downstream action | What changed because of it? | Decision record table written by the agent’s job; write events in audit | You build this table; nothing writes it for you |
| Human review | Who checked it, and what did they decide? | Reviewer and outcome columns on the decision record | Needed wherever an appeal right exists |
Four documented facts sit under that table. The audit log system table carries user_identity, identity_metadata with run_by and run_as, service_name, action_name, request_params, response, event_time, source_ip_address and user_agent, and it is in Public Preview. Unity Catalog captures lineage automatically to the column level and exposes it in two system tables on a rolling one-year window. MLflow tracing reserves the metadata keys mlflow.trace.user and mlflow.trace.session for user and session. And Unity Gateway monitors requests, token usage and latency in system tables and logs requests and responses to Unity Catalog Delta tables.
Why Is the Directing Human So Hard to Capture?
An agent runs under a service principal, an OAuth token or an API key. The platform logs that identity faithfully. What it cannot log is the person who typed the request into a chat window three systems away, unless your application passes that identity down.
The fix is a contract, not a feature. The front end that receives the request authenticates the user, then passes the user ID into every trace as mlflow.trace.user and into the request ID your application attaches to each gateway call. The agent’s service principal stays narrow and separate. Now the audit row says the agent ran as one identity, and the trace says which human asked. Joined on the request ID, you have accountability without giving the agent the human’s permissions.
This matches where standards bodies are heading. OWASP lists identity and privilege abuse among its top ten agentic risks, and NIST’s AI Agent Standards Initiative, launched 17 February 2026, names agent identity and security research as one of its three pillars. Neither has published a field standard yet, so the contract above is yours to define. Our walkthrough of the OWASP Top 10 for agentic applications covers the other nine risks.
How Do the Tables Become One Record per Action?
Each source writes its own table on its own schedule. Gateway payload logs, lineage rows and audit events each land with their own delay. A reviewer does not want five tables. They want one row per agent action.
Build it as a silver table. A scheduled job reads new traces, joins gateway rows on request ID, joins audit events on run_as plus a time window, and attaches lineage rows for the same job or notebook run. The gold layer then holds a decision record with every field in the table above, retained on your schedule rather than the platform default. Our guide to audit log retention covers how long each record type should live.
This is a data engineering problem, not a governance document. The same orchestration discipline that conforms ERP and CRM sources into one governed lakehouse conforms agent telemetry into one record. If the agent’s context assembly is itself undocumented, fix that first; our piece on context engineering for AI agents shows how to make retrieval inputs traceable.
What Breaks an Agent Audit Trail in Practice?
Shared service principals
Three agents run under one principal because it was quicker to set up. Every audit row now says the same identity did everything. The run_as field is accurate and useless. One principal per agent, owned by a group, is the cheapest governance control you will ever ship.
Logging the prompt, not the context
Teams log the user’s question and the model’s answer, then discover the wrong answer came from a stale row the retrieval step pulled. Without the retrieved context and the lineage of the table it came from, the incident cannot be explained. Log what the agent saw, not only what it was asked.
No record of the downstream write
The agent updated a discount field in the CRM through a tool call. The trace shows the call; the CRM shows the new value; nothing ties them together. The decision record table, written by the agent’s own job with the request ID, is what closes that gap. It is the one table in this design you must build yourself.
Retention set by default
The audit table defaults to 365 days and lineage system tables hold one year. A three-year record rule outlives both. Copy what you need into tables you own, on a schedule you can show an auditor.
Automating a process nobody mapped
As a Lean Six Sigma Master Black Belt, I ask the same question of every agent: what did a human do here before, and how often did they get it wrong? Without that baseline, the audit trail records activity but cannot show whether the agent improved anything. Map the decision first, then automate it, then log it.
Are You Ready to Govern Agents This Way? Signals in Both Directions
Ready: your agents call models through one gateway, each agent has its own service principal, and your application already knows which user sent each request. The remaining work is the silver join and the decision record table, a matter of weeks.
Ready: the data your agents read already sits in Unity Catalog with lineage on. Then the data-read field is a query, not an investigation.
Not ready: agents reach external models with personal API keys, or read from exports and shared drives. Consolidate into one governed lakehouse first, or every field in the table above will have holes.
Not ready: nobody can name the business owner of each agent. A trail with no accountable reader is storage, not governance. If you want help designing the record and the controls behind it on Databricks, our AI governance and compliance work starts with the agents you run today. This article is general information, not legal advice; counsel or a qualified assessor makes the compliance call.
Frequently Asked Questions (FAQs)
What is AI agent governance?
It is the set of controls that keep an AI agent within its intended purpose and permissions, and the records that prove it did. In practice that means identity, access policy, monitored tool use and an audit trail per action.
What should an AI agent log?
At minimum: time, agent identity, the human who directed it, session or request ID, model and version, inputs, data read, policy applied, tool calls, output, downstream action and any human review. The directing human and the downstream action are the fields teams most often miss.
How long should agent audit logs be kept?
Set retention by the strictest rule that applies to the decision. Colorado’s replacement AI law, for example, requires covered records for at least three years, longer than the platform defaults on Databricks audit and lineage system tables.
Does Unity AI Gateway replace an audit trail?
No. It became generally available on 4 August 2026 and logs requests, responses, token usage and latency for governed models and tools. It is one source; the audit trail joins it with identity, lineage and decision records.
Should agents use their own service principals?
Yes. One principal per agent, with narrow grants and a group owner, makes the run_as field meaningful and lets you revoke one agent without breaking the others.

