PCI DSS 4.0.1 Log Review: Automating Requirement 10.4.1.1 on a Lakehouse

Analytics AIML is an AI performance firm. We rebuild the three foundations that decide whether an AI investment ships, scales, and shows up on the P&L — a sharper problem, a governed data foundation, and demand that survives the zero-click age.

Frank Shines

September 28, 2026

Pci Dss Log Review: PCI DSS 4.0.1 Log Review: Automating Requirement 10.4.1.1 on a Lakehouse

Quick Answer

PCI DSS log review means examining security-relevant logs every day, and since 31 March 2025 the standard expects automated mechanisms to do it. On a lakehouse you ingest every in-scope log, normalise it to OCSF, run scheduled detections and write each day’s review to an evidence table. The platform produces evidence. Your QSA still decides compliance.

  • Cost: storage and scheduled compute, not per-gigabyte ingest licensing for the full year.
  • Effort: most weeks go into source inventory and parsing, not detection logic.
  • Risk: a source that silently stops sending looks exactly like a quiet day.
  • When it fits: teams whose cardholder data environment logs are scattered across several tools.

Requirement 10.4.1.1 is one short sentence. As it is quoted across PCI compliance guidance, PCI DSS v4.0.1 says: “Automated mechanisms perform audit log reviews.” It was a future-dated requirement, and it became mandatory on 31 March 2025.

The review it automates is not small. Requirement 10.4.1 calls for daily review of all security events, logs of every system component that stores, processes or transmits cardholder data, logs of critical system components, and logs of every server and component that performs a security function. Retention sits in Requirement 10.5.1: at least 12 months of history, with the most recent three months immediately available for analysis.

Read those together and the difficulty is obvious. The logs that count live in data silos: the firewall in one console, the identity provider in another, the database audit trail on a server, the payment application in a vendor portal. Automating review across silos means solving the silos first. That is a data engineering job, and a lakehouse is a good place to do it.

What the Two Requirements Actually Ask of You

Strip away the vendor language and Requirement 10.4.1 with 10.4.1.1 produces three engineering duties.

  • Coverage. Every in-scope source is collected, every day, with no gaps you cannot explain.
  • Review. A machine, not a person scrolling a console, examines those logs daily and surfaces what needs attention.
  • Proof. You can show an assessor, for any sampled day, that the review ran, what it covered and what happened to anything it raised.

Manual review on its own is no longer the control. People still triage what the automation raises, but the daily pass over the logs has to be automated. Most teams underestimate the third duty. A detection that fired correctly but left no durable record of the run is indistinguishable, in an assessment, from one that never ran.

Building Automated Log Review on a Lakehouse

This is the build we use for a security data lake with a PCI scope. Each step leaves an artefact behind, and that artefact is the point.

1. Inventory in-scope sources before touching code

List every system component in the cardholder data environment and every component that performs a security function for it. For each one, record where its logs go today, the format, the volume and the owner. This list becomes a reference table in the lakehouse, and step 3 checks against it every day.

2. Land raw logs in bronze

Ingest each source into its own append-only Delta table using Auto Loader for files and streams, or Lakeflow Connect where a managed connector exists. Put these tables in a dedicated catalog with grants limited to the security engineering role. Keep the raw form exactly as received; it is your evidence of what the source actually sent.

3. Normalise to OCSF in silver

Use Lakeflow Declarative Pipelines to map every source to the Open Cybersecurity Schema Framework: authentication events to the authentication class, network traffic to network activity, and so on. One user, one IP address and one timestamp format across every source is what makes a daily cross-source review possible. Add pipeline expectations that compare each day’s arrivals against the inventory table, so a silent source raises its own alert. Our security data lake guide goes deeper on this layer.

4. Detect in gold on a schedule

Write detections as versioned SQL or Python in a repository and run them daily as Lakeflow Jobs. Jobs support schedules, task dependencies and retries, so the review runs after ingestion completes rather than at a fixed hour that ingestion sometimes misses. Our detection engineering piece covers how to write rules that stay maintainable.

5. Write an evidence row for every run

Each daily job appends one record to an evidence table: run ID, review window, sources expected, sources received, row counts per source, detections executed, alerts raised, and a link to the ticket each alert became. An assessor samples days. This table answers the sample in one query.

6. Retain to the standard, then test a restore you never need

Keep at least 12 months, with the latest three queryable immediately, per Requirement 10.5.1. On Delta that is one table read at any age, so “immediately available” is a compute choice, not a restore project. Protect the archive: access to it is itself logged in the Databricks system.access.audit table, which the docs label Public Preview with a 365-day default retention, so copy those records into the archive as well.

Duty Lakehouse mechanism Evidence artefact What the QSA still judges
Coverage of in-scope sources Inventory table plus daily arrival expectations in the silver pipeline Per-day sources expected vs received Whether your scope and inventory are complete
Automated daily review Detections run as scheduled Lakeflow Jobs over OCSF gold tables Job run history and the evidence table Whether the detections are adequate for your risks
Follow-up on exceptions Alerts written to a table and pushed to the ticketing system Alert-to-ticket links with timestamps Whether exceptions were handled appropriately
Retention Delta tables with deliberate retention and scheduled deletion Oldest and newest queryable dates per source Whether the period and availability meet 10.5.1
Integrity of the log store Restricted catalog, least-privilege grants, platform audit log Grant history and access records for the archive Whether logs are protected from change

 

What Trips Teams Up

The build above is straightforward on paper. These are the places it goes wrong in real environments.

  • The quiet source. A firewall stops forwarding after a firmware update. Detections run, find nothing, and the evidence says the review passed. Without an arrival check against the inventory, a missing source reads as a clean day.
  • Parsing drift. A vendor changes a log format and half of the fields land null in silver. The review runs on incomplete data. Pipeline expectations on required OCSF fields catch this on day one instead of at assessment time.
  • Clock and time zone mismatches. Sources that log in local time without an offset break correlation across sources. Normalise every timestamp to UTC in silver and keep the original in bronze.
  • Review without follow-up. An alert table that nobody works is not a review. Tie every alert to a ticket and record the outcome, or the automation proves only that you noticed and did nothing.
  • A log store the admins can edit. If the people whose actions are logged can alter the archive, its value as evidence collapses. Separate the security catalog and restrict grants to the smallest workable group.
  • Scope creep in reverse. Teams move the easy sources into the lakehouse and leave the awkward mainframe or payment appliance behind. The review is then automated for most of the scope and manual, or absent, for the rest.

Where the SIEM Still Fits

None of this requires ripping out a SIEM. Many teams keep the SIEM for real-time alerting on a short window and use the lakehouse for the full daily review, the 12-month history and cross-source investigation. Our Splunk alternative write-up walks through that split and its trade-offs.

Databricks is also moving into this space itself. In March 2026 it announced Lakewatch, an agentic SIEM, in Private Preview. It is Databricks’ product, not ours, and a preview product is not something to build a PCI control around yet. The architecture above works with or without it, because the evidence lives in your own governed tables.

One principle holds either way: fix the review process before you automate it. If nobody can say today which alerts matter and who owns them, automation will produce a faster, better-documented version of that confusion.

Seven Questions Before You Sign a Contract or Write a Pipeline

Whether you are evaluating a vendor or scoping an internal build, get clear answers to these first.

  1. Is there a complete, owned inventory of every log source in the cardholder data environment and every security function that serves it?
  2. How will the design detect a source that stops sending, and who is paged when it does?
  3. Which schema will every source be normalised to, and who maintains the mappings when a vendor changes a format?
  4. For any day in the last year, can the design show which detections ran, over which sources, and what they raised, in one query?
  5. How are alerts tied to tickets and outcomes, and where is that link stored?
  6. Who can change the log archive, and where is that access itself recorded?
  7. What does the SIEM keep doing, and what moves to the lakehouse, and at what log volume does that split pay for itself?

If the answers point to a governed security layer on the lakehouse, our lakehouse security practice builds the ingestion, OCSF normalisation, scheduled detections and evidence tables described here. This article is general information, not legal or assessment advice; your QSA decides whether your controls meet PCI DSS.

This is general information, not legal advice: your QSA and counsel decide how PCI DSS applies to your environment.

Frequently Asked Questions (FAQs)

What does PCI DSS Requirement 10.4.1.1 require?

As the requirement is widely quoted, “Automated mechanisms perform audit log reviews.” It became mandatory on 31 March 2025 and automates the daily review described in Requirement 10.4.1.

Is manual log review still acceptable under PCI DSS 4.0.1?

Not as the review control on its own. People still investigate and resolve what the automation raises, but the daily pass over the logs is expected to be automated.

How long must PCI audit logs be kept?

Requirement 10.5.1, as commonly quoted, calls for at least 12 months of audit log history with the most recent three months immediately available for analysis.

Does moving logs to a lakehouse make us PCI compliant?

No. A lakehouse produces evidence: coverage checks, run history, alerts and retention. Compliance is assessed by your QSA against the whole standard, including scope, processes and people.

Why normalise PCI logs to OCSF?

Daily review across a firewall, an identity provider and a database only works when the same user, address and time mean the same thing in every source. OCSF gives every source one set of classes and field names.

— Rise above the flood

Build a content engine that gets cited.

AIMGrowth is the discipline for the AI-answer economy. We ship it in 90 days, fixed scope.