Logging That Survives an Audit — Article 12 Traceability with CloudWatch + S3 + Glue

An EU AI Act audit will not start with a polite request for design documents. It will start with a question — show me every decision this system made about citizen X over the last twelve months.

Logging That Survives an Audit — Article 12 Traceability with CloudWatch + S3 + Glue

An EU AI Act audit will not start with a polite request for design documents. It will start with a question — show me every decision this system made about citizen X over the last twelve months — and your answer is whatever your logs can produce in the room. Article 12 is the obligation that turns observability from an SRE concern into a regulatory one, and most teams running production AI on AWS today have the right primitives but the wrong shape.

Why this matters now

High-risk AI systems placed on the EU market on or after 2 August 2026 must demonstrate Article 12 traceability from day one. There is no grandfather clause for the logs you wish you had been writing. If a notified body, a national market surveillance authority, or a downstream deployer asks for a reconstruction of an inference six months in the past, the answer is either in your S3 buckets in a queryable shape, or it is not.

The trap is that an audit is reactive. By the time the request lands, you have weeks — not days — to retrofit a logging schema, backfill what you can, and build the query layer auditors expect. Teams that wait until a request arrives discover that CloudWatch Logs alone cannot answer twelve-month-lookback questions at a price that anyone is willing to sign off on, and that EMF logs written without a schema version are unparseable two model deployments later. The work is not hard, but it is sequential, and the sequence has to start before the request, not after.

What Article 12 actually requires

Article 12 of the EU AI Act requires high-risk AI systems to automatically record events ("logs") over the lifetime of the system, with traceability appropriate to the intended purpose. The text is short on prescription and long on outcome: logs must support post-market monitoring (Article 72), enable identification of situations that may present a risk under Article 79, and feed the transparency obligations of Article 13 and the human oversight obligations of Article 14.

A common mistake is to collapse Article 12 and Article 14 into the same workstream. They are distinct. Article 12 is what was recorded. Article 14 is what a human did with the recording, in time to intervene. Article 12 is your foundation; Article 14 is the dashboard, the alerting, and the escalation runbook built on top. Get Article 12 wrong and Article 14 has nothing to read from. Conflate them and you build a real-time oversight tool that cannot answer historical questions, or a data lake that no operator looks at.

Retention is implied rather than fixed: logs should be kept for a duration appropriate to the intended purpose and at least the lifetime of the system as placed on the market, with national rules and sector-specific obligations (financial services, medical devices, employment) often pushing this to six or seven years. Plan for seven and revise downward only with counsel.

The three layers of logs auditors care about

Auditors do not want one giant log stream. They want three layers, separable, each with a different retention and access profile.

LayerWhat it capturesPrimary consumerTypical volume
Inference-timeEvery model invocation: prompt, response, model version, guardrail decisions, latency, agent/user identity (redacted)Incident replay, drift analysisHigh (per request)
Decision-timeEvery AI-driven decision applied to a person: inputs, model output, business rule applied, human-review flag, final outcomeArticle 12 audit, Article 22 GDPR requestsMedium (per decision)
System-timeDeployment events, guardrail config changes, dataset/version pins, incident recordsChange-control review, root-cause analysisLow (per change)

Inference-time logs are noisy and expensive. Most auditors do not actually want to read them — they want to know they exist and that you can produce a sample. Decision-time logs are the headline asset: every record where an AI output crossed into a real-world consequence for a real person. System-time logs are what lets you answer which model version was live on this date, with which guardrail policy. Without that third layer, the first two are uninterpretable.

AWS implementation

The AWS primitives map cleanly onto the three layers. The work is in the wiring, not the components.

Inference-time — Bedrock Invocation Logging. Bedrock writes invocation logs (request, response, token counts, model ID) to CloudWatch Logs and S3 simultaneously. Enable both. CloudWatch is for the last 7–14 days of operational queries; S3 is the system of record. Use a delivery configuration like:

{
  "loggingConfig": {
    "cloudWatchConfig": {
      "logGroupName": "/aws/bedrock/invocations",
      "roleArn": "arn:aws:iam::123456789012:role/BedrockLoggingRole"
    },
    "s3Config": {
      "bucketName": "ccx-ai-audit-logs",
      "keyPrefix": "bedrock/invocations/"
    },
    "textDataDeliveryEnabled": true,
    "imageDataDeliveryEnabled": false,
    "embeddingDataDeliveryEnabled": false
  }
}

Redact PII before it reaches the model where possible, and apply a second-pass redaction in a Lambda subscriber on the CloudWatch Logs stream to scrub anything that slipped through. The redacted copy is what flows to S3.

System-time — CloudTrail. Every change to a Bedrock guardrail, IAM policy on the model role, or S3 bucket policy on the audit lake is a CloudTrail event. Pipe the management-events trail to the same S3 bucket under a separate prefix. This is the cheapest and most valuable layer: it is small, immutable, and answers the what changed and when question that anchors every incident review.

Decision-time — custom EMF logs. Bedrock invocation logs do not know which of your invocations were merely a chat turn and which were the invocation that produced an adverse decision. Your application has to emit that. Use CloudWatch Embedded Metric Format from the service that owns the decision:

{
  "_aws": {
    "Timestamp": 1746748800000,
    "CloudWatchMetrics": [{
      "Namespace": "ccX/Decisions",
      "Dimensions": [["DecisionType", "ModelVersion"]],
      "Metrics": [{"Name": "DecisionCount", "Unit": "Count"}]
    }]
  },
  "schemaVersion": "2026-05-01",
  "decisionId": "dec_01HXYZ...",
  "subjectId": "sub_hash_a1b2c3...",
  "decisionType": "credit_application",
  "outcome": "rejected",
  "modelVersion": "claude-sonnet-4-7@1m-v2",
  "guardrailVersion": "gr-v14",
  "humanReviewFlag": false,
  "businessRuleApplied": "policy-2026-q2",
  "inputHash": "sha256:...",
  "outputHash": "sha256:..."
}

Version the schema explicitly (schemaVersion). Without it, every model deployment that adds a field silently breaks last year's Athena queries.

Glue + Athena. Point a Glue crawler at the S3 prefixes and partition by year/month/day/source. Athena then becomes the query interface for both day-to-day SRE work and audit response. QuickSight sits on top of Athena for the standing audit narrative — a dashboard a compliance officer can open without writing SQL, showing decision volumes, human-review rates, and guardrail-deny rates over time. That dashboard is also your Article 14 oversight surface, which is why the layering matters.

Retention and immutability

Three tiers, mapped to the three log layers:

AgeStorage classApplies to
0–30 daysS3 StandardAll layers, hot-query window
30 days – 1 yearS3 Standard-IAInference-time, decision-time
1–7 yearsS3 Glacier Instant RetrievalDecision-time, system-time
7+ yearsS3 Glacier Deep Archive or expiryPer legal-hold policy

Apply S3 Object Lock in compliance mode on the decision-time and system-time prefixes. Object Lock is the difference between we keep our logs and we can prove our logs were not altered. The KMS key that encrypts these prefixes should be a separate CMK from your operational keys, with a key policy that denies kms:ScheduleKeyDeletion outright. Inference-time logs typically do not need Object Lock — the decision-time record references them by ID, and that reference is what is immutable.

A baseline lifecycle policy:

Rules:
  - Id: DecisionLogsTiering
    Filter: { Prefix: "decisions/" }
    Status: Enabled
    Transitions:
      - Days: 30
        StorageClass: STANDARD_IA
      - Days: 365
        StorageClass: GLACIER_IR
    Expiration:
      Days: 2555  # 7 years

Anti-patterns to avoid

A sample Athena query an auditor will ask for

The canonical request: show every AI-driven decision applied to subject X over the last twelve months, with the model version, guardrail version, outcome, and whether a human reviewed it.

SELECT
    d.decisionid,
    from_unixtime(d.timestamp / 1000) AS decided_at,
    d.decisiontype,
    d.outcome,
    d.modelversion,
    d.guardrailversion,
    d.humanreviewflag,
    d.businessruleapplied,
    d.inputhash,
    d.outputhash
FROM ai_audit.decisions d
WHERE d.subjectid = 'sub_hash_a1b2c3d4e5f6'
  AND d.year >= CAST(year(current_date - interval '12' month) AS VARCHAR)
  AND from_unixtime(d.timestamp / 1000)
        >= current_timestamp - interval '12' month
ORDER BY d.timestamp DESC;

The query runs in seconds against partitioned Parquet. The same query, against CloudWatch Logs Insights at scale, either does not return or returns a bill.

What this looks like at ccX

The audit-ready logging schema described above is one of the deliverables in our 2-week EU AI Act readiness audit — a focused engagement against your existing AWS estate (Bedrock configuration, CloudTrail coverage, S3 layout, retention policy, KMS posture) that ends with a versioned EMF schema, Glue table definitions, lifecycle policies, and the canonical Athena queries auditors actually ask. ccX is AWS 3× Professional certified across Solutions Architect, DevOps, and Security, with the GenAI Specialty on top, so the schema is the same shape we have shipped on engagements ranging from EU institutions and Big 4 firms to Fortune 500 enterprises — not a reference architecture pulled off a slide. Hungarian-domiciled means your evidence chain stays inside EU jurisdiction, which matters when a notified body asks where the logs physically sit. And the work is senior-only — no juniors interpreting Article 12 or the Article 12 / Article 14 separation on your engagement, and sector-specific retention (financial services, healthcare, employment) is handled by people who have read the underlying sector law, not just the AI Act.

Where to go from here

Logging design for Article 12 compliance is fact-specific; this is a reference pattern, not legal advice. Consult counsel for binding decisions.


Sitemap