Logging That Survives an Audit — Article 12 Traceability with CloudWatch + S3 + Glue
An EU AI Act audit will not start with a polite request for design documents. It will start with a question — show me every decision this system made about citizen X over the last twelve months.
An EU AI Act audit will not start with a polite request for design documents. It will start with a question — show me every decision this system made about citizen X over the last twelve months — and your answer is whatever your logs can produce in the room. Article 12 is the obligation that turns observability from an SRE concern into a regulatory one, and most teams running production AI on AWS today have the right primitives but the wrong shape.
Why this matters now
High-risk AI systems placed on the EU market on or after 2 August 2026 must demonstrate Article 12 traceability from day one. There is no grandfather clause for the logs you wish you had been writing. If a notified body, a national market surveillance authority, or a downstream deployer asks for a reconstruction of an inference six months in the past, the answer is either in your S3 buckets in a queryable shape, or it is not.
The trap is that an audit is reactive. By the time the request lands, you have weeks — not days — to retrofit a logging schema, backfill what you can, and build the query layer auditors expect. Teams that wait until a request arrives discover that CloudWatch Logs alone cannot answer twelve-month-lookback questions at a price that anyone is willing to sign off on, and that EMF logs written without a schema version are unparseable two model deployments later. The work is not hard, but it is sequential, and the sequence has to start before the request, not after.
What Article 12 actually requires
Article 12 of the EU AI Act requires high-risk AI systems to automatically record events ("logs") over the lifetime of the system, with traceability appropriate to the intended purpose. The text is short on prescription and long on outcome: logs must support post-market monitoring (Article 72), enable identification of situations that may present a risk under Article 79, and feed the transparency obligations of Article 13 and the human oversight obligations of Article 14.
A common mistake is to collapse Article 12 and Article 14 into the same workstream. They are distinct. Article 12 is what was recorded. Article 14 is what a human did with the recording, in time to intervene. Article 12 is your foundation; Article 14 is the dashboard, the alerting, and the escalation runbook built on top. Get Article 12 wrong and Article 14 has nothing to read from. Conflate them and you build a real-time oversight tool that cannot answer historical questions, or a data lake that no operator looks at.
Retention is implied rather than fixed: logs should be kept for a duration appropriate to the intended purpose and at least the lifetime of the system as placed on the market, with national rules and sector-specific obligations (financial services, medical devices, employment) often pushing this to six or seven years. Plan for seven and revise downward only with counsel.
The three layers of logs auditors care about
Auditors do not want one giant log stream. They want three layers, separable, each with a different retention and access profile.
| Layer | What it captures | Primary consumer | Typical volume |
|---|---|---|---|
| Inference-time | Every model invocation: prompt, response, model version, guardrail decisions, latency, agent/user identity (redacted) | Incident replay, drift analysis | High (per request) |
| Decision-time | Every AI-driven decision applied to a person: inputs, model output, business rule applied, human-review flag, final outcome | Article 12 audit, Article 22 GDPR requests | Medium (per decision) |
| System-time | Deployment events, guardrail config changes, dataset/version pins, incident records | Change-control review, root-cause analysis | Low (per change) |
Inference-time logs are noisy and expensive. Most auditors do not actually want to read them — they want to know they exist and that you can produce a sample. Decision-time logs are the headline asset: every record where an AI output crossed into a real-world consequence for a real person. System-time logs are what lets you answer which model version was live on this date, with which guardrail policy. Without that third layer, the first two are uninterpretable.
AWS implementation
The AWS primitives map cleanly onto the three layers. The work is in the wiring, not the components.
Inference-time — Bedrock Invocation Logging. Bedrock writes invocation logs (request, response, token counts, model ID) to CloudWatch Logs and S3 simultaneously. Enable both. CloudWatch is for the last 7–14 days of operational queries; S3 is the system of record. Use a delivery configuration like:
{
"loggingConfig": {
"cloudWatchConfig": {
"logGroupName": "/aws/bedrock/invocations",
"roleArn": "arn:aws:iam::123456789012:role/BedrockLoggingRole"
},
"s3Config": {
"bucketName": "ccx-ai-audit-logs",
"keyPrefix": "bedrock/invocations/"
},
"textDataDeliveryEnabled": true,
"imageDataDeliveryEnabled": false,
"embeddingDataDeliveryEnabled": false
}
}
Redact PII before it reaches the model where possible, and apply a second-pass redaction in a Lambda subscriber on the CloudWatch Logs stream to scrub anything that slipped through. The redacted copy is what flows to S3.
System-time — CloudTrail. Every change to a Bedrock guardrail, IAM policy on the model role, or S3 bucket policy on the audit lake is a CloudTrail event. Pipe the management-events trail to the same S3 bucket under a separate prefix. This is the cheapest and most valuable layer: it is small, immutable, and answers the what changed and when question that anchors every incident review.
Decision-time — custom EMF logs. Bedrock invocation logs do not know which of your invocations were merely a chat turn and which were the invocation that produced an adverse decision. Your application has to emit that. Use CloudWatch Embedded Metric Format from the service that owns the decision:
{
"_aws": {
"Timestamp": 1746748800000,
"CloudWatchMetrics": [{
"Namespace": "ccX/Decisions",
"Dimensions": [["DecisionType", "ModelVersion"]],
"Metrics": [{"Name": "DecisionCount", "Unit": "Count"}]
}]
},
"schemaVersion": "2026-05-01",
"decisionId": "dec_01HXYZ...",
"subjectId": "sub_hash_a1b2c3...",
"decisionType": "credit_application",
"outcome": "rejected",
"modelVersion": "claude-sonnet-4-7@1m-v2",
"guardrailVersion": "gr-v14",
"humanReviewFlag": false,
"businessRuleApplied": "policy-2026-q2",
"inputHash": "sha256:...",
"outputHash": "sha256:..."
}
Version the schema explicitly (schemaVersion). Without it, every model deployment that adds a field silently breaks last year's Athena queries.
Glue + Athena. Point a Glue crawler at the S3 prefixes and partition by year/month/day/source. Athena then becomes the query interface for both day-to-day SRE work and audit response. QuickSight sits on top of Athena for the standing audit narrative — a dashboard a compliance officer can open without writing SQL, showing decision volumes, human-review rates, and guardrail-deny rates over time. That dashboard is also your Article 14 oversight surface, which is why the layering matters.
Retention and immutability
Three tiers, mapped to the three log layers:
| Age | Storage class | Applies to |
|---|---|---|
| 0–30 days | S3 Standard | All layers, hot-query window |
| 30 days – 1 year | S3 Standard-IA | Inference-time, decision-time |
| 1–7 years | S3 Glacier Instant Retrieval | Decision-time, system-time |
| 7+ years | S3 Glacier Deep Archive or expiry | Per legal-hold policy |
Apply S3 Object Lock in compliance mode on the decision-time and system-time prefixes. Object Lock is the difference between we keep our logs and we can prove our logs were not altered. The KMS key that encrypts these prefixes should be a separate CMK from your operational keys, with a key policy that denies kms:ScheduleKeyDeletion outright. Inference-time logs typically do not need Object Lock — the decision-time record references them by ID, and that reference is what is immutable.
A baseline lifecycle policy:
Rules:
- Id: DecisionLogsTiering
Filter: { Prefix: "decisions/" }
Status: Enabled
Transitions:
- Days: 30
StorageClass: STANDARD_IA
- Days: 365
StorageClass: GLACIER_IR
Expiration:
Days: 2555 # 7 years
Anti-patterns to avoid
- Logging PII unredacted. This creates a GDPR liability that the AI Act does not override. The redaction layer must run before persistence, not after.
- Rolling indexes that drop history. OpenSearch with a 90-day rollover policy is great for ops and useless for Article 12. Auditors ask twelve-month questions.
- CloudWatch Logs as the only store. At decision-time volume, CloudWatch Logs Insights queries either time out or cost more than the model itself. S3 + Athena is the right shape past 30 days.
- EMF schema drift without versioning. A field rename in May 2026 should not silently invalidate the April 2026 query. Version the schema, keep a registry, and write Athena views that abstract over versions.
- One bucket for everything. Inference-time and decision-time have different retention, different access patterns, and different sensitivity. Separate prefixes, separate lifecycle rules, separate KMS keys.
A sample Athena query an auditor will ask for
The canonical request: show every AI-driven decision applied to subject X over the last twelve months, with the model version, guardrail version, outcome, and whether a human reviewed it.
SELECT
d.decisionid,
from_unixtime(d.timestamp / 1000) AS decided_at,
d.decisiontype,
d.outcome,
d.modelversion,
d.guardrailversion,
d.humanreviewflag,
d.businessruleapplied,
d.inputhash,
d.outputhash
FROM ai_audit.decisions d
WHERE d.subjectid = 'sub_hash_a1b2c3d4e5f6'
AND d.year >= CAST(year(current_date - interval '12' month) AS VARCHAR)
AND from_unixtime(d.timestamp / 1000)
>= current_timestamp - interval '12' month
ORDER BY d.timestamp DESC;
The query runs in seconds against partitioned Parquet. The same query, against CloudWatch Logs Insights at scale, either does not return or returns a bill.
What this looks like at ccX
The audit-ready logging schema described above is one of the deliverables in our 2-week EU AI Act readiness audit — a focused engagement against your existing AWS estate (Bedrock configuration, CloudTrail coverage, S3 layout, retention policy, KMS posture) that ends with a versioned EMF schema, Glue table definitions, lifecycle policies, and the canonical Athena queries auditors actually ask. ccX is AWS 3× Professional certified across Solutions Architect, DevOps, and Security, with the GenAI Specialty on top, so the schema is the same shape we have shipped on engagements ranging from EU institutions and Big 4 firms to Fortune 500 enterprises — not a reference architecture pulled off a slide. Hungarian-domiciled means your evidence chain stays inside EU jurisdiction, which matters when a notified body asks where the logs physically sit. And the work is senior-only — no juniors interpreting Article 12 or the Article 12 / Article 14 separation on your engagement, and sector-specific retention (financial services, healthcare, employment) is handled by people who have read the underlying sector law, not just the AI Act.
Where to go from here
Logging design for Article 12 compliance is fact-specific; this is a reference pattern, not legal advice. Consult counsel for binding decisions.