Article 9 on AWS Bedrock — A Reference Architecture for Risk Management
Article 9 of the EU AI Act asks for something most engineering teams don't build by default: a risk-management system that runs continuously across the lifecycle of a high-risk AI model.

Article 9 of the EU AI Act asks for something most engineering teams don't build by default: a risk-management system that runs continuously across the lifecycle of a high-risk AI model. If your inference path goes through Amazon Bedrock, you already have most of the primitives you need — but only if you wire them together deliberately. This piece is a reference architecture, written for the engineers who will actually be on call when an auditor asks how risk decisions were made on a Tuesday afternoon in March.
Why this matters now
The high-risk obligations of the EU AI Act, including Article 9, become enforceable on 2 August 2026. That's the date by which providers and deployers of high-risk systems must be able to show, in concrete artifacts, that a risk-management system exists and is operating. The General-Purpose AI provisions are already live (since August 2025), and market-surveillance authorities are publishing guidance ahead of the high-risk deadline.
The clause that trips most teams up is in Article 9(2): the risk-management system must be "a continuous iterative process planned and run throughout the entire lifecycle of the high-risk AI system, requiring regular systematic review and updating." It is not a Privacy Impact Assessment with a different cover page. It is not a single document signed at launch. It is a loop — identify, evaluate, mitigate, monitor, feed back into identification — and the auditor will want to see the loop running, not a snapshot of one iteration.
What Article 9 actually requires (engineer's reading)
Strip the legal text down to its operational obligations and you get four phases that need infrastructure backing them:
- Identification of foreseeable risks to health, safety, fundamental rights — including risks that emerge only when the system is used as intended and under reasonably foreseeable misuse.
- Evaluation of those risks against measurable criteria, including risks identified from post-market monitoring data.
- Mitigation through design choices, technical controls, and information provided to deployers — with residual risk explicitly judged acceptable.
- Monitoring in production, with results fed back into (1) so the cycle continues.
Two properties matter for the architecture. First, every decision must be documented and versioned — Article 11 and Annex IV require a technical file that demonstrates the system's reasoning, and Article 9 decisions are part of it. Second, the loop must be causally connected: monitoring output has to influence identification, or you've built four disconnected pipelines and called them a system.
Mapping the four phases to AWS services
The reference architecture uses Bedrock as the inference plane and standard AWS primitives for the surrounding control plane. Nothing exotic — the discipline is in the wiring.
Identification — Bedrock Evaluations plus a custom harness
Bedrock Evaluations gives you a managed framework for model evaluation jobs against built-in or custom datasets. It covers the baseline: accuracy, robustness, toxicity. For Article 9 you'll need more than the baseline, because foreseeable misuse is domain-specific. Wrap Bedrock Evaluations with a custom red-team harness orchestrated by Step Functions: each evaluation run is a state machine execution that calls Bedrock Evaluations for the standard metrics, runs domain-specific adversarial prompts via parallel Lambda invocations, and writes a structured result document to S3. The harness runs on a schedule (EventBridge), on every model or prompt-template change (CodePipeline trigger), and on demand when monitoring surfaces a new risk hypothesis.
The output of this phase is a versioned risk register entry, not a pass/fail. Each identified risk gets an ID, a description, the evidence (eval run ARN, dataset version, prompts that triggered it), and a draft severity classification.
Documentation — versioned risk register with immutable audit subset
This is where most teams under-invest. The risk register lives in two stores:
- S3 holds the canonical artifacts: every evaluation report, every risk decision, every mitigation rationale, in JSON. The bucket has versioning enabled. A subset — anything that becomes part of the technical file or feeds an Article 73 incident report — is copied into a second bucket with S3 Object Lock in compliance mode, with a retention period matching your statute-of-limitations posture (typically 10 years for high-risk systems under the Act). Object Lock matters because once an auditor or a court asks "what did you know on 14 March 2027," nobody on your team can quietly rewrite the answer.
- DynamoDB holds the queryable index: risk ID, current status, owner, last evaluation timestamp, current severity, link to the S3 artifact. This is what the dashboard reads and what the Step Functions workflow updates.
A Lambda function generates the Annex IV technical-file sections from the register on demand, using a Bedrock model with a tightly constrained prompt and a deterministic post-processor — the LLM drafts prose, the post-processor pulls structured fields directly from DynamoDB so the numbers can't drift.
Mitigation — Guardrails, fallback paths, and human-in-the-loop
Bedrock Guardrails are the obvious technical control: content filters across hate, insults, sexual content, violence, misconduct; denied topics defined per use case; sensitive-information filters for PII redaction; word filters for domain-specific blocklists; contextual grounding checks for hallucination on RAG paths. Configure them per use case, not globally — a medical-coding assistant and a customer-support bot have different acceptable-content surfaces.
Guardrails are necessary but not sufficient. The mitigation layer needs three more pieces:
- Fallback paths as Step Functions Choice states. When a Guardrail blocks a response, when grounding confidence is low, or when an evaluation flag fires at runtime, the workflow routes to a safe-completion path (canned response, deterministic template, or escalation) rather than failing open.
- Human-in-the-loop for the residual-risk decisions Article 9(5) cares about. Step Functions
Waitstates with task tokens, paired with SNS/SES escalation to a reviewer queue, give you an auditable handoff. The reviewer's decision — approve, reject, modify — is written back to the risk register with their identity and timestamp. - Information-to-deployer artifacts generated from the register: model cards, usage instructions, known-limitations documents. These are versioned alongside the model.
Monitoring — CloudWatch, EMF, and the feedback edge
Bedrock model invocation logging sends every prompt and completion to CloudWatch Logs (or S3, if you prefer). From there:
- CloudWatch metrics track invocation volume, latency, Guardrail trigger rates, and token usage out of the box.
- Custom EMF metrics emitted from a log-processing Lambda capture what actually matters for Article 9: drift in input distribution (embedding-space distance from the training/evaluation distribution), output-quality proxies (refusal rate, fallback-path rate, human-escalation rate), and bias signals on subgroups when ground truth is available downstream.
- CloudWatch alarms on these metrics feed two sinks: an operational pager and — critically — an EventBridge rule that triggers a new identification cycle in Step Functions. That rule is the feedback edge that makes the system iterative rather than four pipelines pretending to be one.
Serious incidents (Article 73) get their own path: the alarm publishes to an SNS topic that fans out to the incident-reporting workflow, which assembles the notification package from the immutable S3 bucket and DynamoDB and surfaces it to the responsible humans within the 15-day reporting window.
The loop, as a state machine
sequenceDiagram
participant EB as EventBridge (schedule + change + alarm)
participant SF as Step Functions (RiskMgmtLoop)
participant BE as Bedrock Evaluations + Red-team Harness
participant RR as Risk Register (S3 + DynamoDB)
participant GR as Bedrock Guardrails + Runtime
participant CW as CloudWatch + EMF
participant HIL as Human Reviewer (Wait + SNS)
EB->>SF: Trigger iteration (cron / model change / alarm)
SF->>BE: Run identification + evaluation jobs
BE-->>SF: Findings (risks, severities, evidence)
SF->>RR: Write versioned entries (S3 + DynamoDB)
SF->>HIL: Escalate residual-risk decisions
HIL-->>SF: Approve / reject / mitigate
SF->>GR: Update Guardrail config + fallback routes
GR->>CW: Emit invocation + Guardrail metrics
CW->>EB: Alarm on drift / threshold breach
EB->>SF: Re-enter loop with new evidence
The state machine itself is unremarkable ASL — Parallel branches for evaluation jobs, a Map state over the risk findings, a Choice state on residual severity that routes to either auto-mitigation or human review, and a final state that updates DynamoDB and emits an EventBridge event. The point is not the JSON; the point is that the loop is a single execution graph rather than four cron jobs that nobody owns.
Common mistakes
Four patterns we see often enough to call out:
- Treating Article 9 as a one-time PIA. A 40-page document at launch satisfies nothing. The Act asks for a system in operation; auditors will look for execution history.
- Using Guardrails alone as "the" risk system. Guardrails are a mitigation control. They produce no risk register, no evaluation history, no feedback loop. They are necessary, not sufficient.
- Not versioning risk decisions. If you can't reconstruct what was known and decided on a specific date, you cannot defend the decision later. Object Lock on the audit subset is cheap insurance.
- Running monitoring without a feedback edge. Dashboards that nobody reviews and alarms that page humans but don't trigger re-identification break the iterative requirement. The EventBridge edge from CloudWatch back into Step Functions is the load-bearing wire.
What this looks like at ccX
We are an AWS Advanced Tier Services partner with three Professional-level certifications (Solutions Architect, DevOps Engineer, Data Engineer) and the AWS Generative AI Specialty. We are Hungarian-domiciled, which keeps engagements inside the EU data-residency and contracting envelope our clients ask for, and we deliver senior-only — every engineer on a project has shipped this kind of work before. Our standard EU AI Act shape is a two-week audit that produces a gap analysis against Articles 9, 11, 13, 14, 15, and 17, mapped to your existing AWS footprint with a concrete remediation backlog. The team has prior delivery against EU institutions, Big 4 advisory engagements, and Fortune 500 programs, which is why our deliverables tend to survive contact with internal audit.
Where to go from here
If you want the structured shape of the engagement, the EU AI Act service page covers scope, deliverables, and timing. If you want to talk through your specific Bedrock deployment, get in touch and we'll set up a working session with a senior engineer rather than a sales call.
This is a practical reference architecture, not legal advice. Article 9 obligations are evaluated against your specific deployment context — consult counsel for binding determinations.