Article 9 on AWS Bedrock — A Reference Architecture for Risk Management

Article 9 of the EU AI Act asks for something most engineering teams don't build by default: a risk-management system that runs continuously across the lifecycle of a high-risk AI model.

Article 9 on AWS Bedrock — A Reference Architecture for Risk Management

Article 9 of the EU AI Act asks for something most engineering teams don't build by default: a risk-management system that runs continuously across the lifecycle of a high-risk AI model. If your inference path goes through Amazon Bedrock, you already have most of the primitives you need — but only if you wire them together deliberately. This piece is a reference architecture, written for the engineers who will actually be on call when an auditor asks how risk decisions were made on a Tuesday afternoon in March.

Why this matters now

The high-risk obligations of the EU AI Act, including Article 9, become enforceable on 2 August 2026. That's the date by which providers and deployers of high-risk systems must be able to show, in concrete artifacts, that a risk-management system exists and is operating. The General-Purpose AI provisions are already live (since August 2025), and market-surveillance authorities are publishing guidance ahead of the high-risk deadline.

The clause that trips most teams up is in Article 9(2): the risk-management system must be "a continuous iterative process planned and run throughout the entire lifecycle of the high-risk AI system, requiring regular systematic review and updating." It is not a Privacy Impact Assessment with a different cover page. It is not a single document signed at launch. It is a loop — identify, evaluate, mitigate, monitor, feed back into identification — and the auditor will want to see the loop running, not a snapshot of one iteration.

What Article 9 actually requires (engineer's reading)

Strip the legal text down to its operational obligations and you get four phases that need infrastructure backing them:

  1. Identification of foreseeable risks to health, safety, fundamental rights — including risks that emerge only when the system is used as intended and under reasonably foreseeable misuse.
  2. Evaluation of those risks against measurable criteria, including risks identified from post-market monitoring data.
  3. Mitigation through design choices, technical controls, and information provided to deployers — with residual risk explicitly judged acceptable.
  4. Monitoring in production, with results fed back into (1) so the cycle continues.

Two properties matter for the architecture. First, every decision must be documented and versioned — Article 11 and Annex IV require a technical file that demonstrates the system's reasoning, and Article 9 decisions are part of it. Second, the loop must be causally connected: monitoring output has to influence identification, or you've built four disconnected pipelines and called them a system.

Mapping the four phases to AWS services

The reference architecture uses Bedrock as the inference plane and standard AWS primitives for the surrounding control plane. Nothing exotic — the discipline is in the wiring.

Identification — Bedrock Evaluations plus a custom harness

Bedrock Evaluations gives you a managed framework for model evaluation jobs against built-in or custom datasets. It covers the baseline: accuracy, robustness, toxicity. For Article 9 you'll need more than the baseline, because foreseeable misuse is domain-specific. Wrap Bedrock Evaluations with a custom red-team harness orchestrated by Step Functions: each evaluation run is a state machine execution that calls Bedrock Evaluations for the standard metrics, runs domain-specific adversarial prompts via parallel Lambda invocations, and writes a structured result document to S3. The harness runs on a schedule (EventBridge), on every model or prompt-template change (CodePipeline trigger), and on demand when monitoring surfaces a new risk hypothesis.

The output of this phase is a versioned risk register entry, not a pass/fail. Each identified risk gets an ID, a description, the evidence (eval run ARN, dataset version, prompts that triggered it), and a draft severity classification.

Documentation — versioned risk register with immutable audit subset

This is where most teams under-invest. The risk register lives in two stores:

A Lambda function generates the Annex IV technical-file sections from the register on demand, using a Bedrock model with a tightly constrained prompt and a deterministic post-processor — the LLM drafts prose, the post-processor pulls structured fields directly from DynamoDB so the numbers can't drift.

Mitigation — Guardrails, fallback paths, and human-in-the-loop

Bedrock Guardrails are the obvious technical control: content filters across hate, insults, sexual content, violence, misconduct; denied topics defined per use case; sensitive-information filters for PII redaction; word filters for domain-specific blocklists; contextual grounding checks for hallucination on RAG paths. Configure them per use case, not globally — a medical-coding assistant and a customer-support bot have different acceptable-content surfaces.

Guardrails are necessary but not sufficient. The mitigation layer needs three more pieces:

Monitoring — CloudWatch, EMF, and the feedback edge

Bedrock model invocation logging sends every prompt and completion to CloudWatch Logs (or S3, if you prefer). From there:

Serious incidents (Article 73) get their own path: the alarm publishes to an SNS topic that fans out to the incident-reporting workflow, which assembles the notification package from the immutable S3 bucket and DynamoDB and surfaces it to the responsible humans within the 15-day reporting window.

The loop, as a state machine

sequenceDiagram
    participant EB as EventBridge (schedule + change + alarm)
    participant SF as Step Functions (RiskMgmtLoop)
    participant BE as Bedrock Evaluations + Red-team Harness
    participant RR as Risk Register (S3 + DynamoDB)
    participant GR as Bedrock Guardrails + Runtime
    participant CW as CloudWatch + EMF
    participant HIL as Human Reviewer (Wait + SNS)

    EB->>SF: Trigger iteration (cron / model change / alarm)
    SF->>BE: Run identification + evaluation jobs
    BE-->>SF: Findings (risks, severities, evidence)
    SF->>RR: Write versioned entries (S3 + DynamoDB)
    SF->>HIL: Escalate residual-risk decisions
    HIL-->>SF: Approve / reject / mitigate
    SF->>GR: Update Guardrail config + fallback routes
    GR->>CW: Emit invocation + Guardrail metrics
    CW->>EB: Alarm on drift / threshold breach
    EB->>SF: Re-enter loop with new evidence

The state machine itself is unremarkable ASL — Parallel branches for evaluation jobs, a Map state over the risk findings, a Choice state on residual severity that routes to either auto-mitigation or human review, and a final state that updates DynamoDB and emits an EventBridge event. The point is not the JSON; the point is that the loop is a single execution graph rather than four cron jobs that nobody owns.

Common mistakes

Four patterns we see often enough to call out:

What this looks like at ccX

We are an AWS Advanced Tier Services partner with three Professional-level certifications (Solutions Architect, DevOps Engineer, Data Engineer) and the AWS Generative AI Specialty. We are Hungarian-domiciled, which keeps engagements inside the EU data-residency and contracting envelope our clients ask for, and we deliver senior-only — every engineer on a project has shipped this kind of work before. Our standard EU AI Act shape is a two-week audit that produces a gap analysis against Articles 9, 11, 13, 14, 15, and 17, mapped to your existing AWS footprint with a concrete remediation backlog. The team has prior delivery against EU institutions, Big 4 advisory engagements, and Fortune 500 programs, which is why our deliverables tend to survive contact with internal audit.

Where to go from here

If you want the structured shape of the engagement, the EU AI Act service page covers scope, deliverables, and timing. If you want to talk through your specific Bedrock deployment, get in touch and we'll set up a working session with a senior engineer rather than a sales call.

This is a practical reference architecture, not legal advice. Article 9 obligations are evaluated against your specific deployment context — consult counsel for binding determinations.


Sitemap