AASIFWorked example · read-only
STEP 6 OF 7

Plan how you will prove it works

For each measure you kept, confirm how you will prove it works: the method, what counts as evidence and who checks it.

  1. 1Confirm the verification method for each kept measure
  2. 2Accept or edit the evidence proposal
  3. 3Set who reviews the evidence (I0-I3)

M = measure from step 5 · V = verification method · 'Method expected' shows how strongly the method is expected at your AI-SIL

Required independence at AI-SIL 3

  • Review of the AI-SIL classification · independent of the responsible team
  • Review of the AASIF Safety Concept · independent of the responsible team
  • Go-live safety assessment (M38) · independent of the responsible team
  • Periodic control audit · independent of the responsible team

Item definition and operating envelope documented (task, authority, action space, contexts allowed)

  • Design review· Method expected: requiredConfirms: Patterns, constraints and guardrails are specified and fit the AI-SIL
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Least-privilege, scoped and non-transferable agent authority

  • Design review· Method expected: requiredConfirms: Patterns, constraints and guardrails are specified and fit the AI-SIL
  • Control audit· Method expected: requiredConfirms: Governance, roles and operational controls exist and are followed
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Unique agent identity registered in an agent catalog

  • Control audit· Method expected: requiredConfirms: Governance, roles and operational controls exist and are followed
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Hard action limits enforced outside the model (value caps, rate limits, allow-lists)

  • Design review· Method expected: requiredConfirms: Patterns, constraints and guardrails are specified and fit the AI-SIL
  • Scenario-based evaluation (evals)· Method expected: requiredConfirms: Behavior in representative, edge and failure scenarios incl. tool calls
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Reversibility by design: prefer reversible actions; staging or undo for irreversible ones

  • Design review· Method expected: requiredConfirms: Patterns, constraints and guardrails are specified and fit the AI-SIL
  • Scenario-based evaluation (evals)· Method expected: requiredConfirms: Behavior in representative, edge and failure scenarios incl. tool calls
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Oversight mode defined per action class (in / on / out of the loop)

  • Design review· Method expected: requiredConfirms: Patterns, constraints and guardrails are specified and fit the AI-SIL
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Approval gate for irreversible or above-threshold actions

  • Design review· Method expected: requiredConfirms: Patterns, constraints and guardrails are specified and fit the AI-SIL
  • Scenario-based evaluation (evals)· Method expected: requiredConfirms: Behavior in representative, edge and failure scenarios incl. tool calls
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Oversight sufficiency test (Sufficient / Nominal / Insufficient / Theatrical), incl. measured error-detection rate of reviewers

  • Control audit· Method expected: requiredConfirms: Governance, roles and operational controls exist and are followed
  • Runtime monitoring and log review· Method expected: requiredConfirms: Continuous evidence in operation
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Oversight metrics monitored (override rate, response time, outlier reviewers)

  • Runtime monitoring and log review· Method expected: requiredConfirms: Continuous evidence in operation
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Decision context package for reviewers (reasoning summary, evidence, flags)

  • Scenario-based evaluation (evals)· Method expected: requiredConfirms: Behavior in representative, edge and failure scenarios incl. tool calls
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Defined safe state and degradation modes

  • Design review· Method expected: requiredConfirms: Patterns, constraints and guardrails are specified and fit the AI-SIL
  • Scenario-based evaluation (evals)· Method expected: requiredConfirms: Behavior in representative, edge and failure scenarios incl. tool calls
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Escalation on uncertainty (calibrated uncertainty signal plus escalation rule; no raw confidence %)

  • Scenario-based evaluation (evals)· Method expected: requiredConfirms: Behavior in representative, edge and failure scenarios incl. tool calls
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Kill switch: halt and lock autonomous action

  • Scenario-based evaluation (evals)· Method expected: requiredConfirms: Behavior in representative, edge and failure scenarios incl. tool calls
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Behavioral anomaly and drift detection against a baseline

  • Runtime monitoring and log review· Method expected: requiredConfirms: Continuous evidence in operation
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Outcome monitoring beyond the target metric (second-order effects)

  • Runtime monitoring and log review· Method expected: requiredConfirms: Continuous evidence in operation
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Input and context validity check (data quality, staleness, domain match)

  • Scenario-based evaluation (evals)· Method expected: requiredConfirms: Behavior in representative, edge and failure scenarios incl. tool calls
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Incident and near-miss reporting process

  • Control audit· Method expected: requiredConfirms: Governance, roles and operational controls exist and are followed
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Decision logging (inputs, actions, model version, rationale)

  • Control audit· Method expected: requiredConfirms: Governance, roles and operational controls exist and are followed
  • Runtime monitoring and log review· Method expected: requiredConfirms: Continuous evidence in operation
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Tamper-evident logs

  • Control audit· Method expected: requiredConfirms: Governance, roles and operational controls exist and are followed
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Log retention and audit access defined

  • Control audit· Method expected: requiredConfirms: Governance, roles and operational controls exist and are followed
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Scenario-based evaluation before release (representative, edge and failure cases)

  • Scenario-based evaluation (evals)· Method expected: requiredConfirms: Behavior in representative, edge and failure scenarios incl. tool calls
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Known-unsafe scenarios detected and routed to humans

  • Scenario-based evaluation (evals)· Method expected: requiredConfirms: Behavior in representative, edge and failure scenarios incl. tool calls
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Unknown-unsafe exploration (red-teaming, adversarial edge cases)

  • Adversarial testing (red-teaming)· Method expected: requiredConfirms: Resistance to misuse, prompt injection and unauthorized actions
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Staged deployment (shadow → limited → full)

  • Control audit· Method expected: requiredConfirms: Governance, roles and operational controls exist and are followed
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Regression re-test after model, prompt or tool change

  • Periodic re-testing· Method expected: requiredConfirms: Evidence stays valid after changes
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Untrusted content isolation (external content treated as data; provenance marked)

  • Adversarial testing (red-teaming)· Method expected: requiredConfirms: Resistance to misuse, prompt injection and unauthorized actions
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Tool-call validation and safe output handling (schema checks, sandboxed execution)

  • Scenario-based evaluation (evals)· Method expected: requiredConfirms: Behavior in representative, edge and failure scenarios incl. tool calls
  • Adversarial testing (red-teaming)· Method expected: requiredConfirms: Resistance to misuse, prompt injection and unauthorized actions
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Secrets and credential isolation (no secrets in context, short-lived tokens)

  • Control audit· Method expected: requiredConfirms: Governance, roles and operational controls exist and are followed
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Supply-chain vetting of models, tools and agent products

  • Control audit· Method expected: requiredConfirms: Governance, roles and operational controls exist and are followed
  • Independent assessment / certification· Method expected: requiredConfirms: Third-party confirmation
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Data provenance and quality gate

  • Control audit· Method expected: requiredConfirms: Governance, roles and operational controls exist and are followed
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Fundamental-rights screening; formal FRIA where Art. 27 applies

  • Control audit· Method expected: requiredConfirms: Governance, roles and operational controls exist and are followed
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Bias and fairness testing on affected groups

  • Scenario-based evaluation (evals)· Method expected: requiredConfirms: Behavior in representative, edge and failure scenarios incl. tool calls
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Transparency to affected persons (AI disclosure, explanation, appeal path)

  • Control audit· Method expected: requiredConfirms: Governance, roles and operational controls exist and are followed
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Named accountable owner and responsibility matrix (value chain + three lines of defense)

  • Control audit· Method expected: requiredConfirms: Governance, roles and operational controls exist and are followed
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Re-classification triggers defined and monitored

  • Control audit· Method expected: requiredConfirms: Governance, roles and operational controls exist and are followed
  • Periodic re-testing· Method expected: requiredConfirms: Evidence stays valid after changes
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Minimum development process capability

  • Control audit· Method expected: requiredConfirms: Governance, roles and operational controls exist and are followed
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Independent safety assessment before go-live

  • Independent assessment / certification· Method expected: requiredConfirms: Third-party confirmation
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Operator and user training; AI literacy; end-user responsibility

  • Control audit· Method expected: requiredConfirms: Governance, roles and operational controls exist and are followed
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Agent interaction map and trust boundaries (freedom from interference)

  • Design review· Method expected: requiredConfirms: Patterns, constraints and guardrails are specified and fit the AI-SIL
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Authenticated and validated inter-agent messaging

  • Adversarial testing (red-teaming)· Method expected: requiredConfirms: Resistance to misuse, prompt injection and unauthorized actions
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Network circuit breakers (aggregate thresholds regardless of contributing agent)

  • Scenario-based evaluation (evals)· Method expected: requiredConfirms: Behavior in representative, edge and failure scenarios incl. tool calls
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Independence check for decomposed or layered controls (no shared base model, context or memory)

  • Design review· Method expected: requiredConfirms: Patterns, constraints and guardrails are specified and fit the AI-SIL
  • Independent assessment / certification· Method expected: requiredConfirms: Third-party confirmation
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Cross-firm interface contract (assume / guarantee safety obligations)

  • Control audit· Method expected: requiredConfirms: Governance, roles and operational controls exist and are followed
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Answers and statements only from authoritative, versioned sources; otherwise hand over to a human

  • Scenario-based evaluation (evals)· Method expected: requiredConfirms: Behavior in representative, edge and failure scenarios incl. tool calls
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

No commitments, offers or exceptions outside the agent's authority; such requests are routed to a human

  • Design review· Method expected: requiredConfirms: Patterns, constraints and guardrails are specified and fit the AI-SIL
  • Scenario-based evaluation (evals)· Method expected: requiredConfirms: Behavior in representative, edge and failure scenarios incl. tool calls
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Lawful-basis review of decision logic and data items before go-live and after rule changes

  • Control audit· Method expected: requiredConfirms: Governance, roles and operational controls exist and are followed
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Data minimization in agent context and outputs

  • Control audit· Method expected: requiredConfirms: Governance, roles and operational controls exist and are followed
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Output and egress control (no auto-rendered external links or images; egress allow-list)

  • Design review· Method expected: requiredConfirms: Patterns, constraints and guardrails are specified and fit the AI-SIL
  • Adversarial testing (red-teaming)· Method expected: requiredConfirms: Resistance to misuse, prompt injection and unauthorized actions
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Environment separation (sandbox or staging; gated production writes) and tested restore

  • Design review· Method expected: requiredConfirms: Patterns, constraints and guardrails are specified and fit the AI-SIL
  • Scenario-based evaluation (evals)· Method expected: requiredConfirms: Behavior in representative, edge and failure scenarios incl. tool calls
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

Audit by sampling of autonomous decisions

  • Runtime monitoring and log review· Method expected: requiredConfirms: Continuous evidence in operation
Reviewed by suggested: I2
Status
Implemented by (from “How you'll do it”, step 5)

Not filled in yet.

Edit in step 5

50 specified · 0 implemented · 0 verified

Step 6 complete — every selected measure has a status.

Back

Classification is deterministic — AI only adapts wording and suggests, you confirm.