Artificial intelligence is increasingly embedded in safety-critical and regulated systems, from autonomous vehicles and medical devices to industrial robots, aerospace platforms, rail infrastructure, and energy systems.
As AI assumes greater responsibility within these environments, organizations must demonstrate more than strong model performance. They must show that the complete AI-enabled system is acceptably safe for its intended purpose, users, operating conditions, and risk profile.
This is the purpose of an AI safety case.
An AI safety case is a structured, evidence-supported argument that an AI-enabled system is acceptably safe within a defined operational context. It connects safety claims to hazards, requirements, controls, tests, monitoring results, and other engineering artifacts so reviewers can determine whether the system is ready for certification or deployment.
A safety case is therefore not simply a collection of documents.
Certification authorities, regulators, auditors, and internal assurance teams need to understand:
- What the organization is claiming about system safety
- Which hazards and risks those claims address
- Why the selected controls are appropriate
- Whether the evidence is relevant, sufficient, and current
- Which system configuration the evidence applies to
- What assumptions and operational constraints remain
- How future model, data, prompt, interface, or environmental changes affect the argument
AI makes this process more complex because its behavior may be probabilistic, context-dependent, data-driven, and continuously evolving. A model may behave differently after retraining, fine-tuning, prompt changes, retrieval updates, new tool access, or deployment in a different environment.
AI safety case development must therefore combine traditional safety engineering with AI-specific evaluation, lifecycle traceability, continuous monitoring, and change impact analysis.
What Is an AI Safety Case?
An AI safety case is a documented and structured justification that an AI-enabled system does not create unacceptable risk when used for a defined application in a defined environment.
A strong safety case connects three foundational elements:
- Claims — Statements about the safety of the system
- Arguments — Reasoning that explains why those claims should be accepted
- Evidence — Verifiable artifacts supporting the arguments
This structure is commonly known as Claims–Arguments–Evidence, or CAE.
For example:
- Claim: The AI-assisted diagnostic system is acceptably safe for use by trained clinicians within the approved hospital environment.
- Argument: Risks are controlled through validated performance thresholds, human oversight, representative data, cybersecurity protections, and post-deployment monitoring.
- Evidence: Hazard analyses, clinical validation reports, test results, data documentation, usability studies, monitoring plans, traceability records, and approval reports.
The safety case is the reasoning framework that connects these artifacts. It explains why the available information provides sufficient confidence to accept the system within clearly defined boundaries.
AI Safety Case vs. Assurance Case
The terms safety case and assurance case are closely related but are not always identical.
A safety case focuses specifically on demonstrating that safety risks are acceptably controlled.
An assurance case may address a broader range of system qualities, including:
- Security
- Reliability
- Privacy
- Fairness
- Ethics
- Regulatory compliance
- Mission assurance
- Trustworthiness
For an AI-enabled system, the safety case may form one part of a wider assurance case that also addresses data governance, cybersecurity, explainability, human rights, organizational controls, and quality management.
AI Safety Case vs. Risk Assessment
A risk assessment identifies and evaluates hazards, harms, threats, failure modes, and associated controls.
A safety case goes further.
It uses the outputs of the risk assessment as part of an explicit argument that:
- Risks have been identified adequately
- Selected controls are appropriate
- Controls have been implemented
- Controls have been verified
- Residual risks meet defined acceptance criteria
- Operational monitoring can detect degradation or new hazards
The risk assessment is therefore an important source of evidence, but it is not the complete safety case.
AI Safety Case vs. Compliance Report
A compliance report explains how a system or process aligns with regulatory clauses, standards, or organizational policies.
A safety case must demonstrate why the resulting system should be considered acceptably safe.
Compliance can strengthen the argument, but compliance alone does not prove safety. A system may satisfy documented process requirements while still containing:
- Incomplete hazard analysis
- Weak safety assumptions
- Unrepresentative tests
- Missing edge-case coverage
- Ineffective operational controls
- Unsupported safety claims
AI Safety Case vs. Technical File
A technical file or certification package may contain:
- System descriptions
- Design documentation
- Requirements
- Risk assessments
- Test records
- Validation reports
- Compliance matrices
- User instructions
- Configuration records
The safety case provides the logic connecting these artifacts to the overall safety conclusion.
| Artifact | Primary question |
Risk assessment| What can go wrong, and how serious is it? |
|
| Safety case | Why should the system be accepted as sufficiently safe? |
| Compliance report | Which regulatory or standard obligations have been addressed? |
| Technical file | Which engineering and compliance records support approval? |
Why AI Safety Cases Matter for Certification
Certification requires confidence that a system has been designed, developed, evaluated, and controlled appropriately for its intended use.
For AI-enabled systems, this confidence cannot be based solely on accuracy, benchmark performance, or successful functional demonstrations.
Certification stakeholders may need evidence relating to:
- Safety requirements
- Hazard analysis
- Data quality
- Model behavior
- Robustness
- Human oversight
- Cybersecurity
- Misuse prevention
- Explainability
- Configuration management
- Supplier dependencies
- Operational monitoring
- Change management
- Incident response
An AI safety case organizes these elements into a coherent and reviewable structure.
Turning Regulatory Obligations into Technical Claims
Regulations and standards frequently define high-level obligations that must be translated into technical and operational claims.
For example, a regulation may require effective human oversight.
The safety case can translate this obligation into claims concerning:
- Operator authority
- Interface design
- Override mechanisms
- User competence
- Response time
- Alert presentation
- Fallback behavior
- Escalation procedures
Those claims can then be supported by:
- Requirements
- Design specifications
- Human-factors studies
- Simulations
- Usability tests
- Training records
- Operational procedures
This transformation from obligation to claim, requirement, control, and evidence is central to certification readiness.
Demonstrating That Risks Are Controlled
A safety case should show that every significant hazard or risk is connected to:
- A safety objective
- One or more risk controls
- Implemented requirements
- Verification activities
- Acceptance criteria
- Supporting evidence
- A residual-risk decision
This makes the organization’s risk-control strategy visible, reviewable, and defensible.
Supporting Independent Review
Certification bodies and independent assessors need more than assurances from the development team.
They may need to examine:
- The logic of the argument
- The quality of the evidence
- The completeness of traceability
- The validity of assumptions
- The handling of conflicting information
- The independence of reviews
- The configuration to which evidence applies
- The management of unresolved findings
A structured safety case allows assessors to identify gaps without reconstructing the complete engineering lifecycle manually.
Maintaining Accountability
AI systems often involve multiple teams and organizations, including:
- Model providers
- System integrators
- Data providers
- Software suppliers
- Safety engineers
- Quality teams
- Operators
- Regulators
- Independent assessors
The safety case becomes a shared assurance artifact that clarifies responsibility for each claim, control, evidence item, review, approval, and update.
How AI Safety Cases Differ from Traditional Safety Cases
Traditional safety cases were developed primarily for systems whose architectures, functions, interfaces, and failure modes could be specified before deployment.
AI challenges several of these assumptions.
Probabilistic and Nondeterministic Behavior
Traditional software is generally expected to produce predictable outputs for a given input and system state.
AI systems may produce different outputs depending on:
- Model sampling
- Confidence thresholds
- Input wording
- Context windows
- Retrieval results
- Environmental data
- Model version
- Fine-tuning
- Connected tools
The safety argument must therefore address output distributions, uncertainty, rare failures, and behavior across a range of operational conditions.
Emergent Capabilities
AI capabilities may not be fully known before evaluation.
New behaviors can emerge through:
- Greater model scale
- Fine-tuning
- Prompting strategies
- Tool access
- Agent orchestration
- Retrieval augmentation
- Multi-model integration
Assurance cannot rely entirely on pre-defined functional requirements. Capability discovery, adversarial testing, exploratory evaluation, and operational monitoring may be needed to determine what the system can actually do.
Limited Ground Truth
Some AI tasks do not have a single stable or objectively correct answer.
Examples include:
- Medical recommendations
- Risk prioritization
- Engineering design assistance
- Natural-language analysis
- Complex classification
- Autonomous decision-making
- Human preference prediction
When ground truth is incomplete, evidence may rely on:
- Expert agreement
- Comparative performance
- Simulated scenarios
- Proxy metrics
- Operational outcomes
- Statistical confidence
- Error analysis
The resulting arguments may be inductive, comparative, or statistical rather than strictly deductive.
Context-Dependent Behavior
AI behavior is influenced by the complete system context, including:
- Prompts
- User roles
- Interfaces
- Retrieval sources
- Tool permissions
- Workflows
- Environmental conditions
- Human decisions
- Organizational procedures
The underlying model cannot always be assured independently from the system in which it operates.
The safety case must therefore address the complete AI-enabled system.
Continuous Model and Data Change
AI systems may change through:
- Retraining
- Fine-tuning
- Dataset updates
- New model versions
- Threshold adjustments
- Prompt changes
- Policy changes
- Retrieval updates
- New tool integrations
Each change can invalidate existing evidence or alter the risk profile.
Traditional certification often assumes a stable baseline. AI assurance increasingly requires a dynamic or continuously maintained safety case.
Third-Party Model Dependencies
Organizations may use externally supplied foundation models that are updated outside their direct control.
The safety case must answer questions such as:
- Which model version was evaluated?
- Can the model change without notice?
- Which evidence is available from the provider?
- Which claims depend on supplier assertions?
- How are provider changes detected?
- Which compensating controls address evidence gaps?
- What happens if the provider changes the model, interface, or terms?
Human-AI Interaction
Many AI risks arise through interaction between the model and human users.
Examples include:
- Automation bias
- Overreliance
- Misinterpretation
- Poor confidence indicators
- Inadequate override mechanisms
- Insufficient training
- Alert fatigue
- Unsafe task allocation
Human-factors evidence must therefore be integrated into the safety argument.
Agentic AI Risks
Agentic AI systems can plan, invoke tools, access data, modify files, call APIs, and interact with other systems.
Safety claims may need to address:
- Permission boundaries
- Tool restrictions
- Action validation
- Goal misinterpretation
- Unauthorized escalation
- Long-horizon behavior
- Loss of control
- Human intervention
- Fail-safe states
- Emergency shutdown
These characteristics shift assurance from isolated model evaluation toward system-level control and governance.
| Dimension | Traditional engineered system | AI-enabled system |
| Behavior | Primarily specification-driven | Learned, probabilistic, and context-sensitive |
| Failure analysis | Component and logic focused | Includes data, model, interaction, and context failures |
| Requirements | Generally explicit | May be incomplete for emergent behavior |
| Change sources | Controlled engineering changes | Models, data, prompts, tools, users, and environment |
| Validation | Requirement-based testing | Scenario, statistical, adversarial, and operational evaluation |
| Evidence lifecycle | Often stabilizes after approval | May require continuous refresh |
What Are Claims, Arguments, and Evidence?
The CAE structure is the foundation of a well-organized safety case.
Safety Claims
A claim is a statement that must be shown to be true or sufficiently justified.
A top-level claim might state:
The AI-enabled system is acceptably safe for its intended use within the defined operational environment.
This statement is too broad to support directly. It must be decomposed into more specific subclaims.
Possible subclaims include:
- The intended use is clearly defined.
- Foreseeable hazards have been identified.
- Safety requirements are complete and correct.
- Risk controls have been implemented.
- The model operates within defined safety limits.
- Data quality is suitable for the intended population.
- Human oversight is effective.
- Security risks are controlled.
- Operational monitoring detects degradation.
- Changes trigger re-evaluation.
Types of AI Safety Claims
Assertion-Based Claims
These state that a safety property has been achieved.
Example:
The system detects pedestrians with sufficient reliability within the approved Operational Design Domain.
Constraint-Based Claims
These define the conditions under which the system is considered safe.
Example:
The model is used only for decision support and cannot autonomously approve treatment.
Capability-Based Claims
These address whether an AI system possesses or lacks a safety-relevant capability.
Example:
The model cannot independently execute privileged system commands.
Comparative Claims
These compare the AI system with an established baseline.
Example:
The AI-assisted review process does not increase the rate of missed safety requirements compared with the approved manual process.
Process Claims
These concern the development, review, or governance process.
Example:
All model changes undergo documented impact analysis and revalidation.
Operational Claims
These concern runtime behavior and controls.
Example:
Performance degradation is detected before it creates unacceptable operational risk.
Safety Arguments
An argument explains why the evidence supports the claim.
Possible argument strategies include the following.
Requirements-Based Arguments
The claim is supported by demonstrating that:
- Safety requirements were derived correctly.
- Requirements were implemented.
- Requirements were verified.
- Validation confirms intended safety outcomes.
Hazard-Based Arguments
The top-level claim is decomposed according to identified hazards.
Each hazard is connected to:
- Causes
- Consequences
- Controls
- Safety requirements
- Tests
- Residual risk
Risk-Based Arguments
The organization argues that:
- Risks have been identified.
- Risk levels are understood.
- Controls reduce the risks sufficiently.
- Residual risks satisfy acceptance criteria.
Capability-Based Arguments
The organization evaluates whether the AI possesses safety-relevant or dangerous capabilities.
Examples may include:
- Autonomous planning
- Cyber exploitation
- Deception
- High-impact decision-making
- Tool manipulation
- Self-modification
Comparative Arguments
The AI system is compared with:
- Human performance
- A legacy system
- A non-AI baseline
- An earlier approved model
- An alternative technical solution
The baseline must itself be justified as acceptable and relevant.
Defense-in-Depth Arguments
Safety is supported by multiple barriers, such as:
- Model alignment
- Input filtering
- Permission restrictions
- Output validation
- Human approval
- Runtime monitoring
- Emergency shutdown
Operational Arguments
Operational arguments rely on:
- Monitoring
- Incident detection
- Safety performance indicators
- Human intervention
- Feedback loops
- Corrective action
Safety Evidence
Evidence consists of verifiable artifacts used to support the argument.
Evidence should be:
- Relevant to the claim
- Valid for the approved configuration
- Sufficient in coverage
- Credible
- Traceable
- Current
- Reviewable
Evidence families may include:
- Empirical evidence
- Formal evidence
- Comparative evidence
- Expert judgment
- Model-based analysis
- Mechanistic analysis
- Operational data
- Process and governance records
Most AI safety cases will need several evidence families rather than a single form of proof.
Context, Assumptions, and Justifications
Safety claims are rarely universally true.
They depend on context such as:
- Intended users
- Deployment environment
- Model version
- Dataset version
- Geographic region
- User training
- System interfaces
- Infrastructure
- Operational Design Domain
Assumptions should be explicit.
Examples include:
- Operators are trained.
- Network connectivity is available.
- Input sensors operate within specified tolerances.
- Users do not bypass access controls.
- The model provider communicates version changes.
A hidden assumption can become a major weakness in the safety argument.
Defeaters and Counterarguments
A defeater is information that weakens or invalidates a claim.
Examples include:
- A failed test
- An unexplained performance decline
- A newly discovered hazard
- A security vulnerability
- A model update
- Evidence that excludes part of the intended population
- Contradictory evaluation results
- An unresolved audit finding
Strong safety cases do not ignore these issues. They explain how each defeater has been resolved, mitigated, accepted, or escalated.
Goal Structuring Notation for AI Safety Cases
Claims–Arguments–Evidence describes the conceptual structure of a safety case.
Goal Structuring Notation, or GSN, provides a graphical method for representing that structure.
A GSN model can show:
- Goals or claims
- Strategies or argument approaches
- Solutions or evidence
- Context
- Assumptions
- Justifications
- Undeveloped claims
GSN can help reviewers understand how high-level safety goals are supported by lower-level requirements, controls, evaluations, and evidence.
A typical representation may use:
- Rectangles for goals or claims
- Parallelograms for strategies
- Circles for solutions or evidence
- Rounded rectangles for context
- Ovals for assumptions or justifications
GSN does not replace engineering evidence. It makes the reasoning behind that evidence more visible and reviewable.
Core Components of an AI Safety Case
A complete safety case should include more than a top-level claim and supporting files.
Core components include:
- Intended use
- System definition
- Operational context
- Applicable regulations and standards
- Safety claims
- Argument structures
- Hazards and risks
- Safety requirements
- Controls
- Verification and validation evidence
- Data and model evidence
- Assumptions
- Known limitations
- Defeaters
- Residual uncertainty
- Traceability
- Monitoring provisions
- Change-management rules
- Review and approval records
- Configuration information
System Definition and Intended Use
The safety case should define the complete AI macrosystem, including:
- AI models
- Non-AI software
- Hardware
- Data pipelines
- Sensors
- Human operators
- External services
- Tools and APIs
- Deployment infrastructure
- Monitoring systems
Operational Design Domain
The Operational Design Domain, or ODD, defines the boundaries within which the system is intended to operate.
It may specify:
- Environmental conditions
- Geographic constraints
- Input characteristics
- User qualifications
- Hardware assumptions
- Operating speeds
- Network conditions
- Prohibited conditions
- Required human supervision
Safety claims should not extend beyond the conditions supported by evidence.
Hazard and Risk Analysis
The safety case should identify how the AI-enabled system can contribute to harm.
Relevant methods may include:
- Failure Mode and Effects Analysis
- Failure Mode, Effects, and Criticality Analysis
- Fault Tree Analysis
- Hazard and Operability Study
- Systems-Theoretic Process Analysis
- Hazard Analysis and Risk Assessment
- Threat Analysis and Risk Assessment
- Misuse-case analysis
- Human-factors analysis
AI-specific hazards may involve:
- Incorrect predictions
- Hallucinations
- Distribution shift
- Model drift
- Data bias
- Prompt injection
- Unsafe tool use
- Automation bias
- Loss of control
- Inadequate fallback
- Supplier model changes
How to Develop an AI Safety Case Step by Step
Step 1: Define the Intended Use and Operational Context
Start by defining:
- Intended users
- Intended purpose
- Operating environment
- System boundaries
- Inputs and outputs
- Dependencies
- Prohibited uses
- Foreseeable misuse
- Performance limits
- Human responsibilities
- Fallback behavior
A system demonstrated as safe in a laboratory may not be safe in a public, industrial, clinical, or adversarial environment.
Step 2: Identify Applicable Regulations and Standards
Determine which obligations apply based on:
- Industry
- Geography
- Product classification
- AI risk category
- Safety integrity level
- Intended use
- Software and hardware architecture
- Cybersecurity exposure
Create an applicability matrix connecting each obligation to:
- Safety claims
- Requirements
- Controls
- Evidence
- Reviews
Step 3: Identify Hazards and AI-Specific Risks
Analyze:
- Functional failures
- Data failures
- Model failures
- Human-interaction risks
- Security risks
- Misuse
- Environmental risks
- Supplier risks
- Operational degradation
The analysis should include intended use, foreseeable misuse, abnormal conditions, and credible adversarial scenarios.
Step 4: Define the Top-Level Safety Claim
The top-level claim should be:
- Clear
- Bounded
- Conditional
- Version-specific
- Context-specific
- Reviewable
A useful claim pattern is:
AI system version [X] is acceptably safe for [intended use] by [intended users] within [approved environment], subject to [constraints] and [monitoring controls].
This is more defensible than simply stating that the AI system is safe.
Step 5: Decompose the Claim
Break the top-level claim into subclaims.
Decomposition can follow:
- Hazards
- Lifecycle stages
- Components
- Safety functions
- Standards
- Risk controls
- Organizational responsibilities
- Operational conditions
Each subclaim should be narrow enough to support with specific evidence.
Step 6: Select the Argument Strategy
For each claim, define how the evidence will support it.
The safety case may combine:
- Requirements-based reasoning
- Hazard-based reasoning
- Risk-based reasoning
- Capability-based reasoning
- Comparative reasoning
- Defense-in-depth
- Process assurance
- Operational assurance
Step 7: Define Evidence Requirements
Create an evidence plan before testing begins.
For each evidence item, define:
- Evidence description
- Supported claim
- Responsible owner
- Creation method
- Acceptance criteria
- Required independence
- Applicable standard
- Review authority
- Version
- Refresh trigger
This prevents teams from discovering late in the certification process that essential evidence was never generated.
Step 8: Establish Bidirectional Traceability
Build traceability across:
Regulation → Hazard → Safety claim → Requirement → Risk control → Design element → Test case → Test result → Evidence
Traceability should allow reviewers to answer:
- Which requirements address this hazard?
- Which tests verify this control?
- Which claim does this evidence support?
- Which standard clause requires the artifact?
- What is affected by a model change?
- Are any claims unsupported?
- Are any requirements untested?
Step 9: Assess Evidence Quality
Evidence should be evaluated for:
- Relevance
- Completeness
- Credibility
- Coverage
- Independence
- Repeatability
- Statistical confidence
- Configuration consistency
- Operational applicability
- Freshness
A result generated for an outdated model cannot support the current system unless equivalence is justified.
Step 10: Review Gaps and Defeaters
Search for:
- Unsupported claims
- Missing tests
- Unverified controls
- Open hazards
- Conflicting evidence
- Hidden assumptions
- Unresolved audit findings
- Outdated evidence
- Incomplete supplier information
Step 11: Conduct Independent Review
Independent reviewers should evaluate:
- Argument logic
- Evidence sufficiency
- Assumptions
- Risk acceptance
- Traceability
- Configuration
- Remaining uncertainty
The required level of independence depends on the applicable standards and risk level.
Step 12: Prepare the Certification Package
Organize the safety case into a controlled and reviewable submission.
The package should identify:
- Approved baseline
- Applicable configuration
- Evidence index
- Review status
- Approval status
- Open issues
- Residual risks
- Conditions of use
Step 13: Maintain the Safety Case
The safety case should remain active after approval.
Define:
- Monitoring metrics
- Review intervals
- Change triggers
- Incident triggers
- Revalidation thresholds
- Recertification criteria
- Evidence owners
What Evidence Is Needed for AI Certification?
Evidence requirements depend on the industry, jurisdiction, intended use, risk classification, and applicable standards.
Most AI certification safety cases require evidence from several categories.
Requirements and Design Evidence
Requirements evidence may include:
- Intended-use requirements
- System requirements
- Safety requirements
- AI performance requirements
- Data requirements
- Human-oversight requirements
- Interface requirements
- Cybersecurity requirements
- Fallback requirements
- Monitoring requirements
Design evidence may include:
- System architecture
- Model architecture
- Data-flow diagrams
- Interface descriptions
- Safety barriers
- Redundancy strategies
- Fail-safe behavior
- Deployment architecture
Risk Management Evidence
Risk evidence may include:
- Hazard analysis
- FMEA or FMECA
- Fault Tree Analysis
- HAZOP
- STPA
- Misuse analysis
- Threat modeling
- Human-factors analysis
- Residual-risk evaluations
- Risk-acceptance records
Data Evidence
Relevant data evidence may include:
- Data provenance
- Collection methods
- Dataset version
- Representativeness analysis
- Labeling procedures
- Label accuracy
- Missing-data analysis
- Bias assessment
- Data-leakage controls
- Privacy controls
- Data-quality metrics
- Data preprocessing
- Training-validation separation
- Dataset limitations
The safety case should explain why the data is suitable for the intended use and population.
Model Development Evidence
Model development evidence may include:
- Model architecture
- Training objectives
- Training procedures
- Hyperparameters
- Fine-tuning records
- Training environment
- Model version
- Reproducibility records
- Performance limitations
- Known failure modes
- Supplier documentation
- Model cards
- System cards
When third-party models are used, evidence gaps and compensating controls should be documented.
Verification and Validation Evidence
Verification asks whether the system was built correctly.
Validation asks whether the correct system was built for the intended use.
Evidence may include:
- Unit testing
- Integration testing
- System testing
- Requirements-based testing
- Scenario testing
- Simulation
- Stress testing
- Boundary testing
- Robustness testing
- Fault injection
- Adversarial testing
- Performance evaluation
- Subgroup performance analysis
- Human-factors validation
- Operational trials
Security and Misuse Evidence
Security and safety increasingly overlap in AI-enabled systems.
Evidence may include:
- Threat models
- Penetration testing
- Prompt-injection testing
- Access-control testing
- Tool-permission testing
- Data-exfiltration testing
- Adversarial examples
- Abuse-case testing
- Red-team reports
- Incident-response exercises
- Logging validation
- Supply-chain assessments
Explainability and Transparency Evidence
Depending on the system, this may include:
- Model cards
- System cards
- Explanation methods
- User-facing confidence indicators
- Known limitation documentation
- Decision logs
- Instructions for use
- Human-oversight guidance
- Interpretability evaluations
Explainability should be connected to a safety purpose. An explanation that appears convincing but is technically unreliable may weaken the argument.
Operational Evidence
Operational evidence may include:
- Runtime logs
- Safety Performance Indicators
- Drift metrics
- Override events
- User complaints
- Incident reports
- Near-miss reports
- Control failures
- Escalation records
- Corrective actions
- Field performance
- Maintenance records
Operational evidence helps determine whether assumptions made during certification remain valid.
Organizational Evidence
Organizational evidence may include:
- Roles and responsibilities
- Competence records
- Training
- Review approvals
- Quality procedures
- Supplier assessments
- Configuration management
- Change control
- Audit records
- Independence records
- Governance decisions
Positive Evidence, Negative Evidence, and Safety Confidence
A mature safety case should not contain only information that supports the desired conclusion.
It should also show that the organization actively searched for evidence that could weaken or disprove the argument.
Positive Evidence
Positive evidence directly supports a safety claim.
Examples include:
- Passing test results
- Successful simulations
- Verified requirements
- Demonstrated performance
- Approved reviews
- Confirmed risk controls
Negative Evidence
Negative evidence may include:
- The failure of a capable red team to bypass controls
- The absence of unsafe behavior across a justified evaluation set
- The inability of the model to demonstrate a hazardous capability
- A systematic search that did not identify credible counterexamples
Negative evidence must be interpreted carefully.
Failure to discover a problem does not prove that the problem cannot exist.
The safety case should justify:
- Why the evaluation was sufficiently challenging
- Why scenarios were representative
- Why the red team was competent and appropriately incentivized
- Why test duration and coverage were adequate
- Why remaining undetected failures are acceptably unlikely
Red-Team Evidence
Red teams attempt to:
- Bypass safeguards
- Trigger unsafe behavior
- Misuse tools
- Escalate privileges
- Expose data
- Induce deception
- Circumvent controls
- Demonstrate harmful capabilities
A red-team report should document:
- Objectives
- Scope
- Team competence
- Methods
- Findings
- Reproducibility
- Severity
- Corrective actions
- Retesting
Residual Uncertainty
Every AI safety case contains uncertainty.
The objective is not to claim perfect safety. The objective is to show that uncertainty has been:
- Identified
- Evaluated
- Controlled
- Monitored
- Communicated
- Accepted by authorized decision-makers
What Makes Evidence Certification-Ready?
Evidence is certification-ready when it provides justified confidence for the claim within the defined context.
More evidence is not always better. Evidence must be appropriate.
Relevance
Does the evidence directly support the claim?
A high accuracy score, for example, does not automatically support a claim about cybersecurity or effective human oversight.
Coverage
Does the evidence cover:
- Relevant hazards
- Operating conditions
- User groups
- Edge cases
- Failure conditions
- Foreseeable misuse
- Adversarial scenarios
Validity
Does the method actually measure what it claims to measure?
Repeatability and Reproducibility
Can the result be reproduced under controlled conditions?
Independence
Was the evidence created or reviewed with an appropriate level of independence?
Statistical Confidence
Are sample sizes, uncertainty, confidence intervals, and error distributions adequate?
Configuration Consistency
Does the evidence apply to the exact:
- Model
- Dataset
- Prompt
- Software version
- Hardware
- Tool configuration
- Operating environment
Operational Relevance
Does the evaluation reflect real deployment conditions?
Integrity and Provenance
Can reviewers verify:
- Where the evidence originated
- Who created it
- Whether it was changed
- Which approval applies
- Which baseline contains it
Freshness
Has the system changed since the evidence was produced?
| Criterion | Review question |
| Relevance | Does the artifact directly support the stated claim? |
| Coverage | Does it address the required risks and scenarios? |
| Configuration | Was it generated for the approved system version? |
| Integrity | Is provenance and change history available? |
| Independence | Was an appropriate independent review performed? |
| Confidence | Are uncertainty and limitations recorded? |
| Freshness | Is the evidence still valid after recent changes? |
Building End-to-End Safety Traceability
Traceability is the foundation of a defensible AI safety case.
A complete chain may include:
- Regulatory obligation
- Safety objective
- Hazard
- Risk control
- Safety requirement
- Design element
- Verification activity
- Test result
- Evidence item
- Safety claim
- Review approval
Regulatory Requirement to Safety Claim
Each applicable obligation should map to one or more explicit claims.
Example:
- Obligation: Human oversight must be effective.
- Claim: Trained operators can identify and override unsafe recommendations before those recommendations affect the controlled process.
Hazard to Safety Requirement
A hazard should lead to control requirements.
- Hazard: The model recommends an unsafe operating parameter.
- Requirement: The system shall prevent AI-generated parameters from exceeding approved safety limits.
Requirement to Test
The requirement should map to verification.
- Test: Submit recommendations outside the approved range and confirm that the system blocks execution.
Test Result to Evidence
The approved test result becomes evidence.
It should identify:
- System version
- Test environment
- Input data
- Expected result
- Actual result
- Reviewer
- Date
- Status
Detecting Orphaned Artifacts
Traceability analysis should identify:
- Claims without evidence
- Requirements without source obligations
- Hazards without controls
- Controls without tests
- Tests without requirements
- Evidence not connected to claims
Reusable AI Safety Case Patterns
Reusable patterns help organizations address recurring assurance problems.
Hazard-Based Pattern
The safety claim is decomposed according to identified hazards.
Use when:
- Hazards are well understood
- Safety controls are traceable
- Sector standards require hazard-based reasoning
Capability-Based Pattern
The case addresses whether the AI possesses or lacks a dangerous capability.
Use for:
- Frontier AI
- Agentic systems
- Misuse risks
- High-impact tool use
Comparative Pattern
The system is shown to be no worse than an accepted baseline.
Use carefully when:
- Ground truth is incomplete
- Existing human or system performance provides a comparator
- The baseline’s acceptability is justified
Defense-in-Depth Pattern
Multiple barriers reduce the probability or consequence of failure.
Use for:
- Cybersecurity
- Misuse prevention
- Autonomous action
- High-consequence decisions
Human-Oversight Pattern
The argument shows that human intervention effectively reduces risk.
Evidence may include:
- User training
- Interface validation
- Override testing
- Response-time analysis
- Workload evaluation
Continuous-Evolution Pattern
The safety case incorporates monitoring, change detection, evidence refresh, and revalidation.
Use for systems involving:
- Frequent model updates
- Drift
- Learning components
- External model providers
- Changing retrieval content
Threshold-Based Pattern
The argument depends on defined acceptance thresholds.
Examples include:
- Error rates
- Safety scores
- Capability levels
- Drift metrics
- Incident rates
Thresholds must be justified rather than selected arbitrarily.
Dynamic AI Safety Cases and Continuous Assurance
A static safety case may become obsolete quickly when AI components change.
A dynamic AI safety case links claims to system configurations, evidence, monitoring metrics, and change triggers so the argument can be reviewed and updated continuously.
Why Static Safety Cases Become Obsolete
A static case may be invalidated by:
- A new model version
- Retraining
- Fine-tuning
- Dataset updates
- New prompts
- New retrieval sources
- New user groups
- New deployment environments
- Added tools
- Security vulnerabilities
- Incidents
- Regulatory changes
Safety Performance Indicators
Safety Performance Indicators may include:
- Error rate
- Override frequency
- Unsafe-output rate
- Drift score
- Confidence calibration
- Control-failure rate
- Incident rate
- Security-event rate
- Out-of-distribution input rate
Each indicator should connect to:
- A safety claim
- A threshold
- An owner
- An escalation process
- A review trigger
Checkable Safety Arguments
A checkable argument links a safety claim to evidence or a system condition that can be monitored automatically or semi-automatically.
For example:
The model’s error rate remains below the approved threshold.
This claim can be linked to a monitored metric.
If the metric exceeds the threshold, the claim may become unsupported and require review.
Revalidation Triggers
Revalidation may be required after:
- A model update
- A dataset change
- An interface change
- A prompt change
- New tool access
- A supplier change
- A new hazard
- An incident
- A regulatory update
- A monitoring threshold breach
| Change | Likely safety-case impact |
| New model version | Reassess capabilities, robustness, performance, and failure modes |
| Dataset update | Reassess representativeness, bias, drift, and data quality |
| New tool access | Reassess authorization, misuse, and loss-of-control risks |
| Prompt update | Revalidate behavior and operational constraints |
| New environment | Reassess context assumptions and scenario coverage |
| Safety incident | Reopen affected hazards, controls, claims, and risk decisions |
How Model Updates Affect Certification Evidence
AI changes should not be treated as routine software updates without safety analysis.
Foundation Model Changes
A new foundation-model version may affect:
- Capability
- Behavior
- Output format
- Safety filters
- Tool use
- Bias
- Robustness
- Latency
- Explainability
Previous evidence may no longer apply.
Fine-Tuning and Retraining
Fine-tuning can improve intended performance while creating new failure modes.
The safety case should identify:
- Training-data changes
- Objective changes
- Behavioral changes
- Regression testing
- Updated risks
- Updated evidence
Prompt Changes
A system prompt can materially change model behavior.
Prompt changes should be configuration-controlled and evaluated for:
- Safety-policy compliance
- Tool use
- Refusal behavior
- Hallucination
- Role interpretation
- User manipulation
Retrieval Updates
Changes to a retrieval knowledge base may affect:
- Accuracy
- Bias
- Currency
- Source reliability
- Confidentiality
- Unsafe recommendations
Tool and Permission Changes
Giving an AI system access to additional tools can transform a low-risk advisory model into a high-impact agent.
The safety case should reassess:
- Authorization
- Validation
- Action boundaries
- Human approval
- Logging
- Fail-safe behavior
Standards and Regulatory Frameworks Relevant to AI Safety Cases
Applicable standards depend on the system’s industry, jurisdiction, intended use, and risk classification.
ISO/IEC 42001
ISO/IEC 42001 provides requirements for an Artificial Intelligence Management System.
It can support safety-case evidence concerning:
- Governance
- Roles and responsibilities
- Risk management
- Lifecycle controls
- Documentation
- Monitoring
- Continual improvement
ISO/IEC 23894
ISO/IEC 23894 provides guidance on AI risk management.
It can support:
- Risk identification
- Risk analysis
- Risk treatment
- Monitoring
- Communication
NIST AI Risk Management Framework
The NIST AI RMF organizes AI risk activities around:
- Govern
- Map
- Measure
- Manage
It can inform governance and risk evidence within an AI safety case.
EU AI Act
For applicable high-risk AI systems, safety-case evidence may support obligations relating to:
- Risk management
- Data governance
- Technical documentation
- Record keeping
- Transparency
- Human oversight
- Accuracy
- Robustness
- Cybersecurity
- Post-market monitoring
The EU AI Act does not use “safety case” as the required term in every provision. However, a structured safety case can help organizations organize the technical and governance evidence required for conformity assessment.
IEC 61508
IEC 61508 is relevant to functional safety for electrical, electronic, and programmable electronic systems.
An AI safety case may need to explain how AI components interact with safety functions and whether deterministic safety mechanisms provide adequate control.
ISO 26262 and ISO 21448
In automotive systems:
- ISO 26262 addresses functional safety.
- ISO 21448 addresses Safety of the Intended Functionality.
AI safety cases may need to address perception limitations, environmental scenarios, behavioral uncertainty, and the Operational Design Domain.
ISO/SAE 21434
ISO/SAE 21434 addresses automotive cybersecurity.
Where cyber threats can create safety consequences, Threat Analysis and Risk Assessment should be connected to the safety argument.
ISO 14971 and IEC 62304
For medical devices:
- ISO 14971 supports risk management.
- IEC 62304 addresses medical-device software lifecycle processes.
AI-enabled devices may also require clinical evaluation, data governance, usability evidence, and post-market monitoring.
DO-178C, ARP4754A, and ARP4761
In aerospace:
- DO-178C addresses airborne software assurance.
- ARP4754A supports aircraft and system development.
- ARP4761 supports safety assessment.
AI components may require additional assurance arguments where traditional objectives do not fully address learned behavior.
EN 50126, EN 50128, and EN 50129
Rail applications may require evidence concerning:
- Reliability
- Availability
- Maintainability
- Safety
- Software assurance
- Safety-related electronic systems
- Safety-case development
UL 4600
UL 4600 supports structured safety arguments for autonomous products and may be relevant for systems operating without a human driver or operator.
IEC 62443
IEC 62443 may be relevant where industrial cybersecurity threats can create safety consequences.
Standards should not be presented as interchangeable. Applicability depends on the system, sector, architecture, market, and certification scheme.
What Should an AI Safety Case Certification Package Include?
A comprehensive certification package may include:
- Executive summary
- Intended use
- System scope
- Operational context
- System architecture
- AI model description
- Applicable regulations and standards
- Safety claims
- Argument structure
- Hazard analysis
- Risk controls
- Safety requirements
- Data documentation
- Model development records
- Verification results
- Validation results
- Human-factors evidence
- Security and misuse testing
- Traceability matrix
- Residual-risk decisions
- Known limitations
- Monitoring plan
- Change-management plan
- Incident-response process
- Configuration baseline
- Evidence index
- Review and approval records
- Open issues and conditions of use
Using AI to Support Safety Case Development
AI can accelerate parts of the safety-case workflow.
Potential applications include:
- Generating initial claim structures
- Suggesting subclaims
- Classifying evidence
- Detecting missing traceability links
- Mapping standards to requirements
- Generating test cases
- Summarizing reports
- Identifying inconsistent terminology
- Supporting change impact analysis
- Monitoring evidence status
AI assistance can reduce manual effort, but AI-generated safety arguments may contain:
- Hallucinations
- Unsupported assumptions
- Incorrect standard mappings
- Missing evidence
- Weak reasoning
- Fabricated references
A hybrid approach is therefore essential.
AI can assist with drafting, classification, and analysis, while qualified humans validate:
- Claims
- Argument logic
- Evidence
- Assumptions
- Standard interpretation
- Residual-risk decisions
Automation should support traceability and analysis, not replace accountable approval.
AI Safety Case Development by Industry
Automotive
Key focus areas include:
- Perception
- Object detection
- Operational Design Domain
- Driver interaction
- Fallback behavior
- Scenario coverage
- Cybersecurity
- SOTIF
Aerospace and Defense
Focus areas may include:
- Mission assurance
- Software and system assurance
- Autonomous behavior
- Human authority
- Explainability
- Configuration control
- Independent verification
Medical Devices
Focus areas include:
- Clinical safety
- Patient populations
- Bias
- Diagnostic performance
- Human oversight
- Usability
- Post-market surveillance
Rail
Focus areas include:
- Signaling
- Traffic management
- Predictive maintenance
- Safety integrity
- Operational constraints
- Fail-safe behavior
Industrial Automation and Robotics
Focus areas include:
- Human-robot interaction
- Safe motion
- Workspace boundaries
- Sensor uncertainty
- Emergency stop
- Cyber-physical security
Critical Infrastructure
Focus areas include:
- Resilience
- Cybersecurity
- Human control
- System availability
- Incident response
- Supplier risk
Frontier and Agentic AI
Focus areas include:
- Dangerous capabilities
- Loss of control
- Tool use
- Deception
- Privilege escalation
- Monitoring
- Containment
- Emergency shutdown
Example: AI-Enabled Medical Device Safety Case
Consider an AI system that assists clinicians in identifying abnormalities in medical images.
Intended Use
The system provides diagnostic support to trained radiologists. It does not make autonomous clinical decisions.
Top-Level Claim
The AI-assisted imaging system is acceptably safe for use by trained radiologists within the approved clinical workflow.
Key Hazards
- False negatives
- False positives
- Performance bias
- Poor-quality input
- Automation bias
- Model drift
- Cybersecurity compromise
Safety Controls
- Minimum image-quality checks
- Confidence thresholds
- Human review
- Warning messages
- Restricted use
- Performance monitoring
- Access controls
- Periodic revalidation
Supporting Evidence
- Clinical validation
- Subgroup performance analysis
- Data-quality assessment
- Human-factors study
- Cybersecurity testing
- Traceability records
- Monitoring plan
- Change-control records
Dynamic Review Triggers
The safety case should be reviewed after:
- A model update
- Expansion to a new patient population
- Integration with a new imaging device
- Performance drift
- A serious incident
- A cybersecurity vulnerability
Common AI Safety Case Development Challenges
Claims That Are Too Broad
“The AI is safe” is not a certifiable claim.
Claims should specify:
- System
- Version
- Purpose
- Environment
- Conditions
- Constraints
Evidence That Does Not Match the Claim
A high accuracy score does not support every type of safety claim.
Each evidence item must directly support the relevant argument.
Missing Operational Context
Laboratory evidence may not apply to real deployment conditions.
Hidden Assumptions
Assumptions involving users, data, suppliers, infrastructure, and environmental conditions should be documented explicitly.
Incomplete Hazard Coverage
Traditional failure analysis may overlook:
- Bias
- Emergent behavior
- Misuse
- Prompt injection
- Tool misuse
- Automation bias
- Drift
Overreliance on Benchmarks
Benchmarks may be:
- Unrepresentative
- Contaminated
- Too narrow
- Easily optimized
- Weakly connected to real risk
Broken Traceability
Disconnected documents and spreadsheets make it difficult to maintain consistency across claims, requirements, tests, and evidence.
Outdated Evidence
Evidence may become invalid after changes to models, data, prompts, interfaces, or deployment conditions.
Treating Compliance as Proof of Safety
Compliance supports the argument, but the organization must still demonstrate that risks are acceptably controlled.
AI Safety Case Development Best Practices
Start Early
Begin developing the safety case during concept and requirements engineering, not after testing.
Define Evidence Before Testing
Create the evidence plan alongside the claims and requirements.
Use Conditional Claims
State the context and conditions under which every claim is valid.
Control Configurations
Version:
- Models
- Data
- Prompts
- Requirements
- Tests
- Evidence
- Safety arguments
Maintain Bidirectional Traceability
Reviewers should be able to navigate from regulation to evidence and from evidence back to the source claim.
Include Negative Evidence
Search actively for failures, weaknesses, and counterexamples.
Track Uncertainty
Document known limitations and evidence gaps.
Use Independent Review
Critical claims should receive appropriate independent challenge.
Automate Traceability, Not Approval
AI and automation can support analysis, but authorized humans remain responsible for safety acceptance.
Maintain the Case Continuously
Update the safety case as the system, operating environment, and regulatory context evolve.
How Visure Supports AI Safety Case Development
Developing and maintaining an AI safety case requires control over requirements, hazards, risks, tests, evidence, standards, changes, and approvals.
The Visure Requirements ALM Platform helps organizations establish a connected source of truth for safety and certification information.
Centralized Safety Claims and Requirements
Teams can manage:
- Safety objectives
- Top-level claims
- Subclaims
- System requirements
- Safety requirements
- Assumptions
- Constraints
- Evidence references
within a controlled requirements environment.
This helps prevent safety claims and supporting requirements from becoming fragmented across documents, spreadsheets, and disconnected tools.
Bidirectional Traceability
Visure enables traceability across:
- Regulations
- Standards
- Hazards
- Safety goals
- Requirements
- Design elements
- Tests
- Results
- Evidence
This helps teams identify:
- Unsupported claims
- Untested requirements
- Uncontrolled hazards
- Missing evidence
- Incomplete risk controls
- Broken compliance mappings
Risk and Hazard Management
Safety teams can connect hazards and risks directly to:
- Mitigations
- Requirements
- Controls
- Verification activities
- Residual-risk decisions
This makes the risk-control strategy easier to review and maintain.
Verification and Validation Management
Requirements can be linked to:
- Test cases
- Test procedures
- Test results
- Review status
- Defects
- Evidence
This creates visible coverage for certification and independent assessment.
AI-Assisted Requirements Quality
AI-assisted capabilities can support:
- Requirement-quality analysis
- Ambiguity detection
- Consistency review
- Requirement generation
- Test-case generation
- Traceability suggestions
- Impact analysis
Human review remains essential for approval and safety decisions.
Change Impact Analysis
When a model, requirement, hazard, standard, or system configuration changes, teams can identify affected:
- Claims
- Controls
- Tests
- Evidence
- Reviews
- Baselines
This is particularly important for dynamic AI safety cases.
Baselines and Configuration Management
Visure helps teams establish controlled baselines for:
- Requirements
- Safety claims
- Evidence
- Tests
- System versions
- Certification releases
This makes it easier to demonstrate which approved artifacts apply to a specific product configuration.
Review and Approval Workflows
Formal workflows can support:
- Technical review
- Safety review
- Quality approval
- Independent assessment
- Risk acceptance
Multi-Standard Compliance Reuse
Organizations working under multiple frameworks can reuse requirements, controls, evidence, and traceability while maintaining standard-specific mappings.
Audit Trails
Audit histories help demonstrate:
- Who changed an artifact
- What changed
- When it changed
- Why it changed
- Who reviewed it
- Which baseline contains it
Secure and On-Premises Deployment
For regulated or security-sensitive organizations, on-premises deployment can help maintain control over sensitive requirements, risks, evidence, and certification information.
AI Safety Case Development Checklist
Before submitting or approving an AI safety case, verify that:
- Intended use and operational context are defined.
- System boundaries are documented.
- Applicable regulations and standards are identified.
- Hazards and AI-specific failure modes are analyzed.
- The top-level safety claim is approved.
- Claims are decomposed into reviewable subclaims.
- Argument strategies are documented.
- Evidence requirements are defined.
- Acceptance criteria are approved.
- Verification and validation are complete.
- Model, data, prompt, and configuration versions are recorded.
- Bidirectional traceability is established.
- Evidence quality has been reviewed.
- Assumptions and limitations are documented.
- Negative evidence and countercases have been considered.
- Residual risks are accepted by authorized stakeholders.
- Independent review is complete.
- The safety-case baseline is approved.
- Monitoring and update triggers are established.
Conclusion
AI safety case development is not simply the production of a compliance document. It is the disciplined construction of a defensible safety argument.
A credible AI safety case explains:
- What the organization claims
- Why those claims should be accepted
- Which evidence supports them
- Which assumptions apply
- Which uncertainties remain
- How the case will be maintained as the system evolves
Because AI behavior depends on models, data, prompts, tools, human interaction, and operating context, safety evidence must extend beyond traditional software testing. It must incorporate risk analysis, data assurance, model evaluation, human factors, cybersecurity, operational monitoring, organizational controls, and change management.
The strongest safety cases connect regulations, hazards, requirements, controls, tests, results, evidence, and approvals through complete traceability.
They also recognize that certification is not the end of assurance.
Model updates, data changes, drift, incidents, new tools, and changing operating conditions may alter the validity of previously accepted claims. For AI-enabled systems, safety assurance must therefore become a continuous, evidence-driven engineering process.
Take the first step toward revolutionizing your product engineering lifecycle management—try Visure Requirements ALM Platform free and experience the difference AI-driven solutions can make!