Standards

The frameworks the evidence has to satisfy.

Aplark Assure does not certify anything and is not itself certified. What it does is produce evidence in the shape these frameworks expect, with the provenance and reproducibility properties their reviewers ask for.

The argument, in full

Why testing alone cannot demonstrate extremely low failure rates.

ARP4761 sets a target failure condition probability of 10⁻⁹ per flight hour for a catastrophic hazard on civil aircraft. That figure did not originate as a statistical requirement — it is a design assurance target, met historically through architecture (redundancy, independence, monitoring) rather than through directly testing a component down to that rate.

Demonstrating a rate that low by testing alone would require sample counts, at realistic confidence levels, that exceed what any flight-test or simulation programme can fly — commonly cited estimates run to the order of a billion independent trials for a single hazard. No credible programme budget or schedule accommodates that, for a component of any kind, learned or otherwise.

A learned component makes this worse, not better: it has no specification to trace evidence back to, and correlated failure modes across a test set (near-duplicate frames, clustered scenarios) reduce the number of independent, effective samples far below the raw count collected.

The conclusion is not "give up." It is that the argument has to move to where it can actually be carried — architectural mitigation, redundancy and independence, or a declared operational domain narrow enough that the residual claim is honest. Aplark Assure computes and states that shortfall in orders of magnitude, on the same screen as the readiness assessment, rather than presenting a test campaign as though it could close a gap it structurally cannot.

ARP4754A / ARP4761

Origin of the 10⁻⁹ per flight hour catastrophic target.

DO-178C / DO-331

Assurance objectives for the software and model artefacts involved.

EASA AI Concept Paper

Published guidance acknowledging the testing-sufficiency gap for learning assurance.

Airborne systems

The most developed assurance culture, and the hardest fit for learned components.

DO-178C

Software Considerations in Airborne Systems and Equipment Certification

Sets the objectives an argument is structured against, and the design assurance levels that determine how much evidence is enough.

DO-330

Software Tool Qualification Considerations

Governs whether a tool’s output can be relied on without independent verification. Tool qualification is in permanent scope for Aplark Assure.

DO-331

Model-Based Development and Verification

Relevant where the item under assessment is specified as a model.

ARP4754A / ARP4761

Development of Civil Aircraft and Systems / Safety Assessment Process

Where hazards, severity classifications and target failure rates originate — including the 10⁻⁹ per flight hour figure that learned components cannot reach by testing.

EASA AI Concept Paper

Guidance for Level 1 and 2 Machine Learning Applications

The clearest published statement of what a learning assurance argument is expected to contain.

Functional safety and autonomy

Where the intended-functionality problem is named directly.

IEC 61508

Functional Safety of Electrical/Electronic Safety-related Systems

The root standard for safety integrity levels across sectors.

ISO 21448

Safety of the Intended Functionality (SOTIF)

Addresses hazards arising from performance limitations rather than faults — which is the failure mode of a perception component.

UL 4600

Standard for Safety for the Evaluation of Autonomous Products

Goal-based and argument-centric, with explicit treatment of the operational design domain.

MIL-STD-882E

System Safety

Provides the severity taxonomy used throughout the platform — Catastrophic, Critical, Marginal, Negligible — mapped to the hazard register.

Space

Assurance under the constraint that intervention is not available.

ECSS-Q-ST-80C

Space Product Assurance — Software Product Assurance

Product assurance objectives for flight software.

ECSS-E-ST-40C

Space Engineering — Software

Engineering process requirements for space software.

NPR 7150.2

NASA Software Engineering Requirements

Classification-driven software requirements for flight systems.

Evidence integrity

How the record is made checkable by someone who does not trust us.

RFC 6962

Certificate Transparency

The Merkle log construction the evidence ledger follows, giving inclusion and consistency proofs that verify without the server.

RFC 8785

JSON Canonicalization Scheme

Canonical bytes for signing. Two independent implementations must agree on them, checked against a conformance vector suite.

in-toto Attestation / DSSE

Dead Simple Signing Envelope and Attestation Statement

The envelope and predicate structure carried by every signed record.

PDF/A-3b

Long-term archiving format with embedded files

The dossier format, chosen so the package and its evidence remain readable for decades.

Argument and governance

The shape of the argument itself.

GSN Community Standard

Goal Structuring Notation

The assurance case notation, with an explicit defeater register rather than goals alone.

ISO/IEC 42001

AI Management Systems

Organisational context around the technical evidence.

NIST AI RMF

AI Risk Management Framework

A common vocabulary for risk framing in mixed-sector programmes.

Listing a standard here means the platform is built with its objectives in view. It does not imply certification, compliance, qualification or endorsement by any authority or standards body.

Standards · APLARK