Standards
The frameworks the evidence has to satisfy.
Aplark Assure does not certify anything and is not itself certified. What it does is produce evidence in the shape these frameworks expect, with the provenance and reproducibility properties their reviewers ask for.
The argument, in full
Why testing alone cannot demonstrate extremely low failure rates.
ARP4761 sets a target failure condition probability of 10⁻⁹ per flight hour for a catastrophic hazard on civil aircraft. That figure did not originate as a statistical requirement — it is a design assurance target, met historically through architecture (redundancy, independence, monitoring) rather than through directly testing a component down to that rate.
Demonstrating a rate that low by testing alone would require sample counts, at realistic confidence levels, that exceed what any flight-test or simulation programme can fly — commonly cited estimates run to the order of a billion independent trials for a single hazard. No credible programme budget or schedule accommodates that, for a component of any kind, learned or otherwise.
A learned component makes this worse, not better: it has no specification to trace evidence back to, and correlated failure modes across a test set (near-duplicate frames, clustered scenarios) reduce the number of independent, effective samples far below the raw count collected.
The conclusion is not "give up." It is that the argument has to move to where it can actually be carried — architectural mitigation, redundancy and independence, or a declared operational domain narrow enough that the residual claim is honest. Aplark Assure computes and states that shortfall in orders of magnitude, on the same screen as the readiness assessment, rather than presenting a test campaign as though it could close a gap it structurally cannot.
ARP4754A / ARP4761
Origin of the 10⁻⁹ per flight hour catastrophic target.
DO-178C / DO-331
Assurance objectives for the software and model artefacts involved.
EASA AI Concept Paper
Published guidance acknowledging the testing-sufficiency gap for learning assurance.
Airborne systems
The most developed assurance culture, and the hardest fit for learned components.
DO-178C
Software Considerations in Airborne Systems and Equipment Certification
Sets the objectives an argument is structured against, and the design assurance levels that determine how much evidence is enough.
DO-330
Software Tool Qualification Considerations
Governs whether a tool’s output can be relied on without independent verification. Tool qualification is in permanent scope for Aplark Assure.
DO-331
Model-Based Development and Verification
Relevant where the item under assessment is specified as a model.
ARP4754A / ARP4761
Development of Civil Aircraft and Systems / Safety Assessment Process
Where hazards, severity classifications and target failure rates originate — including the 10⁻⁹ per flight hour figure that learned components cannot reach by testing.
EASA AI Concept Paper
Guidance for Level 1 and 2 Machine Learning Applications
The clearest published statement of what a learning assurance argument is expected to contain.
Functional safety and autonomy
Where the intended-functionality problem is named directly.
IEC 61508
Functional Safety of Electrical/Electronic Safety-related Systems
The root standard for safety integrity levels across sectors.
ISO 21448
Safety of the Intended Functionality (SOTIF)
Addresses hazards arising from performance limitations rather than faults — which is the failure mode of a perception component.
UL 4600
Standard for Safety for the Evaluation of Autonomous Products
Goal-based and argument-centric, with explicit treatment of the operational design domain.
MIL-STD-882E
System Safety
Provides the severity taxonomy used throughout the platform — Catastrophic, Critical, Marginal, Negligible — mapped to the hazard register.
Space
Assurance under the constraint that intervention is not available.
ECSS-Q-ST-80C
Space Product Assurance — Software Product Assurance
Product assurance objectives for flight software.
ECSS-E-ST-40C
Space Engineering — Software
Engineering process requirements for space software.
NPR 7150.2
NASA Software Engineering Requirements
Classification-driven software requirements for flight systems.
Evidence integrity
How the record is made checkable by someone who does not trust us.
RFC 6962
Certificate Transparency
The Merkle log construction the evidence ledger follows, giving inclusion and consistency proofs that verify without the server.
RFC 8785
JSON Canonicalization Scheme
Canonical bytes for signing. Two independent implementations must agree on them, checked against a conformance vector suite.
in-toto Attestation / DSSE
Dead Simple Signing Envelope and Attestation Statement
The envelope and predicate structure carried by every signed record.
PDF/A-3b
Long-term archiving format with embedded files
The dossier format, chosen so the package and its evidence remain readable for decades.
Argument and governance
The shape of the argument itself.
GSN Community Standard
Goal Structuring Notation
The assurance case notation, with an explicit defeater register rather than goals alone.
ISO/IEC 42001
AI Management Systems
Organisational context around the technical evidence.
NIST AI RMF
AI Risk Management Framework
A common vocabulary for risk framing in mixed-sector programmes.
Listing a standard here means the platform is built with its objectives in view. It does not imply certification, compliance, qualification or endorsement by any authority or standards body.