AI assurance and evaluation

Aplark Assure

Test intelligence before deployment.

An AI assurance and verification platform for organisations developing safety- and mission-critical AI. It generates, preserves, evaluates, traces and independently verifies the evidence that demonstrates an AI system satisfies its defined safety, robustness, operational and deployment requirements.

AVAILABLE NOW
Request a Demo

The problem

Accuracy is not readiness

A model that scores 94.2% on a held-out test set has told you almost nothing about whether it is fit to fly. The test set does not describe the operational domain, the samples inside it are not independent, and the aggregate hides exactly the conditions that matter — the low-light cell, the sensor-degraded cell, the cell an adversary will choose.

Certification authorities do not accept scores. They accept arguments: a structured claim that a hazard has been mitigated to a stated residual rate, backed by evidence whose provenance, reproducibility and sufficiency can each be checked independently. Producing that argument is currently manual, adversarially reviewed, and takes months.

Aplark Assure makes the argument a build artefact. It does not make your model safer and it does not certify anything — it records, evaluates and proves what your evidence actually supports, including when the honest answer is that it supports nothing.

How it works

One workflow, nine stages

Every screen, endpoint and CLI verb in Aplark Assure sits at a named position on a single workflow spine. A capability that cannot be placed on the spine is out of scope.

The spine runs Define → Test → Capture → Verify → Assess → Explain Gaps → Build Assurance Case → Generate Dossier → Gate Deployment. Work enters at Define, where hazards, requirements, the operational design domain and the mitigation chain are declared, and the system computes the residual rate per hazard along with the sample count needed to demonstrate it.

That sufficiency calculation happens before a single test executes. Discovering that a residual rate is statistically undemonstrable is worth a great deal on day one and nothing at all after four months of GPU time.

Capabilities

What it does

Evidence ledger

An append-only, tamper-evident Merkle log in the manner of RFC 6962. Every artefact is hashed, every environment fingerprinted, every record signed. Inclusion and consistency proofs are checkable without the server.

Statistical sufficiency on effective samples

Coverage is assessed on effective sample count, not raw count. 40,000 frames drawn from 22 sorties are 22 effective samples for a per-sortie hazard, because consecutive video frames are correlated at roughly ρ = 0.9. The platform states the correlation assumption alongside every figure that depends on it.

ODD definition and coverage

Operational design domain declared as versioned dimensions, ranges and bins, with operational exposure weights traced to their CONOPS provenance. Coverage is reported per reachable cell, and a cell that is out of declared domain is reported as such rather than silently excluded.

Deterministic execution and reproducibility class

Runs are executed under pinned determinism controls and classified R0–R3 by how reproducibly they replay. R2 and R3 runs receive an automated diagnosis. Evidence cited at an inadmissible class is refused, with the reason named.

Readiness vector

Six dimensions, fixed order, assessed against a named criticality level, never summed into a composite score. Each dimension decomposes in one click into the gap entries that constitute it — a readiness dimension you cannot turn into work is decoration.

Gap report as a work queue

The default landing screen. Ordered by hazard exposure × severity, each entry naming the missing ODD cell, the hazard and requirement it blocks, the additional effective samples required, the estimated wall-clock and storage cost, and a single action that launches the campaign closing it.

Assurance case and dossier

A GSN argument graph with an explicit defeater register, bound to ledger evidence, rendering to a deterministic PDF/A-3b certification package with the evidence embedded. Two generations of the same input produce a byte-identical digest.

Offline verification

aplark-verify is a free, unrestricted static binary. A reviewer with no account, no documentation and no network can verify an export bundle on a machine that has never contacted the server. This is deliberately the easiest thing in the product to do without us.

Deployment gate

Policy evaluated against readiness and edge conformance for a specific artefact digest on a specific platform, producing a signed decision. Runs in the customer pipeline, offline. On refusal every failing rule is named — never a bare denial.

Who it is for

  • Safety engineers who own the hazard register and have to defend a residual rate
  • ML engineers who need to know which cell failed and reproduce it
  • Certification leads assembling a package they will have to defend line by line
  • Independent reviewers and certification authorities, who receive the dossier and the verifier and nothing else
  • Deployment engineers who need to confirm that the exact artefact shipping is the one that was cleared

See what Aplark Assure finds in your evidence.

Request a Demo
Aplark Assure · APLARK