Synthetic data
Aplark Forge
Build the data your AI needs.
Aplark Forge is intended to produce synthetic datasets targeted at declared coverage gaps, carrying the correlation structure and provenance record that determine whether the data can be cited as evidence at all.
The problem
Volume is not evidence
Synthetic data generation is not difficult to do badly. A million generated samples drawn from one underlying scene distribution contribute roughly the statistical weight of the number of independent scenes, not the number of files — and a dataset delivered without that correlation structure declared will be treated by any competent reviewer as unquantifiable.
The hard part is not generation. It is generating with an honest account of how much independent information the result actually contains.
How it works
How it is intended to work
Aplark Forge is being designed to treat the correlation structure as part of the deliverable. Generated sets would carry their cluster structure, the parameters varied and held fixed, and a declared default intra-cluster correlation that a downstream sufficiency calculation can consume directly.
The intended posture is conservative: where the effective sample count is uncertain, the pessimistic assumption would be the default and relaxing it would require evidence.
Trust boundary
No privileged status inside Aplark Assure
A vendor that both generates the test data and adjudicates the sim-to-real gap in the assurance argument holds a conflict of interest that a competent certification authority will identify immediately. We have resolved it in the architecture rather than in the sales conversation.
Data produced by any Aplark generation tool enters Aplark Assure as a declared, provenance-tracked input carrying exactly the same evidentiary burden as a third-party source. It takes the same declared source type and the same default intra-cluster correlation — simulation runs sharing a scene default to ρ = 0.3, and overriding that requires evidence plus a ledger record. It carries the same mandatory sim-to-real-gap defeater in the assurance case. It meets the same provenance requirements on the generating environment as any external dataset.
No code path, configuration flag or licence tier relaxes any of these because the data came from us. Aplark Assure will never produce the thing it evaluates.
Capabilities
What it does
Gap-directed synthesis
Planned to take a coverage gap as input and generate against it, rather than producing bulk data and measuring coverage afterwards.
Declared correlation structure
Intended to ship cluster structure with every dataset so that effective sample size is computable rather than assumed.
Provenance-tracked outputs
Designed to emit the same manifest format Aplark Assure ingests, with no shortcut for being first-party.
Who it is for
- Programmes with coverage gaps that field collection cannot close economically
- Teams that need generated data to survive independent review