Methods
An evaluation is only reproducible if the instrument is available, so we publish the specification we report against, the harnesses we build, the protocols we adjudicate to, and the rules that bound what any of it can establish.
Released methods
Harnesses are released with the first finding that uses them.
What we can claim depends on what we were given
Access is determined and recorded before method design, since it bounds the available conclusions more tightly than any other factor, and the tier we worked under is stated on the front of every report and in the index entry for every finding.
| Claim type | Black box | Grey box | White box |
|---|---|---|---|
| What we have | Queries and outputs. | Plus documentation, confidence scores, a description of the training data, and logs. | Plus weights and internals. |
| We may claim | That the system fails, where, and how often. | Plausible mechanisms, and calibration analysis. | Mechanistic explanation. |
| We may not claim | Why it fails. | Causal attribution to specific design choices. | That the system we were given is the system in production. Provenance is established separately, and usually cannot be established by us. |
No tier of access allows us to certify that a system is safe, since an evaluation establishes how one version behaved on one test set at one point in time. Every report sets out what would have to be true to generalise from that.
The protocol is published before the study is run
Before execution begins we publish a timestamped method file covering the unit of analysis, the metrics, the harm threshold, the ground truth source, disaggregation, sample construction, and the uncertainty budget, and stating what result would falsify the hypothesis. A reader can then set what we said we would measure against what we reported measuring.
We establish what a test detects before we rely on it
For each method we document what it detects, its sensitivity to specification choices such as prompt phrasing and detection thresholds, its known false positive and false negative behaviour, and the range of systems over which it is valid. A new method is validated against a system with known ground truth before it is applied to a study subject.
A validated method is valid over a stated range of systems and conditions. Every method file names that range. Applying a method outside it produces a number with no interpretable meaning.
Instruments and results are released by default
Harnesses, scoring code, and adjudication protocols are released as a matter of course, and every published chart ships with the machine-readable results file behind it, so a reader can replot or reanalyse without asking us for anything. Reference datasets we collect ourselves are released with the finding that first uses them, together with the collection protocol and the annotation guidelines.
Where anyone reproduces our work and arrives at a different answer, we publish that as well.
Datasets containing personal data, imagery licensed on restricted terms, and test sets whose release would let a subject train against them are withheld or redacted. Where we withhold, the finding states what was withheld and why, and an independent party may request access under the terms set out in the report.