Deep Field Labs · published 2026-10-04 · v1, frozen on publication · source: docs/dfl-prediction-grading-standard-v1.md. The live numbers it governs are on the track record; how we compute them is on methods.

DFL Prediction Grading Standard — v1

Status: v1 — FROZEN on publication, 2026-10-04 (DecayGuard public launch). Later changes create v2; v1 is never edited. Published at https://decayguard.deepfieldlabs.dev/grading-standard.html

The problem this solves: every vendor quotes an accuracy number; almost none say what would have counted as a miss, when a prediction is allowed to be judged, or whether the history behind the number can be checked. This standard defines those terms precisely enough that any vendor's re-entry (or comparable event) predictions can be graded the same way. We hold ourselves to it in public; buyers are invited to ask every vendor the questions at the bottom.

1. The prediction record

A gradable prediction is a dated, immutable claim:

FieldMeaning
issuedUTC date the claim was made. Never edited.
norad_idthe object
claimevent type (re-entry, manoeuvre)
window_endlast date the claim covers
statusopen → one of confirmed / expired / void

One open claim per object: re-issuing daily while a claim is open would inflate the sample. The earliest call is the graded one.

2. Grading states

Grading is causal: evidence dated after the grading clock is invisible to that grading run. Nothing may be graded with information that did not exist at grading time.

3. Evidence grace

Evidence publishes late (measured on Space-Track Historical decay messages: median ≈ 6 days, p90 ≈ 26 days after the event). A window that closed yesterday is therefore not yet a miss. A prediction only matures — becomes judgeable — once window_end + grace has passed. DFL uses grace = 28 days, calibrated from that lag distribution and recalibrated as the archive grows. Publishing precision over unmatured cohorts is survivorship bias in the favourable direction; this standard forbids it in the headline.

4. The three precision figures (all published, one headline)

MetricDefinitionRole
precision_closedconfirmed / (confirmed + expired), lifetime rawfull-disclosure floor
precision_steadysame, restricted to post-backlog cohortstrend
precision_maturesame, restricted to matured post-backlog predictionsheadline

"Backlog cohort": predictions issued while first draining a pre-existing alert backlog measure the backlog, not the system; they are excluded from steady/mature and included in closed. The cutoff date is fixed and published (backlog_cohort_end).

n_pending_grace (issued, window closed, not yet matured) must be published next to any precision figure.

5. Capture (the other half)

Precision without capture rewards silence. capture_decay = of the observed events in the period that the system could have seen (object in its candidate universe before the event), the fraction an issued prediction warned about, with median_lead_days alongside. Both numbers published; a precision quoted without its capture is not compliant with this standard.

6. Verifiability (what makes any of this checkable)

A track record that cannot be independently re-hashed is a claim, not a record.

7. Pre-registration

Go/no-go launch criteria, metric definitions, and any candidate change to the interpretation are frozen in writing before the evaluation runs. A negative result (a rejected candidate) is published to the same standard as a positive one.

8. Questions to ask any prediction vendor

1. What is your evidence-mature precision — and what grace period and cohort rules produced it? 2. What is your capture rate over the same period, and its median lead time? 3. How many predictions are currently pending grace? 4. Can I re-verify your published history cryptographically, without trusting your servers? 5. Were your evaluation criteria written down before or after the results existed?

DFL's answers are on the DecayGuard track-record page, recomputed nightly, hash-anchored daily.

Exact-text fingerprint of this standard (SHA-256 of the source markdown): 6efc07595a3ff3552b37dc6c34de08fca03b4e7ef85ca180bddc15fa10e01103. If this page ever differs from that hash, it is not v1.


DFL Prediction Grading Standard — implementation notes

Companion to v1. The standard is frozen and its fingerprint never changes; this file records how DFL's own code tracks it, including where the code was wrong. Rendered beneath the standard on https://decayguard.deepfieldlabs.dev/grading-standard.html

2026-10-07 — Correction to the published headline

From 2026-10-04 to 2026-10-07 DFL's implementation applied the §3 maturity test to expired predictions only; confirmed predictions from the same unmatured cohorts entered the headline at once. That is the survivorship bias §3 forbids, in the favourable direction. On launch day the published figure was 0.955 ("155 matured"); the figure under the rule as written was 0.22 on 9 matured predictions — too few to support any claim. The figure fell nightly as cohorts matured (0.639 on 2026-10-07, when the cohort-consistent value was 0.349 on 129).

Corrected together on 2026-10-07: the implementation (core/decayguard/prediction_ledger.py), the dashboard's client-side reproduction, the published metrics and the explanatory copy. n_matured is now published next to n_pending_grace, and n_pending_grace counts every closed prediction (confirmed, miss or expired) that has not yet matured. The ledger and its hash chain are unchanged — every call, hit and miss is where it was. The standard's text needed no change.

The current issuance regime (from 2026-08-17) matures from mid-November 2026; until then the headline describes the July–August cohorts.

Lesson for anyone adopting this standard: encode §3/§4 as a test that holds both statuses to the same maturity rule, and require n_matured beside every headline figure. A precision that moves while n_matured stands still is the signature of this error.