# Prediction register: method and present status

**Six draft candidates. Zero registered predictions. Zero resolved predictions. No prospective accuracy score or calibration record exists for this cohort.** Source checks performed in September 2026 are retrospective lineage, not prediction successes. These are proposed operating rules; their publication here as a local candidate does not demonstrate their operation.

This is a **single-elicitor candidate record**, assisted by research tools. The six probabilities were adopted from one drafting lineage. A named accountable human has not yet accepted them; no crowd, consensus or independently calibrated forecast is claimed. Hapax Research Labs has an interest in showing its method useful. Readers may inspect or reject the forecasts without adopting that position.

## The two ledgers

Ledger A contains retrospective source checks and attached corrections. The original September 22 review contains graded findings, including attribution and wording corrections. Its “12/12” summary is an inherited review claim, not twelve prospective trials and not a fresh replication. The later accepted narrative corrects the unconditional/conditional survey-median confusion and limits productivity and acceleration inferences. Those corrections stay attached to the lineage.

Ledger B contains the six draft forecasts P1–P6. They concern, in order, EU delegated adoption, public laboratory acknowledgment, a joint network report, a California annual report, a benchmark estimate and a final standards document. P1 assigns 40% to **no qualifying act**. P2/P3/P4/P5/P6 assign 22%/42%/92%/68%/76% to the stated positive events. The exact cards, not this shorthand, control.

Draft counts: 6. Registered: 0. Due: 0. Resolved YES/NO: 0. Ambiguous: 0. Annulled: 0. Brier, skill, calibration and resolvable coverage: **not applicable**, not zero. A future coverage measure divides YES+NO resolutions by all registered claims reaching their evidence cutoff; overdue, censored, ambiguous and annulled rows remain in that denominator and visible.

## Registration and version discipline

Registration requires an accountable human elicitor, external checker, accepted exact wording and reference-class/baseline packet, the operator’s D4 pass, and admitted publication of the exact version through the existing publication route with an anonymous receiving-surface readback. The registration time is that actual publication time, never a task-claim time. The first forecast closes at that time. A hash without a public timestamp or receipt is only a local byte binding.

The candidate cards preserve both the adopted final row and the detailed source rule. P1’s local timezone, P2’s final minute, P5/P6’s unspecified timezones, P3’s membership/subset interpretation and P6’s landing-page/PDF discrepancy rule are exposed for pre-registration disposition. An operator or checker cannot be represented as having approved an unresolved operationalization. If accepted, a new registration edition binds it explicitly; this candidate remains preserved.

After registration, preserve original propositions and probabilities. Corrections, clarifications, withdrawals and updates are new dated records with predecessor hashes and public diffs. A changed event is a new question, not a corrected winning score. Publish adverse scores and corrections at the original’s prominence. Unresolved entries never disappear.

## Evidence representation and prior art

Cards use the receiving portal’s existing record projection and source references. The body supplies the eight assessor fields: claimant, claim, target, horizon, resolution, scoring, obligation and history. The existing `inference_strength` carrier is interpreted non-ordinally as `epistemicRelation`; probability is separate. `sourceKind`, `methodKind`, citation, byte integrity, support, verifier and structured scope describe different facts, not evidence levels. Source reports about a forecast’s context do not establish its outcome. The projection is a compatibility artifact, not a new schema or a HACA conformance claim.

Prior art is **BACKED at method-source tier** for explicit resolution criteria and separate indeterminate/malformed question states in [Metaculus’s FAQ](https://www.metaculus.com/faq/). Direct automated capture was denied with HTTP 403 in this build; the inherited source reference remains available for checker review, without a fresh full-source verification claim. The proposed practices remain HRL’s requirements independently of that attribution.

The inherited methods citation map has defects: reference 0 is absent from the saved search mapping, and reference 2 points to a Metaculus checklist rather than OECD. Consequently this edition does **not** claim fresh GJP or OECD source verification. Proper scoring and transparent denominators are stated as the adopted method, without laundering the polluted trailing bibliography. New primary-source backing must be attached before quoting those attributions publicly.

## Elicitation and baselines

Frame the event first. Publish the reference-class universe, exclusions, data cutoff, n, k and uncertainty before using a rate. Have the external checker review a masked universe packet before seeing final probabilities. Test alternative plausible classes and changes of at least five percentage points. Then record every source, its access time, dependence, direction, log-odds movement and posterior; complete an opposite-case search, pre-mortem and leave-strongest-item-out recalculation. Suggested movement classes are ±0.10/0.25/0.50/1.00 log-odds, judgmental rather than measured likelihood ratios.

The inherited probabilities preceded this protocol. Their complete elicitation history cannot be reconstructed as fact. Cards retain the numbers but mark the absent worksheets and recalculations as blockers. Every candidate carries a numeric 0.50 ignorance baseline with a disclosed rationale, pending checker acceptance. Empirical frequency takes precedence if an admissible class can be enumerated; otherwise inspect persistence and trend before approving ignorance. Replacing a baseline requires a pre-registration version, never a favorable post-outcome adjustment.

No extremizing at n=1. Any future transformation requires at least 50 prospectively resolved questions, a held-out evaluation showing improvement, and publication of raw and transformed probabilities. Six correlated outcomes cannot demonstrate calibration.

## Dependence matrix

Freeze this candidate matrix before outcomes. It is a qualitative driver hypothesis, not measured correlation. U11’s illustrative table called P4 “NIST” and P6 “state law”; here the adopted identities are P4 California and P6 NIST. P1 is a negative event, so the original generic positive-direction labels cannot be copied literally. A faster common governance/capability driver tends to decrease P1’s no-act probability and increase the paired positive event; these direction judgments require independent acceptance.

| Pair | Shared driver | Direction for the actual propositions | Strength |
|---|---|---|---|
| P1–P2 | Regulatory pressure and disclosure | Negative | Medium |
| P1–P3 | International coordination pace | Negative | High |
| P1–P4 | Governance implementation momentum | Negative | Medium |
| P1–P5 | Capability growth prompting regulation | Negative | Low |
| P1–P6 | Standards and regulatory coordination | Negative | Medium |
| P2–P3 | Public evaluation evidence | Positive | Medium |
| P2–P4 | Reporting duties and lab disclosure | Positive | Medium |
| P2–P5 | Frontier capability progress | Positive | High |
| P2–P6 | Measurement standards and disclosures | Positive | Medium |
| P3–P4 | General governance momentum | Positive | Low |
| P3–P5 | Capability evidence motivating evaluation | Positive | Low |
| P3–P6 | International measurement practice | Positive | High |
| P4–P5 | Capability growth and implementation | Positive | Low |
| P4–P6 | US standards and statutory practice | Positive | High |
| P5–P6 | Measurement demands | Positive | Low |

Clusters: EU/international=P1+P3; frontier capability/lab=P2+P5; US standards/law=P4+P6. Clustering does not remove all dependence, notably P3–P6. Annual sensitivity reports also show (a) all six as one descriptive cluster and (b) institutional outputs P1/P3/P4/P6 versus frontier P2/P5. No binomial significance, naïve n=6 confidence interval or multiplication of independent probabilities is permitted.

## Scoring

For each resolved proposition, y=1 when its exact proposition is true, y=0 when false. Brier=(p−y)²; baseline Brier=(b−y)². P1 **no act** therefore has y=1. Secondary log loss=−[y ln(p)+(1−y) ln(1−p)] using natural logs; no post-outcome clipping of probabilities. The candidate values lie strictly inside (0,1).

Report every row, ordinary mean and cluster-balanced mean. Average within clusters, then equally across the three cluster means; with all six resolved these two means happen to be equal because cluster sizes are equal. For interim reports show rows and available-cluster descriptive scores with coverage and denominators; withhold a full-cohort headline until every cluster is represented. Brier skill=1−mean forecast Brier/mean baseline Brier using the same questions and weighting; undefined when the baseline denominator is zero. Ambiguous and annulled outcomes are excluded from accuracy arithmetic and included in resolution-quality reporting. Never impute y=0.5.

Do not collapse predictive accuracy, measurement quality, disclosure, response compliance and transparent rule revision into one score. A final report or framework does not itself establish safety effectiveness. Reviewers should separately track output completion, evidence availability, promised actions and outcomes. Revision taxonomy distinguishes anticipatory/evidence-responsive/reactive tightening and evidence-based/unsupported/adverse-timing loosening. “Capture” is a hypothesis, not a verdict from timing alone. Retain original scores under retroactive changes; publish classification replay and the proportion of old cases changed by revised rules. A benchmark replacement without bridge breaks the series.

## Adjudication and appeals

The proposer prepares the question. A named case resolver proposes an outcome. An independent **person** with no authorship or financial stake checks contested, ambiguous, source-substituted and claimant-reputational cases and has final appeal authority. Review at least two randomly selected ordinary annual resolutions as well. An independent AI artifact review is useful but cannot fill this human external-checker role.

Default source order: named primary; its archive; an official successor with materially equivalent definitions; two independent secondary reproductions; contemporaneous public records; AMBIGUOUS if still indeterminate. A card’s stricter official-publication predicate overrides this general hierarchy. Substitution cannot make a private statement count as public or lower a numeric threshold. Search snippets alone never resolve a claim.

Open evidence within three business days of a conclusive event or evidence cutoff; propose within ten business days; accept challenges for fourteen calendar days; obtain required external review within ten business days; finalize within five business days thereafter. Publish a delay notice every thirty days if unresolved. A final result includes evidence archives, original criteria, resolver/checker identities and conflicts, challenge replies, timestamps, substitutions and dissent.

Silence yields a negative event conclusion only after a complete predeclared authoritative-source search. If publication ceases or evidence is confidential, record transparency failure/censoring or inability to resolve, with the exact public-event exceptions stated in each card. Intervention that prevents an event is scored as registered and flagged, not counted as predictive vindication. AMBIGUOUS means the factual state or criterion application cannot be established; ANNULLED means the question is defective although reality is sufficiently clear. Ambiguity resolves against awarding HRL credit, not by inventing an adverse binary outcome.

## Decision timing and recurrence

Every card is visibly **LATE FOR PRIMARY DECISION** under its illustrative year-end planning date. That date is a scenario, not a claim about a real reader’s irreversible decision. Named owners, costs, avoided losses and effectiveness remain to be accepted. The illustrative 1/5 preparation ratio must never be applied directly to a publication probability as if it were a probability of catastrophic harm. Monitor primary indicators now; comply with current obligations and act on deployment evidence without waiting for 2027 scoring.

Proposed cadence after actual registration: archive material events within seven days; quarterly link/evidence/resolution review; annual claim-level report with Brier/skill, disclosure and resolution coverage, overdue commitments, revision replay and checker findings; biennial method reassessment retaining original series. No monitoring service or future report delivery is activated by this document.
