Prediction register / draft slate
A proposition.
A deadline.
A result to answer to.
The register separates predictions of assessor behavior (what bodies that assess or regulate AI, including developers assessing their own systems, will decide or publish) from capability measurements. Neither is, by itself, a safety finding.
- P1
EU delegated act
The European Commission will not adopt, on or before 2 August 2027, a delegated act under Article 51(3) of Regulation (EU) 2024/1689 that either changes the numerical 10^25-FLOP threshold in Article 51(2) or supplements Article 51 with a benchmark or indicator for determining high-impact capabilities. Resolve by Commission adoption date, allowing 30 days for documentary publication; drafts, consultations, guidance, designations, implementing acts, and ordinary legislation do not count.
Assessor behavior · Not registered
40%candidate p - P2
Public lab threshold attainment
By 2027-09-30, Anthropic will expressly conclude that the “Automated R&D in key domains” threshold in RSP v3.4 has been met, or Google DeepMind will expressly conclude that a named model reached ML R&D Automation Level 1 or Acceleration Level 1 under FSF v3.1. Only explicit public attainment counts; alert thresholds, precautionary safeguards, or later framework renaming do not.
Assessor behavior · Not registered
22%candidate p - P3
Joint evaluation report
By 2027-09-30, NAAIMES will publish or formally co-issue a joint evaluation report identifying at least three member institutes, naming at least three model versions from at least two developers, reporting a common numeric model-level metric for every named model, and disclosing the task set, scoring rule, model/version, and principal inference or elicitation settings.
NAAIMES, as named by its source: “the International Network for Advanced AI Measurement, Evaluation and Science (NAAIMES) – formerly the International Network of AI Safety Institutes” (UK AI Security Institute, “International evaluation best practice and open questions in AI measurement”, retrieved ). This expansion is not part of the proposition.
The resolver for this forecast is not yet bound. The linked source page is the U.S. CAISI centre, not the network’s archive.
Assessor behavior · Not registered
42%candidate p - P4
California statutory report
By 2027-01-31, Cal OES will file or publicly post its first report required by California Business and Professions Code §22757.13(g). A zero-incident report counts; a portal, press release, or legislative description without the statutory report does not.
The proposition’s date, 31 January 2027, is read as 23:59:59 Pacific time, which is 2027-02-01T07:59:59Z. The statute sets a start date (‘Beginning January 1, 2027, and annually thereafter’) and transmission to the Legislature and the Governor, not a public filing deadline. The 31 January window is this forecast’s own choice.
Assessor behavior · Not registered
92%candidate p - P5
A 24-hour task horizon
By 2027-09-30, METR will publish a task-completion-time-horizon estimate for a named publicly available model with a P50 time horizon of at least 24 hours under TH1.1 or an explicitly bridged successor methodology. Internal or unreleased models and non-estimate lower bounds do not count.
Capability measurement · Not registered
68%candidate p - P6
Final evaluation practices
By 2027-09-30, NIST or CAISI will publish a final, non-draft edition of NIST AI 800-2, Practices for Automated Benchmark Evaluations of Language Models, or an officially renumbered successor with substantially the same scope. Revised drafts and requests for comment do not count.
Assessor behavior · Not registered
76%candidate p
Read the record directly.
The human view and machine exports come from the same Markdown records. Candidate probabilities are marked as drafts in both. A content digest identifies this snapshot; it does not prove a proposition true.