AI Trend Notifier
EN
← wiki

$ cat wiki/concepts/preparedness-framework.md

Preparedness Framework

Definition

OpenAI's internal policy for deciding what a model is allowed to be — trained further, deployed, or made available to whom — based on how it scores against capability thresholds in specific risk domains. The framework names tiers; a model's assessed tier determines which safeguards must be in place before work continues (source).

Two tiers appear in the sources this wiki holds:

TierWhat it has meant in practice
HighThe assessment given to every OpenAI frontier model evaluated for cyber capability before August 2026, including GPT-5.6 Sol (and Terra, Luna)
CriticalFirst invoked 2026-08-07, for Astra. Triggers a development slowdown, not only a deployment gate
The Critical cyber threshold as OpenAI states it: a model that can identify and
develop functional zero-day exploits of all severity levels in many hardened
real-world critical systems without human intervention, or **devise and execute
end-to-end novel strategies for cyberattacks** against hardened targets **given only
a high-level desired goal**
(source).

Why It Matters

  • It is a commitment that costs something when honoured. On 2026-08-07 OpenAI said it could not rule out Critical cyber capability in Astra and would slow development of the model in response. This wiki has recorded frontier labs publishing safety frameworks since it began; this is the first entry in which a published framework is cited as the reason a lab's own flagship is delayed.
  • The trigger is "cannot rule out", not "confirmed". OpenAI states testing is ongoing and that it has not confirmed Astra crossed the threshold. The framework fires on an inability to exclude a capability, which is a materially lower bar than a positive finding — and a bar that gets harder to clear as evaluations get harder to run.
  • A tier is a claim about evaluations, and evaluations are contested. See Eval Harness Configuration: what a model scores depends on the scaffold it is scored in. A threshold expressed in capability terms inherits every ambiguity in the measurement.
  • It is unilateral. The framework is OpenAI's own document, the assessment is OpenAI's own, and the remedy is OpenAI's own choice. Nothing in the sources read here makes any part of it externally reviewable, though the 2026-08-07 response brings government agencies and selected AI safety organisations in to test.

State of the Art (2026-08-08)

The framework's first Critical designation is six days old and the response is published but not observable from outside. What is on the record:

  • 2026-08-07 — Astra assessed as unable to rule out Critical for cybersecurity. Response: development slowed; isolated test environments; restricted network and tool access; sandboxed execution; hardened weight storage; monitoring of every agentic run; robustness testing of safeguards scaled up; external testing with government agencies and safety organisations; recommended security controls supplied to third-party testing partners (source).
  • Before this, High-tier cyber capability was handled by gating access rather than slowing work: the Daybreak programme sells controlled access to cyber-tuned models including GPT-5.5-Cyber for authorised red teaming, penetration testing and exploit validation, behind Trusted Access verification (source), and GPT-5.5-Cyber's EU availability was itself staged (source).

The distinction worth holding: High produced a distribution policy, Critical produced a development policy. Those are different kinds of commitment, and only the second one is visible in a shipping schedule.

Open Problems

  • No published exit condition. OpenAI says development slows "until it has the right safeguards in place." No source read here states what would satisfy that, who decides, or on what evidence — so from outside there is no way to distinguish a framework working from a delay that ends when it becomes commercially inconvenient.
  • "Cannot rule out" has no floor. The same phrase covers a model that is one evaluation short of a confirmed Critical result and a model whose evaluations are merely inconclusive. Nothing in the sources read distinguishes them.
  • Self-assessment. The evaluator, the framework author and the party bearing the delay cost are the same organisation. External testing is announced as part of the response, not as part of the assessment that triggered it.
  • No cross-lab vocabulary. Anthropic, Google DeepMind and OpenAI each publish their own thresholds under their own names. A "Critical" model at one lab has no defined relationship to any tier at another, so the industry cannot say whether two models are at comparable risk levels — only whether each vendor says so.

Key Papers

No paper. This is a corporate policy document, and the wiki holds it through OpenAI's own posts rather than through literature (source, source).

Referenced by

Sources