Then in LinkedIn: Write article → click into the body → paste (Ctrl+V). Headings, links and images come with it. The title usually pastes as the first line — cut it into LinkedIn's title field. back to the article

Astra's Safety Sign-Off Has a Chain of Custody, and Every Link Is OpenAI's Own

OpenAI's system card says a Safety Advisory Group recommended and leadership determined Astra's safeguards "sufficient" for the first model the company has called Critical. Five outside groups touched the evidence; none held the pen.

A paper chain-of-custody tag threaded through the pull-tie of a sealed clear evidence bag on a plain intake counter, the tag's cord caught mid-thread and not yet lying flat, a routing folder's corner lifting at the edge of the frame

A record can move from hand to hand without any of those hands holding the authority to open what it describes — the distinction this article traces through OpenAI's own Preparedness Framework [@openai-preparedness-framework-v2]. — Image prompt and art direction by Brecht Corbeel; image generated to that direction.

Abstract

OpenAI's Preparedness Framework states that an internal Safety Advisory Group recommends, OpenAI Leadership approves or rejects, and the Board's Safety and Security Committee only provides oversight. This article follows that chain for one sentence — the System Card's statement that Astra's safeguards were "sufficient" for release as the first model OpenAI has designated Critical — from the document that defines the process, through OpenAI's public statements in the three weeks before launch, to the five external organizations the System Card names. It stays entirely inside OpenAI's own documents and does not evaluate whether Astra's cyber or biological capabilities are themselves dangerous, whether the Framework should already have been rewritten, or how this compares to Anthropic's Responsible Scaling Policy for Mythos 5.1 — those are other pieces in this series. What it establishes, using only OpenAI's own text, is where the authority to say "safe enough" sits, how far short every outside name falls, and where a circulating description of OpenAI's August commitments doesn't match what OpenAI wrote.

One Sentence Ends the Review, and Every Noun in It Is OpenAI’s

Buried in section 10 of the GPT-6 Astra System Card, after eighty-some pages of capability tables and red-team transcripts, sits the sentence that actually matters: “The internal report informed our Safety Advisory Group’s recommendation and OpenAI leadership’s determination that these safeguards are sufficient for Astra’s public launch” [2]. That sentence is doing the only load-bearing work in the whole document. Everything above it — the benchmark scores, the red-team findings, the named outside organizations — is evidence. That sentence is the verdict, and it is worth reading slowly, because every noun in it names an OpenAI body. The internal report is OpenAI’s. The Safety Advisory Group is OpenAI’s. The leadership is OpenAI’s. Astra is the first model OpenAI has ever designated as meeting the Critical cybersecurity threshold under its own Preparedness Framework [3] — the tier the Framework itself defines as posing “a qualitatively new threat vector for severe harm with no ready precedent” [1] — and the sentence that certifies it safe enough to ship was written, reviewed, and signed entirely by the company that built it.

That is not, on its own, a scandal. Every company that ships a product decides internally whether to ship it; expecting otherwise misunderstands how corporate decision-making works anywhere. What makes this specific sentence worth tracing is that OpenAI’s own governing document — the Preparedness Framework, first published in December 2023 and revised to its current “Version 2” in April 2025 — describes, in granular and fairly unusual detail, exactly which office is allowed to say yes, exactly which offices are allowed to weigh in without being able to say yes, and exactly what conditions have to be met before anyone outside the company touches the evidence at all [1]. A document that specific creates an opportunity most companies never give you: you can check its terms against what it actually names doing the checking. This article does that, for one document and one sentence, and nothing else.

The Framework Names Three Bodies; Only One of Them Can Say No

Start with the Framework’s own account of who does what, stated once, early, in plain declarative sentences: “An internal, cross-functional group of OpenAI leaders called the Safety Advisory Group (SAG) oversees the Preparedness Framework and makes expert recommendations on the level and type of safeguards required for deploying frontier capabilities safely and securely. OpenAI Leadership can approve or reject these recommendations, and our Board’s Safety and Security Committee provides oversight of these decisions” [1]. Three bodies, three distinct verbs. SAG “recommends.” Leadership “can approve or reject.” The Board’s committee “provides oversight.” Only the middle verb has the power to change what happens next; the other two describe input and observation.

A bound intake ledger open on a counter, a ruled column of numbered signature lines with only the first two rows filled in and a desk pen resting at an angle beside the next blank line

Figure 1. OpenAI's own Preparedness Framework names exactly which office fills which line: a Safety Advisory Group recommends, and only OpenAI Leadership's line can approve or reject the recommendation [@openai-preparedness-framework-v2]. — Image prompt and art direction by Brecht Corbeel; image generated to that direction.

The Framework does not leave that reading to inference. Further into the document, in the section spelling out SAG’s governance role in full, it adds a clause that removes any ambiguity about whether SAG’s recommendation could ever function as a binding requirement: “For the avoidance of doubt, OpenAI Leadership can also make decisions without the SAG’s participation, i.e., the SAG does not have the ability to ‘filibuster’” [1]. That is the document explicitly ruling out the one reading that would make SAG’s recommendation a check on Leadership rather than an input to it. It also addresses who staffs SAG in the first place: “The members of the SAG and the SAG Chair are appointed by the OpenAI Leadership,” serving one-year terms that Leadership can choose to renew [1]. The body that reviews Astra’s safeguards is not an outside check on the executives who decide whether to ship it; it is a panel those same executives appoint, and its recommendation is explicitly non-binding on them by the document’s own admission.

The fairest complication to that reading sits in the Framework’s own appendix on decision-making practice, and it deserves stating in full rather than left out because it cuts against a cleaner story: the Board’s Safety and Security Committee “will be given visibility into processes, and can review decisions and otherwise require reports and information from OpenAI Leadership,” and, critically, “where necessary, the Board may reverse a decision and/or mandate a revised course of action” [1]. That is real power, not ceremonial language, and “provides oversight” in the Framework’s summary sentence understates what the appendix actually grants. But the power described is still bounded in a way worth naming precisely: it is reactive rather than a precondition — the Committee reviews a decision Leadership has already made, rather than clearing one before it takes effect — and its exercise is conditioned on the same kind of discretionary trigger (“where necessary”) that governs every external door in this document. It is also, and this is the point the rest of this article turns on, a committee of OpenAI’s own Board of Directors: an internal governance body reviewing another internal body’s decision, not an outside authority entering the chain at all. Astra’s System Card records no Board review or reversal in connection with its launch, only the SAG-to-Leadership determination quoted above [2]. Whether that silence means the Committee reviewed and concurred, or was never asked to look, is not something any document fetched for this article settles either way — a genuine gap, not a resolved one.

A slim review binder on a small wheeled cart approaching the records-intake counter from a distance, still short of arriving alongside a nearby ledger whose stamp impression is already visible and dry

Figure 4. The Board's Safety and Security Committee "provides oversight" and may "reverse a decision... where necessary" — a real power, but a reactive one: it reviews a determination OpenAI Leadership has already made, not one it clears in advance [@openai-preparedness-framework-v2]. — Image prompt and art direction by Brecht Corbeel; image generated to that direction.

An Outsider Enters This Process Only When OpenAI Decides One Should

If the internal chain runs entirely through OpenAI’s own appointees, the natural next question is what role the Framework leaves for anyone outside the company. It leaves three, all gathered under a section titled “Transparency and external participation,” and all three are written in conditional language rather than commitment language. “Third-party evaluation of tracked model capabilities: If we deem that a deployment warrants deeper testing of Tracked Categories of capability… then when available and feasible, OpenAI will work with third-parties to independently evaluate models” [1]. “Third-party stress testing of safeguards: If we deem that a deployment warrants third party stress testing of safeguards and if high quality third-party testing is available, we will work with third parties to evaluate safeguards” [1]. “Independent expert opinions for evidence produced to SAG: The SAG may opt to get independent expert opinion on the evidence being produced to SAG… If provided, these opinions will form part of the analysis presented to SAG” [1].

Read those three bullets as a set and a pattern emerges that the document never states outright but that its own word choices make plain: “if we deem,” “when available and feasible,” “may opt to.” Every door an outsider could walk through opens only because OpenAI decided, in that instance, to open it, and every one of those doors feeds evidence toward SAG — the appointed, overruled-if-necessary body — rather than toward a decision point with any independent authority of its own.

Even inside that body, the menu of possible outcomes is narrower than “approve or block.” The Framework spells out exactly three decision points available to SAG once it has reviewed the evidence for a given deployment: it “can find that it is confident that the safeguards sufficiently minimize the associated risk of severe harm… and recommend deployment”; it “can request further evaluation of the effectiveness of the safeguards”; or it “can find the safeguards do not sufficiently minimize the risk of severe harm and recommend potential alternative deployment conditions or additional or more effective safeguards” [1]. All three options are recommendations. The document says so directly in the next line: “All of SAG’s recommendations will go to OpenAI Leadership for final decision-making” [1]. There is no fourth option on that list — no path by which SAG, on its own, halts a release outright — which is consistent with, not contrary to, everything the Framework says elsewhere about where its authority actually sits.

The Framework does commit, separately, to a form of transparency that is not conditional in the same way: “We will release information about our Preparedness Framework results in order to facilitate public awareness of the state of frontier AI capabilities for major deployments,” including “the scope of testing performed” and “our reasoning for the deployment decision,” while noting that “such disclosures… may be redacted or summarized where necessary” [1]. That is a real commitment, and Astra’s System Card mostly honors it — which is precisely why this article can trace anything at all. But publishing your reasoning after a decision is a different act than sharing the authority to make it, and the Framework’s own text keeps those two acts in separate paragraphs.

Five Names Enter the System Card. None of Them Signs the Sentence.

Astra’s System Card names five outside organizations that touched some piece of the underlying evidence: UK AISI (the UK’s government AI Security Institute), Apollo Research, Irregular, SecureBio, and Gray Swan. Reading what each was actually asked to do — and, as important, what each says it was not able to do — shows the discretionary pattern above playing out exactly as the Framework’s own language predicts.

Five slim manila routing folders in a wall-mounted out-tray, each carrying a small blank external-reviewer tab, with the topmost folder's corner lifted as though mid-handoff

Figure 2. Five outside organizations are named in Astra's own System Card as having touched pieces of the evidence; the Framework's own text places every one of their engagements inside a clause OpenAI itself controls [@openai-gpt-6-astra-system-card]. — Image prompt and art direction by Brecht Corbeel; image generated to that direction.

None of these five engagements is trivial, and crediting their specificity matters before this article turns to what they don’t establish. Irregular solved 86 of 226 challenges on its FrontierCyber suite against 34 of 226 for the prior model, including zero-day vulnerabilities in browsers, mobile phones, and a widely used cloud database, and reported that neither model broke any of its seven hardest “Elite” challenges or attacked a fully hardened target successfully — a finding that cuts directly against treating Astra’s capability jump as unbounded [2]. OpenAI’s own disclosed benchmark numbers point the same two directions at once: the company reported Astra saturating its internal ExploitBench suite at 100%, up from 78.5% for GPT-5.6 Sol, while also disclosing that the testing process itself surfaced two new zero-day vulnerabilities along the way [8]. SecureBio’s numbers are similarly two-sided: Astra’s pre-release checkpoint posted the highest virology-capability scores SecureBio had observed from any model, and was the first checkpoint SecureBio tested that could generate designable proteins within the task’s time and compute limits, while its performance on a separate pathogen-knowledge benchmark actually came in below the prior model, 61.6% against 67.6% [2]. These are real, granular, sometimes unflattering findings, produced by organizations with no obvious incentive to flatter OpenAI, and they are the reason this article can name specific numbers at all rather than gesturing at “external review” in the abstract. What none of the five engagements did, and what none was structured to do, is decide anything.

A sixth channel touched the evidence that the System Card itself never names. Astra was submitted for review by the U.S. government ahead of launch, under the voluntary safety commitments OpenAI and other frontier labs made to the current administration, according to reporting published alongside the release [7]. Pressed on what that review actually covered, OpenAI’s president would say only that there were “nuances on exactly how” the process works and that he did not want to “misstate anything” [7] — a hedge that fits the pattern the rest of this article traces exactly: a door opened because OpenAI decided, in this one instance, to open it, on terms the public record does not yet specify.

Two of these five sit closest to the piece of evidence that matters most for a “sufficient to release” determination on a model whose defining risk is autonomous cyber capability: whether OpenAI’s own monitoring can actually catch the model doing something it shouldn’t. That is UK AISI’s assignment, and UK AISI’s own sentence about it is the sharpest scope-limit in the entire document. Astra shows “a substantial increase in CoT Controllability over GPT 5.6 Sol, following the constraint on 93% of samples compared to 48%,” and a no-CoT math time horizon of 30.9 minutes against 3.6 minutes for the prior model — both real, both concerning for anyone trying to watch what the model is “thinking” [2]. But the sentence that follows those findings is the one that matters for this article’s question, not for the capability question: “UK AISI did not directly test if Astra evades monitors successfully and makes no claims about the overall monitorability of the model” [2]. The organization nearest the sharpest concern in the whole card explicitly declines to render the verdict the card’s headline sentence needs. Apollo Research’s own hedge, on strategic deception, runs the same direction: three days, two with chain-of-thought visibility, on a near-final rather than final checkpoint, is a real engagement and a genuinely narrow one, and Apollo says so itself rather than letting the number stand unqualified [2]. OpenAI’s own chief scientist conceded a version of the same limit in his own words at launch, telling reporters that “current techniques for monitoring and observing models’ behavior may not hold” as systems grow more capable, and that the company would “not accept degradation in our ability to monitor model alignment beyond a certain level” [6] — a commitment made in press remarks rather than in the Framework’s own text, but one that concedes the exact limit UK AISI’s own findings describe.

What OpenAI Actually Committed To in August Is Narrower Than It Sounds

Astra’s path to this System Card ran through a compressed and unusually public three weeks. On August 7, 2026, OpenAI published a short post reporting that its “latest internal evaluations of Astra… indicate significant advancements in agentic coding and cybersecurity” — findings that, combined with “expert assessments” the post never identifies as internal or external, had “led us to conclude last night that we cannot rule out critical cyber capabilities under our Preparedness Framework” [4], alongside a list of immediate steps: stricter security controls for higher-capability workloads, a pause on “internal activities involving Astra that do not yet meet these strengthened security control requirements,” and a specific commitment that “we will work with relevant government agencies and select AI safety organizations to test the capabilities for this model.” Roughly two weeks later, a longer follow-up post detailed what that pause had actually involved: “a two-week pause in reinforcement learning (RL) training on our latest models intended for deployment while we further hardened and red-teamed our research environments,” with the company’s “largest planned frontier RL run” remaining on hold “while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding” [5]. That larger RL run was restarted on August 28, once the new security requirements were in place, according to OpenAI’s own account published alongside Astra’s launch [3].

It would be easy to read OpenAI’s two August posts as a pledge that a cyber-critical capability threshold like this one will trigger mandatory, binding outside review before any future release — the posts read, at a skim, like exactly that kind of commitment. Reading OpenAI’s own two posts directly, that specific commitment is not there. The August 7 post promises to “work with” government agencies and “select” safety organizations of OpenAI’s own choosing, and to “provide recommended security controls to third-party testing partners” — OpenAI setting the terms of engagement, not ceding a checkpoint [4]. The August 18 post is, if anything, vaguer about externality specifically: its closing section promises only that “we intend to involve external organizations and share more of what we learn as our approach develops,” alongside a separate, undated intention to “evolve our Preparedness Framework” [5]. “Intend to involve” and “will work with” are not synonyms for “mandatory,” and neither post uses that word, or any word that binds a future release to an outside party’s sign-off. This matters for the piece’s own argument, not against it: OpenAI’s real-time language, written under the pressure of its own first Critical-threshold determination, tracks the Preparedness Framework’s existing discretionary register — “if we deem,” “when available and feasible” — rather than departing from it. The company did not quietly walk back a mandatory-evaluation pledge between August and September. On the available primary text, it never made one.

A self-inking date-and-approval stamp mounted in its own press fixture on a counter, the stamp arm caught partway down toward the ink pad but not yet touching it

Figure 3. One office holds the mechanism that actually turns a recommendation into a release: OpenAI Leadership can approve or reject the Safety Advisory Group's finding, and even the Board's committee can only review or reverse that decision after it is made [@openai-preparedness-framework-v2]. — Image prompt and art direction by Brecht Corbeel; image generated to that direction.

What the August posts do establish, reliably, is that OpenAI slowed itself down for real: a two-week RL pause, new isolation and monitoring requirements, workloads held back until they met a stricter bar, all before anyone outside the company had rendered a finding on anything [5]. That is a genuine, verifiable act of institutional caution, and it happened entirely inside OpenAI’s own walls, using OpenAI’s own authority, on OpenAI’s own timeline, exactly as the Framework’s own governance chain says it should.

The Fair Reading: A Company That Publishes Its Own Bad News Is Not Nothing

The strongest case against everything above is not that this reconstruction is wrong; it is that it asks the wrong document to do a job it was never built for. UK AISI is a government body operating without statutory power to block a private company’s product launch — expecting its evaluation to function as a veto misreads what an institute like this currently is, under any government’s existing law, anywhere. Judged by that standard, OpenAI publishing UK AISI’s and Apollo’s findings verbatim — including a finding that Astra “has capabilities that could enable it to evade monitoring” and an evaluator’s own admission that it could not test the thing that would matter most — is a more transparent posture than the available alternative, which is running no external evaluation at all, or running one and declining to publish what it found [2]. The August 7 promise to “work with” outside organizations was, on the evidence of the System Card that followed it, kept: five named organizations did engage, across cyber capability, alignment propensity, biological risk, and adversarial robustness, and their findings — flattering and unflattering alike — made it into the published document rather than staying internal [2]. A reader should hold that fact next to everything above it before concluding that “advisory, not binding” is the same as “absent.” It plainly is not; the five rows in the table above are five real engagements, not five names borrowed for cover.

The Determination Is Public; the Authority Behind It Is Not

None of that changes what the opening sentence actually says, or who is named in it. “The internal report informed our Safety Advisory Group’s recommendation and OpenAI leadership’s determination that these safeguards are sufficient for Astra’s public launch” is, on the Framework’s own terms, an accurate description of how the decision was structured: a body OpenAI appoints, whose recommendation OpenAI’s own document says cannot bind anyone, informing a determination OpenAI’s own leadership made — reviewable only afterward, and only if OpenAI’s own Board committee judges review “necessary,” by a body that is itself OpenAI’s Board rather than anyone outside the company [2] [1]. Every external name attached to that determination touched a bounded piece of the evidence, under terms OpenAI set, for a window OpenAI chose, and the organization nearest the single most consequential open question — can anyone actually tell if this model is hiding something in its own reasoning — said plainly that it did not test the thing that would answer that question and made no claim about the answer [2].

The Framework itself anticipates that its own terms will change: it commits to reviewing “the Preparedness Framework for continued sufficiency at least once a year,” with SAG reviewing any proposed changes and Leadership deciding on them through the same process described above, and it separately provides for a “fast-track” if “a risk of severe harm rapidly develops,” under which the SAG Chair “should also coordinate with OpenAI Leadership for immediate reaction as needed” [1]. Both provisions describe internal review triggering internal action — a company updating its own rulebook by its own schedule, using its own emergency channel when one is needed. Nothing in the documents fetched for this piece describes either mechanism as having produced a revised Framework, a fast-tracked external referral, or a Board reversal in connection with Astra specifically. That is not evidence any of these things should have happened; it is only the observation that the governance chain this article has traced ran, start to finish, exactly once, through exactly the bodies the Framework names, and produced exactly the sentence this article opened with.

What would change this account is specific and checkable: a revision to the Preparedness Framework that gives an external body binding authority over a release decision rather than advisory input into one, or a government statement establishing that UK AISI or a comparable institute holds formal power over a private launch that this document’s own language does not currently describe. Neither exists in anything OpenAI, the UK government, or any of the five named evaluators has published as of this writing. Until one does, the chain-of-custody tag on Astra’s safety sign-off has OpenAI’s name on every line that matters, and the document that defines the process was written, entirely deliberately, to keep it that way.

Sources

  1. OpenAI. Preparedness Framework, Version 2. OpenAI (2025).
  2. OpenAI. GPT-6 Astra System Card. OpenAI (2026).
  3. OpenAI. Path to Astra: Critical Capabilities and Frontier Safeguards. OpenAI (2026).
  4. OpenAI. Responding to the Next Frontier of Critical Cyber Capabilities. OpenAI (2026).
  5. OpenAI. Pacing Model Development in an Era of Cyber-Critical Capabilities. OpenAI (2026).
  6. Jared Perlo. OpenAI debuts GPT-6 Astra, says it triggered security measures. NBC News (2026).
  7. Emily Forlini. OpenAI launches GPT-6 Astra, its most powerful model yet, and touts its ability to use your computer. Fortune (2026).
  8. Gyana Swain. OpenAI launches GPT-6 Astra, its first model to cross a critical cybersecurity threshold. CSO Online (2026).

Originally published at https://absolutedigitalpublishers.com/articles/astras-safety-sign-off-chain-of-custody.