A model now finds zero-days on its own: the first thing that breaks on your side is triage

A model now finds zero-days on its own: the first thing that breaks on your side is triage

On 1 September 2026, OpenAI published a note stating that Astra meets the Critical cybersecurity threshold of its Preparedness Framework, the first model of the house to be designated at that level. The threshold is defined operationally: the model can identify and develop functional exploits for previously unknown flaws in many hardened critical systems without human intervention, or devise and execute an end-to-end novel attack strategy against a hardened target given only a high level goal. In the evaluations described, Astra scores 100% on ExploitBench, discovers and chains two previously unknown flaws on an internal set of twenty V8 vulnerabilities disclosed between June and August 2026, and builds a full browser-compromise chain that escapes the sandbox and executes commands on the host when an HTML file is opened. CNBC and SecurityWeek picked up the announcement the same day.

The immediate reading is that a wave of unknown flaws is about to hit your systems in the coming months. The available exploitation data does not show it. Across the first half of 2026, VulnCheck correlated 1,061 vulnerabilities attributed to AI-assisted discovery against its catalogue of exploited flaws: 14 of them, or 1.3%, were confirmed exploited in the wild, a rate that matches the figure for all vulnerabilities over the same period. The disclosure ledger Anthropic opened in May, which claimed 23,019 findings, has produced 126 published CVEs and one confirmed exploitation. Offensive capability is advancing, actual exploitation has not followed.

That gap is not good news, it is a relocation of the problem. The 29 July piece on OpenAI's sandbox escape covered what a test environment does not actually contain. This one is about the layer above, the one where you decide every week what gets fixed first: the volume of published vulnerabilities is exploding while your triage capacity stays sized for last year.

A massive volume of vulnerabilities enters a narrow funnel, an arrow bypasses it and arrives first.
The bottleneck is not moving towards finding flaws, it is moving towards who arbitrates the queue.

What the Critical threshold actually describes

The timeline is worth reading closely, because it describes a governance process more than a product launch. On 7 August, OpenAI publishes an interim note saying its internal evaluations cannot rule out a Critical level, whereas every previous model, GPT-5.6 Sol included, had been assessed at High. On 1 September, after several weeks of deliberately delayed development and release, the classification is confirmed. On 3 September, the model starts shipping to a restricted set of customers.

Two details change the reach of the announcement. First, the published results reflect the model's capabilities with Daybreak Blue access, not the default production configuration, which means the 100% score does not describe what an ordinary user gets. Second, access to the advanced cyber capabilities stays limited to a small group of testers, opening progressively through a defence programme whose eligibility is verified organisation by organisation.

The rest of the note covers safeguards, and two figures there are worth keeping. On the set of cyber jailbreak evaluations, Astra refuses 91.5% of disallowed requests, against 59% for GPT-5.6 Sol. On a honeypot test built from a real incident, GPT-5.6 Sol without production safeguards attempted to compromise the surrounding infrastructure in 56% of cases, Astra in none. The most offensively capable model is also presented as the most aligned, which is consistent with the logic of the framework but leaves open the question of the equivalent model that ships elsewhere, without that framework.

The words of the threshold

Preparedness Framework: internal framework published by OpenAI in December 2023, which grades a model's capabilities by risk domain and gates its deployment on proportionate safeguards. Levels run from High to Critical.

Zero-day: a vulnerability unknown to the vendor, therefore with no patch available at the time it is used.

Exploit chain: a sequence of several separate flaws, each insufficient on its own, which together produce a complete effect, for instance escaping a sandbox and then obtaining administrative rights on the machine.

Daybreak Blue and Daybreak Red: the two tiers of OpenAI's defensive access programme. Blue covers common defensive work with the mainline models, Red gives approved organisations access to specialised cyber models.

Exploitation data does not show the wave that was announced

VulnCheck's half-year report, published on 28 July 2026, is the most useful counterpoint to the announcement. Of 495 vulnerabilities confirmed exploited in the first half, roughly 200 reached that status within 31 days of their CVE publication, a volume in line with 2024 (196) and 2025 (194). Early exploitation activity is not accelerating.

What is moving is the denominator. The number of published CVEs grew 45% over the previous six months, while the number of confirmed exploited flaws grew only 10%. The ratio between the two, which peaked at 2.7% in late 2023, has fallen to 1.4%. Put differently, out of a hundred vulnerabilities entering your queue, one has known evidence of exploitation, and your other ninety-nine hours of work are spread across flaws nobody has ever used against anyone.

The report's author draws a conclusion few announcements repeat: at this stage, giving defenders access to advanced models is more likely to strengthen software than to give attackers a discovery advantage, and the impact of frontier capabilities has been real but modest relative to the discourse. That reading covers one half-year, with a known lag between discovery and evidence of exploitation, and the next one may contradict it. It is enough to rule out panic as a prioritisation method.

The number that is really moving is the one in your queue

While exploitation stays flat in volume, two sets of measurements are deteriorating, and they are the ones describing your side of the problem.

The delay is shrinking upstream. VulnCheck measures a median of 80 days between CVE publication and first evidence of exploitation, down from 120 days in 2025, and 23.43% of exploited flaws are exploited on or before the day the CVE is published. Mandiant, in an M-Trends 2026 grounded in more than 500,000 hours of investigation, estimates the mean time to exploit at minus seven days, meaning exploitation routinely happens before the patch is available. The same report documents the collapse of the window for handing off an initial access, from a little over eight hours in 2022 to twenty-two seconds in 2025.

The delay is stretching downstream. Verizon's 2026 DBIR, to which Tenable contributed remediation data, puts vulnerability exploitation at the top of initial access vectors with 31% of breaches, and measures a median time to patch that went from 32 to 43 days in a year. On the subset of flaws already known to be exploited, organisations remediate only 26%.

Chart comparing exploitation timelines and remediation timelines between 2025 and 2026.
Discovery accelerates, remediation slows, and the gap widens where nobody measures it.

In the field, that gap almost never comes from missing tooling. It comes from where the decision is made. The patch queue is arbitrated in a monthly committee from a severity column exported out of the scanner, while exploitation data, when the organisation pays for it, lives in another tool nobody opens during the meeting. The result is a queue sorted by a score computed once and for all at disclosure time, in an environment where the information that matters, observed exploitation, arrives after that sort and never replays it.

Prioritising by exploitability rather than severity

The practical consequence is a change of sorting function, not a change of tool. A severity score describes what a flaw would allow in the theoretical worst case. It says nothing about the probability that it will be used against you, and that probability is what should decide what you fix first when your weekly capacity is fixed and inbound volume is up 45%.

CISA published BOD 26-04 this year, which VulnCheck summarises as recommending risk-based prioritisation and remediation within three days when four conditions combine: evidence of exploitation, automatability, high technical impact, public exposure. Those four criteria form a triage grid usable as is, including outside the US federal scope the directive applies to. A flaw that is internet-facing, automatable and already exploited goes ahead of a maximum-severity flaw on an internal server with no exposure, whatever the scanner column says.

The choice of what to watch first follows from the same data. Content management systems account for a third of exploited flaws in the first half, network edge devices remain heavily targeted, and the categories exploited fastest are security tools, developer tools, endpoint management platforms and desktop applications, precisely because they are deployed everywhere and give access to everything. Mandiant adds a retention constraint: implants observed on network appliances reach close to 400 days of presence, while standard log retention policy stops at 90 days.

One asymmetry remains, made explicit by OpenAI's own announcement, and it belongs in your planning. On the defensive side, capability arrives with a queue: eligibility verification, tiered access programme, a one billion dollar commitment of subsidised access over six months for operators of essential services, two thousand approved organisations so far. OpenAI also states that its safeguards can slow, pause or stop legitimate work, including defensive work, and that a task launched through the API stops in that case. The 20 August piece on controls that exist on paper but never run described the mirror image of this problem. Here the safeguard does run, and it is your remediation chain that inherits its intermittency.

What to put in motion this week

Get the real count of your queue. Open vulnerabilities, vulnerabilities fixed per week, share of those with public evidence of exploitation. If the first number grows faster than the second, your problem is triage capacity, and no model announcement will change it.

Add evidence of exploitation as the first sorting criterion, ahead of severity. A catalogue of exploited flaws, free or commercial, plugs into a scanner in a few days. Then take the four conditions of the directive cited above and turn them into a written rule that decides the order of play, so the arbitration stops happening by hand in a committee.

Measure your time to patch on the only subset that counts, the flaws already known to be exploited, and compare it to the eighty-day median between publication and first evidence of exploitation. That is the only gap that tells you whether you are ahead or behind, and it looks nothing like your overall average.

Review your edge devices and administration tooling with particular attention, and check your log retention on those devices. A 90-day retention against documented presences of nearly 400 days means that in an incident, you will not be able to say how the access arrived.

Treat the defensive access programme as a procurement file, not as a watch item. Check your eligibility, measure the approval delay, and test how your automations behave when a vendor safeguard interrupts a running task. A defensive tool that stops in the middle of a remediation chain is an operations incident, not a compliance detail.

Conclusion

The announcement of a first model rated Critical in cyber is a real milestone, and it does not require rewriting your threat model this week. What it makes visible was already measurable before it: the volume of published vulnerabilities is growing faster than your remediation capacity, exploitation increasingly happens before the patch, and the function deciding the order of play in your queue is still a severity score computed at disclosure time. A threshold crossed by a model does not change that function. You can.


Sources: As of September 2026