According to OpenAI, its forthcoming Astra model has achieved “Critical” status on its cyber risk scorecard, earning top marks in the company’s internal preparedness guidelines.
Naturally, these terrifyingly advanced capabilities will only be handed to a trusted few dubbed Daybreak Blue, which reportedly includes heavyweights such as Cisco, Cloudflare and Palo Alto Networks. OpenAI has also flagged that its safety monitoring systems could prove restrictive, warning users that legitimate activity might be mistakenly blocked as suspicious behaviour.
That’s an unusual amount to disclose before a product launch. It also raises an important question: is this responsible transparency, or a more sophisticated way of launching something nobody outside OpenAI fully controls?
Defining “Critical” In OpenAI’s Lexicon
This “Critical” designation exists solely within OpenAI’s proprietary safety scale, carrying no weight with external regulators.
By OpenAI’s definition, a model reaches this level when it can autonomously write zero-day exploit code against protected targets, or dream up and execute a multi-step cyber campaign from scratch. It’s a measure of what the tool can do, not what it wants to do. The company isn’t warning of an impending AI attack, but simply admitting the software can automate tasks that once required a seasoned hacker.
During testing, OpenAI says Astra found two previously unknown vulnerabilities and used them as part of an exploit chain, including a browser-compromise chain that escaped a sandbox and executed commands on a host. The company is still passing those details along to relevant parties to be patched.
The big caveat is that these tests were run under the elevated Daybreak Blue setup, meaning everyday users are unlikely to see this level of performance when the model actually launches.
The Transparency Argument, And Its Limits
The case for going public right now comes down to three motives: flagging that a key risk threshold has been breached, giving governments and security teams room to brace for impact and setting expectations around blocked tasks before frustrated users run into them. It’s a much more meaningful style of transparency than bragging about test results. It outlines the model’s actual power while highlighting the points where security measures will cause friction.
All this openness is still carefully filtered through the company’s PR department. OpenAI set the baseline metrics, controlled the testing environment, selected the public disclosures and chose the favoured partners. While a launch-day system card may offer further technical background, it falls short of an independent audit.
More from Artificial Intelligence
- How Does Machine Learning Work? The Five Algorithms You See Every Day
- Stop Losing Money By Ignoring AI: Advice From $1B Company Founder Shane Morand
- Lunar Cyber Launches Token Exposure Monitoring As Infostealers Target Developer And AI Credentials
- Malware Is Now Stealing Claude Sessions To Drain Paid AI Usage – How Does That Work?
- 5 Of The Most Promising Humanoid Robotics Companies Changing The Industry
- What Does It Mean To Be AI Agnostic?
- AI Is Hungry For Power, Is Nuclear The Answer?
- AI Was Meant To Fix Drug Discovery – So Why Do Nine In Ten Drugs Still Fail?
Who Gets Inside The Club?
Limiting advanced cyber capabilities to Daybreak Blue members makes sense on the surface, given that high-stakes managers need powerful tooling. In practice, though, it creates an unequal playing field. A favoured inner circle gets a head start on finding security flaws, long before smaller teams, independent researchers or standard IT departments can access anything similar.
The legitimacy of this setup comes down to some missing details. OpenAI hasn’t explained who qualifies for Daybreak Blue, nor what proof is required to confirm a protective mission. It’s also unclear if independent researchers or smaller organisations will get any seat at the table at all. Should a member find a zero-day flaw, does the community hear about it?
A head start for security and a corporate walled garden can look the same. Which one this becomes rests on facts OpenAI is withholding.
What This Costs Ordinary Users
OpenAI is candid about the fact that its safety controls will trigger false alarms.
Safety monitoring can delay, pause or stop valid operations, particularly benign security tasks that look like malicious activity. Where ChatGPT or Codex might prompt a user to verify an action, API processes simply stop without warning.
This creates an operational problem, beyond just a minor nuisance. A security team investigating a live intrusion often requires an agent to execute aggressive commands, which an automated safety check could easily flag as hostile behaviour. An intervention during a time-sensitive emergency introduces a costly delay. Conversely, removing those controls would leave the target exposed to the model’s capabilities. OpenAI is effectively giving customers a higher-performing tool while restricting their ability to see why execution stopped or bypass the automated refusal.
OpenAI highlights that Astra rejected 91.5% of malicious prompts in cyber-jailbreak tests, up from 59% in the prior model, while avoiding host networks entirely during honeypot trials. These statistics sound impressive, but they fall short of a full security evaluation.
Without visibility into the testing methodologies, refusal definitions, false-positive frequencies for valid work and endurance against evolving attack methods, independent verification isn’t possible. Genuine transparency is certainly questionable when the entity profiting from the technology controls every variable, writes the evaluation criteria, selects the reviewers, designs the controls and certifies its own safety.
