OpenAI launches GPT-6 Astra with hacking risks in check
- GPT-6 Astra is the first OpenAI model to hit the "critical" cybersecurity threat level, but deployment is moving forward with significant safeguards in place.
- It is currently limited to defensive tasks like secure code review. Advanced capabilities like exploit creation are blocked for now, but trusted users in the Daybreak program will eventually get access to complex workflows.
- Astra drastically outperforms GPT-5.6 Sol, hitting a perfect 100% on ExploitBench and 42.4% on ExploitGym.
OpenAI just released GPT-6 Astra, and it's the company's first model to reach the critical" threat level in the company's Preparedness Framework for cybersecurity. Despite these unprecedented security risks, OpenAI is moving ahead with the deployment.
That doesn't mean GPT-6 Astra is being rolled out as an unrestricted hacking tool. The version rolling out now will do defensive work such as secure code review and patching, but will not entertain more advanced requests like creating proof-of-concept exploits, OpenAI says.
The company adds that it intends to relax some of those restrictions with its Daybreak program for trusted users, ultimately supporting workflows such as vulnerability validation, malware analysis and detection engineering.
So, how did we arrive at this point? OpenAI had to severely lock down its development pipelines after the recent mess where rogue AI agents hacked the Hugging Face platform. Astra is a huge step up from GPT-5.6 Sol in capabilities.
Massive performance leapIn OpenAI's testing, the new model scored 100% on ExploitBench, compared to 78.5% for GPT-5.6 Sol. It also scored 42.4% on ExploitGym against 30.3% for its predecessor. More worrying, Astra uncovered and took advantage of two new zero-day vulnerabilities in the course of testing, which OpenAI says it is reporting to the appropriate maintainers.
(Image credit: OpenAI)OpenAI explains that GPT-6 Astra is built on improvements to pre-training, reinforcement learning, alignment, and computer use. It can run software, surf the web, write code, and complete multistep tasks. OpenAI says its OSWorld 2.0 testing shows it's about 47% faster per task than GPT-5.6 Sol, while scoring 72.6% vs 65.7%.
The company says it has responded with stronger jailbreak resistance, broader monitoring, encrypted model checkpoints, tighter access controls, and misalignment monitoring across tool-using deployments. Astra also outperformed Sol in remaining within authorized boundaries.
But one big weakness that OpenAI does acknowledge is that it's actually a lot harder to monitor. Internal testing indicated that GPT-6 can actively influence its own "chain of thought" reasoning to avoid leaving incriminating evidence. In adversarial situations, the model has even sandbagged its performance to fly under the radar and avoid internal monitors.
For developers, Astra costs $10 per million input tokens and $50 per million output tokens, with an accelerated API at twice the normal price. It's rolling out now to a limited number of people before it's available to ChatGPT Plus, Pro, Business and Enterprise users and the API and Amazon Bedrock.
