OpenAI officially launches GPT-6 Astra: How to try it


A year after the launch of GPT-5, OpenAI has announced the arrival of GPT-6 Astra, which will begin rolling out immediately. The AI company called Astra “the world’s most intelligent and aligned model.”

OpenAI Astra marks a new generation of models for the AI company, but also carries new risks.

The day before Astra’s launch, OpenAI confirmed that Astra had been deemed a “critical” cybersecurity risk, but that the company would move ahead with a launch with safeguards in place. The company’s Preparedness Framework tracks risk in three domains (chemical/biological, self-improvement, and cybersecurity), and an OpenAI spokesperson confirmed to Mashable that Astra is the first OpenAI model to be designated a critical risk — the highest threat level — in any domain.

Strangely, in its blog post announcing Astra, OpenAI did not call Astra its most advanced model, but only “the most capable model we have ever broadly deployed.” The company also released a detailed system card outlining the model’s training and safety risks.

How to try GPT-6 Astra

OpenAI will be rolling out GPT-6 starting today, though it may not be available to everyday users right away.

OpenAI said that GPT-6 Astra will be available at launch to select organizations, but that it will be available to ChatGPT Plus, Pro, Business, and Enterprise users “over the coming days.” It will also be available via the OpenAI API and AWS.

The API pricing will be:

You can learn more about pricing and availability at the GPT-6 Astra webpage.

GPT-6 Astra: Cybersecurity and safety

OpenAI wrote that “with the right tools and access, GPT‑6 Astra can find previously unknown security flaws and develop new ways to exploit them across many well-protected systems without a person guiding each step.” As a result, Astra’s most advanced cybersecurity skills will only be available to a select group of testing partners, and OpenAI said it’s tightened security around the model as well. Per OpenAI, these safeguards are designed to prevent two outcomes:

  • Stopping malicious actors from using Astra to develop new exploits or carry out cyberattacks against critical systems

  • Stopping the “model itself causing cyber harm when taking an unauthorized (or misaligned) action”

On the ExploitBench cybersecurity benchmarking test, OpenAI reports that GPT-6 achieves a perfect 100 percent score, even at its lowest-tested reasoning level. On the ExploitGym Honeypot benchmarking test, OpenAI reported a perfect 0.0 percent score.

table showing gpt-6 performance on exploitbench benchmarking test

GPT-6 shows significant improvement over GPT 5.6 Sol on the ExploitBench benchmarking test.
Credit: OpenAI

Despite posing existential cybersecurity risks, OpenAI said in its announcement that GPT-6 is generally safer than its recent predecessors. In particular, the company said, “GPT‑6 Astra is significantly safer in higher-risk scenarios.”

The GPT-6 Astra system card reports better scores on safety evaluations measuring risks such as promoting emotional reliance in users, engaging with self-harm requests, and inappropriate responses to users under 18.

table showing safety evaluations for gpt-6 astra and other recent openAI models

OpenAI says Astra achieved higher scores than other recent models in safety evaluations (higher scores are better).
Credit: OpenAI

At the same time, OpenAI admitted that “GPT‑6 Astra’s monitorability has decreased relative to GPT‑5.6 Sol,” meaning that the model may be more difficult to understand and control.

“We have found that GPT‑6 Astra is more capable of controlling its own CoT than GPT‑5.6 Sol, and less likely to include incriminating information in its CoT,” reads an OpenAI blog post. “In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks.”

GPT-6 Astra: Hallucination rate

OpenAI provided fewer details on the hallucination and factuality rates of GPT-6. In previous system cards, OpenAI provided specifics about its models’ propensity for inventing facts or repeating users’ own misinformation.

That being said, the system card states that GPT-6 has made strides in this regard, generating fewer mistruths and hallucinations than GPT-5.6 Sol and other recent models.

chart showing hallucination rates for gpt-6 astra


Credit: OpenAI

“We find that Astra makes substantially fewer factual errors than GPT-5.6 Sol and is significantly less likely to reproduce user-reported hallucinations. These improvements are particularly pronounced at very low latency and reasoning settings.”

Hallucinations remain one of the most stubborn problems with AI models.

GPT-6 Astra: Benchmark performance

OpenAI released benchmark scores for GPT-6 Astra, showing marked improvement on most common benchmarking tests compared to both GPT-5.6 Sol and Fable 5.1, which Anthropic released earlier this week. Notably, GPT-6 Astra scored lower on Humanity’s Last Exam (with tools).

Highlights include:

  • Terminal-Bench Science 0.1: 64.6 percent

  • FrontierMath Tier 4 (v2): 97.6 percent

  • GPQA Diamond: 96.0 percent

  • Humanity’s Last Exam (w/ tools): 57.2 percent

  • ExploitBench: 100 percent

  • Exploit Gym: 42.4 percent

  • ARC-AGI-1: 98.5 percent

    ___________________________________________________________________________________________________________

    Disclosure: Ziff Davis, Mashable’s parent company, in April 2025 filed a lawsuit against OpenAI, alleging it infringed Ziff Davis copyrights in training and operating its AI systems.



Source link