GPT-6 Astra officially debuted. OpenAI positions it as currently the most powerful model, focusing on computer operations, professional work, science, programming, and cybersecurity. FrontierMath Tier 4 v2 achieves 97.6%, ARC-AGI-3 peaks at 99.9%, and ExploitBench even scores 100%. What truly deserves attention is that model capabilities are shifting from "solving problems" quickly to "completing tasks."
From answering questions to completing complete tasks
Astra supports over 1 million token contexts and enhances web search, computer use, code execution, and multi-tool collaboration. OpenAI's direction is clear: the model does not just generate answers, but continuously completes multi-step tasks across browsers, terminals, and professional software.
This is especially critical for AI Agents. Complex workflows used to be interrupted due to lost context and failed tool calls; Astra aims to turn models into true "executors," further reducing manual handover in software development, research, and enterprise automation.
99.9% does not equal AGI
The most eye-catching is ARC-AGI-3. ARC Prize independent tests show Astra reached a maximum of 99.9% in the OpenAI Provider Adapter environment but 62.7% in the unified Standard harness. The runtime framework and context management have a huge impact on the final score.
Even so, this is still a clear leap forward. Astra's ability to explore rules, build internal models, and plan actions in unfamiliar environments shows that cutting-edge models are enhancing continuous learning and task execution, but the ARC Prize also clearly states that a high score alone cannot prove AGI has been achieved.
Cybersecurity has reached the Critical level for the first time
Astra scored 100% on ExploitBench, and OpenAI also rated its model's cybersecurity capability as Critical in the Preparedness Framework for the first time. The model is already capable of finding unknown vulnerabilities, developing exploitation chains, and performing complex security tasks.
However, OpenAI acknowledged that historical vulnerabilities could have contaminated ExploitBench data, so it used new vulnerabilities disclosed between June and August 2026 for internal testing. During testing, Astra also discovered and exploited two previously unknown zero-day vulnerabilities, with related issues currently being responsibly disclosed.
What GPT-6 Astra truly changes is not a single leaderboard, but the boundaries of AI capabilities: when models can continuously operate software, conduct research, and handle complex engineering tasks, the next round of AI entrepreneurship competition will shift from "integrating large models" to "who can truly take over the entire workflow."