GPT-6 Astra Launched: Autonomous Corporate Reasoning
OpenAI debuts GPT-6 Astra with Test-Time Compute and Computer Use, achieving an 80.8% score on SWE-bench and reducing agentic errors by 74.7%.

Stock photo for illustration only, not from the actual event
- OpenAI officially introduces the GPT-6 Astra flagship AI model
- Features native Test-Time Compute and Computer Use capabilities
- Achieves 80.8% on the SWE-bench Verified benchmark
- Reduces catastrophic agentic execution failures by 74.7%
The competition for frontier model supremacy reaches a new milestone with OpenAI officially releasing GPT-6 Astra. Positioned at the absolute peak of the developer's model hierarchy, Astra is engineered to transition away from static predictive models toward autonomous cognitive systems equipped with dynamic test-time compute allocation, native computer use capabilities, and long-duration corporate task execution without intent drift.
Targeting demanding global operations such as formal corporate code auditing, large-scale legal synthesis, quantitative risk modeling, and deterministic sub-agent fleet orchestration, Astra has spent months in early access with select enterprise clients across finance, aerospace, and advanced technology sectors.
The architectural core of GPT-6 Astra is driven by a combination of complementary innovations:
- A dynamic compute scheduler utilizing latent Monte Carlo tree search (Latent MCTS) combined with process reward models
- Native Computer Use allowing real-time pixel-precise interaction with web DOM elements and Linux terminal windows
- Rigorous reinforcement learning alignment with formal verification feedback to prevent goal drift
"In internal evaluations and independent audits, the model produced unintended results and side effects 74.7% less frequently than its primary direct rival, Claude Fable 5.1."
OpenAI
In independent benchmarks, GPT-6 Astra achieved a record 64.6% on the Artificial Analysis index, outperforming Claude Fable 5.1 at 52.6%. On SWE-bench Verified, Astra secured 80.8% in autonomous problem resolution, while achieving 66.4% on the extreme Humanity's Last Exam suite, trailing only Claude Mythos 5.1.
The integration of Test-Time Compute and advanced Computer Use marks a critical shift toward software execution agents capable of handling real enterprise workflows. Robust alignment frameworks and formal verification are becoming vital as models transition from passive assistants to autonomous operators.

Stock photo for illustration only, not from the actual event
Source: Dev.to
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment