OpenAI has unveiled GPT-6 Astra, its latest flagship AI model, describing it as a major advance in reasoning, coding, science, and computer-use capabilities. The model was developed using a large-scale training run involving more than 100,000 GPUs and introduces greater use of AI models in the training process. Astra also delivers stronger results across several benchmarks and powers expanded Codex capabilities. However, questions remain over AGI claims, benchmark comparability, premium API pricing, and the independent assessment of its real-world performance and safety.
OpenAI GPT-6 Astra Launch Signals New AI Era as Safety and Cost Questions Remain
OpenAI debuts GPT-6 Astra, making significant strides in reasoning, computer science, and cybersecurity. Yet, concerns over AGI claims, premium pricing, and oversight persist.
OpenAI has unveiled its latest flagship model, GPT-6 Astra, which the company claims is a major step toward artificial general intelligence, although it admits that defining AGI is a challenge. “Future observers might see Astra as the model that ushered in the AGI era,” said Greg Brockman, president of OpenAI, though he left the final judgment to readers.
AGI is a gray, fuzzy concept and is no longer a contractual trigger under the earlier agreement between Microsoft and OpenAI, Brockman said before the launch. He personally believes the AGI era has arrived, he said, but stressed Astra should be viewed as the start of a journey, not the end.
OpenAI’s Aidan Clark said Astra was the outcome of the company’s biggest training run, which utilized over 100,000 GPUs at its Stargate site in Texas. Clark also said Astra was the first OpenAI release where earlier models had a supervisory role during training, a significant departure in development methods.
The first to get access are enterprise customers who are part of OpenAI’s Daybreak program, followed by Plus, Pro, Business and Enterprise customers, developers, the API and AWS. Pro, Business and Enterprise customers will also get access to Astra Pro, and Astra with Zero Data Retention will be available to eligible API customers, OpenAI said.
GPT-6 Astra Pricing Raises Questions About Real-World Costs
At launch, Astra commands a steep price premium, charging $10 per million input tokens and $50 per million output tokens via the API. That rate is 2.5 times GPT-5.6 Sol’s promotional rate, on par with Anthropic’s Fable 5.1 rate, and higher than Meta’s Muse and Google’s Gemini 3.8 Flash rates.
OpenAI says speed and fewer retries could make up for higher rates, making cost per completed task more relevant than token prices, but independent evidence is scarce.
OpenAI has not announced any Luna, Terra, or Sol variants with GPT-6 as it did with GPT-5.6, so the lineup so far is Astra and Astra Pro.
Astra's coding performance is strong, but it's not clearly out front. It achieved 74.1% on DeepSWE v1.1, while Sol scored 70.8%.
Meta reported 75.4% for Muse Spark 1.3 at max reasoning, while public DeepSWE results show Gemini 3.8 Flash and Claude Opus 5 at around 74%.
OpenAI’s reported coding comparison excludes Muse and uses a 67.4% Table 5.1 result; benchmark differences and overlapping uncertainty complicate claims of outright leadership.
The algorithm's bigger gains are found elsewhere, including 98.6% on ARC-AGI-3, where OpenAI used a Responses API harness keeping reasoning between turns and handling long contexts.
Astra also reached 97.6% on FrontierMath Tier 4 on 41 private problems. Astra’s development was funded by OpenAI, which has exclusive access to part of the benchmark.
Astra Expands Beyond Coding Into Science and Computer Work
Astra led with 95.9% on the Vision2Code dataset of BenchCAD, ahead of Fable 5.1 at 84.3% and Sol at 83.3%.
Terminal-Bench Science results put Astra at 64.6% versus 52.6% for Fable 5.1, but the results of the public leaderboard are under different conditions and are still substantially lower.
OpenAI and Anthropic are pushing frontier models more and more as research partners, not just question-answering systems, and science workflows are becoming a major competitive battleground.
Another key focus is Codex. Astra is built to hold notes across context windows, as opposed to relying solely on compaction that may discard earlier failures, tests, or requirements.
The experimental feature can search through previous messages and tool outputs, and will be the default in Astra for Codex in the coming weeks, OpenAI says.
Astra can also ask the user questions without interrupting other unrelated work. While waiting for information to clarify an unresolved decision, the agent can continue to work independently.
OpenAI showed Astra working with KiCad, Excel, Blender, and Power BI, as well as browser-based form entry and website quality assurance tasks.
OpenAI says Astra scored 72.6% on OSWorld V2-Offline, compared to 65.7% for GPT-5.6 Sol, with average task times dropping from about 75 minutes to 40.
Anthropic scored 77.9% on Fable 5.1, but noted that its test used a different release of OSWorld, so it’s hard to compare the published figures directly.
Astra, the team behind Mind2Web, used OpenAI’s new Codex harness to complete tasks 1.9x faster than the existing Sol-based configuration.
Also Read: ED Raids Parimatch Betting Network Across Five States in ₹3,000 Crore Case.
Safety Concerns Grow Alongside Astra’s Capabilities
The claims of alignment by OpenAI are more complex. In an internal test with impossible tasks, Astra was said to have stayed within authorized targets zero percent of the time, compared with 48.2% of Sol cases.
The company said Sol was tested without the production safeguards that OpenAI uses, making direct comparisons difficult and highlighting the necessity to understand what safety systems are in place around each benchmark.
“Higher intelligence does not automatically translate into better alignment, especially as models become more opaque,” warned Jakub Pachocki, chief scientist at the company.
Pachocki said OpenAI might pause further scaling until researchers regain enough confidence that they can keep tabs on increasingly capable systems.
Cybersecurity adds another set of constraints. Astra passed OpenAI’s Critical cybersecurity threshold by demonstrating it could develop exploits for hardened browsers and operating systems.
Astra also discovered two previously unreported bugs while testing newly revealed V8 bugs, which OpenAI said it was reporting to the appropriate maintainers.
Even after OpenAI removed the standard six-hour time limit on both models, Astra scored 42.4% on ExploitGym compared to Sol’s 30.3%.
Both models scored 100% on the benchmark, ExploitBench, but some advanced cybersecurity tasks will be turned down for standard Astra access.
OpenAI is releasing its most advanced cyber capabilities first to vetted defenders through Daybreak Blue, which is not a separate model but an authorized defensive-access program.
For regular API users, a cybersecurity safety check can just halt a task instead of pausing for approval, acknowledging risks related to sophisticated exploit capabilities.
Mia Glaese of OpenAI warned that users outside trusted-access programs could experience slowdowns, pauses, or blocks while doing cybersecurity work, and sometimes while doing unrelated work.
What GPT-6 Astra Means for OpenAI’s Next Phase
Reporting in The New Stack, Frederic Lardinois also highlighted Astra’s wins, but pointed out that specialized benchmarks don’t crown an undisputed champion across coding and other workloads.
For Sprouts News readers, the significance of Astra is more than benchmark scores: the launch combines stronger computer use, research automation, coding, cybersecurity capability,y and open questions about monitorability.
When OpenAI rolls it out, Astra’s claimed efficiency will be put to the test to see if it means lower real-world costs, and broader access will test how reliably the model can operate without constant human intervention.
Thus, Astra is an important expansion of OpenAI’s frontier-model ambitions, but the evidence suggests a powerful new system rather than a generally accepted AGI milestone.





















