Morning Edition · Thursday, September 10, 2026Published at 2:24 AM EDT · New York
The model reports 72.6 percent on OSWorld 2.0 and more than doubles its predecessor on AutomationBench, with all published figures coming from OpenAI's own evaluations.
OpenAI followed the September 3 launch of GPT-6 Astra with a business-focused rollout, positioning the model for work through reasoning, computer use, and writing and design judgment. Access reaches Pro, Business and Enterprise plans and the application programming interface (API), with enterprise administrators required to switch it on because it is off by default in managed workspaces.
The published figures focus on the model's ability to operate software. Astra reports 72.6 percent on OSWorld 2.0 with roughly 47 percent less time per task than GPT-5.6 Sol, and AutomationBench rises from 18.1 percent to 41.4 percent, indicating that longer sequences of actions complete without human intervention. On security capability, OpenAI reports 100 percent on ExploitBench against 78.5 percent for the prior model, and 42.4 percent on ExploitGym against 30.3 percent. API pricing is 10 dollars per million input tokens and 50 dollars per million output tokens.
Every one of those figures is vendor-reported, and a more accurate reading is narrower than the marketing suggests. Independent write-ups conclude Astra is a clear efficiency and computer-use gain over OpenAI's own prior models rather than a decisive move past rival models. A paper posted to arXiv this week adds a specific caution: fixed-rollout pass@k evaluations, a method that tests success across only a fixed number of attempts per task, capture just a limited set of quantities from the samples actually collected, so extrapolating success rates far beyond that sampled budget is not supported by the data.
Demonstrations are circulating faster than measurements. One widely shared clip shows Astra directing a robot arm holding a paintbrush and improving across attempts, an impressive demonstration that lacks a controlled baseline for comparison.
What this means
Computer-use scores are now the dimension OpenAI is competing on, which shifts enterprise buying from chat quality to how many steps of internal software a model can complete unattended. Software vendors whose value sits in the user interface lose leverage when an agent operates the interface instead of a person, while OpenAI gains a reason to charge work-tier prices. The unresolved question is whether independent evaluations on real enterprise environments reproduce the OSWorld gap or show it narrowing under messier conditions, and enterprise pilots will settle that within a quarter.
Part of a tracked trend
Agentic AI Moves Into Enterprise and Government Workflows
Over the next 3-9 months, AI agents move from demos into real enterprise and public-sector workflows, with deployment success tied to domain and task understanding more than raw model capability.
Start a discussion in Townsquare.
More from this edition
What to watch
Observations to monitor, not financial advice.
Synthesized from: OpenAI News · Polylog editors · arXiv stat.ML
Comments
0No comments yet.