Morning Edition · Monday, September 14, 2026Published at 2:22 AM EDT · New York
OpenAI's account describes Astra writing communications, changing code, and monitoring live systems with far fewer human check-ins, weeks after the same model was rated Critical for cyber capability under OpenAI's own risk framework.
OpenAI published a customer account on September 14 describing how Perplexity uses GPT-6 Astra. According to the write-up, the model drafts communications, modifies software, and monitors production systems, and Perplexity's engineers check in on it considerably less often than they did with earlier models. OpenAI also says Astra builds its own small test programs that imitate the responses a dependent service would return, so it can exercise a workflow end to end before a human reviews the result.
The claim is a vendor claim. OpenAI publishes no error rate, no rollback count, and no independent measurement of how often Astra's unattended changes needed correction. Perplexity is also a close commercial partner, so the account is best read as evidence that one sophisticated customer is willing to grant that autonomy, not as a measurement of how reliable that autonomy is.
The autonomy claim coincides with a separate risk classification. CSO Online reported that Astra is the first OpenAI model to reach the Critical cybersecurity tier under the company's Preparedness Framework, meaning it can autonomously find and exploit previously unknown vulnerabilities under the right conditions. OpenAI gates those offensive capabilities behind a vetted-access program called Daybreak, while the generally available model refuses the tasks. Astra is priced at $10 per million input tokens and $50 per million output tokens with roughly a one million token context window, placing it at the same headline price as Anthropic's Claude Fable 5.1.
Separately, an AI channel on Telegram circulated a demonstration in which Astra built a working simulation of a violin and then trained itself to play it, later attempting original composition. That claim rests on a single unattributed post and has no primary source behind it, unlike the Bach chorale and Ableton Live demonstrations that accompanied Astra's launch and were reproduced by multiple outlets.
Part of a tracked trend
Agentic AI Moves Into Enterprise and Government Workflows
Over the next 3-9 months, AI agents move from demos into real enterprise and public-sector workflows, with deployment success tied to domain and task understanding more than raw model capability.
Start a discussion in Townsquare.
More from this edition
OpenAI, which needs buyers to accept that a model rated Critical for cyber capability is safe enough to edit production code, since the revenue case for Astra at $10 and $50 per million tokens rests on replacing supervised engineering hours rather than on benchmark scores.
The OpenAI write-up and the quoted claim from Perplexity co-founder Johnny Ho about trusting the model with end-to-end systems are genuine, and the Critical rating under the Preparedness Framework is independently reported, but "checks in much less frequently" is a customer testimonial with no error rate, rollback count, or audit behind it, and both parties have a commercial interest in the impression it creates.
An open-source-intelligence read of how likely this story is true with its real nuance, not a judgment of any outlet. It assesses the claim, weighing independent and adversarial reporting. How we label confidence.
What this means
The economically significant number in agent deployment is not a benchmark score, it is supervision frequency. If a customer of Perplexity's sophistication genuinely reduces check-ins on production changes, the labor cost per unit of software work falls in a way benchmark suites do not capture, and the buyer of AI shifts from a per-seat tool to a metered replacement for on-call engineering time. The counterweight is that the same model carries a Critical cyber rating from OpenAI, so the software that can be trusted to edit production is also the software with the strongest demonstrated ability to break into it, which puts the burden on customer-side permission scoping rather than model-side refusal.
What to watch
Observations to monitor, not financial advice.
Synthesized from: OpenAI · Polylog editors
Comments
0No comments yet.