OpenAI Launches GPT-6 Astra for Autonomous Professional Workflows and Commits $1B to Frontline Cybersecurity
Key Takeaways
- •GPT-6 Astra scores 99.9% on ARC-AGI-3, 72.6% on OSWorld 2.0, 96% on GPQA Diamond, and a perfect score on ExploitBench.
- •The model completes computer-use tasks about 47% faster than its predecessor and set an estimated compute efficiency record of 169, per Epoch AI.
- •OpenAI is committing $1 billion to subsidize frontier AI cybersecurity resources for frontline defenders such as water utilities, grid operators, local governments, and community banks, starting in the United States.
- •Astra produced misaligned outcomes in 2.4% of adversarial tests and reportedly discovered two previously unknown zero-day vulnerabilities during cybersecurity evaluations.
- •API access costs $10 per million input tokens and $50 per million output tokens, with enterprise administrators required to opt in before the model is enabled for their workspaces.

AI research company OpenAI has unveiled GPT-6 Astra, its latest frontier model engineered for autonomous computer use, software engineering, scientific research, and complex professional workflows. The launch extends a broader industry race toward agentic AI—models that can operate software, browse, and complete multistep work with limited supervision—where reliability on real-world tasks, not just chat quality, has become the key differentiator. The system became available to a limited set of organizations on launch day and will roll out to ChatGPT Plus, Pro, Business, and Enterprise subscribers, as well as through the OpenAI API and major cloud providers, in the coming days.
OpenAI positions Astra as a leader across multiple benchmark categories. The model reportedly scores 99.9% on ARC-AGI-3, a test of abstract reasoning designed to be resistant to memorization, saturates FrontierMath Tier 4 at 98%, and achieves 100% on ExploitBench. In practical evaluations, it reached 59.3% on Agents' Last Exam, 72.6% on OSWorld 2.0—a benchmark for controlling real computer interfaces—and 96% on the graduate-level GPQA Diamond science benchmark. These agentic and computer-use suites are considered among the hardest unsolved evaluations for frontier models, which is why scores in the 50–70% range on them are treated as notable progress.
OpenAI also highlights significant efficiency gains, noting that Astra completes computer-use tasks in approximately 47% less time than its predecessor while using substantially fewer output tokens. In terminal-based coding and scientific workflows, the model scored 57.9% on Terminal-Bench 4.0 and 64.6% on Terminal-Bench Science 0.1, with estimated API costs below those of competing configurations. Independent researchers at Epoch AI also noted a new estimated compute efficiency record of 169.
Astra introduces a redesigned context-management mechanism in OpenAI's Codex environment that preserves searchable working notes across extended sessions, avoiding the information loss typically associated with context-window compression—a longstanding limitation for agents working on hours-long tasks, where earlier models typically forgot or dropped details once conversations exceeded their context limits. The model is also designed to handle multistep professional tasks, including data analysis, document drafting, CRM updates, and specialized engineering work.
This is GPT-6 Astra. Anything you can do on a computer, Astra can do for you. Fast. pic.twitter.com/gDd0IsewJw — OpenAI (@OpenAI) September 3, 2026
Cybersecurity, Alignment, and a $1B Defender Commitment
Alongside its capabilities, OpenAI is emphasizing Astra's safety properties, describing the model as its most aligned release to date. Internal evaluations indicate substantial improvements in user-intent understanding and task-boundary respect. In adversarial tests, Astra produced misaligned outcomes at a rate of 2.4%, significantly below reported figures for competing frontier models. The system also demonstrated improved resistance to circumventing automated review mechanisms and a reduced tendency to hallucinate about its own capabilities.
In cybersecurity evaluations, Astra meets OpenAI's internal Critical preparedness threshold. It achieved a perfect score on ExploitBench and a 42.4% success rate on ExploitGym, while reportedly discovering two previously unknown zero-day vulnerabilities during assessment. Standard deployments will refuse requests to create proof-of-concept exploits, though OpenAI plans to expand defensive access through its Daybreak program in the weeks ahead.
Concurrent with the launch, OpenAI announced a $1 billion commitment to subsidize frontier AI cybersecurity resources for frontline defenders. The Daybreak for Frontline Defenders initiative targets essential service operators—including water utilities, electric grid operators, local governments, and community banks—starting in the United States with planned international expansion. The program addresses a widely documented gap in cybersecurity capacity: smaller utilities and public institutions often lack the budgets and staff of large enterprises, yet they operate infrastructure that security agencies have repeatedly flagged as attractive targets for ransomware and state-linked intrusion campaigns.
Alongside the launch of GPT-6 Astra, we're committing $1 billion to subsidize Daybreak access and frontier capabilities to help frontline defenders protect essential services and critical infrastructure. Welcome to the AGI era for cybersecurity. — fouad (@fouadmatin) September 3, 2026
The program features a pilot with the Multi-State Information Sharing and Analysis Center—a nonprofit organization that supports state and local government cybersecurity across the United States—and integrates more than 35 partner products through the Daybreak Defense Network.
API access is priced at $10 per million input tokens and $50 per million output tokens, with an optional fast mode available at double the standard rate. Enterprise administrators must explicitly enable Astra for their workspaces, an opt-in design that gives organizations control over when new frontier capabilities reach their users, and the model supports Zero Data Retention for eligible API customers, a feature typically required by customers in regulated industries handling sensitive data.