⚡ Uncle Cat AI Radar
ModelsAgentsSafety

OpenAI Launches GPT-6 Astra for Long-Horizon Work

Astra combines a 1.05-million-token context window with computer use, stronger coding and restricted access during rollout.

A model built to act

OpenAI has begun rolling out GPT-6 Astra, its new flagship model for extended software, research and computer-use tasks. Initial access is limited to selected organizations, including enterprise customers in a trusted-access program; availability through the API and ChatGPT Plus, Pro, Business and Enterprise plans is scheduled to expand over the following days.

Astra accepts text and images, supports a 1.05-million-token context window and can produce up to 128,000 output tokens. Its toolset includes computer use, web and file search, hosted shell access, code execution, patch application, MCP connections and image generation. OpenAI also introduced asynchronous tool calls and mid-turn steering, allowing applications to add instructions or receive unrelated work while an external tool is still running.

The model costs $10 per million input tokens and $50 per million output tokens, with higher effective rates for prompts exceeding 272,000 tokens. Independent early results are not uniformly favorable on value: Artificial Analysis found gains in some agentic and factuality evaluations, but said Astra was 75% more expensive than GPT-5.6 Sol at maximum effort and often behind it on intelligence per dollar.

Capability and containment

OpenAI describes Astra as a major advance in coding, professional workflows and computer operation. It reports 75.2% on DeepSWE v1.1 and 72.6% on the offline OSWorld 2.0 computer-use benchmark. The company is restricting rollout because Astra is also its first model assessed at the “Critical” cybersecurity-capability level. OpenAI says production safeguards, monitoring and access controls differ from the configurations used in its strongest cyber evaluations.

Why it matters

Astra shifts the frontier-model contest toward systems that can sustain work across browsers, terminals and professional applications rather than merely answer prompts. Its significance will depend on whether those benchmark gains survive long, messy production tasks—and whether the security controls permit useful access without creating an unmanageable cyber risk.

Sources