OpenAI는 백만 토큰 컨텍스트를 지원하며 자율 컴퓨터 작동이 가능한 새로운 모델 GPT-6 Astra를 출시했다.
We are introducing GPT‑6 Astra, the world’s most intelligent and aligned model. GPT‑6 Astra synthesizes years of research and strategic investments across pre-training, reinforcement learning, and alignment. It achieves state-of-the-art performance in computer use, web browsing, software engineering, cybersecurity, scientific research, and professional workflows. On FrontierMath Tier 4, Astra scores 98%, having already contributed to solving longstanding open problems in mathematics. It also saturates ARC-AGI-3 with a 99.9% score and ExploitBench with a perfect 100%. Furthermore, it establishes a new benchmark for computer and browser automation, executing the most demanding professional tasks with unprecedented speed, accuracy, and judgment.
GPT‑6 Astra is launching today to a select group of organizations and will roll out to all ChatGPT Plus, Pro, Business, and Enterprise users over the coming days. It will also be accessible via the OpenAI API, Microsoft Azure, and AWS Bedrock.
In Terminal-Bench Science 0.1, which evaluates agents’ ability to execute scientific research workflows using code and terminal tools—including data analysis, simulation execution, and model fitting—GPT‑6 Astra achieves a leading score of 64.6%. This outperforms Claude Fable 5.1’s 52.6% while operating at an estimated API cost roughly 31% lower. Under a lower-cost configuration, Astra scores 61.1%, significantly surpassing GPT‑5.6 Sol’s peak of 22.4% at an estimated API cost approximately 27% lower.
Astra represents our most rigorously aligned model, featuring substantial improvements in interpreting user intent and governing model behavior. Users can now delegate complex tasks with greater confidence in Astra’s decision-making. To validate this alignment, we developed a new evaluation inspired by the Hugging Face incident, designed to measure whether a model will exceed its operational boundaries when confronted with difficult or impossible tasks. While GPT‑5.6 Sol breached its authorized target 48% of the time without production safeguards, GPT‑6 Astra maintained strict compliance in 100% of test cases.
GPT‑6 Astra establishes a new standard for the speed, accuracy, and safety of computer interaction. It automates routine tasks such as completing online forms, updating CRM customer records, and managing calendars. It conducts web research, drafts email summaries, and populates document editors. Beyond administrative work, Astra analyzes scientific datasets, generates visualizations, builds websites, and executes frontend quality assurance checks to verify functionality. It also autonomously installs, tests, and troubleshoots software, resolving on-screen issues independently. These capabilities are validated by our latest benchmark results.
On Agents’ Last Exam, which assesses agent performance on complex professional tasks within real-world software environments spanning financial modeling, engineering, and media production, GPT‑6 Astra achieves a record 59.3%. This surpasses Claude Opus 5 (55.5%) and GPT‑5.6 Sol (53.6%). At peak performance settings, Astra consumes approximately 65% fewer output tokens than Opus 5.
These advancements deliver substantial efficiency gains for knowledge-work applications. In latency simulations on OSWorld 2.0, Astra completes computer-automation tasks in approximately 47% less time than GPT‑5.6 Sol, achieving a 72.6% success rate in roughly 40 minutes per task, compared to Sol’s 65.7% in approximately 75 minutes.³
GPT‑6 Astra’s computer-automation capabilities span multiple disciplines, including game development, electrical engineering, and daily professional workflows. For example, a 15-second accelerated demonstration shows Astra executing printed circuit board (PCB) layout in KiCad. By placing components and routing copper traces, it converts an electronic schematic into a manufacturable PCB. Since PCB layout remains a manual bottleneck in electronics design, accelerating this process allows engineers to iterate, optimize, and test new concepts at a substantially higher velocity.
Concurrently, we have updated the Codex harness to dramatically accelerate computer-automation speeds. When paired with Astra’s inherent efficiency, task completion on the Mind2Web benchmark runs 1.9 times faster than the current GPT‑5.6 Sol implementation.⁴ This leap in speed enables the model to handle numerous time-intensive personal and professional tasks more quickly than manual execution.
GPT‑6 Astra integrates advanced computer automation with specialized training for enterprise environments, enabling it to navigate complex professional challenges. It merges the analytical depth required for intricate problem-solving with the capacity to execute multi-step workflows and generate polished documents, spreadsheets, and presentations.
On BenchCAD, which evaluates a model’s ability to reconstruct 3D objects from multi-view renders by generating CAD code, GPT‑6 Astra achieves a leading geometric-overlap score of 95.9% when equipped with tool access. This outperforms GPT‑5.6 Sol (83.3%) and Claude Fable 5.1 (84.3%).⁵ In the tested configurations, Astra’s estimated API cost is approximately 43% lower than Sol’s and 86% lower than Fable 5.1’s.
Astra excels at maintaining template fidelity, delivering slides that are professionally formatted and concisely communicate key insights through structured narratives. It generates clear documents, presentations, spreadsheets, and analytical reports that strictly adhere to user templates while mirroring individual writing and visual preferences. Additionally, Astra is optimized to extract only relevant contextual information for each output, eliminating redundant repetition. This ensures the generation of immediately actionable artifacts that align precisely with organizational standards and business contexts.
As demonstrated in the reference file, GPT‑6 Astra created a presentation on GPT‑Gaia—a fictional model—using only a handful of slides from OpenAI’s corporate template. The output maintains consistent tone and layout throughout, guaranteeing that generated slide decks meet exact business formatting standards.
Astra also demonstrates enhanced visual judgment when designing websites, games, applications, and 3D renderings. Through the Sites feature in ChatGPT, users can generate, host, and distribute websites, web applications, and interactive games directly from natural language prompts.
In architectural visualization, Astra models a residential structure in Blender and exports it as a fully navigable environment in Unreal Engine 5, enabling designers and clients to evaluate spatial layouts and immersive experiences prior to construction.
The model animates games with detailed graphics, compelling mechanics, and precise physics, empowering non-technical users to design and play custom gaming experiences that transcend basic prototypes within minutes. (Credit: Pietro Schirano)
When prompts contain ambiguity, GPT‑6 Astra outperforms its predecessors in making sound editorial and technical judgments. It leverages contextual cues to resolve routine uncertainties and poses targeted clarifying questions when outcomes hinge on specific details. Within Codex, Astra operates asynchronously: it continues independent work while awaiting responses, applies reasonable defaults for non-critical items, and pauses exclusively for user approval on consequential decisions.
The following examples illustrate how Astra collaborates on routine tasks where incomplete information could significantly alter the final output.
Astra also maintains superior task orientation as workflows evolve. Previous iterations occasionally misinterpreted corrective feedback as entirely new objectives, causing them to abandon initial parameters. Astra seamlessly integrates new requirements, pivots when directed, and addresses tangential queries without losing sight of the primary objective.
GPT‑6 Astra stands as the premier model for software engineering to date.
On Terminal-Bench 4.0, which evaluates agent proficiency in complex terminal-driven tasks encompassing software engineering, system configuration, and data analysis, GPT‑6 Astra achieves a record 57.9%. This surpasses GPT‑5.6 Sol 2 (37.3%) and Claude Fable 5.1 (55.8%), while reducing estimated API costs per task by approximately 9% and 63%, respectively.
Astra introduces a novel context-management architecture for Codex that preserves and retrieves information once the context window reaches capacity. Traditionally, models relied on session compaction to summarize progress during extended workflows like debugging or large-scale refactoring, often discarding critical nuances regarding failed patches or component behaviors. Astra circumvents this limitation by maintaining persistent notes across context windows, retaining granular details without repetitive compression. Crucially, earlier context segments remain fully searchable, allowing Astra to locate historical requirements or tool outputs even if they were omitted from active notes. This experimental capability is currently available via the config.toml file in Codex and will be enabled by default for Astra in the coming weeks.
GPT‑6 Astra represents a significant leap forward in scientific research, mathematics, and healthcare. Today, we are releasing two additional findings on the ga