OpenAI Publishes Developer Guide for Its GPT-6 Model Family
Key Takeaways
- •OpenAI's GPT-6 family is split into three tiers—GPT-6 Astra for the hardest reasoning, GPT-6.1 Sol for complex coding and computer-use tasks, and GPT-6 Luna for high-volume work such as extraction, classification, and structured summaries.
- •Pricing spans roughly 100-fold across the lineup: Astra costs $10 per million input tokens and $50 per million output tokens, Sol is priced at $2 and $10 respectively, and Luna costs $0.10 and $0.50.
- •The GPT-6 generation supports tasks lasting hours or days through mid-run steering, asynchronous tool calls, and parallel work, allowing developers to update instructions without canceling already-running tool calls.
- •GPT-6.1 Sol offers beta multi-agent workflows through the Responses API by delegating portions of tasks to subagents, while GPT-6 Astra in Codex can request user clarification while continuing work that does not depend on the answer.
- •OpenAI recommends prompt caching and context compaction as core cost-management measures—cached input can cost up to 95% less than uncached tokens—and advises that simpler prompts may work better than detailed instruction libraries on the newer models.

OpenAI has published a new guide outlining how developers should use its GPT-6 model family, along with new tooling designed to let AI agents operate across tasks that can stretch over hours or even days. The document consolidates how OpenAI intends the family to be used, covering model selection, agent design, prompting style, and cost control in one place.
The company divides the lineup into three primary models—GPT-6 Astra, GPT-6.1 Sol, and GPT-6 Luna—each targeting a different balance of intelligence, cost, and speed. Tiered families of this kind have become a common structure among major AI providers, letting teams route routine work to cheaper models while reserving flagship capacity for the hardest jobs.
GPT-6 Astra is positioned as OpenAI's most intelligent model and is intended for the hardest reasoning workloads. GPT-6.1 Sol is aimed at complex coding, research, and computer-use tasks, while Luna is designed for high-volume workloads with clearer objectives, such as extraction, classification, and structured summaries. OpenAI's current API documentation likewise describes Astra as its highest-intelligence option and Sol as a lower-cost alternative with near-Astra performance.
Pricing varies sharply across the three tiers. Astra costs $10 per million input tokens and $50 per million output tokens. GPT-6.1 Sol is priced at $2 and $10 respectively, while Luna costs $0.10 for input and $0.50 for output—a roughly 100-fold spread between the top and bottom tiers on both input and output. Tokens, the units of text that models read and generate, are the standard billing unit for AI APIs, so where a workload sits in the lineup directly shapes what it costs to run.
Alongside the model tiers, OpenAI is emphasizing longer-running autonomous work with the GPT-6 generation. According to the company, the models can now take on tasks spanning hours or days, supported by features including mid-run steering, asynchronous tool calls, and parallel work. Developers can update instructions while a task is running without canceling existing tool calls, and agents can continue independent work while slower tools operate in the background. For tasks that run for hours, mid-run steering reduces the need to cancel and restart work that has already consumed tokens.
GPT-6.1 Sol also supports multi-agent workflows in beta through the Responses API. The model can delegate independent portions of a task to subagents and combine their findings into a final response. The beta label means the workflow is still in development and subject to change. GPT-6 Astra extends that approach in Codex, OpenAI's coding agent, where it can ask users for clarification while continuing portions of a task that do not depend on the answer. OpenAI recommends that developers explicitly define which decisions an agent can make independently and which require approval.
All three models can also use computer-control capabilities to interact with websites and desktop applications when direct APIs or connected tools are unavailable. OpenAI said this can allow an agent to investigate a software bug, modify code, and then open the product in a browser to verify the fix.
To control the cost of increasingly long agent workflows, the guide recommends greater use of prompt caching and context compaction. Cached input can cost as much as 95% less than uncached tokens, depending on the model. Because agents running for hours read and generate large volumes of text under per-token billing, the guide treats these optimizations as core cost controls rather than optional extras.
OpenAI also noted that developers may need to simplify how they prompt the newer models. The company argues that increasingly capable systems are better at handling nuance and ambiguity, meaning overly prescriptive instructions can sometimes hurt performance rather than improve it. For teams carrying over detailed prompt libraries written for earlier models, the guidance suggests a lighter touch may now work better.