
GPT-6 Astra Announced: Capabilities, Pricing, and What Changes
OpenAI's Astra launch brings stronger scientific and computer work, new agent APIs, and premium pricing. Here is the evidence and a practical adoption guide.
Read MoreA guide to OpenAI's July–September 2026 releases: GPT-6 Sol and Luna, Astra, image and voice models, with API prices, access limits, and migration choices.

OpenAI released GPT-6 Sol and GPT-6 Luna on September 22, 2026. They join Astra as lower-priced options for coding, reasoning, and agent workflows. Sol is the candidate to evaluate for demanding everyday work; Luna targets focused work at much higher volume. Those roles reflect OpenAI's positioning, not an independent ZharfAI performance ranking.
This guide covers the model launches and major access changes recorded from July 1 through September 25, 2026. It also explains the image, voice, transcription, and specialist models that a text-only comparison would miss. We checked the official API release record and current model documentation on September 25. Recommendations and cost examples below are our analysis; we did not run comparative model benchmarks.
Cover: an original ZharfAI editorial map of the model families discussed here. It is not official OpenAI artwork or a performance chart.
The chronology matters because a new API release, a product rollout, and an expansion of restricted access are different events.
| Date in 2026 | Models or access change |
|---|---|
| July 6 | GPT-Realtime-2.1 and GPT-Realtime-2.1 Mini |
| July 9 | GPT-5.6 Sol, Terra, and Luna |
| July 28 | GPT Transcribe and GPT Live Transcribe |
| August 7 | Daybreak Blue/Red access tiers, including separately approved GPT-5.6 Cyber |
| September 3 | GPT-6 Astra |
| September 8 | GPT Image 2.5 Sunburst/Flare; GPT-Rosalind general availability through trusted access |
| September 10 | GPT-Live 1 general availability in the API |
| September 22 | GPT-6 Sol and GPT-6 Luna |
Our GPT-5.6 family coverage explains the earlier generation. The Astra launch analysis covers that model in depth. This update connects those releases to the newer options without treating every catalog entry as a September launch.
The Sol model card positions it for complex coding and agent work. The Luna card emphasizes efficient, focused tasks at scale. Both accept text and images and produce text; neither is a native image generator or a voice-conversation endpoint. Each lists a 1,050,000-token context window and a 128,000-token maximum output. These are capacity limits, not a guarantee of accurate recall across a million-token document.
Astra remains OpenAI's highest-capability general-purpose tier. A practical evaluation could assign routine classification to Luna, a repository change to Sol, and an unusually difficult investigation to Astra. That routing is a proposed experiment. Keep the model that meets your acceptance criteria at the lowest total cost, including repair time.
There is no GPT-6 Terra in the documented family as of this review. Terra remains part of GPT-5.6. Similarly, a reasoning setting or product label should not be mistaken for a separately launched API model.
The following are US dollars per million text tokens, using Standard processing for prompts with at most 272,000 input tokens. The current API price sheet is the source for this comparison; subscription credits are a separate billing system.
| Model | Input | Cached input | Cache write | Output |
|---|---|---|---|---|
| GPT-6 Astra | $10.00 | $1.00 | $12.50 | $50.00 |
| GPT-6 Sol | $2.00 | $0.20 | $2.50 | $10.00 |
| GPT-6 Luna | $0.10 | $0.01 | $0.125 | $0.50 |
| GPT-5.6 Sol | $4.00 | $0.40 | $5.00 | $20.00 |
| GPT-5.6 Luna | $0.20 | $0.02 | $0.25 | $1.20 |
Sol's listed input and output rates are half its predecessor's current rates. Luna's input rate is also half, while its output rate falls from $1.20 to $0.50. Avoid applying one headline discount to every line item. These are token-price comparisons; a model that uses different amounts of reasoning or requires another attempt can produce a different saving per completed task.
For GPT-6 Sol and Luna, prompts above 272,000 input tokens use twice the input/cache rates and 1.5 times the output rate for the whole request, according to their model cards. The threshold is not a surcharge applied only to the excess tokens. Batch and Flex are listed at half Standard rates; Fast mode is twice the applicable API rate. EU data residency is limited to Standard processing for these models.
Consider an illustrative request with 100,000 uncached input tokens and 10,000 billable output tokens. The text bill is $1.50 for Astra, $0.30 for Sol, or $0.015 for Luna. At 300,000 input tokens and the same output count, the long-context rates make those totals $6.75, $1.35, and $0.0675 respectively.
These calculations exclude tools, cache writes, retries, and regional surcharges. Count billable reasoning output as well as visible text when budgeting. For a document-heavy application, measure the real distribution of input lengths: a workflow that repeatedly crosses the threshold may need better retrieval or compaction before it needs a different model.
The September 22 product changelog describes Sol and Luna rolling out to Plus, Pro, Business, Enterprise, and Edu users in Work and Codex, not ordinary Chat. Free and Go users can access Luna in the desktop app. Rollout state and workspace controls still apply; Enterprise administrators must enable the models. This article does not verify access for any particular account.
For API applications, use the exact IDs gpt-6-sol and gpt-6-luna. A model appearing in a desktop picker does not establish your API organization's entitlement, quota, or regional configuration. Confirm those in the environment that will run the workload.
There is also a dated migration issue: the product model guide says GPT-5.5 retires from ChatGPT, Work, and Codex on October 14, 2026. That notice explicitly excludes the OpenAI API. Teams should inspect saved model defaults and scheduled tasks without turning a product retirement into an unsupported claim of API shutdown.
Sol and Luna support none, low, medium, high, xhigh, and max reasoning effort, with medium as the API default. Their Chat Completions function calling works only with reasoning_effort: "none"; use Responses for reasoning with tools. Astra differs: it does not support none and requires Responses for tool calling.
The GPT-6 integration guide also describes asynchronous tool execution and mid-turn steering. These features can let an agent continue independent work or incorporate a correction while a task is running. Your application still has to execute tools, associate results with calls, and handle pending work correctly.
Before migration, record the endpoint, reasoning setting, tools, output schema, and acceptance checks of the existing application. Replay the same representative inputs against the proposed replacement. Include an interrupted task and a failed tool call. A correct final answer should reflect the latest user instruction and the actual saved state, even when an earlier tool result arrives late.
GPT Image 2.5 Sunburst is positioned for precise editing, while GPT Image 2.5 Flare targets fast everyday generation. Both accept text/image inputs and return images, with additional xhigh and max quality settings. Select them through the Images API or the Responses image-generation tool, rather than treating them as general text models.
Both list $5 per million text input tokens, $8 per million image input tokens, and $30 per million image output tokens. Shared token rates do not establish an identical price per picture: token consumption and selected quality still matter. The documentation specifically says the GPT Image 2 calculator does not estimate GPT Image 2.5 consumption.
For an e-commerce editing trial, compare whether packaging, logos, material appearance, and requested edits survive review. For Persian artwork, check every glyph and joining form at final size. These are suggested acceptance tests; the release documentation is not evidence of flawless Persian typography or product fidelity.
GPT-Live 1 handles full-duplex conversation: it can listen and speak concurrently while delegating reasoning and tools to a backend agent. It uses the Live sessions API. Voice-session pricing is $0.05 per minute, billed per second, with backend model and tool usage charged separately. A twenty-minute session therefore has a $1 voice charge before its backend work.
The earlier Realtime 2.1 and Realtime 2.1 Mini are reasoning voice models with tool use in the Realtime API. Mini is the lower-cost, distilled option. They should not be conflated with the Live API's separate voice-and-backend arrangement.
Choose an architecture before comparing unit prices. A support assistant that keeps speaking while an order lookup runs has different requirements from a bounded voice command. Test interruptions, correction of names and numbers, and what the assistant says when the backend has not finished. Smooth speech should never count as proof that an order changed successfully.
GPT Transcribe covers audio files and finalized Realtime turns at $0.0045 per minute. GPT Live Transcribe produces low-latency streaming text at $0.017 per minute. Both document context and vocabulary/language hints. Twenty minutes costs $0.09 or $0.34 respectively at those listed duration rates.
These are speech-to-text models, not conversational voice agents. For meeting notes, decide whether the user needs live captions or a checked final transcript. Grade names, dates, amounts, omitted words, and corrections on representative recordings. A word-error average can hide the one wrong account number that makes a transcript unusable. Persian and mixed-language calls need their own evaluation set; we have not measured those models on one.
GPT-Rosalind's September event is general availability within trusted access for approved internal life-sciences research. The price sheet lists $5 input, $0.50 cached input, and $25 output per million tokens, with billing beginning October 5, 2026. This is not unrestricted access or evidence of suitability for clinical decisions.
GPT-5.6 Cyber likewise requires separate approval and provisioning for authorized security work. Daybreak Blue and Red are access tiers, not two additional general-purpose GPT-6 models. A team considering either specialist route should establish eligibility and the intended research scope before including it in a production dependency plan.
Start with one recurring task and an acceptance rubric. For invoice extraction, specify required fields, allowed missing values, currency handling, and reconciliation rules. For coding, specify the behavior change and checks that demonstrate it. For an image, identify which visual details must remain unchanged. A single broad score cannot substitute for these different contracts.
Compare completion rate, correction time, total token or duration cost, and elapsed time. Record endpoint, model ID, effort, and cache conditions alongside each result. Escalate from Luna to Sol or Astra only where the observed quality gain justifies it; choose the dedicated media model when the required output is speech or an image. This produces evidence for a purchasing or migration decision instead of a ranking based on names and launch enthusiasm.
Reviewed September 25, 2026. Release dates come from dated official changelogs; specifications, rates, and access conditions come from living documentation checked on that date. The coverage window is July–September 2026 through September 25. No independent benchmark, account-access verification, or Persian-language performance study was conducted for this article.

OpenAI's Astra launch brings stronger scientific and computer work, new agent APIs, and premium pricing. Here is the evidence and a practical adoption guide.
Read More
OpenAI's July 9 family spans flagship, balanced, and efficient models, with state-of-the-art terminal, coding, browsing, and science results.
Read More
How Iranian teams can reach LLM APIs lawfully: domestic providers, gateways, self-hosted open models, Persian evaluation, a cost method, and a client that swaps providers by config.
Read MoreIf this note maps to a real system in your organisation, start with the services page or a shipped case study.