GPT-6.1 Sol: Pricing, vs Astra and Opus 5.5, Where It Is in ChatGPT, API and Amazon Bedrock

Logeshwaran
—

GPT-6.1 Sol is the upgrade to GPT-6 Sol that OpenAI released at DevDay on September 29, 2026. It gets close to GPT-6 Astra on coding, computer use and document work at one-fifth of Astra's price: $2 per million input tokens, $0.10 cached and $10 output on the API. It also arrived on Amazon Bedrock the same day, as openai.gpt-6.1-sol. Here is the part the launch posts bury: if you open ChatGPT and look for it in a normal chat, you will not find it. GPT-6.1 Sol lives only in ChatGPT Work and Codex, and only on paid plans. The other surprise is the price tag. It costs exactly what GPT-6 Sol cost, except that cached input is now half price, so the upgrade itself is free.

Jake runs a phone repair shop, and over the summer a freelancer built him a booking site plus a small tool that reads supplier invoices into a spreadsheet. Jake pays for ChatGPT Plus, and on the Tuesday after DevDay he wanted the new model to read the invoice tool's code and explain why it kept skipping lines with a discount on them. He opened ChatGPT, opened the model menu, and saw no GPT-6.1 Sol. He refreshed, logged out and back in, and then called Ethan, who looks after AWS for a few local businesses. "Is my account broken? Every article says it's out." Ethan was already using it through Amazon Bedrock for a client. "Your account's fine. You're just in the wrong room. Click Work, not Chat." This page is the rest of that conversation: where GPT-6.1 Sol really is, what it costs on every route, how it compares with Astra, GPT-6 Sol and Claude, working code for the API and Bedrock, how to point Codex at it, and the mistakes that make it look broken when it is not.

⚡ Quick Answer

• What it is → OpenAI's mid-tier GPT-6 model, an upgrade to GPT-6 Sol released September 29, 2026, with near-Astra results at a fifth of Astra's token price. What it is.

• Where it is in ChatGPT → ChatGPT Work and Codex on Plus, Pro, Business, Enterprise and Edu. Not in Chat, not on Free or Go. Where to find it.

• GPT-6.1 Sol pricing → $2 input, $0.10 cached input, $2.50 cache writes, $10 output per million tokens. Prompts over 272K input tokens cost more for the whole request. Full table.

• API name → gpt-6.1-sol, with reasoning effort from low to max. Use the Responses API for tools. Code.

• Amazon Bedrock → us.openai.gpt-6.1-sol or global.openai.gpt-6.1-sol on bedrock-runtime, or openai.gpt-6.1-sol on Mantle in us-east-1. Each output token counts as ten against your quota. Bedrock details.

Astra still wins the very hardest work, and dots, the new always-on agents, run on Astra, not on Sol.

If the model names are blurring together, that is normal. OpenAI shipped three GPT-6 models in September, then replaced one of them a week later, and every launch came with its own benchmark charts. The short version is that Astra is the expensive flagship, Luna is the cheap fast one, and Sol sits in the middle. GPT-6.1 Sol is the new middle. It took GPT-6 Sol's place exactly seven days after GPT-6 Sol launched.

🧭 NEW HERE? READ THESE FIRST

If any of the AWS or model words on this page feel unfamiliar, these five make the rest easy:

📌 Bookmark this; the Bedrock half of this page leans on all five.

What GPT-6.1 Sol is, in one minute

OpenAI's GPT-6 family has three tiers. GPT-6 Astra is the most capable and the most expensive. GPT-6 Luna is the fastest and cheapest, meant for focused, high-volume jobs like extraction and classification. Sol is the balance between them. OpenAI introduced GPT-6 Sol and Luna on September 22, 2026, and released GPT-6.1 Sol on September 29 at DevDay in San Francisco.

OpenAI describes GPT-6.1 Sol as near-Astra performance for complex coding, computer use and professional work at a lower cost. In practice that means three kinds of jobs: reading and changing real codebases, operating software through a screen the way a person would, and working through long documents and multi-step business tasks. It reads text and images and writes text. It does not take audio or video, and it does not generate images itself.

The specifications that matter day to day:

DetailGPT-6.1 Sol
ReleasedSeptember 29, 2026 (API, ChatGPT Work, Codex and Amazon Bedrock the same day)
API model namegpt-6.1-sol (one snapshot, no dated versions yet)
Context window1,050,000 tokens on the OpenAI API, of which up to 922,000 can be input; 1M on Bedrock
Maximum output128,000 tokens (Bedrock lists 131,072)
Knowledge cutoffApril 30, 2026 (GPT-6 Sol's was April 20)
Input and outputText and images in, text out
Reasoning effortlow, medium (the default), high, xhigh, max. No none or minimal
Standard API price$2 input, $0.10 cached input, $2.50 cache writes, $10 output per million tokens

One name you may see in searches is GPT-6.1 Astra. There is no such model. OpenAI said on September 28, the day before DevDay, that it would not release it, because it fell short on staying within the scope of what it was asked to do and on how it reported back about its own work. GPT-6 Astra remains the flagship.

Where GPT-6.1 Sol is in ChatGPT, and why you cannot find it in Chat

This is where most people get stuck, so it comes before the pricing. ChatGPT now has more than one place to work. Chat is the familiar conversation screen. ChatGPT Work is the agent workspace, where the model can work on files, apps and longer tasks, on the web, in the desktop app and on mobile. Codex is OpenAI's coding agent, in the ChatGPT desktop app, a command-line tool and editor extensions. GPT-6.1 Sol, GPT-6 Sol and GPT-6 Luna are available in Work and Codex. None of them are available in Chat.

So when Jake opened a normal chat and searched the model menu, he was looking in a room the model does not live in. Clicking Work and starting a new task, then opening the model control beneath the box where you type, is what finally showed it.

ChatGPT planPrice per monthGPT-6.1 Sol?
Free$0No, not at launch
Go$8No, not at launch
Plus$20Yes, in Work and Codex
Pro$100, $200 or $500Yes, in Work and Codex
Business$20 per user billed annually, $25 monthly (2+ users)Yes, in Work and Codex
Enterprise and EduContractYes, but off until a workspace admin turns it on

If you are on Enterprise or Edu and still see nothing in Work, that last row is almost always the reason. GPT-6.1 Sol starts switched off in those workspaces, and an administrator has to enable it. Selecting a model in your own settings does not grant access to it, so there is nothing to fix on your side.

Inside Work and the desktop app, OpenAI has simplified the controls into a Power slider. Start at the default for your account, slide toward Smarter for harder reasoning or Faster for quicker, cheaper work, and open Advanced when you want to pick GPT-6.1 Sol by name and set its reasoning effort yourself. The efforts in ChatGPT are named Light, Medium, High and Extra High, with Max and Ultra on top. Max gives one model more time on a single problem. Ultra splits a big task across subagents working in parallel. Most tasks need neither.

Two usage details are worth knowing before you burn through a week's allowance in an afternoon:

  • Sol uses a fifth of Astra's allowance. ChatGPT counts work in credits per million tokens. GPT-6.1 Sol costs 50 credits per million input tokens, 2.5 cached and 250 output. GPT-6 Astra costs 250, 25 and 1,250. If Astra keeps running you out of messages, switching heavy, repetitive work to Sol stretches the same plan about five times further.
  • Fast mode drains included limits 2.5 times faster. Standard and Fast speed are both available for GPT-6.1 Sol. Fast uses your included subscription usage at 2.5 times the Standard rate, and purchased credits at 2 times. Ultrafast, the newest speed tier, with up to eight times faster token generation in Codex, is promised for Sol but had not arrived as of October 7; for now it is Astra only.

One more date belongs here because it affects the same menus. GPT-5.5 retires from ChatGPT, ChatGPT Work and Codex on October 14, 2026, on every plan. It stays on the API. OpenAI's suggested replacement on paid plans is GPT-6 Sol, but if your saved settings, scheduled tasks or Codex configuration still name gpt-5.5, GPT-6.1 Sol is the better swap where your plan offers it. Change them before the 14th so nothing stops working quietly on a weekend.

Ethan's tip for Jake was simpler than any of that. "Put the invoice tool's folder in a Work task, pick GPT-6.1 Sol under Advanced, leave it on Medium, and ask it to find why discount lines go missing. If it fumbles, move up one effort level, not to Astra." It found the bug on Medium: the parser treated any line containing a minus sign as a page footer and skipped it.

GPT-6.1 Sol pricing: the full table, and the one number that changed

GPT-6.1 Sol's standard API price is $2.00 per million input tokens, $0.10 per million cached input tokens, $2.50 per million cache-write tokens and $10.00 per million output tokens, for prompts up to 272,000 input tokens. Put next to the rest of the family, the interesting line is the one that did not change:

Model (OpenAI API, Standard)InputCached inputCache writesOutput
GPT-6.1 Sol$2.00$0.10$2.50$10.00
GPT-6 Sol$2.00$0.20$2.50$10.00
GPT-6 Astra$10.00$1.00$12.50$50.00
GPT-6 Luna$0.10$0.01$0.125$0.50
GPT-5.6 Sol$4.00$0.40$5.00$20.00

Prices per million tokens, prompts of 272K input tokens or fewer.

Input and output prices are identical to GPT-6 Sol. The only change is cached input, which dropped from $0.20 to $0.10, so cached tokens now cost 5% of the normal input price instead of 10%. For anything that resends the same long instructions, documents or code on every call, which describes most agents, the newer model is both better and slightly cheaper. Against Astra, every column is one-fifth, which is where OpenAI's "one-fifth of the cost" line comes from.

Four rules change the bill more than the table suggests:

  1. Long prompts reprice the whole request. When a prompt goes over 272,000 input tokens, input and cache prices double and output costs 1.5 times as much, for the entire request, not just the tokens past the line. That makes the long-context price $4.00 input, $0.20 cached, $5.00 cache writes and $15.00 output. A 300,000-token prompt costs $1.20 in input, where a 270,000-token one costs $0.54. If you can trim a prompt under 272K, it is often the single biggest saving available.
  2. Reasoning is billed as output. GPT-6.1 Sol always reasons; there is no none setting. The thinking it does before answering is counted as output tokens at $10 per million, even though you do not see it. Higher effort means more of those tokens, so the cheapest reliable effort for a job is worth finding. Medium is the default.
  3. Batch and Flex are half price on the OpenAI API. Batch and Flex processing cost 50% of Standard: $1.00 input, $0.05 cached, $1.25 cache writes and $5.00 output. Fast mode costs twice Standard. Neither discount exists on Amazon Bedrock, where GPT-6.1 Sol supports the Standard tier only.
  4. Regional processing costs 10% more. On the OpenAI API, data residency endpoints add a 10% uplift for models released since March 5, 2026, which includes GPT-6.1 Sol. Bedrock has the same 10% pattern for keeping traffic in the US, covered in the Bedrock section below.

The Bedrock pricing guide walks through how tokens, cached tokens and tiers add up across every model on AWS, if you want the general picture before the worked example below.

What a month of GPT-6.1 Sol really costs: a worked example

Prices per million tokens are hard to feel, so here is one realistic job worked through. Ethan's client runs a review bot over a small codebase. Every request sends the same 20,000-token prefix (instructions plus the files that matter), 4,000 tokens of new material, and gets about 2,000 tokens back, reasoning included. It runs 100 times a working day, about 3,000 requests a month. Assume the cached prefix is rewritten four times a day as it expires, 120 writes a month.

RoutePer request3,000 requests120 cache writesMonth
GPT-6.1 Sol, OpenAI API$0.030$90.00$6.00$96.00
GPT-6.1 Sol, Bedrock global routing$0.030$90.00$6.00$96.00
GPT-6.1 Sol, Bedrock US routing$0.033$99.00$6.60$105.60
GPT-6 Sol, OpenAI API$0.032$96.00$6.00$102.00
GPT-6 Astra, OpenAI API$0.160$480.00$30.00$510.00
GPT-6.1 Sol, OpenAI Batch$0.015$45.00$3.00$48.00

How the GPT-6.1 Sol line adds up: the cached 20,000 tokens cost $0.002, the 4,000 new tokens $0.008, and the 2,000 output tokens $0.020, which is $0.030 a request. Each cache write is 20,000 tokens at $2.50 per million, or $0.05.

Three lessons fall out of that table. Output is two-thirds of the bill, so the reasoning effort you choose matters more than the input price. Astra costs five times as much for the same traffic, so it should earn its place on the jobs Sol fails, not run everything by default. And if the work does not need an answer within seconds, OpenAI's Batch processing halves the whole thing. A nightly review of the day's commits is a perfect Batch job.

Whatever route you choose, set a budget alarm before the first real day of traffic. Our guide to AWS billing alerts that actually fire covers the Bedrock side, and OpenAI's dashboard has its own monthly limits.

GPT-6.1 Sol vs GPT-6 Astra: is Sol good enough?

For most work, yes, and the numbers OpenAI published at launch make the case. These are OpenAI's own evaluations, so treat them as the vendor's best foot forward and test on your own tasks before moving production traffic:

  • DeepSWE v1.1 (real software-engineering tasks in real codebases): GPT-6.1 Sol matches GPT-6 Astra at roughly one-fifth of the cost, and beats GPT-6 Sol's best score by 6.4 percentage points while using a lower reasoning effort.
  • OSWorld 2.0, offline set (demanding computer-use workflows): 7 points better than GPT-6 Sol at maximum effort, at less than half the cost, and within 2.1 points of Astra at roughly one-seventh of the cost per task.
  • GDP.pdf (professional questions answered from complex PDFs, tables and fine print): approaches Astra at roughly one-fifth of the cost per task.
  • Factual errors on hard prompts: at low effort, the share of answers with a factual error fell from 11.4% with GPT-6 Sol to 7.7%, a cut of about a third, and across settings it stays within 1.9 points of Astra.

Where Astra still clearly wins is the hardest scientific work. On Terminal-Bench Science 0.1, which covers data analysis, simulation and theorem proving, Astra posted the top score among the models OpenAI tested, 68.1%, and OpenAI itself says Astra should be used for the most difficult research tasks. GPT-6.1 Sol more than doubled GPT-6 Sol's score there, but it is not Astra.

Astra also keeps a few things Sol does not have yet: the Ultrafast speed tier, and the job of powering dots, OpenAI's new always-on agents. If you need either, the choice is made for you.

The sensible pattern is the one Ethan uses with clients. Default to GPT-6.1 Sol at medium effort. When a task fails, raise the effort one step and try again. Only when Sol fails at high effort does the request go to Astra. Most requests should never need to leave Sol, which is the whole point of a model that costs a fifth as much.

GPT-6.1 Sol vs GPT-6 Sol: what changed, and what breaks when you switch

On paper GPT-6.1 Sol is a drop-in replacement: the same price, the same context window, the same output limit. In practice a few differences can break working code or a working Bedrock setup, so read this list before you change one string and deploy.

WhatGPT-6 SolGPT-6.1 Sol
Reasoning effortnone through maxlow through max; none and minimal are not supported
Tools on Chat CompletionsFunction calling only with effort noneNo tool calling on Chat Completions; use the Responses API
Cached input price$0.20 per million$0.10 per million
Knowledge cutoffApril 20, 2026April 30, 2026
Server-side tools on Bedrock MantleSupportedNot supported
Bedrock end-of-life promiseNo sooner than September 22, 2027Not announced; at least 6 months' notice under OpenAI's own terms

The first two rows catch the most people. If your GPT-6 Sol code sets reasoning.effort to none to get fast, cheap answers, GPT-6.1 Sol will not accept it. Use low, or keep those calls on GPT-6 Sol or Luna, both of which still support none. And if you call tools through Chat Completions, the move to GPT-6.1 Sol is also a move to the Responses API.

The Bedrock rows matter for teams on AWS. An app that runs GPT-6 Sol on Mantle and leans on server-side tools, such as hosted web search, loses them on GPT-6.1 Sol today. And the lifecycle row is a quiet one: GPT-6 Sol on Bedrock carries AWS's usual promise of at least a year before end of life, while GPT-6.1 Sol follows OpenAI's first-party terms of at least six months' notice. For most teams that changes nothing. For anyone signing a long contract on top of it, it is worth one line in the risk register.

The Bedrock side of GPT-6 Sol, with every ID and price, is in our GPT-6 on Amazon Bedrock guide, which now links here for 6.1.

GPT-6.1 Sol vs Claude Opus 5.5 and Sonnet 5.5

The comparison most people searching for this model actually want is against Anthropic's two current models. On price, GPT-6.1 Sol lands in a very specific spot:

Model (first-party API)InputCache readOutputContext
GPT-6.1 Sol$2.00$0.10$10.001.05M
Claude Sonnet 5.5$2.00$0.20$10.001M
Claude Opus 5.5$4.00$0.20$20.001M
GPT-6 Astra$10.00$1.00$50.001.05M

Prices per million tokens. Each company counts tokens with its own tokenizer, so the same text is not always the same number of tokens on both.

Against Claude Sonnet 5.5, the list prices are identical: $2 in and $10 out. GPT-6.1 Sol's cached input is half of Sonnet's cache-read price, so a heavily cached agent leans slightly toward Sol. That is close enough that the real decision is which model does your task better and which ecosystem you already live in.

Against Claude Opus 5.5, GPT-6.1 Sol is half the price per token, and OpenAI compared the two directly. On AutomationBench, multi-step business workflows, OpenAI reports GPT-6.1 Sol 2.2 points above Opus 5.5 at medium effort, at roughly a third of the cost. On GDP.pdf, it reports Sol scoring higher than Opus 5.5 at less than half the cost per task. On the Terminal-Bench Science tasks, the average cost per task at maximum effort was $5.47 for GPT-6.1 Sol, $23.21 for Opus 5.5 and $23.80 for Astra. Those are OpenAI's tests and OpenAI's chosen settings. Anthropic publishes its own numbers for Opus 5.5 on tasks it chose, and our Claude Opus 5.5 guide covers them.

The honest summary: every lab picks benchmarks it wins. Run twenty of your own real tasks through two models at the effort you would actually use, and count passes and dollars. That afternoon of testing tells you more than any launch chart, and both companies sell their models on Amazon Bedrock, so one AWS account can run the comparison.

How to call GPT-6.1 Sol through the OpenAI API

On OpenAI's own API, the model name is gpt-6.1-sol, and the Responses API is the main way in. Install or update the OpenAI SDK, set your key, and send a request:

pip install --upgrade openai
export OPENAI_API_KEY="sk-..."
from openai import OpenAI

client = OpenAI()

response = client.responses.create(
    model="gpt-6.1-sol",
    reasoning={"effort": "medium"},
    input="This Python function skips invoice lines that contain a discount. "
          "Find the bug and show the fix:\n\n" + open("parser.py").read(),
)
print(response.output_text)

A few rules make the difference between a clean first call and an afternoon of confusing errors:

  • Reasoning effort is low, medium, high, xhigh or max. Medium is the default if you leave it out. none and minimal are rejected.
  • Tools need the Responses API. Chat Completions works with GPT-6.1 Sol, but only for plain requests without tool calling. Function calls, web search, file search, code interpreter, computer use, MCP and the rest go through Responses.
  • Batch works. The Batch endpoint supports GPT-6.1 Sol, at half the Standard price, for anything that can wait.
  • Not supported: Realtime, audio, embeddings, image generation endpoints, fine-tuning and predicted outputs. Image generation is still available as a tool inside a Responses request.
  • Multi-agent is in beta. GPT-6.1 Sol can delegate parts of a request to subagents inside a single Responses call.
  • Data residency. GPT-6.1 Sol supports US and EU data residency. Fast mode is not available with EU residency.

Rate limits depend on your account tier, which OpenAI simplified on October 6 from five tiers to three. For GPT-6.1 Sol on Standard processing, Build allows 5,000 requests and 1,000,000 tokens per minute, Launch 10,000 requests and 4,000,000 tokens, and Grow 15,000 requests and 40,000,000 tokens. Accounts move up automatically as total credit purchases reach each tier's minimum.

If you already call GPT-6 Sol, the switch is the model name plus the two checks from the section above: no none effort, and no tools on Chat Completions. Run your test suite against the new name before you flip production, because the newer model may phrase answers differently even where it is more accurate.

GPT-6.1 Sol on Amazon Bedrock: model IDs, Regions and the 10x quota rule

GPT-6.1 Sol became generally available on Amazon Bedrock on September 29, 2026, the same day as the OpenAI launch. That matters for teams whose data has to stay inside their AWS account: Bedrock runs the model on AWS infrastructure, prompts and responses are not used for training, and you do not opt into sharing data with OpenAI. Access is governed by IAM, calls are recorded in CloudTrail, and PrivateLink can keep traffic inside your network. If Bedrock itself is new to you, start there; the rules below are Bedrock's, not OpenAI's.

There are two front doors to Bedrock, and GPT-6.1 Sol uses each one differently:

EndpointWhat you send as the modelWhere it runsAPIs
bedrock-runtime, US routingus.openai.gpt-6.1-solUS Regions (callers in the US or Canada)Responses, Chat Completions, Converse, Invoke
bedrock-runtime, global routingglobal.openai.gpt-6.1-solSupported commercial Regions worldwideResponses, Chat Completions, Converse, Invoke
bedrock-mantle, in-Regionopenai.gpt-6.1-solUS East (N. Virginia), us-east-1, onlyResponses, Chat Completions

The rule people miss: on bedrock-runtime there is no in-Region option. Sending the bare openai.gpt-6.1-sol to bedrock-runtime does not work; you must send the us. or global. profile. The bare ID is for Mantle, and Mantle serves this model only in us-east-1. Neither endpoint supports the Anthropic-style Messages API for this model. The full difference between the two endpoints is in our Mantle vs bedrock-runtime guide.

For US routing, the calling Region decides where requests can land:

You call fromRequests can run in
us-east-1, us-east-2 or us-west-2us-east-1, us-east-2, us-west-2
us-west-1 (N. California)us-east-1, us-east-2, us-west-1, us-west-2
ca-central-1 (Canada)ca-central-1, us-east-1, us-east-2, us-west-2
ca-west-1 (Calgary)ca-west-1, us-east-1, us-east-2, us-west-2

Note what is missing: there is no EU, UK or Asia Pacific profile. A team in Europe or Asia can use GPT-6.1 Sol on Bedrock only through global routing, which means requests can be processed outside their home region. If you have promised a customer that their data stays in the EU, GPT-6.1 Sol on Bedrock cannot keep that promise today.

GPT-6.1 Sol Bedrock pricing matches OpenAI's own price for global routing and adds 10% for US routing or Mantle:

Bedrock optionInputCache writeCache readOutput
Global routing (bedrock-runtime)$2.00$2.50$0.10$10.00
US routing (bedrock-runtime)$2.20$2.75$0.11$11.00
Mantle in us-east-1$2.20$2.75$0.11$11.00
Global, over 272K input tokens$4.00$5.00$0.20$15.00
US or Mantle, over 272K input tokens$4.40$5.50$0.22$16.50

Prices per million tokens, Standard tier. The long-context rate applies to the whole request once input passes 272,000 tokens.

Priority, Flex and Reserved tiers are not offered for GPT-6.1 Sol on Bedrock. If you were counting on Flex's discount for background work, the OpenAI API's Batch or Flex processing is the only half-price route for this model right now.

And now the rule that surprises everyone the first time they hit a throttling error: the output-token burndown rate for GPT-6.1 Sol on Bedrock is 10. Every output token consumes ten tokens of your tokens-per-minute quota. A request that returns 2,000 tokens, reasoning included, uses 20,000 tokens of quota for its output alone. You are still billed for 2,000 output tokens; it is only the quota counter that runs ten times faster. If your quota looked generous on paper and you are being throttled at a fraction of it, this is why. Plan quota requests around output, and check your current limits in the Service Quotas console before launch, because quotas vary by account and Region.

One quota rule works in your favor. Cached input tokens read from the prompt cache do not count against the input-tokens-per-minute quota, so caching the stable part of your prompt eases throttling as well as the bill.

Calling GPT-6.1 Sol on Bedrock: working code

Bedrock speaks the OpenAI API's language for this model, so the OpenAI SDK works unchanged once it points at AWS. Create a Bedrock API key in the Bedrock console (short-term keys are safer for anything beyond a test), then set two environment variables. For Mantle:

export OPENAI_API_KEY="<your Bedrock API key>"
export OPENAI_BASE_URL="https://bedrock-mantle.us-east-1.api.aws/openai/v1"

For bedrock-runtime, only the base URL changes:

export OPENAI_API_KEY="<your Bedrock API key>"
export OPENAI_BASE_URL="https://bedrock-runtime.us-east-1.amazonaws.com/openai/v1"

Note the path: /openai/v1, not /v1. Then the same Python code calls either endpoint; only the model string differs:

from openai import OpenAI

client = OpenAI()   # reads OPENAI_API_KEY and OPENAI_BASE_URL

response = client.responses.create(
    model="us.openai.gpt-6.1-sol",      # bedrock-runtime; use "openai.gpt-6.1-sol" on Mantle
    reasoning={"effort": "medium"},
    input="Summarize the risks in this supplier contract in five bullet points:\n\n"
          + open("contract.txt").read(),
    max_output_tokens=4000,
)
print(response.output_text)

If your code already uses the Converse API with boto3, GPT-6.1 Sol works there too, on bedrock-runtime only, with your normal AWS credentials instead of an API key:

import boto3

brt = boto3.client("bedrock-runtime", region_name="us-east-1")

resp = brt.converse(
    modelId="us.openai.gpt-6.1-sol",
    messages=[{"role": "user", "content": [{"text": "Explain prompt caching in two sentences."}]}],
    inferenceConfig={"maxTokens": 2000},
)
blocks = resp["output"]["message"]["content"]
print(next(b["text"] for b in blocks if "text" in b))

The last line looks for the first text block rather than assuming it comes first, because a reasoning model's reply can include other block types ahead of the text.

Explicit prompt caching is where Bedrock users save the most. Mark the end of the stable part of your prompt with a breakpoint, and turn on explicit mode with a 30-minute lifetime, which is the only lifetime offered. The cached prefix must be at least 1,024 tokens, and one request can create up to four cache writes. With the OpenAI SDK, the Bedrock-specific option travels in extra_body:

stable = open("review_instructions_and_files.txt").read()   # 20,000 tokens that never change

response = client.responses.create(
    model="us.openai.gpt-6.1-sol",
    input=[{
        "role": "user",
        "content": [
            {"type": "input_text", "text": stable,
             "prompt_cache_breakpoint": {"mode": "explicit"}},
            {"type": "input_text", "text": "Review today's change:\n" + diff_text},
        ],
    }],
    extra_body={"prompt_cache_options": {"mode": "explicit", "ttl": "30m"}},
)
print(response.usage)   # look for cached tokens on the second call

Send the same request twice within half an hour and the usage details on the second call should show cached tokens. If they stay at zero, something in the stable part is changing between calls, often a timestamp or a reordered list. On Converse, only implicit caching works for this model; the cachePoint field is not supported for explicit caching.

IAM for bedrock-runtime. The calling identity needs permission on the inference profile and the model, plus bedrock:CallWithBearerToken if you use an API key. AWS's GPT-6 Sol card also calls out the account's default project, and the GPT-6.1 Sol card lists the same default-project support, so include it rather than chase an access-denied error later:

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "Gpt61SolRuntime",
      "Effect": "Allow",
      "Action": [
        "bedrock:InvokeModel",
        "bedrock:InvokeModelWithResponseStream"
      ],
      "Resource": [
        "arn:aws:bedrock:*:111122223333:inference-profile/us.openai.gpt-6.1-sol",
        "arn:aws:bedrock:*:111122223333:inference-profile/global.openai.gpt-6.1-sol",
        "arn:aws:bedrock:*::foundation-model/openai.gpt-6.1-sol",
        "arn:aws:bedrock:::foundation-model/openai.gpt-6.1-sol",
        "arn:aws:bedrock:*:111122223333:project/default"
      ]
    },
    {
      "Sid": "ApiKeyAccess",
      "Effect": "Allow",
      "Action": "bedrock:CallWithBearerToken",
      "Resource": "*"
    }
  ]
}

Replace 111122223333 with your account ID. The model ARN appears with and without a Region so either routing scope matches. Mantle uses its own IAM actions, bedrock-mantle:CreateInference and, with API keys, bedrock-mantle:CallWithBearerToken; a policy that allows bedrock:* does not cover them. If IAM policy JSON still reads like a foreign language, that guide decodes it line by line.

What works with GPT-6.1 Sol on Bedrock, and what does not

Bedrock wraps the model in its own features, and not all of them apply. The two endpoints differ more than most people expect:

Featurebedrock-runtimebedrock-mantle
StreamingYesYes
Structured outputs (JSON schema)Yes, all four APIsYes, Responses and Chat Completions
Implicit prompt cachingYes, all four APIsYes
Explicit prompt cachingResponses, Chat Completions, Invoke; not ConverseYes
GuardrailsChat Completions, Invoke and Converse; not ResponsesNo
Model invocation loggingYesNo
Application inference profiles (cost tags)Invoke and Converse onlyNo
ProjectsDefault project onlyYes
Server-side tools, Knowledge Bases, prompt routing, token counting, prompt optimizationNoNo

The row that causes real trouble is Guardrails. If your company requires a Bedrock guardrail on every model call, GPT-6.1 Sol's Responses API cannot have one, and Mantle cannot have one at all. Use Chat Completions, Invoke or Converse on bedrock-runtime, where guardrails work. Logging is the other compliance row: model invocation logs exist only on bedrock-runtime, so audit teams usually prefer it.

For structured JSON output, use an object schema, list every field in required, and set additionalProperties to false. On Converse, the schema goes in outputConfig.textFormat, and you must also set additionalModelRequestFields.text.format.strict to true, or the schema is not enforced the way you expect. Always check for a refusal or an incomplete response before parsing.

Because Bedrock cannot count tokens for this model, estimate before you send very long prompts. A long document that lands just over 272,000 input tokens costs roughly double what it would just under, so leave yourself a margin rather than aiming at the line.

Using GPT-6.1 Sol in Codex, including Codex on Amazon Bedrock

Codex is where OpenAI expects most GPT-6.1 Sol work to happen, and OpenAI recommends it for complex coding and agent work when it is available to your account. In the Codex command-line tool, pick it for one run with a flag:

codex --model gpt-6.1-sol
codex exec -m gpt-6.1-sol "Review the current changes"

To make it the default for the desktop app, the CLI and the editor extensions at once, add one line to ~/.codex/config.toml, which all three read:

model = "gpt-6.1-sol"

Inside an interactive session, /model switches models or effort, and "More reasoning…" shows Max and Ultra where your model supports them.

Codex on Amazon Bedrock is the part most guides skip. Codex can send its model requests to Bedrock instead of OpenAI, authenticated with AWS rather than a ChatGPT sign-in or an OpenAI key. You choose the endpoint with a provider name in the same config.toml. For bedrock-runtime, which is the recommended choice for new setups:

model_provider = "amazon-bedrock-runtime"
web_search = "disabled"
model = "us.openai.gpt-6.1-sol"

Hosted web search is not available on bedrock-runtime, which is why it is switched off. The amazon-bedrock provider selects Mantle instead. Then give Codex AWS credentials, either a Bedrock API key with a Region:

export AWS_BEARER_TOKEN_BEDROCK="<your Bedrock API key>"
export AWS_REGION=us-east-1

or any normal AWS SDK credential source: aws configure, environment variables, aws login, or aws sso login --profile codex-bedrock with AWS_PROFILE set. The API key is checked first, so an old key left in your environment silently wins over your SSO profile. The desktop app and editor extensions often do not see your shell's variables; put them in ~/.codex/.env and restart.

To check the route, open /status in the CLI and confirm the provider, then send Reply with exactly: bedrock-ok in a new task. The reply alone does not prove which route answered, so confirm the provider and model in the client as well.

Two honest caveats. First, OpenAI's Codex page for Bedrock still listed only GPT-6 Astra, GPT-6 Sol and GPT-6 Luna as supported models on October 7, while AWS's launch post says Codex can use GPT-6.1 Sol on Bedrock. If your Codex version's model picker does not offer 6.1 Sol on the Bedrock provider, set the model line explicitly as above; if your version rejects it, use us.openai.gpt-6-sol on Bedrock for now, or GPT-6.1 Sol through your ChatGPT plan. Second, several Codex features need OpenAI's cloud and do not work on the Bedrock route: Codex cloud tasks, ChatGPT Work on the web, Fast mode, image generation, GitHub delegation and pull-request reviews, and Slack and Linear integrations. Local work, the CLI, the desktop app, editor extensions, MCP, skills, subagents and local code review all work.

If you would rather use an open-source agent, OpenCode also talks to Bedrock; our OpenCode guide has that setup, and our Kimi Code guide covers the main alternative from Moonshot.

GPT-6.1 Sol not working? The mistakes behind most failures

Almost every "GPT-6.1 Sol is broken" moment traces back to one of these. Work down the list in order; the early ones are the common ones.

  1. Looking for it in ChatGPT Chat. It is in Work and Codex only, on paid plans. On Enterprise and Edu, an admin must turn it on.
  2. Using a Free or Go plan. Neither includes GPT-6.1 Sol at launch. The API and Bedrock are separate, pay-per-token routes.
  3. Sending none or minimal reasoning effort. GPT-6.1 Sol accepts low through max only. Code copied from GPT-6 Sol or Luna often sets none.
  4. Calling tools through Chat Completions. Chat Completions works only without tools for this model. Move tool calls to the Responses API.
  5. Sending the bare ID to bedrock-runtime. On bedrock-runtime, send us.openai.gpt-6.1-sol or global.openai.gpt-6.1-sol. The bare openai.gpt-6.1-sol is for Mantle.
  6. Calling Mantle outside us-east-1, or with the wrong path. Mantle serves this model in us-east-1 only, at /openai/v1, not /v1.
  7. Calling the US profile from outside the US and Canada. The us. profile accepts calls from the six US and Canadian source Regions listed above. From Europe or Asia, use global. or call from a supported Region.
  8. An IAM policy that names only the model. Include the inference profiles, the model ARN, the default project and, for API keys, bedrock:CallWithBearerToken. For Mantle, the bedrock-mantle: actions. The Bedrock AccessDeniedException guide walks through the same checks for Claude, and they apply here.
  9. Throttling far below your quota. Remember the burndown rate of 10: output tokens count ten times against the per-minute quota. Request more quota, lower the effort, or cap max_output_tokens.
  10. Asking Bedrock for Priority or Flex. Only the Standard tier exists for this model on Bedrock. Leave service_tier out, or set it to default.
  11. A guardrail on the Responses API. Guardrails do not apply to Responses on bedrock-runtime, or to anything on Mantle. Use Chat Completions, Invoke or Converse if a guardrail is required.
  12. Explicit caching on Converse. The cachePoint field is not supported for this model. Use Responses, Chat Completions or Invoke for explicit caching.
  13. Codex configured for Chat Completions. Current Codex releases do not support wire_api = "chat" or a Chat Completions-only endpoint. Point custom providers at a Responses-compatible endpoint.

If none of those fit, test the smallest possible request in the Bedrock playground or the OpenAI dashboard. If the playground answers and your code does not, the problem is in your code or credentials, not in the model or your account's access.

What about dots? The DevDay agent that runs on Astra, not Sol

The other headline from DevDay was dots, and plenty of people searching for GPT-6.1 Sol are really trying to understand those. A dot is an always-on agent inside ChatGPT that takes ongoing responsibility for work and keeps making progress between conversations. It has its own cloud computer and browser, connects to more than 4,000 apps through OpenAI's plugins, and can be reached in ChatGPT on desktop, web and mobile, or in Slack and Teams.

The detail that answers the most common confusion: dots run on GPT-6 Astra, not on GPT-6.1 Sol. They are a separate product, not a feature of the new model.

The rest of what you need to know, briefly:

  • Who gets them: ChatGPT Pro and Business Premium, in eligible markets. Enterprise, Edu and Healthcare workspaces can try a beta once an admin enables it; it is off by default. Plus is not included at launch.
  • Cost: your first dot is included in Pro or Business Premium at no extra charge. Talking to your dot does not count toward ChatGPT usage limits, but tasks it starts in Codex or ChatGPT Work count as usual.
  • Control: you choose which apps it can reach, and Custom Rules let you allow an action, require approval for it, or block it. Some sensitive tasks, such as changing a password, always stay with you. When you are not working with it, its background research uses read-only tools that cannot send messages, change app content or control your computer. Your own computer stays separate unless you choose to connect it.
  • Getting started: create your first dot in the ChatGPT desktop app or a desktop browser, connect your apps, and it introduces itself. After setup, you can message it from the mobile app.

Dots are an interesting product, but they are a ChatGPT subscription feature, not something you call from an API or run on AWS. For building your own agents on Sol, the Responses API and Codex sections above are the place to start.

Can you run GPT-6.1 Sol locally?

No. GPT-6.1 Sol is a closed model. OpenAI has not released its weights, so there is nothing to download into Ollama, LM Studio or llama.cpp, and anything claiming to be a downloadable GPT-6.1 Sol is not the real model. The closest you can get to "your own copy" is Amazon Bedrock, where the model runs inside AWS under your account's IAM, logging and network rules, with prompts kept out of training.

If what you really want is a strong model on your own hardware, look at the open-weight families instead. Our guides to running GLM-5.3 locally and running Qwen3.8 on Windows and Kali have the real memory numbers, and the Kimi K3 guide explains why some big open models are still impractical at home.

Which GPT-6.1 Sol route should you pick?

Every route runs the same model. What changes is who bills you, where your data goes, and which features you get. Pick by your situation:

  • You want to use it yourself, today: ChatGPT Plus at $20 a month, in Work or Codex. No code, no AWS. Our Is ChatGPT Plus worth it? guide helps if you are on Free or Go and wondering whether to upgrade.
  • You are building an app and want the lowest price: the OpenAI API, with Batch or Flex at half price for anything that can wait, and explicit caching for your stable prompt text.
  • Your company's data has to stay in its AWS account: Amazon Bedrock, bedrock-runtime with the us. profile, and Chat Completions or Converse if you need guardrails. Budget 10% more than OpenAI's list price for US routing, and size your quota for the burndown rate.
  • You want the OpenAI-native features on AWS: Mantle in us-east-1, with Projects and explicit caching on the Responses API, accepting no guardrails and no invocation logs.
  • Your customers need EU or Asia Pacific processing: not GPT-6.1 Sol on Bedrock today. The OpenAI API's EU data residency is the option for EU traffic.
  • Your developers live in Codex: sign in with ChatGPT for the full feature set, or point Codex at Bedrock if the bill and the audit trail have to sit in AWS.

Ethan ended up running both. Jake uses GPT-6.1 Sol through Plus for the shop's odd jobs. Ethan's clients reach it through Bedrock on bedrock-runtime, with a guardrail on Chat Completions and a quota request sized for ten tokens per output token.

Frequently asked questions

What is GPT-6.1 Sol?

GPT-6.1 Sol is OpenAI's mid-tier GPT-6 model, an upgrade to GPT-6 Sol released on September 29, 2026. It approaches GPT-6 Astra on coding, computer use and document work at one-fifth of Astra's token price, and is available in ChatGPT Work, Codex, the OpenAI API and Amazon Bedrock.

When was GPT-6.1 Sol released?

September 29, 2026, at OpenAI DevDay. It reached ChatGPT Work, Codex, the OpenAI API and Amazon Bedrock the same day, one week after GPT-6 Sol and GPT-6 Luna launched on September 22.

How much does GPT-6.1 Sol cost?

On the OpenAI API, $2 per million input tokens, $0.10 per million cached input tokens, $2.50 per million cache-write tokens and $10 per million output tokens, for prompts up to 272K input tokens. Larger prompts cost 2x for input and 1.5x for output on the whole request. Batch and Flex are half price.

Is GPT-6.1 Sol better than GPT-6 Astra?

Not overall, but it is close for most work. OpenAI reports it matching Astra on the DeepSWE v1.1 coding benchmark at about one-fifth of the cost, and coming within 2.1 points on OSWorld 2.0. Astra still leads on the hardest scientific tasks and is the model behind dots and Ultrafast mode.

What is the difference between GPT-6.1 Sol and GPT-6 Sol?

GPT-6.1 Sol scores higher on coding, computer use and factual accuracy at the same input and output price, with cached input halved to $0.10. It drops the none reasoning effort, has no tool calling on Chat Completions, and on Bedrock Mantle it does not support server-side tools.

Why can't I find GPT-6.1 Sol in ChatGPT?

Because it is not in Chat. GPT-6.1 Sol is available in ChatGPT Work and Codex on Plus, Pro, Business, Enterprise and Edu. Open Work, start a task, and choose it under Advanced. On Enterprise and Edu it stays off until a workspace admin enables it.

Is GPT-6.1 Sol free?

No. It is not included in ChatGPT Free or Go at launch. The cheapest way to use it is ChatGPT Plus at $20 a month in Work or Codex, or pay-per-token through the OpenAI API or Amazon Bedrock.

What is the GPT-6.1 Sol context window?

1,050,000 tokens on the OpenAI API, of which up to 922,000 can be input, with up to 128,000 output tokens. Amazon Bedrock lists a 1M context and 131,072 output tokens. Prompts over 272K input tokens are billed at the higher long-context rate.

What is the GPT-6.1 Sol API model name?

gpt-6.1-sol on the OpenAI API. Use the Responses API for tools; Chat Completions works only without tool calling. Reasoning effort accepts low, medium (the default), high, xhigh and max.

Is GPT-6.1 Sol available on Amazon Bedrock?

Yes, generally available since September 29, 2026. Use us.openai.gpt-6.1-sol or global.openai.gpt-6.1-sol on bedrock-runtime, or openai.gpt-6.1-sol on bedrock-mantle in us-east-1. Direct in-Region calls on bedrock-runtime are not supported.

How much does GPT-6.1 Sol cost on Bedrock?

With global routing, the same as OpenAI: $2.00 input, $0.10 cache read, $2.50 cache write and $10.00 output per million tokens. US routing and Mantle add 10%, at $2.20, $0.11, $2.75 and $11.00. Only the Standard tier is offered; there is no Priority or Flex.

Why is GPT-6.1 Sol on Bedrock throttling me so early?

Its output-token burndown rate is 10, so every output token uses ten tokens of your per-minute quota, although you are billed for one. Plan quota around output, lower the reasoning effort, cap the output length, or request a higher quota in Service Quotas.

Can I use GPT-6.1 Sol in Codex with Amazon Bedrock?

Set model_provider to amazon-bedrock-runtime and model to us.openai.gpt-6.1-sol in ~/.codex/config.toml, with AWS credentials. OpenAI's Codex Bedrock page still listed only GPT-6 models on October 7, so if your version rejects 6.1 Sol, use us.openai.gpt-6-sol for now.

Is GPT-6.1 Sol better than Claude Opus 5.5?

On OpenAI's own tests, it beats Opus 5.5 on AutomationBench and GDP.pdf at a third to half of the cost per task. It costs half as much per token: $2 and $10 against $4 and $20. Benchmarks chosen by one lab favor that lab, so compare both on your own tasks.

GPT-6.1 Sol vs Claude Sonnet 5.5: which is cheaper?

Their list prices are the same, $2 per million input and $10 per million output tokens. GPT-6.1 Sol's cached input is $0.10 against Sonnet 5.5's $0.20 cache read, so heavily cached agents cost slightly less on Sol. The two use different tokenizers, so the same text can count differently.

Does GPT-6.1 Sol have Ultrafast mode?

Not yet. OpenAI promised GPT-6.1 Sol Ultrafast at launch, but as of October 7 it was available only for GPT-6 Astra. Standard and Fast speeds are available for Sol now; Fast costs twice the Standard price on the API.

Can I run GPT-6.1 Sol locally?

No. It is a closed model and OpenAI has not released its weights, so it cannot run in Ollama, LM Studio or llama.cpp. Amazon Bedrock is the closest option to running it inside your own environment.

Are dots powered by GPT-6.1 Sol?

No. Dots, the always-on ChatGPT agents announced at DevDay, run on GPT-6 Astra. They are available on ChatGPT Pro and Business Premium in eligible markets, with a beta for Enterprise, Edu and Healthcare once an admin enables it.

Is there a GPT-6.1 Astra?

No. OpenAI said on September 28, 2026 that it would not release GPT-6.1 Astra because it did not meet the bar on staying within scope and on reporting back about its work. GPT-6 Astra remains OpenAI's flagship model.

πŸ“š ALSO READ

Where to go next, from the Claude models on the other side of this comparison to the plan that unlocks Sol in ChatGPT:

📌 Bookmark this if you are choosing between Sol, Astra and Claude this month.

A model launch always feels like a race you are already losing, with new names every week and charts that all point up. You are not behind. GPT-6.1 Sol is a cheaper way to get very good work done, sitting in two rooms of ChatGPT, one API and one corner of AWS, and you now know which door leads where. Jake's invoice tool counts discount lines again, and his freelancer was a little embarrassed about the minus sign. Ethan's clients got the same model with their own AWS bill and audit trail. Start with one real task, at medium effort, and let the results tell you whether you ever need Astra.

📌 If you keep one line from this page

GPT-6.1 Sol is near-Astra work at a fifth of the price, but it lives in Work, Codex, the API and Bedrock, never in a plain chat.

On Bedrock, size your quota for ten tokens per output token and keep prompts under 272K input tokens.

Revision note. Written October 7, 2026, a week after GPT-6.1 Sol reached ChatGPT, the API and Amazon Bedrock. If you went looking for it in a normal chat and came back empty-handed, nothing is wrong with your account; it simply lives in Work and Codex for now.

Related