GPT-6 on Amazon Bedrock: Astra, Sol, Luna IDs, Prices, Code

Logeshwaran
—

OpenAI's GPT-6 models are on Amazon Bedrock: GPT-6 Astra since September 8, 2026, and GPT-6 Sol and GPT-6 Luna since September 22. All three take text and images, return text, and have a 1,050,000-token context window with up to 128,000 output tokens. You call them with the ordinary OpenAI SDK pointed at a Bedrock URL, or through Bedrock's own Converse API, and you pay AWS, not OpenAI. The model IDs are global.openai.gpt-6-sol or us.openai.gpt-6-sol (and the same pattern for Astra and Luna), and the first thing that surprises everyone is this: the plain model ID does not work on the main Bedrock endpoint. You must call a cross-Region profile. The second surprise costs money: keeping your requests inside the United States costs 10% more than letting Bedrock route them anywhere in the world. This guide covers which GPT-6 model to pick, every model ID and endpoint, the full price table, where each model is available, working code, and the errors you will hit first.

Ethan's shop had been paying OpenAI directly for a year, on a company card, with an API key that lived in three developers' laptops. When a client with an AWS contract asked whether the same features could run "inside AWS, on our bill, with our IAM," Jake assumed it would be a week of rewriting. It was an afternoon: the OpenAI SDK he already used, a different base URL, a Bedrock API key, and a model name with a global. in front of it. What took the rest of the week was a pricing question Ethan asked on day two: "Why is the US option more expensive than the whole-world option?" The answer is below, and it matters more than the code.

⚡ Quick Answer

• Which model → Astra for the hardest work, Sol for everyday coding and multistep tasks, Luna for high-volume summarizing, extraction and classification. Compare.

• Model IDs → on bedrock-runtime use global.openai.gpt-6-sol or us.openai.gpt-6-sol; the bare openai.gpt-6-sol only works on the separate bedrock-mantle endpoint in one Region. All IDs.

• Price (per million tokens, global) → Astra $10 in / $50 out, Sol $2 / $10, Luna $0.10 / $0.50. US-only routing and the mantle endpoint add 10%. Full table.

• First error you will see → AccessDenied because your IAM role can invoke the profile but not your account's default project. Fix.

From EU and Asia Pacific Regions, these models are reachable only through global routing. If you have data-residency rules, read that section first.

New to Bedrock itself? Our plain-English guide to Amazon Bedrock explains the idea: one AWS service that rents you models from many companies, billed on your AWS account and governed by IAM. OpenAI joining that shelf is the news; everything else about how Bedrock works still applies.

GPT-6 Astra vs Sol vs Luna: which one to use

The three models share a context window, an output limit and the same input types. What differs is how much thinking each is built to do, and how much each costs. AWS's own descriptions are short and useful, so here they are in plain terms.

GPT-6 Astra GPT-6 Sol GPT-6 Luna
Built forThe hardest end-to-end work: complex reasoning, coding, computer use, research, document creationEveryday demanding work: building features, debugging, refactoring, code review, multistep tasks across toolsRepeatable work at scale: summarizing, extracting, classifying, routing, focused questions
On Bedrock sinceSeptember 8, 2026September 22, 2026September 22, 2026
Context / max output1,050,000 / 128,0001,050,000 / 128,0001,050,000 / 128,000
Input / outputText, image / textText, image / textText, image / text
Global price per 1M tokens$10 in, $50 out$2 in, $10 out$0.10 in, $0.50 out
Structured outputs on bedrock-runtimeNot supportedNot supportedSupported
bedrock-mantle Regionus-west-2 onlyus-east-1 onlyus-east-1 only

Read the price row as a ratio, because that is how you should choose: Astra costs five times Sol, and Sol costs twenty times Luna. A task that Luna does well enough costs one hundredth of the same task on Astra. The honest pattern for most teams is Luna for the high-volume plumbing (sort this ticket, pull the invoice number out of this email, summarize this call), Sol as the daily workhorse for anything that writes or changes code, and Astra reserved for the few jobs where a better answer is worth fifty dollars per million output tokens. Sol and Luna also let you set reasoning effort to none, low, medium, high, xhigh or max, with medium as the default; turning Luna down to none or low for classification is the cheapest speed-up you will ever make.

One detail from the model cards that trips people who build extraction pipelines: structured outputs, the feature that forces a reply to match a JSON schema, is listed as supported on the main bedrock-runtime endpoint for Luna, and not supported there for Sol or Astra. If your pipeline depends on guaranteed JSON, that is one more reason Luna is the extraction model.

What about GPT-6.1 Astra?

If you searched for "Bedrock GPT-6 Astra" after late-September headlines, here is the short version. Reports on September 28, 2026 said OpenAI had held back a planned successor, GPT-6.1 Astra, before launch. That does not affect the GPT-6 Astra you can call on Bedrock today, which has been generally available there since September 8, with an end of life no sooner than September 8, 2027. When a newer Astra ships on Bedrock, it will appear as a new model ID; your existing ID keeps working through its lifecycle.

Endpoints and model IDs: the part everyone gets wrong first

Bedrock now has two ways in for OpenAI models, and each accepts different model IDs. This is the single most confusing thing about the launch, so it gets a table of its own.

bedrock-runtime (recommended) bedrock-mantle
Base URLhttps://bedrock-runtime.{region}.amazonaws.com/openai/v1https://bedrock-mantle.{region}.api.aws/openai/v1
Model ID to sendglobal.openai.gpt-6-sol or us.openai.gpt-6-sol (same pattern: -astra, -luna)openai.gpt-6-sol (bare ID)
Bare model ID?Not supported; in-Region calls are not available on this endpointRequired
APIsResponses, Chat Completions, ConverseResponses, Chat Completions
RegionsMany source Regions (next section)One: us-east-1 for Sol and Luna, us-west-2 for Astra
Guardrails, application inference profilesYes, through the Converse API onlyApplication inference profiles not supported
Server-side tool callingNot supportedSupported
Prompt cachingSol and Luna, Responses API only; not listed for AstraAll three, Responses API only

Three rules fall out of that table. First, on bedrock-runtime, always send a profile ID, starting global. or us.; the bare openai.gpt-6-sol is refused there, because these models are served only through cross-Region inference on that endpoint. Second, the path is /openai/v1, not /v1; forgetting the openai segment is the most common 404 of launch week. Third, pick the endpoint by feature: bedrock-runtime is the one AWS recommends for new applications, and the only one with Converse, Guardrails and application inference profiles for cost tagging; bedrock-mantle is the one with server-side tool calling, pinned to a single Region per model. None of the three supports the Anthropic-style Messages API or Bedrock's older InvokeModel request format.

These models run only on the Standard service tier, pay per token with no commitment. Priority, Flex and Reserved capacity are not offered for them, so leave service_tier out or set it to default.

Where GPT-6 runs: global, US and the data-residency catch

Bedrock offers three ways to route a request. In-Region keeps it in one Region. Geo cross-Region inference spreads it across Regions within one geography, respecting data residency for that geography. Global cross-Region inference routes it to any Region in the world, and AWS's own description says to use it when you have no data residency needs. If the difference between a Region and a geography is fuzzy, our Regions versus Availability Zones map sorts it out.

For the GPT-6 family on bedrock-runtime, the pattern is simple and consequential:

  • US and Canada source Regions (us-east-1, us-east-2, us-west-1, us-west-2, ca-central-1, and for Sol and Luna also ca-west-1) can use either the us. geo profile or the global. profile.
  • Every other listed Region, including Frankfurt, Ireland, London, Paris, Stockholm, Tokyo, Seoul, Mumbai, Singapore, Sydney and SΓ£o Paulo, can use only the global. profile. There is no EU or Asia Pacific geo profile for these models, and no in-Region option on this endpoint anywhere.

That is the catch for regulated teams outside the US. Calling GPT-6 from eu-central-1 today means global routing, which by AWS's definition can process the request in any Region. If your organization requires EU-only processing, these models do not meet that bar on Bedrock at launch, and you should not let a developer discover that after a pilot has gone to production. Watch the model cards; a geo profile for the EU would appear there as a new row.

GPT-6 on Bedrock pricing, and the 10% for staying home

The base prices on Bedrock are OpenAI's own first-party Standard rates. Global cross-Region inference charges exactly those rates. The us. geo profile and in-Region calls on the mantle endpoint add a 10% premium. Here is the full short-context table, per million tokens, for prompts of 272,000 input tokens or fewer.

Model and routing Input Cache write Cache read Output
Astra, global$10.00$12.50$1.00$50.00
Astra, US geo or mantle in-Region$11.00$13.75$1.10$55.00
Sol, global$2.00$2.50$0.20$10.00
Sol, US geo or mantle in-Region$2.20$2.75$0.22$11.00
Luna, global$0.10$0.125$0.01$0.50
Luna, US geo or mantle in-Region$0.11$0.1375$0.011$0.55

Above 272,000 input tokens, the long-context rates apply to the whole request, not just the part above the line, and they roughly double input and add half to output: Astra goes to $20 in and $75 out globally ($22 and $82.50 on the US or mantle options), Sol to $4 and $15 ($4.40 and $16.50), Luna to $0.20 and $0.75 ($0.22 and $0.825). A prompt of 272,001 tokens therefore costs about twice as much as a prompt of 272,000 tokens. If you are stuffing whole codebases into context, it is worth trimming to stay under the line.

Now Ethan's question, answered. The us. profile promises your request stays in US Regions; the global. profile promises nothing about geography. Bedrock charges 10% more for the promise. For most workloads that is a fair trade and the right default in the US. For a startup with no residency obligations, it is a straight 10% off the model bill for changing a few characters: $1 saved on every million Sol output tokens, $5 on every million from Astra. Decide it on purpose; do not let a copied code sample decide it for you. The AWS sample code for Luna uses us., the sample for Sol uses global., which is exactly how teams end up paying two different rates for the same kind of call.

Prompt caching is where the real savings are. A cache read costs one tenth of fresh input on all three models, so a long system prompt or a large shared document that you send on every call should be cached. On bedrock-runtime, caching is listed for Sol and Luna through the Responses API only; on the mantle endpoint it is available for all three, again through Responses. Chat Completions and Converse calls do not get it. For the bigger picture on keeping a Bedrock bill sane, see our Bedrock pricing guide.

Calling GPT-6 on Bedrock: keys, working code and the migration checklist

The good news for anyone with existing OpenAI code is that the OpenAI SDK works unchanged. You swap the base URL and the key, and use a Bedrock model ID. The key itself deserves a paragraph, because Bedrock has two kinds and the docs are firm about which one belongs in production.

  • Short-term API key: lasts up to 12 hours or the length of your session, inherits the permissions of the IAM identity that generated it, and is the one AWS recommends for production. Generate one in the Bedrock console under API keys, or in code with the aws-bedrock-token-generator package, whose provide_token() refreshes automatically when called before each request.
  • Long-term API key: lasts until an expiry date you set, creates a dedicated IAM user behind the scenes, and is documented as "for exploration only." It is the fastest way to a first request, and the wrong thing to leave in a deployed app.
  • No key at all: the OpenAI-compatible paths accept AWS SigV4 signatures, so anything running with an IAM role on EC2, ECS, Lambda or AgentCore can call them with the role it already has.

Here is the whole first request for Sol through the recommended endpoint, using a short-term key from the console.

  1. In the Amazon Bedrock console, open API keys and generate a short-term key (switch the console to the Region you will call first).
  2. Install the SDK: pip install openai.
  3. Set two environment variables, then run the script.
export OPENAI_API_KEY="your-short-term-bedrock-api-key"
export OPENAI_BASE_URL="https://bedrock-runtime.us-east-1.amazonaws.com/openai/v1"

# first_request.py
from openai import OpenAI
client = OpenAI()   # reads the two variables above

resp = client.responses.create(
    model="global.openai.gpt-6-sol",
    input="Summarize what Amazon Bedrock does in two sentences.",
)
print(resp.output_text)

# streaming, Chat Completions style
stream = client.chat.completions.create(
    model="global.openai.gpt-6-luna",
    messages=[{"role": "user", "content": "Classify this ticket: my card was charged twice"}],
    stream=True,
)
for chunk in stream:
    if chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="")

The same call with no key, signed by your AWS credentials, looks like this in curl (7.75 or later), which is also the quickest way to prove a role's permissions from a server:

curl -X POST "https://bedrock-runtime.us-east-1.amazonaws.com/openai/v1/chat/completions" \
  -H "Content-Type: application/json" \
  --aws-sigv4 "aws:amz:us-east-1:bedrock" \
  --user "$AWS_ACCESS_KEY_ID:$AWS_SECRET_ACCESS_KEY" \
  -d '{"model":"global.openai.gpt-6-luna","messages":[{"role":"user","content":"Hello"}]}'

To use the mantle endpoint instead, change the base URL to https://bedrock-mantle.us-east-1.api.aws/openai/v1 (us-west-2 for Astra) and the model to the bare openai.gpt-6-sol. And if you would rather stay inside the AWS SDK, the Converse API accepts the same profile IDs, which is the route when you want Bedrock Guardrails on the output or an application inference profile to tag costs by team.

Moving from OpenAI direct: the checklist

  1. Base URL becomes the Bedrock one, and the path must include /openai/v1.
  2. Key becomes a Bedrock short-term key or SigV4; the token generator package replaces your key-rotation code.
  3. Model name becomes a profile ID: global.openai.gpt-6-sol, not gpt-6-sol.
  4. client.models.list() does not exist on bedrock-runtime; use Bedrock's ListInferenceProfiles or the model cards. It works on the mantle endpoint only.
  5. Prompt caching moves to the Responses API; Chat Completions calls do not get cache pricing here.
  6. Quotas count output tokens ten to one, so size your tokens-per-minute request on output, not input.
  7. Drop service_tier other than default; Priority and Flex are not offered.
  8. Structured outputs are only on Luna via bedrock-runtime; keep your own JSON validation for Sol and Astra.
  9. Logging and cost tags are Bedrock features: turn on model invocation logging, and use Converse with an application inference profile where you need per-team cost lines.

The errors you will hit first, and the fixes

AccessDeniedException even though the role can invoke the model. This is the launch-week classic, and the cause is one line in the model card: your identity needs bedrock:InvokeModel on two resources, the inference profile you are calling, and your account's default project, whose ARN looks like arn:aws:bedrock:us-east-1:111122223333:project/default. Most policies written for older Bedrock models grant the first and not the second. Add the project ARN to the policy's resources and the error goes away. If you need a refresher on reading those strings, see what an ARN is and IAM policy JSON, line by line.

{
  "Version": "2012-10-17",
  "Statement": [{
    "Effect": "Allow",
    "Action": "bedrock:InvokeModel",
    "Resource": [
      "arn:aws:bedrock:*:111122223333:inference-profile/global.openai.gpt-6-sol",
      "arn:aws:bedrock:*::foundation-model/openai.gpt-6-sol",
      "arn:aws:bedrock:us-east-1:111122223333:project/default"
    ]
  }]
}

Treat that policy as a starting shape, not a copy-paste answer: cross-Region profiles route to foundation models in several Regions, so the resource list must cover the Regions the profile can use, and your own account ID and Region replace the examples.

ValidationException or "model not supported" with openai.gpt-6-sol. You sent the bare ID to bedrock-runtime. Use global.openai.gpt-6-sol or us.openai.gpt-6-sol, or switch to the mantle endpoint where the bare ID is correct.

404 on client.models.list(). The main endpoint does not implement the OpenAI models list; it is not a permissions problem. List models with Bedrock's own APIs, or call the mantle endpoint, which does support it.

404 Not Found. Your base URL ends in /v1 instead of /openai/v1, or you pointed the mantle endpoint at the wrong Region for that model: us-east-1 for Sol and Luna, us-west-2 for Astra.

The us. profile is not available in your Region. Outside the US and Canada only global. exists for these models. That is not a bug; it is the residency catch described above.

ThrottlingException far earlier than your token math suggests. On bedrock-runtime, the tokens-per-minute quota counts each output token as ten tokens. A job that generates long answers burns quota ten times faster than its input suggests. Plan quota increases around output, and request them through Service Quotas before launch, not after the first busy morning.

service_tier rejected. Only Standard is supported for these models. Remove the field or set it to default.

JSON schema ignored on Sol or Astra. Structured outputs are not supported for those two on bedrock-runtime. Use Luna for strict extraction, or validate and retry on your side.

The OpenAI-compatible API on Bedrock, explained

"Does Amazon Bedrock support OpenAI models?" used to have a narrow answer: only OpenAI's open-weight models. It now has a full one, and the way AWS delivered it is worth understanding, because it changes how you build. Bedrock exposes an OpenAI-compatible API: the same request and response shapes as OpenAI's own Responses API and Chat Completions API, served from an AWS hostname under the /openai/v1 path. Any library or tool that lets you set a base URL, the official OpenAI SDKs for Python and JavaScript, most agent frameworks, many IDE assistants, can talk to Bedrock without a Bedrock-specific plugin.

What you give up by staying on the compatible API instead of Bedrock's native Converse API is the AWS-only extras: Bedrock Guardrails and application inference profiles for cost tagging work through Converse only. What you gain is portability: the same code can point at OpenAI, at Bedrock, or at another OpenAI-compatible provider by changing two environment variables. For most teams the right split is the compatible API for application code that may move, and Converse for the few services that need Guardrails or per-team cost tags. Both reach the same GPT-6 models at the same prices.

Reasoning effort travels through the compatible API as well. Sol and Luna accept none, low, medium, high, xhigh and max, with medium as the default, and effort is the biggest lever you have on both latency and cost, because reasoning tokens are billed as output.

Watching GPT-6 calls: logs, traces and cost tags

Once GPT-6 is in production, three AWS features keep it honest. Model invocation logging, supported on bedrock-runtime for all three models, writes every request and response to CloudWatch Logs or S3 so you can answer "what did we send, and what came back" after the fact; decide your retention and who can read those logs before you switch it on, because they contain your prompts. Application inference profiles let you create a named wrapper around a model per team or product, so Cost Explorer shows "support-bot" and "code-review" as separate lines instead of one GPT-6 total; they work only through the Converse API. And tracing, if your calls come from an agent, lands in CloudWatch, where CloudWatch Omni can show each model call inside the agent's trace and score the answers. Together those three turn "we use GPT-6 somewhere" into "this team spent this much on this model for this job, and the answers scored this well."

Amazon Bedrock vs Azure OpenAI vs OpenAI direct

The most searched question about this launch is which of the three to use, and the honest answer is that the model is the same; what differs is whose cloud wraps it. OpenAI directly gets new models and features first and bills you on OpenAI's account. Azure OpenAI puts OpenAI's models inside Microsoft's cloud, governed by Microsoft Entra ID and billed on an Azure agreement. Amazon Bedrock puts them inside AWS, governed by IAM, logged in CloudTrail, reachable over VPC endpoints, billed on your AWS account, and sitting on the same shelf as models from Anthropic, Meta, Mistral, Amazon and others behind one API. On Bedrock, global routing is priced at OpenAI's own rates, so the choice is not about price. It is about where your identity, your data and your other workloads already live. If your company runs on AWS, keeping model calls inside AWS removes a separate vendor, a separate key and a separate bill from the security review. Our AWS vs Azure vs Google Cloud comparison covers the wider decision.

Two AWS-side facts help the case. Inference data sent through Bedrock is not used to train the models, and using these models does not require opting into sharing your data with OpenAI. And AWS says OpenAI's workplace chat product and its Codex coding agent can be configured to use GPT-6 Astra on Bedrock, which matters to companies that want the chat product and the API on the same governed pipe.

For teams and IT: adopting GPT-6 on Bedrock cleanly

  1. Decide routing once, in writing. us. for US data-residency needs at 10% more, global. otherwise. Outside North America, confirm with compliance that global routing is acceptable before any pilot.
  2. Grant the default project in IAM. Add project/default to every role that invokes these models, and prefer roles with SigV4 over long-term API keys in production.
  3. Route by model tier. Luna for classification and extraction, Sol for coding and multistep work, Astra by exception. Make the cheap model the default in shared code.
  4. Cache long, repeated prompts through the Responses API, and keep inputs under 272,000 tokens where you can.
  5. Use Converse where you need Guardrails or cost tags; application inference profiles only work there.
  6. Size quotas for output at ten to one, and request increases before launch.
  7. Log and watch. Turn on model invocation logging, and read the model line in Cost Explorer weekly for the first month.

Is GPT-6 available on Amazon Bedrock?

Yes. GPT-6 Astra has been generally available on Amazon Bedrock since September 8, 2026, and GPT-6 Sol and GPT-6 Luna since September 22, 2026. All three accept text and image input, return text, and support a 1,050,000-token context window with up to 128,000 output tokens.

What is the model ID for GPT-6 Sol on Bedrock?

On the bedrock-runtime endpoint, use the inference profile ID global.openai.gpt-6-sol or us.openai.gpt-6-sol. On the bedrock-mantle endpoint in us-east-1, use the bare ID openai.gpt-6-sol. The same pattern applies to gpt-6-astra and gpt-6-luna, except that Astra's mantle endpoint is in us-west-2.

How much does GPT-6 cost on Amazon Bedrock?

With global cross-Region inference, per million tokens: Astra $10 input and $50 output, Sol $2 and $10, Luna $0.10 and $0.50, matching OpenAI's own Standard rates. The US geo profile and mantle in-Region calls cost 10% more. Prompts above 272,000 input tokens are billed at higher long-context rates for the whole request.

Why is the US inference profile more expensive than global?

Bedrock adds a 10% premium to US geographic cross-Region inference and to in-Region calls on the mantle endpoint, because those options keep processing inside a defined geography. Global cross-Region inference can route to any Region and is charged at OpenAI's base rates.

Can I use the OpenAI SDK with Amazon Bedrock?

Yes. Set the base URL to https://bedrock-runtime.{region}.amazonaws.com/openai/v1 (or the bedrock-mantle URL), use a Bedrock API key as the OpenAI API key, and pass a Bedrock model ID. Both the Responses API and Chat Completions work. Requests can also be signed with AWS SigV4 instead of a key.

Why do I get AccessDeniedException calling GPT-6 on Bedrock?

Usually because your IAM identity can invoke the inference profile but not your account's default project. Grant bedrock:InvokeModel on both the inference profile and arn:aws:bedrock:{region}:{account-id}:project/default. Also check that you are using a profile ID, not the bare model ID, on bedrock-runtime.

Is GPT-6 on Bedrock available in Europe?

From European source Regions such as Frankfurt, Ireland, London, Paris and Stockholm, the GPT-6 models are available only through global cross-Region inference, which can process requests in any Region. There is no EU geographic profile or in-Region option for them at launch, which matters if you have EU data-residency requirements.

What is the difference between bedrock-runtime and bedrock-mantle?

Both serve OpenAI models through OpenAI-compatible paths under /openai/v1. bedrock-runtime is the recommended endpoint, available from many Regions, supports Converse, Guardrails and application inference profiles, and requires a global or US profile ID. bedrock-mantle runs in a single Region per model, takes the bare model ID, and supports server-side tool calling.

Which is better on Bedrock, GPT-6 Sol or GPT-6 Luna?

They are built for different work. Sol handles demanding tasks such as building features, debugging, code review and multistep workflows. Luna handles high-volume focused tasks such as summarizing, extracting and classifying, at one twentieth of Sol's price, and supports structured outputs on bedrock-runtime. Many teams use Luna by default and Sol where quality demands it.

Does GPT-6 on Bedrock support prompt caching?

Yes, through the Responses API only. A cache read costs one tenth of the normal input price. On bedrock-runtime caching is listed for Sol and Luna; on bedrock-mantle it is available for Astra, Sol and Luna. Chat Completions and Converse calls do not use it.

Why am I throttled so quickly with GPT-6 on Bedrock?

On bedrock-runtime the tokens-per-minute quota counts every output token as ten tokens. Jobs that generate long responses exhaust quota far faster than input size suggests. Request quota increases through Service Quotas based on expected output volume.

Does Amazon Bedrock support Priority or Flex tiers for GPT-6?

No. The GPT-6 models on Bedrock support only the Standard tier, billed per token with no commitment. Priority, Flex and Reserved tiers are not available for them; omit service_tier or set it to default.

Is Amazon Bedrock cheaper than Azure OpenAI or OpenAI direct for GPT-6?

On Bedrock, global cross-Region inference is priced at OpenAI's own first-party Standard rates, and US-only routing adds 10%. The practical difference between the three is less about per-token price and more about which cloud governs identity, logging, networking and billing for your workloads.

Does OpenAI see my data if I use GPT-6 through Bedrock?

AWS states that inference data is not used to train the models and that using GPT-6 Astra on Bedrock does not require opting into sharing data with OpenAI. Your requests are billed and governed through your AWS account, with IAM, CloudTrail and VPC endpoint controls.

Does Amazon Bedrock support OpenAI models?

Yes. Besides OpenAI's open-weight models, Bedrock now serves OpenAI's GPT-6 Astra, Sol and Luna, callable through an OpenAI-compatible Responses and Chat Completions API under the /openai/v1 path, or through Bedrock's native Converse API.

Is Amazon Bedrock OpenAI-compatible?

For these models, yes. Bedrock exposes OpenAI-compatible Responses and Chat Completions endpoints, so the official OpenAI SDKs and most tools that accept a custom base URL work by changing the base URL, the key and the model ID. Bedrock-only features such as Guardrails and application inference profiles require the Converse API.

What reasoning effort settings does GPT-6 support on Bedrock?

GPT-6 Sol and Luna accept none, low, medium, high, xhigh and max, with medium as the default. Lower effort is faster and cheaper because reasoning tokens are billed as output tokens.

Should I use a short-term or long-term Bedrock API key?

Short-term for anything real: it lasts up to 12 hours, inherits your IAM identity's permissions and refreshes automatically through the aws-bedrock-token-generator package. Long-term keys create an IAM user and are documented for exploration only. Workloads on AWS can skip keys entirely and sign with SigV4 using their role.

Why does client.models.list() fail on Bedrock?

The bedrock-runtime endpoint does not implement the OpenAI-compatible models list operation. Use ListFoundationModels or ListInferenceProfiles in the Bedrock API, or the model cards. The bedrock-mantle endpoint does support the models list.

The code change was an afternoon, and it usually is. The decisions around it are what deserve a meeting: which of the three models is the default, whether your requests may leave your geography, and whether the 10% for staying home is worth it to you. Jake settled on Luna for the ticket triage, Sol for the code review bot, and a us. profile for the client who asked where their data goes. Ethan settled on one rule he now repeats to every client: "Choose the routing on purpose. The sample code chose it at random."

📌 If you keep one line from this page

On Bedrock, us. costs 10% more than global.; the difference is where your prompt may travel.

Send a profile ID, grant the default project, and pick Luna, Sol or Astra by the job, not by the name.

Revision note. Written September 29, 2026, from the Amazon Bedrock model cards for GPT-6 Astra, Sol and Luna, AWS's launch announcements of September 8 and September 22, 2026, and OpenAI's published API rates, all as they stood that day. The API-key guidance follows the Bedrock API keys reference. The IAM policy is an illustrative shape to adapt, not a tested least-privilege policy for your account. Regions, prices and supported features will change as the models mature; this page gets a dated update when they do.

Related