Claude Sonnet 5.5 on Amazon Bedrock: Model ID, Pricing, Code
Claude Sonnet 5.5 has been on Amazon Bedrock since September 28, 2026, the same day Anthropic released it. It costs the same as Sonnet 5, $2 per million input tokens and $10 per million output tokens at global routing, takes a 1,000,000-token context window and returns up to 128,000 tokens. You call it with the inference profile ID global.anthropic.claude-sonnet-5-5, not the bare model ID. Here is the part nobody tells you on launch day: there are now three different addresses for "Claude on AWS", each with its own model ID spelling, its own signing name and its own feature list, and the ID that works on one of them returns a 400 on another.
Ethan's shop has run its support bot on Sonnet 5 through Bedrock since July. When 5.5 landed, Jake expected a one-line change: swap the model ID, redeploy, done. The swap took thirty seconds. The next hour went on three errors that had nothing to do with the model being new and everything to do with the request Jake had been sending for three months. This page is that hour, written down so yours takes ten minutes: which door to use, the exact IDs, what it costs, working code for each API, and the errors that mean your Sonnet 5 request needs to change before Sonnet 5.5 will answer it.
New to Bedrock itself? The plain-English guide to Amazon Bedrock explains the idea in ten minutes: one AWS service that rents you models from many companies, billed on your AWS account and governed by IAM. Sonnet 5.5 arriving on that shelf is the news. The rules of the shelf have not changed, and that is what makes the three-doors problem below so easy to walk into.
The three doors to Claude Sonnet 5.5 on AWS, and which one you are standing at
Until this summer there was one way to call Claude on AWS: the Bedrock runtime, using the InvokeModel or Converse API, with a model ID that had a date and a version baked in. Two more doors have opened since, and Sonnet 5.5 is served through all three. They look alike from the outside. They are not.
| Door 1: Bedrock runtime | Door 2: Bedrock Mantle | Door 3: Claude Platform on AWS | |
|---|---|---|---|
| Who runs the servers | AWS | AWS | Anthropic, inside AWS |
| Hostname | bedrock-runtime.REGION.amazonaws.com | bedrock-mantle.REGION.api.aws | aws-external-anthropic.REGION.api.aws |
| Model ID for Sonnet 5.5 | global.anthropic.claude-sonnet-5-5 | anthropic.claude-sonnet-5-5 | claude-sonnet-5-5 |
| APIs | Messages (/anthropic/v1/messages), Converse, InvokeModel | Messages only | The whole Claude API: Messages, Batches, Files, Skills, beta headers |
| Signing name (SigV4) | bedrock | bedrock-mantle | aws-external-anthropic |
| Where Sonnet 5.5 is served | 35 Regions through global routing | GovCloud West only, at launch | Any commercial Region for the workspace; inference pinned US or Global |
| Bill shows up as | AWS Marketplace, under Anthropic | AWS Marketplace, under Anthropic | AWS Marketplace, in Claude Consumption Units |
| Guardrails, Knowledge Bases, Agents | Yes | No | No (Anthropic's own tooling instead) |
For almost everyone reading this, the answer is door 1. It is where the Bedrock console sends you, where Guardrails and Knowledge Bases live, and where Sonnet 5.5 is actually available in commercial Regions. Door 2, the Mantle endpoint, exists so that the Anthropic SDK can talk to Bedrock with the exact request shape it uses against Anthropic's own API, but for Sonnet 5.5 it is only switched on in GovCloud West today, so most of you cannot use it yet even if you want to. Door 3 is a different product with a different contract, and it gets its own section near the end.
Ethan's rule for Jake, which turned the hour into ten minutes: "Write down which hostname you are sending to. Everything else, the ID, the signing name, the features, follows from that one line."
Why the Bedrock runtime has a Messages API now
This is the quiet change that makes the old tutorials wrong. The Bedrock runtime still speaks Converse and InvokeModel, but for Opus 4.7 and every model after it, including Sonnet 5.5, it also serves Anthropic's native Messages API at the path /anthropic/v1/messages. Same request body as Anthropic's own API, same response, same streaming format. That means the official Anthropic Python package can point at Bedrock with a base_url and a Bedrock API key and just work, with no boto3 in sight. The Sept 28 model card ships exactly that sample, and it is the shortest working code on this page.
Claude Sonnet 5.5 Bedrock model ID: the one line that decides whether you get a 400
Every Bedrock model has a base model ID, and newer Claude models cannot be invoked with it on the runtime endpoint. You have to send an inference profile ID instead: the base ID with a routing prefix in front. Send the bare ID and you get this, before the model ever sees your prompt:
ValidationException: Invocation of model ID anthropic.claude-sonnet-5-5 with on-demand
throughput isn't supported. Retry your request with the ID or ARN of an inference profile
that contains this model.
The profile IDs for Sonnet 5.5, and where each one is allowed to send your request:
| ID to send | Routing | Price | Works from |
|---|---|---|---|
global.anthropic.claude-sonnet-5-5 | Anywhere in the world with capacity | List price | 35 commercial Regions |
us.anthropic.claude-sonnet-5-5 | Stays in US and Canada Regions | +10% | The 6 US and Canada Regions, plus GovCloud |
eu.anthropic.claude-sonnet-5-5 | Stays in EU Regions | +10% | The 8 EU Regions |
anthropic.claude-sonnet-5-5 | None on the runtime (refused); the only form on Mantle and on Claude Platform's cousin | List price | Mantle in GovCloud West only |
Notice what is missing. Sonnet 5 has au., in. and jp. profiles and can even be pinned to a single Region in Seoul and Singapore. Sonnet 5.5 launched with us. and eu. only. If you searched for jp.anthropic.claude-sonnet-5-5 or in.anthropic.claude-sonnet-5-5, they do not exist yet. That gap is the single most important fact on this page for anyone outside North America and Europe, and it gets its own section below.
You can also send the full ARN of the profile instead of the ID, in the form arn:aws:bedrock:REGION:ACCOUNT:inference-profile/global.anthropic.claude-sonnet-5-5. Teams do that when an IAM policy or an SCP has to name the exact resource. If ARNs still look like line noise, our ARN syntax guide takes them apart field by field.
Why "global" is the default and not a shortcut
Global cross-Region inference sends each request to whichever Region has capacity, so it rides out a busy hour in one Region without you doing anything. It is also the cheapest form. The 10% premium on us. and eu. buys one thing: a promise about geography. Pay it when a contract or a regulator asks where the prompt may travel. Do not pay it because the word "global" sounds risky.
How to opt for Claude Sonnet 5.5 in AWS Bedrock
Enabling the model is three clicks, and the two reasons it fails are both about your account, not the model. Here is the console route, then what to check when the first call comes back with AccessDeniedException.
- Open the Amazon Bedrock console in a Region that serves Sonnet 5.5 through global routing. Any of the US or EU Regions is a safe pick; the Regions table below lists all 35.
- In the left navigation, choose Model access. Find Anthropic, select Claude Sonnet 5.5, and submit the request. For most accounts access is granted within a few minutes. The first time you enable any Anthropic model in an account, a short use-case form appears; fill it in once and it never asks again.
- Go to Test → Playground, pick Claude Sonnet 5.5, and run one prompt. If the playground answers, the subscription and the Marketplace side are done, and any error you hit from code is yours to fix in IAM or in the request.
Two account-level things cause the long version of this story. First, Claude models are billed through AWS Marketplace, so the account needs a working payment method and no Marketplace restriction from an organization policy. Second, your IAM identity needs bedrock:InvokeModel and, for streaming, bedrock:InvokeModelWithResponseStream, on both the foundation model and the inference profile you call. A policy that names only the model ARN will be refused when you send a profile ID. The minimal statement looks like this:
{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Action": ["bedrock:InvokeModel", "bedrock:InvokeModelWithResponseStream"],
"Resource": [
"arn:aws:bedrock:*::foundation-model/anthropic.claude-sonnet-5-5",
"arn:aws:bedrock:*:111122223333:inference-profile/global.anthropic.claude-sonnet-5-5"
]
}]
}
Global routing has one more IAM wrinkle: the request may be fulfilled in a Region other than the one you called, so the foundation-model resource uses a wildcard Region, as above. Lock that down to a list of Regions and you have quietly re-created the data-residency promise of a us. profile, except without the routing to match, so requests fail instead of staying home. If your first call still comes back denied after this, work through every cause of AccessDeniedException on Claude, in order; it covers the Marketplace subscription, the first-use form, SCPs that block cross-Region profiles, and the zero-quota trap on new accounts.
Working code: Converse, InvokeModel, the Anthropic SDK, and the CLI
Four ways to make the same request. Pick by what your codebase already speaks, not by which one looks newest. All four send the same prompt to the same profile and print the same answer.
1. boto3 Converse: the AWS-native shape
import boto3
client = boto3.client("bedrock-runtime", region_name="us-east-1")
response = client.converse(
modelId="global.anthropic.claude-sonnet-5-5",
messages=[{"role": "user", "content": [{"text": "Explain Amazon Bedrock in three sentences."}]}],
inferenceConfig={"maxTokens": 1024},
)
print(response["output"]["message"]["content"][0]["text"])
Converse is the API to use when you want Guardrails, when you switch between model vendors, or when you are already in boto3. Anthropic-specific knobs that Converse has no field for, such as the thinking mode or the effort level, go in additionalModelRequestFields as a plain dictionary.
2. boto3 InvokeModel: the raw Anthropic body
import boto3, json
client = boto3.client("bedrock-runtime", region_name="us-east-1")
response = client.invoke_model(
modelId="global.anthropic.claude-sonnet-5-5",
body=json.dumps({
"anthropic_version": "bedrock-2023-05-31",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "Explain Amazon Bedrock in three sentences."}],
}),
)
print(json.loads(response["body"].read())["content"][0]["text"])
The anthropic_version line is required and its value has not changed since 2023; it is a format marker, not a model version, so leave it exactly as written. InvokeModel is the door through which mid-conversation system messages reach Sonnet 5.5 on Bedrock, which Converse cannot do because Converse owns the system field.
3. The Anthropic SDK against Bedrock: no boto3, one base URL
pip install -U "anthropic[bedrock]" aws-bedrock-token-generator
from anthropic import Anthropic
from aws_bedrock_token_generator import provide_token
token = provide_token(region="us-east-1") # short-term Bedrock API key from your AWS credentials
client = Anthropic(
base_url="https://bedrock-runtime.us-east-1.amazonaws.com/anthropic",
api_key=token,
)
response = client.messages.create(
model="global.anthropic.claude-sonnet-5-5",
max_tokens=1024,
messages=[{"role": "user", "content": "Explain Amazon Bedrock in three sentences."}],
)
print(next(block.text for block in response.content if block.type == "text"))
This is the one Jake ended up keeping. The request and response are byte-for-byte the Anthropic shape, so code written against Anthropic's API, or against Claude Platform on AWS, moves to Bedrock by changing two strings. The token generator mints a Bedrock API key from whatever AWS credentials you already have, valid for up to 12 hours, so no long-lived key lives in the repo. If you would rather sign with SigV4 and skip the key, the AnthropicBedrock client from the same package does that and accepts the same profile IDs.
4. AWS CLI: one line for a smoke test
aws bedrock-runtime converse \
--region us-east-1 \
--model-id global.anthropic.claude-sonnet-5-5 \
--messages '[{"role":"user","content":[{"text":"Say hello in five words."}]}]' \
--query 'output.message.content[0].text' --output text
Run that before you touch application code. If it prints five words, access, IAM and the profile ID are all correct, and anything that fails afterward is inside your request body.
Thinking, effort, and the Sonnet 5 request that Sonnet 5.5 refuses
This is where Jake's hour went. Sonnet 5.5 keeps the same prices and the same tokenizer as Sonnet 5, so the natural assumption is that the request body carries over untouched. Three parts of it do not.
| What you sent to Sonnet 5 | What Sonnet 5.5 does | Send this instead |
|---|---|---|
"thinking": {"type": "disabled"} | 400. Thinking cannot be switched off. | Omit thinking and set "output_config": {"effort": "low"}. If a route must stay thinking-free, "thinking": {"type": "between_tools"} at effort high or below, with no other field inside thinking. |
"tool_choice": {"type": "any"} or {"type": "tool", "name": ...} | 400. Forced tool use is gone. | {"type": "auto"} plus a sentence in the prompt saying when to call the tool, and "strict": true on the tool so the arguments match the schema. Check that a call happened; retry if not. |
| Editing an earlier turn and replaying the thinking blocks | 400 on accounts created on or after August 31, 2026; silently dropped blocks on older ones | Keep history append-only. Pass thinking blocks back unchanged. Do not prune old screenshots or tool results on the client. |
"thinking": {"type": "enabled", "budget_tokens": 8000} | 400 (already gone on Sonnet 5) | Adaptive thinking, which is the default when you omit the field, plus an effort level. |
Ethan's bot had "thinking": {"type": "disabled"} in it because the Sonnet 5 docs offered it as the way to get fast chat replies. The exact error Sonnet 5.5 returns is worth pasting here, because it tells you the fix inside the message:
"thinking.type.disabled" is not supported for this model. Use "thinking.type.between_tools"
for the lowest thinking setting, or "thinking.type.adaptive" and "output_config.effort"
to control thinking behavior.
Effort is the dial that replaced the on-off switch. Sonnet 5.5 takes low, medium, high, xhigh and max, defaults to high, and the levels are recalibrated from Sonnet 5, so a setting you tuned in August produces a different amount of thinking now. The starting points that hold up: medium for agentic coding and multistep tool use, low for chat, classification, extraction and search. At low the model skips thinking on most simple requests, which is what the old disabled flag was really for. On Converse, send the effort through additionalModelRequestFields:
response = client.converse(
modelId="global.anthropic.claude-sonnet-5-5",
messages=[{"role": "user", "content": [{"text": "Classify this ticket: 'my invoice shows two charges'"}]}],
inferenceConfig={"maxTokens": 512},
additionalModelRequestFields={"output_config": {"effort": "low"}},
)
One more thing changed in the response rather than the request. Text that Sonnet 5 wrote between tool calls came back as text blocks. On Sonnet 5.5, anything longer than a sentence or two between tool calls comes back as a thinking block whose text is empty by default. A chat window that renders only text blocks goes quiet mid-task. The fix is either between_tools, which returns those notes with their text, or adaptive thinking with "display": "summarized" so the summaries arrive. Ethan's bot now shows a one-line status between tool calls instead of a spinner, and the support team stopped asking whether it had crashed.
✅ Size max_tokens for thinking as well as the reply
Thinking counts toward max_tokens even when its text is not returned. A limit that fit Sonnet 5 at disabled can cut Sonnet 5.5 off mid-answer at high. For chat, 4,096 is comfortable at low; for long coding turns, start at 64,000 and stream.
Claude Sonnet 5.5 Bedrock pricing: the table and a worked month
Anthropic kept Sonnet 5.5 at Sonnet 5's price, and Bedrock bills it at the provider's rate. The charge appears on your AWS bill and in Cost Explorer under Anthropic, not under Amazon Bedrock, which is the first place people look and the reason "my Bedrock line is tiny" is such a common false alarm. Our Cost Explorer walkthrough shows exactly where the Marketplace line sits.
| Per million tokens | global. | us. or eu. (+10%) |
|---|---|---|
| Input | $2.00 | $2.20 |
| Output (thinking tokens included) | $10.00 | $11.00 |
| Prompt cache write, 5-minute | $2.50 | $2.75 |
| Prompt cache read | $0.20 | $0.22 |
| Batch (50% off) | Not offered for Sonnet 5.5 on Bedrock at launch | |
| Flex / Priority / Reserved tiers | Not offered at launch; Standard only | |
Read the Batch row twice. On Anthropic's own API and on Claude Platform on AWS, a nightly job can run through the Message Batches endpoint at half price. On Bedrock, Sonnet 5.5's model card lists Batch as unsupported, and the only service tier is Standard. If half-price batch work is a large share of your bill, that single row decides which door you use, and it is not the door the console points you at.
The cache rows are where the real savings live for a chatbot. Sonnet 5.5 lowered the smallest cacheable prompt from 1,024 tokens on Sonnet 5 to 512, so a system prompt that was too short to cache in August now qualifies. A cache read costs a tenth of a normal input token. Bedrock supports both the implicit cache and explicit checkpoints, up to four per request, on system, messages and tools, with a 5-minute or 1-hour lifetime. One thing to know before you copy code from Anthropic's docs: the top-level cache_control shortcut that caches the last block automatically is not available on Bedrock. Put the breakpoint on the block itself.
Ethan's support bot, costed
Numbers from the shop so the arithmetic is on the table. The bot handles 9,000 conversations a month. Each one carries a 2,400-token system prompt and product sheet, about 600 tokens of customer text across the exchange, and returns about 400 tokens including the short thinking Sonnet 5.5 does at low effort.
| Line | Tokens per month | No caching | System prompt cached |
|---|---|---|---|
| System prompt, 2,400 × 9,000 | 21.6M | 21.6 × $2 = $43.20 | 21.6 × $0.20 = $4.32 (+ a few cents of writes) |
| Customer text, 600 × 9,000 | 5.4M | 5.4 × $2 = $10.80 | $10.80 |
| Replies, 400 × 9,000 | 3.6M | 3.6 × $10 = $36.00 | $36.00 |
| Month, global routing | 30.6M | $90.00 | about $51.20 |
Same month on us. | $99.00 | about $56.30 |
Caching the system prompt takes $39 off a $90 month, which is more than the whole 10% residency premium three times over. Jake's reaction: "So the cache matters more than the Region." Ethan's: "The cache matters more than the model. We would have saved that on Sonnet 5 too. We just never read the usage block." The usage block is the proof: if cache_read_input_tokens is zero on the second request, something in the prefix changes every call, usually a timestamp in the system prompt, and the discount never happens.
The example above is illustrative, with round token counts, so treat the shape as the lesson and plug in your own numbers. For the rest of the Bedrock bill, the quotas, and the three ways to pay, the Bedrock pricing guide walks through every line.
Regions and data residency: who can keep Sonnet 5.5 at home
Bedrock serves Sonnet 5.5 in 35 commercial Regions plus the two GovCloud Regions, but "serves" means three different things depending on where you are. The question to ask is not "is it in my Region" but "can my request stay in my country".
| Where you call from | Routing available for Sonnet 5.5 | Can the request stay in your geography? |
|---|---|---|
| N. Virginia, Ohio, N. California, Oregon, Canada Central, Calgary | global. and us. | Yes, within US + Canada, at +10% |
| Frankfurt, Zurich, Stockholm, Milan, Spain, Ireland, London, Paris | global. and eu. | Yes, within the EU, at +10% |
| GovCloud West, GovCloud East | Geo profile only | Yes, within GovCloud. No global option. |
| Mumbai, Hyderabad | global. only | No. Sonnet 5 has an in. profile; Sonnet 5.5 does not yet. |
| Tokyo, Osaka | global. only | No. No jp. profile for 5.5 yet. |
| Sydney, Melbourne | global. only | No. Sonnet 5 has au.; Sonnet 5.5 does not yet. |
| Seoul, Singapore | global. only | No. Sonnet 5 can be pinned in-Region here; 5.5 cannot. |
| Taipei, Jakarta, Malaysia, Thailand, New Zealand, Tel Aviv, UAE, Bahrain, Cape Town, SΓ£o Paulo, Mexico | global. only | No |
This matters most for Indian readers of this site, so let me be plain about it. On September 29, 2026, the day after Sonnet 5.5 launched, AWS announced in-country inference for Claude in Mumbai and Hyderabad. The models in that announcement are Opus 5, Sonnet 5 and Haiku 4.5. Sonnet 5.5 is not on the list. If a client or a regulator requires the prompt to stay in India, Sonnet 5 through in.anthropic.claude-sonnet-5 is the model you can promise today, and Sonnet 5.5 is the model you can use once an in. profile appears. The same logic holds for Japan and Australia. The pattern with Sonnet 5 was that the extra geo profiles arrived in the weeks after launch rather than on launch day, so check the model card's data residency list before you sign anything.
Regions and Availability Zones confuse this further for people who are new to AWS, because a Region in the console is not the same thing as where a global profile may run. If that sentence did not land, the Regions vs Availability Zones map is the ten-minute fix.
⚠️ What this actually breaks
An organization policy that restricts aws:RequestedRegion to Mumbai, or a VPC endpoint policy that allows only ap-south-1 resources, will deny every Sonnet 5.5 call, because the only profile that serves Mumbai is global and global may fulfil the request elsewhere. The error says AccessDenied; the cause is geography. Sonnet 5 on the in. profile is the workaround until 5.5 gets one.
What works on Bedrock and what does not: the feature list that decides your architecture
Sonnet 5.5 is the same model on every door, but the door decides which features you can reach. This is the table to read before you design anything that depends on a server-side tool.
| Feature | Bedrock runtime | Claude Platform on AWS |
|---|---|---|
| Messages API, streaming, tool use, vision, PDF input, citations | Yes | Yes |
| 1M context, adaptive thinking, effort levels | Yes | Yes |
| Prompt caching (explicit breakpoints) | Yes; top-level auto-cache shortcut no | Yes, both |
| Guardrails, Knowledge Bases, Bedrock Agents, Flows, prompt management | Yes | No |
| Mid-conversation system messages | Yes, through InvokeModel only | Yes |
| Computer use | The beta computer_20251124 tool still works; the newer toolset does not | Toolset |
Structured outputs (output_config.format) | No | Yes |
| Count tokens endpoint | No on the runtime (yes on Mantle) | Yes |
| Web search, web fetch, code execution, advisor | No | Yes |
| Agent Skills, MCP connector, programmatic tool calling | No | Yes |
| Message Batches, Files API, Models API | No | Yes |
| Server-side refusal fallback | No (do it client-side) | Yes |
Beta features via anthropic-beta header | No | Yes |
| Request payload limit | 20 MB | Claude API limits |
Three rows change designs. Structured outputs are not on Bedrock for Sonnet 5.5, so a pipeline that relies on the model being forced to return schema-valid JSON has to do it the older way: define a tool whose input schema is the JSON you want, mark it strict, tell the model in the prompt to call it, and read the tool call's input. No web search or code execution means an agent that browses or runs code has to bring its own tools, which is exactly the job of Bedrock AgentCore, with its browser and code interpreter sandboxes. And no Batches is the pricing row from above, seen from the architecture side.
Two things that do work on Bedrock and get missed: the 1M context (the request size cap is 20 MB, so a stack of PDFs can hit the byte limit before the token limit), and refusals. Sonnet 5.5 declines in five categories, and a decline is a normal 200 with stop_reason: "refusal". Bedrock does not retry those on another model for you, so if a cyber or frontier-LLM decline would break your flow, catch the stop reason and send the request to Sonnet 5 yourself.
Sonnet 5.5 vs Sonnet 5 vs Opus 5.5 on Bedrock: which one Ethan's shop runs
| Sonnet 5.5 | Sonnet 5 | Opus 5.5 | |
|---|---|---|---|
| On Bedrock since | September 28, 2026 | June 30, 2026 | September 23, 2026 |
| Price in / out per 1M | $2 / $10 | $2 / $10 | $4 / $20 |
| Context / max output | 1M / 128K | 1M / 128K | 1M / 128K |
| Knowledge cutoff | June 2026 | January 2026 | See the Opus post |
| Thinking off switch | No (between_tools is the floor) | Yes, disabled | No; effort low is the floor |
| Default effort | high | high | medium |
| Smallest cacheable prompt | 512 tokens | 1,024 tokens | See the Opus post |
| Geo profiles on Bedrock | us. eu. | us. eu. au. in. + Seoul and Singapore in-Region | us. eu. au. jp. |
| EOL no sooner than | September 28, 2027 | June 30, 2027 | See the Opus post |
Same price, newer knowledge, cheaper caching, and the thinking changes. That is the whole trade for most teams, and it is why the shop moved the support bot to 5.5 and left the overnight report generator on Sonnet 5: the report job runs under an Indian data-residency clause and needs the in. profile. When 5.5 gets one, that job moves too. Opus 5.5 at twice the price stayed where it was, on the one task that earns it, the weekly code review; the Opus 5.5 pricing and benchmarks post covers when that doubling pays for itself.
A note on the EOL row, because "bedrock claude sonnet eol" is a search people make in a panic. Bedrock promises Sonnet 5.5 will not reach end of life before September 28, 2027, with a six-month legacy period after any announcement. Sonnet 5's floor is June 30, 2027. Neither date is an announcement; they are the earliest possible ones. Pin the model ID you deploy and read the model lifecycle page once a quarter, and you will never be surprised.
Claude Platform on AWS vs Bedrock: the third door, and when it is the right one
Door 3 deserves a plain description, because the name invites confusion. Claude Platform on AWS is Anthropic's own API, operated by Anthropic, reached through an AWS hostname, signed with your IAM credentials, and billed through AWS Marketplace. It is not Bedrock. Anthropic is the data processor; AWS handles identity and the invoice. Sonnet 5.5 was available there from launch day, as claude-sonnet-5-5, with no provider prefix.
What you gain by using it:
- Every feature in the table above that said "No" for Bedrock: structured outputs, Batches at half price, the Files API, web search and fetch, code execution, Agent Skills, the MCP connector, and beta features through the
anthropic-betaheader. - Same-day access to new models and features, instead of waiting on Bedrock's release schedule.
- A choice of SigV4 or a short-term API key for places where signing is awkward, such as a browser or a quick script.
- Data residency by request:
inference_geoset touskeeps inference in US data centers at a 1.1x price;globalis standard price.
What you give up, and the setup steps that catch people:
- Guardrails, Knowledge Bases, Bedrock Agents and the rest of the Bedrock toolbox. You use Anthropic's equivalents or build your own.
- AWS as the sole data processor. Regulated workloads that need FedRAMP High, IL4, IL5 or a HIPAA-ready boundary with AWS as the operating party belong on Bedrock. Anthropic's own guidance says the same.
- Zero data retention is opt-in and arranged through an Anthropic representative. On Bedrock, Anthropic never sees the inference data at all.
- Signing up creates a new Anthropic organization tied to your AWS account. Existing Anthropic API keys and workspaces do not carry over, and a Bedrock private offer does not transfer either; arrange commercial terms before the first request, because discounts are not applied retroactively.
- Before the first call, the account must enable outbound web identity federation once, and every request needs a workspace ID header. The SDK reads
ANTHROPIC_AWS_WORKSPACE_IDandAWS_REGIONfrom the environment; there is no default Region, so a missing one fails at client construction. - Rate limits start on Anthropic's Start tier with a monthly spend cap, and the self-service increase button is not available; you ask an Anthropic representative.
pip install -U "anthropic[aws]"
export ANTHROPIC_AWS_WORKSPACE_ID='wrkspc_01AbCdEf23GhIj'
export AWS_REGION='us-west-2'
from anthropic import AnthropicAWS
client = AnthropicAWS() # SigV4 from your AWS credentials; region + workspace from the environment
response = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=1024,
messages=[{"role": "user", "content": "Explain Amazon Bedrock in three sentences."}],
)
Billing is the other difference people feel. Claude Platform on AWS meters usage in Claude Consumption Units, invoiced monthly in arrears on your AWS bill, with per-model token rates behind them. There is no prepaid balance. Spend limits exist at the organization and workspace level, and they are computed at list price with about two hours of lag, so a burst can overshoot a limit before requests start failing, and the overshoot is billed.
Ethan's way of choosing: "If the sentence starts with 'our compliance team requires AWS to', it is Bedrock. If it starts with 'we need the feature Anthropic shipped last week', it is the Platform. If it starts with 'the console told me to', it is Bedrock and you have already decided." The two use separate capacity pools, so some teams run both and fail over between them, which is a legitimate design once the IDs and signing names are in a config file rather than in someone's head.
Claude Code Bedrock setup for Sonnet 5.5, and other tools
Claude Code talks to Bedrock with three environment variables, and the model string is the same profile ID as everywhere else on this page:
claude update
export CLAUDE_CODE_USE_BEDROCK=1
export AWS_REGION=us-east-1
export ANTHROPIC_MODEL='global.anthropic.claude-sonnet-5-5'
claude
Run claude update first. Each new Claude model on Bedrock has needed a minimum Claude Code version (Opus 5.5 needed 2.1.280), and an old build fails with a model-not-found error that looks like an access problem. If it is an access problem, the Claude Code section of the AccessDenied guide shows the exact IAM shape it needs, including the cross-Region profile permission that catches people on day one.
For frameworks that already support Bedrock, the only change is the model string. Anything that accepts an Anthropic base URL can use the token-generator route from the code section instead. One trap worth naming: a router or retry layer that re-sends a Sonnet 5.5 request body to another model must strip "thinking": {"type": "between_tools"} first, because every other model rejects it.
The first errors you will see, and what each one means
In the order Jake met them, with the fix beside each.
| Error text | What it means | Fix |
|---|---|---|
ValidationException: ... with on-demand throughput isn't supported. Retry your request with the ID or ARN of an inference profile | You sent the bare model ID to the runtime | Prefix it: global., us. or eu. |
AccessDeniedException on the first call | Model access not granted, first-use form not filled, IAM missing the profile ARN, or an SCP blocking cross-Region profiles | Playground test first; then the nine causes in the AccessDenied guide, in order |
"thinking.type.disabled" is not supported for this model | Sonnet 5 request body | Remove it; set effort low, or between_tools |
tool_choice: type "tool" and "any" are not supported for this model | Forced tool use | auto + prompt + strict: true |
output_config.effort 'xhigh' is not supported when thinking is disabled on this model | between_tools combined with xhigh or max | Drop to high, or switch to adaptive thinking |
| 400 after you edited an earlier message in a long conversation | Preserved thinking: the history no longer matches the thinking blocks (accounts created on or after August 31, 2026) | Append-only history; never rewrite a turn the model has already seen |
ThrottlingException: Too many tokens per minute | Account quota for the model in that Region; new accounts can start low | Retry with backoff; request a quota increase in Service Quotas; cache the prompt so fewer input tokens count |
ValidationException: Input is too long or a 413 on a big request | The 20 MB payload cap, usually from base64 images or PDFs | Resize images to 2,000 px or less per side; split the documents |
stop_reason: "refusal" with a normal 200 | A safety decline in one of five categories, named in stop_details | Branch on the stop reason before reading content; fall back client-side if the task allows it |
Signature error against bedrock-mantle | Wrong signing name, or Sonnet 5.5 not served on Mantle in your Region | Use the runtime with the /anthropic base URL unless you are in GovCloud West |
Jake's hour, for the record: the first error was the bare ID, because he had copied anthropic.claude-sonnet-5-5 straight from the model card's heading. The second was thinking.type.disabled. The third was a quiet one, not an error at all: the bot went silent between tool calls, because its notes were now arriving as empty thinking blocks. Ethan's summary, once all three were fixed: "None of that was the model being new. That was our request being old."
Moving from Sonnet 5 to Sonnet 5.5 on Bedrock: the checklist
- Model ID. Replace the profile string with
global.anthropic.claude-sonnet-5-5, or theus./eu.form if you were on a geo profile. If you were onau.,in.orjp., stop: there is no 5.5 equivalent yet. Stay on Sonnet 5 for that workload. - Thinking. Delete
"thinking": {"type": "disabled"}. Try effortlowfirst and measure time to first token; usebetween_toolsonly if the route must stay thinking-free. - Tool choice. Replace
anyandtoolwithauto, a prompt instruction, andstrict: true. Add a check that the call happened. - Response parsing. Read content blocks by type, not position. A response may begin with a thinking block whose text is empty.
- History handling. Make the conversation store append-only. Pass thinking blocks back unchanged. Remove any client-side pruning of old tool results or screenshots.
- max_tokens. Raise it to cover thinking plus the reply; stream anything long.
- Caching. Re-check which prompts now qualify at the 512-token minimum, and put explicit breakpoints on the blocks (no top-level shortcut on Bedrock).
- Effort sweep. Re-run it; the levels moved. Start at
mediumfor agents andlowfor chat, and judge cost per completed task, not per token. - UI. If the interface showed text between tool calls, render non-empty thinking blocks or use
between_tools. - Quotas. Sonnet 5.5 has its own rate-limit pool, separate from Sonnet 5's. Check Service Quotas for the new model before moving volume.
- Bill. Same prices, but re-baseline anyway: a higher default effort than you had means more output tokens unless you set it.
Everything in that list is a request change or a config change. Nothing in it is a rewrite. The team that budgets a day for it and finishes before lunch is the team that read the thinking section first.
Frequently asked questions
Is Claude Sonnet 5.5 available on Amazon Bedrock?
Yes. Claude Sonnet 5.5 has been generally available on Amazon Bedrock since September 28, 2026, the day Anthropic released it, through the global cross-Region inference profile in 35 commercial Regions plus GovCloud. It is also available on Claude Platform on AWS from the same day.
What is the Claude Sonnet 5.5 model ID on AWS Bedrock?
The base model ID is anthropic.claude-sonnet-5-5, but on the bedrock-runtime endpoint you must send an inference profile ID: global.anthropic.claude-sonnet-5-5 for worldwide routing, us.anthropic.claude-sonnet-5-5 to stay in US and Canada Regions, or eu.anthropic.claude-sonnet-5-5 to stay in the EU. The bare ID is refused with a ValidationException on the runtime.
How much does Claude Sonnet 5.5 cost on AWS Bedrock?
$2 per million input tokens and $10 per million output tokens at global routing, the same as Sonnet 5 and the same as Anthropic's own list price. The us. and eu. geo profiles add 10%. Prompt cache reads are $0.20 per million and 5-minute cache writes $2.50. Batch, Flex and Priority tiers are not offered for Sonnet 5.5 on Bedrock at launch.
How to opt for Claude Sonnet 5.5 in AWS Bedrock?
Open the Bedrock console in a supported Region, choose Model access, select Claude Sonnet 5.5 under Anthropic and submit the request; fill in the one-time Anthropic use-case form if it appears. Then open Test, Playground, and run one prompt. Your IAM identity needs bedrock:InvokeModel on both the foundation model and the inference profile ARN.
Why does Bedrock say on-demand throughput isn't supported for Sonnet 5.5?
Because you sent the bare model ID. Newer Claude models on Bedrock are served only through cross-Region inference profiles, so the request must name global.anthropic.claude-sonnet-5-5, us.anthropic.claude-sonnet-5-5 or eu.anthropic.claude-sonnet-5-5, or the full ARN of one of those profiles.
Does Claude Sonnet 5.5 on Bedrock support extended thinking and effort levels?
Yes. Adaptive thinking is on by default and cannot be fully disabled; the lowest setting is thinking type between_tools at effort high or below. Effort levels low, medium, high, xhigh and max are supported, with high as the default. On the Converse API, pass thinking and output_config through additionalModelRequestFields.
Why does my Sonnet 5 request fail on Sonnet 5.5 with thinking.type.disabled is not supported?
Sonnet 5.5 removed the disabled thinking mode. Delete the thinking field and set output_config.effort to low for fast replies, or send thinking type between_tools at effort high or below if the route must do no extended thinking. Forced tool_choice values any and tool also return a 400 on Sonnet 5.5 and need auto plus a prompt instruction.
Can I keep Claude Sonnet 5.5 requests inside India, Japan or Australia on Bedrock?
Not yet. At launch Sonnet 5.5 has only us. and eu. geo profiles; from Mumbai, Hyderabad, Tokyo, Osaka, Sydney and Melbourne it is reachable through global routing only. Sonnet 5 has in., au. and jp. profiles and is the model to use for in-country requirements until Sonnet 5.5 gains them.
Does Claude Sonnet 5.5 on Bedrock support batch inference?
No. The Sonnet 5.5 model card lists Batch as not supported and Standard as the only service tier. Half-price batch processing of Sonnet 5.5 is available through Claude Platform on AWS, which exposes Anthropic's Message Batches API, or through Anthropic's own API.
Does Claude Sonnet 5.5 on Bedrock support structured outputs?
No. Structured outputs are listed as not supported for Sonnet 5.5 on both the bedrock-runtime and bedrock-mantle endpoints. To get schema-valid JSON, define a tool whose input schema is the JSON you want, mark it strict, and read the tool call's input. Structured outputs are available on Claude Platform on AWS.
Is Claude hosted on AWS, and does Claude use AWS?
Both, in different senses. Anthropic runs Claude on several clouds, and on AWS it is offered two ways: Amazon Bedrock, where AWS operates the inference stack and is the data processor, and Claude Platform on AWS, where Anthropic operates the stack inside AWS and you reach it through an AWS hostname with IAM and Marketplace billing. Sonnet 5.5 is on both since September 28, 2026.
Claude Platform on AWS vs Bedrock: what is the difference?
Bedrock is operated by AWS, billed as a native AWS service through Marketplace, managed in the Bedrock console, and includes Guardrails, Knowledge Bases and Agents, with a feature subset of the Claude API. Claude Platform on AWS is Anthropic's full API operated by Anthropic, reached through an AWS hostname with IAM, billed in Claude Consumption Units through Marketplace, with same-day features but none of the Bedrock toolbox.
Can I use the Anthropic Python SDK with Claude Sonnet 5.5 on Bedrock?
Yes. Install anthropic[bedrock] and either use the AnthropicBedrock client with SigV4 and a profile ID, or construct the plain Anthropic client with base_url set to https://bedrock-runtime.REGION.amazonaws.com/anthropic and a short-term Bedrock API key from the aws-bedrock-token-generator package. Both send the native Messages API shape.
Does Claude Sonnet 5.5 on Bedrock support prompt caching?
Yes, both implicit caching and explicit checkpoints, up to four per request on system, messages and tools, with 5-minute and 1-hour lifetimes and a 512-token minimum, down from 1,024 on Sonnet 5. The top-level automatic cache_control shortcut is not available on Bedrock; put the breakpoint on the content block.
What is the Claude Sonnet 5.5 context window and max output on Bedrock?
A 1,000,000-token context window and up to 128,000 output tokens, with text and image input and text output. Bedrock limits a single request payload to 20 MB, so large base64 images or PDFs can hit the byte limit before the token limit.
When will Claude Sonnet 5.5 reach end of life on Bedrock?
No date has been set. Bedrock states the EOL will be no sooner than September 28, 2027, with a legacy period of at least six months after any announcement. Sonnet 5's floor is June 30, 2027. Pin the model ID you deploy and check the Bedrock model lifecycle page each quarter.
How do I use Claude Sonnet 5.5 on Bedrock with Claude Code?
Run claude update, then set CLAUDE_CODE_USE_BEDROCK=1, AWS_REGION, and ANTHROPIC_MODEL to global.anthropic.claude-sonnet-5-5 (or a us. or eu. profile), and start claude. Your IAM identity needs InvokeModel permission on both the foundation model and the inference profile.
The swap really is one string. The afternoon is in the request that string rides on: a thinking flag from the summer, a forced tool call, a chat window that only renders text, and a Region that cannot promise to stay home. Fix those four in the order on this page and Sonnet 5.5 on Bedrock is the quiet upgrade it was meant to be. Jake's bot answers faster now, costs the same, and shows a line of progress where there used to be a spinner. The report job stays on Sonnet 5 until an in. profile arrives, and that is not a compromise, it is the contract being kept.
If you keep one line from this page
Send global.anthropic.claude-sonnet-5-5, delete the disabled thinking flag, and read the usage block before you blame the price.
Pay the 10% only for a geography you have to promise, and check the geo-profile list before you promise one.
Revision note. Written October 1, 2026, three days after launch, from the Bedrock model cards for Sonnet 5.5 and Sonnet 5, AWS's launch and India in-country announcements, and Anthropic's Bedrock and Claude Platform on AWS references as they stood that day. The geo-profile table is the part most likely to be out of date first: Sonnet 5's extra profiles arrived in the weeks after its launch, and I expect 5.5's to follow. The prices are Anthropic's list rates with the published 10% regional premium applied; confirm them on the Bedrock pricing page before you quote a client. Ethan's bot numbers are round on purpose, so the arithmetic is easy to redo with yours.
