Bedrock "Invocation of Model ID With On-Demand Throughput Isn't Supported": The Fix

Logeshwaran
—

The Amazon Bedrock error "Invocation of model ID ... with on-demand throughput isn't supported. Retry your request with the ID or ARN of an inference profile that contains this model" means the model you called cannot be reached by its plain model ID in your Region. The fix is one line: send the model's inference profile ID instead, the same ID with a prefix such as us., eu. or global.. For example, change anthropic.claude-sonnet-5-5 to us.anthropic.claude-sonnet-5-5. Here is the catch that the older answers miss: the prefix that works depends on the model and on the Region you call from. For Claude Sonnet 5.5 on bedrock-runtime, no commercial Region offers the plain ID at all, and if you call from Tokyo, Mumbai, Sydney or SΓ£o Paulo, there is no apac. profile either. Only global.anthropic.claude-sonnet-5-5 works there.

Ethan builds small AI tools on AWS for a few local businesses, and one of them is a support bot for Jake's phone repair shop that answers "how long does a screen repair take?" at midnight. It ran for a year on an older Claude model through LangChain without a single error. On Monday, Ethan changed one string, the model ID, to move the bot to Claude Sonnet 5.5, and every request came back with a 400 and the message above. Jake only knew that the bot had gone quiet. "I didn't break anything," Ethan told him, "I just asked for the new model the old way." This page is the fix he applied, and everything around it: what the error means, how to find the right inference profile ID for any model and Region, the code for boto3, the AWS CLI, LangChain, LiteLLM and the OpenAI SDK, the IAM permissions that the fix quietly requires, and the next errors that tend to appear once this one is gone.

⚡ Quick Answer

• What it means → The model has no in-Region on-demand access where you called it. It is only offered through cross-Region inference profiles. The meaning.

• The fix → Put the inference profile ID in modelId: us., eu., apac., jp., au. or global. plus the model ID. You do not need to create anything. Code for every SDK.

• Which prefix? → The model's page in the Bedrock docs lists its profile IDs per Region, or run aws bedrock list-inference-profiles in your Region. Find yours.

• Then AccessDenied? → IAM must allow the profile and the model in every destination Region, plus the Region-less model ARN for global.. The policy.

• Cost → No extra charge for cross-Region routing. Global profiles are about 10% cheaper than geographic ones. Geo vs global.

Inference profiles do not support Provisioned Throughput. If you bought provisioned capacity, call the provisioned model's ARN instead.

If you landed here from a developer forum thread, you will have seen the short version of the fix already, and it is correct as far as it goes. What it does not tell you is which prefix your model and Region need in 2026, why your IAM policy may need to change, and how to keep data in the geography you promised a customer. Those are the parts that turn a one-line change into a working app.

🧭 NEW HERE? READ THESE FIRST

If Bedrock or the AWS words on this page are new to you, these five make the rest easier:

📌 Bookmark this; the IAM part below leans on the third and fourth.

What "on-demand throughput isn't supported" actually means

On-demand throughput is Bedrock's normal pay-per-token way of calling a model, with no capacity reserved in advance. Every model on Bedrock is offered in up to three ways, and each model's documentation page shows which ones exist in each Region:

  • In-Region: you call the plain model ID, such as amazon.nova-2-lite-v1:0, and the request runs in the Region you called.
  • Geographic cross-Region inference: you call a profile such as us.amazon.nova-2-lite-v1:0, and Bedrock runs the request in one of several Regions inside that geography, the US, the EU or Asia Pacific, for example.
  • Global cross-Region inference: you call global.amazon.nova-2-lite-v1:0, and Bedrock runs the request in any supported commercial Region worldwide.

The error appears when you call the plain model ID for a model that has no in-Region option where you are. Bedrock is not saying the model is unavailable to you; it is saying "not this way". The second sentence of the error says exactly what to do: retry with the ID or ARN of an inference profile that contains this model.

Most newer, popular models are offered this way, because routing across Regions lets AWS spread demand over more capacity. A few real examples from the Bedrock model pages, on the bedrock-runtime endpoint:

ModelPlain ID works in-Region?Profile IDs offered
Claude Sonnet 5.5No, in any commercial Regionus.anthropic.claude-sonnet-5-5, eu.anthropic.claude-sonnet-5-5, global.anthropic.claude-sonnet-5-5
GPT-6.1 SolNo (on bedrock-runtime)us.openai.gpt-6.1-sol, global.openai.gpt-6.1-sol
Amazon Nova 2 LiteYes, where in-Region is listedus., eu., jp. and global.amazon.nova-2-lite-v1:0
Claude 3 Haiku (older)Yes, in some Regionsus. and eu.anthropic.claude-3-haiku-20240307-v1:0

That table explains Ethan's Monday. His old model worked with its plain ID in his Region, so his code never needed a prefix. The new one has no plain-ID option anywhere on bedrock-runtime, so the same code broke the moment the string changed.

The error message variants people search for

The wording is always the same, but the model ID inside it changes, and so do the wrappers around it. These are all the same problem with the same fix:

  • ValidationException: Invocation of model ID anthropic.claude-3-5-haiku-20241022-v1:0 with on-demand throughput isn't supported. Retry your request with the ID or ARN of an inference profile that contains this model.
  • Error: Error 400: Invocation of model ID ... with on-demand throughput isn't supported, as LangChain reports it.
  • "converse API call using bedrock-runtime via boto3 ... throwing ValidationException on-demand throughput isn't supported", the way it is often described on developer forums.
  • The same message for many other models, including Claude 3.7 Sonnet, and GPT-6 Sol and GPT-6.1 Sol, wherever a model is profile-only in the Region you call from.

A related search is "max tokens for Claude models are much lower when using on-demand throughput". That one is about quotas rather than the model ID, and the quota section below explains why a large max_tokens can throttle you sooner than the model's real output would.

The fix: use the inference profile ID (code for every SDK)

You do not need to create an inference profile. AWS defines the cross-Region ones for you, and their IDs are the model ID with a geography prefix. Put that ID where the model ID used to go. With boto3 and the Converse API:

import boto3

brt = boto3.client("bedrock-runtime", region_name="us-east-1")

resp = brt.converse(
    modelId="us.anthropic.claude-sonnet-5-5",      # was: "anthropic.claude-sonnet-5-5"
    messages=[{"role": "user", "content": [{"text": "How long does a screen repair take?"}]}],
)
blocks = resp["output"]["message"]["content"]
print(next(b["text"] for b in blocks if "text" in b))   # skip any reasoning block

InvokeModel takes the same ID in modelId. You can also pass the full ARN of the profile, such as arn:aws:bedrock:us-east-1:111122223333:inference-profile/us.anthropic.claude-sonnet-5-5; for AWS-defined profiles, the ID and the ARN both work.

With the AWS CLI:

aws bedrock-runtime converse \
  --region us-east-1 \
  --model-id us.anthropic.claude-sonnet-5-5 \
  --messages '[{"role":"user","content":[{"text":"Hello"}]}]'

With LangChain, the error in the original forum question came from exactly this layer. Pass the profile ID as the model:

from langchain_aws import ChatBedrockConverse

llm = ChatBedrockConverse(
    model="us.anthropic.claude-sonnet-5-5",
    region_name="us-east-1",
)
print(llm.invoke("How long does a screen repair take?").content)

With LiteLLM, add the bedrock/ prefix in front of the profile ID:

from litellm import completion

response = completion(
    model="bedrock/us.anthropic.claude-sonnet-5-5",
    messages=[{"role": "user", "content": "Hello"}],
)

With the OpenAI SDK against bedrock-runtime, for an OpenAI model such as GPT-6.1 Sol, the profile ID is the model name, and the base URL ends in /openai/v1:

from openai import OpenAI

client = OpenAI(
    base_url="https://bedrock-runtime.us-east-1.amazonaws.com/openai/v1",
    api_key="<your Bedrock API key>",
)
r = client.responses.create(model="us.openai.gpt-6.1-sol", input="Hello")
print(r.output_text)

In the Bedrock console playground, choose the model, then under Inference choose Inference profiles and pick one, such as the US profile, before you send a prompt.

Ethan's fix for Jake's bot was the LangChain change above: one string, nine characters longer. The bot answered its first midnight question again within a minute of the deploy. What took him the rest of the hour was the IAM policy, which comes two sections down.

How to find the right inference profile ID for your model and Region

Prefixes are not interchangeable, and a profile only works from the source Regions it lists. Three reliable ways to find the right one:

  1. The model's page in the Bedrock user guide. Every model now has a detail page with a Programmatic access table listing the model ID, the geographic inference IDs and the global inference ID, and a Regional availability table showing, Region by Region, whether In-Region, Geo and Global are supported.
  2. The console. In the Bedrock console's model catalog or playground, the inference profiles available to your Region are listed for each model.
  3. The AWS CLI, from the Region you will call from. List the AWS-defined profiles, then check where one routes:
# every active system-defined profile for one model, from your source Region
aws bedrock list-inference-profiles \
  --region us-east-1 \
  --type-equals SYSTEM_DEFINED \
  --query "inferenceProfileSummaries[?status=='ACTIVE' && contains(inferenceProfileId, 'claude-sonnet-5-5')].[inferenceProfileId,inferenceProfileName]" \
  --output table

# which Regions a profile can send your requests to
aws bedrock get-inference-profile \
  --region us-east-1 \
  --inference-profile-identifier us.anthropic.claude-sonnet-5-5 \
  --query "models[].modelArn"

The Region inside each model ARN that get-inference-profile returns is a destination Region. Note the bedrock command, not bedrock-runtime: listing profiles is a control-plane call.

The geography prefixes you will meet include us, eu, apac and in, plus newer ones such as jp and au, and global for worldwide routing. Not every model has every prefix, which is where the 2024 advice to "just use apac." stops working. Here is Claude Sonnet 5.5 on bedrock-runtime, from its model page:

You call fromPlain IDGeo profileGlobal profile
US Regions (us-east-1, us-east-2, us-west-1, us-west-2) and CanadaNous.anthropic.claude-sonnet-5-5Yes
EU Regions (Frankfurt, Zurich, Stockholm, Milan, Spain, Ireland, London, Paris)Noeu.anthropic.claude-sonnet-5-5Yes
Asia Pacific (Tokyo, Seoul, Osaka, Mumbai, Hyderabad, Singapore, Sydney, Jakarta, Melbourne and others)NoNoneYes, the only option
Middle East, Africa, SΓ£o Paulo, Mexico, Tel AvivNoNoneYes, the only option

So a developer in Sydney who copies apac.anthropic.claude-sonnet-5-5 from an old answer gets a different error, because that profile does not exist. The working ID there is global.anthropic.claude-sonnet-5-5, which also means requests can be processed outside Asia Pacific. If that matters to your customers, read the next section before you ship.

A small helper that picks the right profile automatically

If your app runs in more than one Region, hard-coding us. will break the day someone deploys it in Frankfurt. This helper asks Bedrock which AWS-defined profiles exist for a model in the current Region and picks one in your order of preference:

import boto3

def pick_profile(model_id, region, prefer=("us", "eu", "jp", "au", "apac", "global")):
    """Return an active system-defined inference profile ID for model_id in region."""
    bedrock = boto3.client("bedrock", region_name=region)   # control plane, not bedrock-runtime
    found, token = {}, None
    while True:
        kwargs = {"typeEquals": "SYSTEM_DEFINED", "maxResults": 100}
        if token:
            kwargs["nextToken"] = token
        page = bedrock.list_inference_profiles(**kwargs)
        for p in page["inferenceProfileSummaries"]:
            pid = p["inferenceProfileId"]
            prefix, _, rest = pid.partition(".")
            if p["status"] == "ACTIVE" and rest == model_id:
                found[prefix] = pid
        token = page.get("nextToken")
        if not token:
            break
    for prefix in prefer:
        if prefix in found:
            return found[prefix]
    return model_id   # no profile: the plain ID may work in-Region

print(pick_profile("anthropic.claude-sonnet-5-5", "ap-southeast-2"))   # global. in Sydney

Two cautions. Put global last only if you are comfortable with worldwide processing; if residency matters, remove it from the list and let the code fail loudly instead. And call this once at startup and cache the result rather than on every request, because it is a control-plane call with its own rate limits. The calling role also needs bedrock:ListInferenceProfiles.

Geographic or global profile? Data residency, price and the SCP catch

Both kinds of profile fix the error. They differ in where your prompts are processed:

QuestionGeographic (us., eu., apac. ...)Global (global.)
Where requests runRegions inside that geography onlyAny supported commercial Region worldwide
PriceStandard pricingAbout 10% lower
Destination listFixed; never changes for that profileCan grow as AWS adds Regions
Region-deny SCP needsAllow every destination Region in the profileAllow "aws:RequestedRegion": "unspecified"
Best forData residency rulesLowest cost, no residency rule

A few facts that settle the usual worries:

  • There is no extra charge for cross-Region routing. You pay the model's price for the Region you call from, not the Region that processed the request.
  • You do not have to enable the destination Regions. Cross-Region inference can route to Regions that are not enabled in your account. Opt-in Regions can appear as destinations, which matters only if your organization's policies deny them.
  • Traffic stays on AWS. Data moving between Regions stays on the AWS network, encrypted in transit, and never crosses the public internet.
  • You can see where each request ran. CloudTrail records cross-Region requests in your source Region, and the additionalEventData.inferenceRegion field shows the Region that processed it.
  • Quotas are separate. Geographic profiles have their own "Cross-region model inference requests per minute" and "tokens per minute" quotas per model, in Service Quotas.

Ethan used the us. profile for Jake's bot, not global.. Jake's customers are all local, and Ethan had promised the shop that customer messages would be processed in the US. The 10% saving was not worth breaking that sentence.

Which quota applies once you switch to a profile

Switching to a profile changes more than the model string: it changes which quota your traffic counts against. On bedrock-runtime, each model has separate tokens-per-minute quotas for the two ways of calling it:

Quota name in Service QuotasCounts
Cross-Region tokens per minute for modelCalls through a geographic profile (us., eu., apac. ...), per Region
Global cross-Region model inference tokens per minute for modelCalls through a global. profile; one global limit
On-demand InvokeModel tokens per minute for modelCalls by the plain model ID, in one Region
InvokeModel requests per minute for modelRequests per minute, for models that have an RPM quota
Cross-Model Max Tokens Per DayAll models together, per account, per Region

AWS pages word some of these names slightly differently, so search Service Quotas for the model name. Despite the "InvokeModel" in some names, these quotas cover every API you call the model with on that endpoint, including Converse, Responses and Chat Completions. Both token quotas count input and output together, and some models count each output token several times over: Claude Opus 5.5 and GPT-6.1 Sol, for example, use a burndown rate of 10, so one output token consumes ten tokens of quota while you are billed for one. Bedrock also reserves your input plus the full max_tokens at the start of every request and returns the unused part at the end, which is why an oversized max_tokens throttles you early.

To raise any of the token quotas for a model, AWS asks you to request an increase on the Cross-Region InvokeModel tokens per minute quota for that model, which covers the others in the same request. AWS gives priority to accounts already using their current allocation, and does not grant increases for models in Legacy or Deprecated status. You can list a model's current quotas from the CLI:

aws service-quotas list-service-quotas \
  --service-code bedrock \
  --region us-east-1 \
  --query "Quotas[?contains(QuotaName, 'Claude Sonnet 5.5')].[QuotaName,Value]" \
  --output table

The exact model name inside the quota names can differ slightly from the marketing name, so if the query returns nothing, search for a shorter fragment such as Sonnet.

After the fix: AccessDeniedException and the IAM policy you need

The most common next error after switching to a profile is an AccessDeniedException, and the reason is that IAM now checks more than one resource. For a geographic profile, the policy must allow three things: the inference profile itself, the foundation model in your source Region, and the foundation model in every destination Region of that profile. For a global profile, add the Region-less model ARN, arn:aws:bedrock:::foundation-model/..., which has no Region or account in it on purpose.

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "UseTheProfile",
      "Effect": "Allow",
      "Action": ["bedrock:InvokeModel", "bedrock:InvokeModelWithResponseStream"],
      "Resource": [
        "arn:aws:bedrock:us-east-1:111122223333:inference-profile/us.anthropic.claude-sonnet-5-5",
        "arn:aws:bedrock:us-east-1:111122223333:inference-profile/global.anthropic.claude-sonnet-5-5"
      ]
    },
    {
      "Sid": "UseTheModelOnlyThroughTheProfile",
      "Effect": "Allow",
      "Action": ["bedrock:InvokeModel", "bedrock:InvokeModelWithResponseStream"],
      "Resource": [
        "arn:aws:bedrock:*::foundation-model/anthropic.claude-sonnet-5-5",
        "arn:aws:bedrock:::foundation-model/anthropic.claude-sonnet-5-5"
      ],
      "Condition": {
        "StringLike": {
          "bedrock:InferenceProfileArn": "arn:aws:bedrock:us-east-1:111122223333:inference-profile/*anthropic.claude-sonnet-5-5"
        }
      }
    }
  ]
}

Replace the account ID with yours. The condition lets the role use the model only through those profiles, which is a tidy way to stop anyone calling it some other way. The Converse API uses the same bedrock:InvokeModel permissions. If you call Bedrock with an API key rather than IAM credentials, also allow bedrock:CallWithBearerToken.

If the policy looks right and access is still denied, look one level up. A Region-deny service control policy in AWS Organizations, the kind that allows only a few approved Regions, blocks a geographic profile if any destination Region is not on its list, and blocks a global profile unless it allows aws:RequestedRegion with the value unspecified. That one catches teams who did everything else right. Our Bedrock AccessDeniedException guide walks through the rest of the checks, including the Anthropic first-use form, which surfaces as FTUFormNotFilled with a 404.

Application inference profiles: when you do want to create one

The forum answer that tells you to create an inference profile is not wrong, just unnecessary for this error. There is one good reason to create your own: tracking cost and usage per application. An application inference profile wraps a model or a system-defined profile, and you can tag it, so the Cost Explorer bill shows exactly what the support bot spent:

aws bedrock create-inference-profile \
  --region us-east-1 \
  --inference-profile-name shop-support-bot \
  --model-source copyFrom=arn:aws:bedrock:us-east-1:111122223333:inference-profile/us.anthropic.claude-sonnet-5-5 \
  --tags key=project,value=shop-support-bot

Point copyFrom at a system-defined cross-Region profile to keep cross-Region routing, or at a foundation model ARN to track a single Region. The command returns an inferenceProfileArn; use that ARN as the model ID in your calls. With application profiles you must use the ARN, not a short ID. In LangChain, also pass base_model="anthropic.claude-sonnet-5-5" so the library knows which model family it is talking to, since the ARN does not say.

Two limits to know: application inference profiles work with InvokeModel and Converse, but not with every API on every model (GPT-6.1 Sol, for example, supports them only on InvokeModel and Converse), and none of these profiles work with Provisioned Throughput.

Inside Knowledge Bases, Flows, Prompt management and evaluations

The error is not limited to your own code. Bedrock features that call a model on your behalf need the profile too, and the error can appear when you test a Knowledge Base or run a Flow with a model that is profile-only in your Region. AWS documents inference profiles as supported in these places:

  • Knowledge Bases: for generating answers after a query, through RetrieveAndGenerate, and for parsing non-text content in a data source, through CreateDataSource. Pass the profile ARN as the model.
  • Model evaluation: submit an inference profile as the model to evaluate in CreateEvaluationJob.
  • Prompt management: choose an inference profile when a stored prompt generates a response, through CreatePrompt.
  • Flows: use an inference profile for an inline prompt in a prompt node, through CreateFlow.

In the console, each of these has the same model selector as the playground, with an Inference profiles option under the model. In the API, the field that asks for a model ARN accepts the inference profile ARN. If a Knowledge Base worked last month and fails after you picked a newer model for answers, this is almost always why.

The same error in Claude Code, Codex, OpenCode and other tools

Developer tools that talk to Bedrock hit this error in exactly the same way, because they send whatever model ID you configure. The fix is the same: give them the profile ID.

Claude Code turns on Bedrock with environment variables. The model must be an inference profile ID, or an application inference profile ARN:

export CLAUDE_CODE_USE_BEDROCK=1
export AWS_REGION=us-east-1
export ANTHROPIC_MODEL='us.anthropic.claude-sonnet-5-5'

# or an application inference profile you created
# export ANTHROPIC_MODEL='arn:aws:bedrock:us-east-1:111122223333:application-inference-profile/abc123'

Anthropic's setup guide adds two details worth knowing. In AWS GovCloud Regions the prefix is us-gov.. And if you do not pin a model at all, recent Claude Code versions default to Opus 5.5 on Bedrock, which costs more per token than Sonnet, so pinning ANTHROPIC_MODEL is also a billing decision, not only a fix for this error.

Codex uses a provider setting in ~/.codex/config.toml. For bedrock-runtime, use model_provider = "amazon-bedrock-runtime" and a profile ID such as model = "us.openai.gpt-6.1-sol"; Codex sends requests to the runtime endpoint's /openai/v1 path in your Region. Our GPT-6.1 Sol guide has the full Codex setup.

OpenCode reads model IDs under its amazon-bedrock provider in opencode.json; a model missing from its built-in list is added there with the profile ID as its id. Our OpenCode guide walks through that file.

Any other SDK or framework: wherever the code asks for a "model ID" or "model ARN", the inference profile ID or ARN is accepted in its place for InvokeModel and Converse. If a library validates the string against a fixed list of known models and rejects the prefix, update the library first; older releases predate cross-Region inference.

Defining an application inference profile in CloudFormation or Terraform

If your infrastructure lives in code, create the per-app profile there too, so the tag that splits your bill cannot be forgotten. In CloudFormation, the resource is AWS::Bedrock::ApplicationInferenceProfile:

Resources:
  SupportBotProfile:
    Type: AWS::Bedrock::ApplicationInferenceProfile
    Properties:
      InferenceProfileName: shop-support-bot
      Description: Claude Sonnet 5.5 for the shop support bot
      ModelSource:
        CopyFrom: !Sub "arn:aws:bedrock:${AWS::Region}:${AWS::AccountId}:inference-profile/us.anthropic.claude-sonnet-5-5"
      Tags: [{Key: project, Value: shop-support-bot}]

In Terraform, the AWS provider's resource is aws_bedrock_inference_profile:

data "aws_caller_identity" "current" {}

resource "aws_bedrock_inference_profile" "support_bot" {
  name        = "shop-support-bot"
  description = "Claude Sonnet 5.5 for the shop support bot"

  model_source {
    copy_from = "arn:aws:bedrock:us-east-1:${data.aws_caller_identity.current.account_id}:inference-profile/us.anthropic.claude-sonnet-5-5"
  }

  tags = {
    project = "shop-support-bot"
  }
}

Either way, the output you want is the profile's ARN, which you pass to the application as its model ID, and which the IAM policy for that application should name.

Fixing the model ID often uncovers the next problem in line. These are the ones that tend to follow, with AWS's own meaning for each:

ErrorHTTPWhat it means
ValidationException (on-demand throughput)400This page: use an inference profile ID or ARN
AccessDeniedException403Permissions: IAM, an SCP, or expired credentials
FTUFormNotFilled404Submit the Anthropic use-case form for the account
ResourceNotFound404The model or profile ID is wrong, or does not exist in that Region
ThrottlingException429You exceeded an account quota for the model
ServiceUnavailable503Temporary capacity pressure, not your quota; retry with backoff
overloaded_error529The model is temporarily over capacity; retry with backoff and jitter

Cross-Region profiles help with the last two as well, because spreading requests over several Regions is exactly what they are for. AWS's advice for 503 and 529 is retries with exponential backoff and random jitter, honoring any Retry-After header, and avoiding many clients retrying at the same instant.

One more that looks unrelated but is not: if long requests or streaming responses fail with connection resets when your app runs behind a NAT gateway, an interface VPC endpoint or a Network Load Balancer, those have a fixed 350-second idle timeout. AWS's fix needs two settings together: Config(tcp_keepalive=True) on the boto3 client, and a kernel keep-alive below 350 seconds, such as sysctl -w net.ipv4.tcp_keepalive_time=45. Either one alone does nothing.

What about Bedrock Mantle and the OpenAI-compatible endpoint?

Bedrock has a second endpoint, bedrock-mantle, which speaks the OpenAI and Anthropic APIs. It takes plain model IDs, such as openai.gpt-6.1-sol, but only in the Regions where that model is offered in-Region on Mantle, and Mantle does not support cross-Region profiles at all. For GPT-6.1 Sol that means Mantle in us-east-1; for Claude Sonnet 5.5, Mantle in-Region is listed only in AWS GovCloud (US-West). So switching endpoints is not a general workaround for this error. If a model is profile-only on bedrock-runtime in your Region, the profile ID on bedrock-runtime is the fix. Our Mantle vs bedrock-runtime guide has the full comparison.

Provisioned Throughput: the one case where profiles do not apply

If you bought Provisioned Throughput, you call the provisioned model's ARN, not an inference profile. Cross-Region inference profiles do not support Provisioned Throughput, so the two paths never mix: provisioned capacity lives in one Region, and you address it by the ARN Bedrock gave you when you bought it.

That also explains a confusing moment for teams that use both. The same model can work by ARN for the provisioned capacity and fail with "on-demand throughput isn't supported" when someone calls the plain model ID for a quick on-demand test. Nothing is broken; the on-demand test needs the profile ID like everyone else's code.

For steady, predictable volume, AWS points to Provisioned Throughput or the Reserved tier, where a model supports them; each model's page lists which apply. For bursty or modest traffic, on-demand through a profile is simpler, needs no commitment, and is usually cheaper.

Upgrading to a profile-only model without downtime

Ethan's Monday would have been a non-event with this order of work. If you are about to move an app from an older model to a newer one that is profile-only, do it in this sequence:

  1. Read the new model's page and note its profile IDs for your Region, its In-Region support, its burndown rate and any APIs it does not support.
  2. Update IAM first. Add the profile, the model in every destination Region and, for global., the Region-less model ARN. Deploying a policy early breaks nothing.
  3. Check the quota. Look up the Cross-Region tokens-per-minute quota for the new model; it is a separate number from the old model's, so check it rather than assume it carried over.
  4. Check any SCP in AWS Organizations against the profile's destination Regions.
  5. Test from the real Region with one Converse call using the profile ID, from the same role the app uses.
  6. Switch the model ID in configuration, not in code, so rolling back is changing one value.
  7. Watch the first hour of CloudWatch metrics for throttles and errors, and look at one CloudTrail record to confirm where requests are being processed.

None of these steps takes long. Skipping them is what turns a model upgrade into an outage.

Monitoring the switch: metrics, logs and where requests ran

Once the profile is in place, three places tell you whether it is working the way you intended.

CloudWatch metrics. Bedrock publishes runtime metrics per model, including invocation counts, latency, client and server errors and throttles, plus the token counts that drive your quota: InputTokenCount, OutputTokenCount, CacheReadInputTokenCount and CacheWriteInputTokenCount. A spike in throttles right after the switch usually means the cross-Region quota for the new model is smaller than the traffic, not that the profile is wrong.

CloudTrail. Every cross-Region request is recorded in your source Region, and the additionalEventData.inferenceRegion field shows which Region processed it. Checking a few records is the quickest way to prove to yourself, or to an auditor, that a us. profile kept requests in the US.

Application inference profiles. If you need usage and cost per application rather than per model, route each app through its own application profile. AWS designed them for exactly this: tag the profile, and the cost shows up under that tag in Cost Explorer; turn on model invocation logging, and the usage is attributable to that profile. Our Cost Explorer guide shows where those tagged costs appear.

Ethan set up one CloudWatch alarm on throttles for Jake's bot and tagged its application profile with the shop's name. The bill for the bot now has its own line, which made the conversation about whether the upgrade was worth it a short one: it was.

The checklist before you deploy the fix

  1. The model's page confirms which profile IDs exist for your source Region.
  2. Your code reads the profile ID from configuration, not a hard-coded string, so the next model change is a config change.
  3. You chose geographic or global deliberately, and the choice matches what you have told customers about where data is processed.
  4. IAM allows the profile, the model in every destination Region, and the Region-less model ARN if you use global..
  5. Any Region-deny SCP allows the destination Regions, or unspecified for global.
  6. Retries use exponential backoff with jitter for 429, 503 and 529.
  7. If you need per-app cost tracking, an application inference profile is tagged and its ARN is in use.
  8. You have checked the cross-Region quotas for the model in Service Quotas before launch day.

Frequently asked questions

What does "on-demand throughput isn't supported" mean in Bedrock?

It means the model cannot be called by its plain model ID with on-demand pricing in the Region you used. The model is offered only through cross-Region inference profiles there, so you must send a profile ID such as us.anthropic.claude-sonnet-5-5 or its ARN.

How do I fix "Retry your request with the ID or ARN of an inference profile that contains this model"?

Replace the model ID in your request with an inference profile ID, which is the model ID with a prefix such as us., eu., apac., jp., au. or global. You do not need to create anything; AWS defines these profiles.

What is an inference profile in Amazon Bedrock?

An inference profile is a Bedrock resource that defines a model and the Regions requests can be routed to. AWS-defined cross-Region profiles spread requests across Regions; application inference profiles are ones you create to track cost and usage.

How do I find the inference profile ID for a model?

Check the model's detail page in the Bedrock user guide, which lists its geographic and global inference IDs per Region, or run aws bedrock list-inference-profiles --type-equals SYSTEM_DEFINED in the Region you call from.

Why does apac. not work for Claude Sonnet 5.5?

Claude Sonnet 5.5 has no APAC geographic profile on bedrock-runtime. From Asia Pacific Regions such as Tokyo, Mumbai or Sydney, only global.anthropic.claude-sonnet-5-5 is available, which can process requests outside Asia Pacific.

Should I use the us. or the global. profile?

Use a geographic profile such as us. when data must stay in that geography. Use global. when you have no residency requirement and want the lower price; global profiles cost about 10% less than geographic ones.

Does cross-Region inference cost more?

No. There is no extra routing charge. You pay the model's price for the Region you call from, and global profiles are about 10% cheaper than geographic ones.

How do I fix this error in LangChain?

Pass the inference profile ID as the model, for example ChatBedrockConverse(model="us.anthropic.claude-sonnet-5-5", region_name="us-east-1"). If you use an application inference profile ARN, also pass base_model with the plain model ID.

How do I fix this error with boto3?

Set modelId to the profile ID in your converse or invoke_model call, for example modelId="us.anthropic.claude-sonnet-5-5", using a bedrock-runtime client in a source Region that the profile supports.

Why do I get AccessDeniedException after switching to an inference profile?

IAM must allow the inference profile and the foundation model in every destination Region, plus the Region-less model ARN for global profiles. A Region-deny SCP can also block destination Regions.

Do I need to enable other AWS Regions for cross-Region inference?

No. Cross-Region inference can route to Regions that are not enabled in your account. Only an organization policy that denies those Regions, such as an SCP, would block it.

Can I see which Region processed my request?

Yes. CloudTrail logs cross-Region inference requests in your source Region, and the additionalEventData.inferenceRegion field shows the Region that processed each request.

Do inference profiles work with Provisioned Throughput?

No. Inference profiles do not support Provisioned Throughput. With provisioned capacity, call the provisioned model ARN instead.

Does Bedrock Mantle have the same error?

Mantle takes plain model IDs but only in Regions where the model is offered in-Region on Mantle, and it does not support cross-Region profiles. For most profile-only models, the fix is the profile ID on bedrock-runtime.

What is an application inference profile?

An inference profile you create with CreateInferenceProfile, copying a model or a system-defined profile, so you can tag it and track one application's cost and usage. Call it by its ARN.

Why did my code work with an older model but not a newer one?

Older models such as Claude 3 Haiku offer in-Region on-demand access in some Regions, so the plain ID worked. Newer models such as Claude Sonnet 5.5 are offered only through profiles on bedrock-runtime.

What is the us-gov. prefix?

In AWS GovCloud (US) Regions, geographic inference profiles use the us-gov. prefix, for example in Claude Code configurations. Check the model's page for GovCloud support, since not every model is offered there.

How do I get the ARN of an inference profile?

Run aws bedrock get-inference-profile with --inference-profile-identifier set to the profile ID in your source Region; the response includes inferenceProfileArn and the destination model ARNs. For AWS-defined profiles, the ID alone also works in calls.

Which quota does an inference profile use?

Geographic profiles count against the model's Cross-Region tokens per minute quota in your Region, global profiles against its Global cross-Region quota, which is one global limit, and plain model IDs against the On-demand quota. All cover Converse and the other APIs too.

Why does Claude Code on Bedrock show this error?

ANTHROPIC_MODEL is set to a plain model ID. Set it to an inference profile ID such as us.anthropic.claude-sonnet-5-5, or an application inference profile ARN, along with CLAUDE_CODE_USE_BEDROCK=1 and AWS_REGION.

Is cross-Region inference secure?

Requests routed between Regions stay on the AWS network, are encrypted in transit, and never cross the public internet. Each request is logged in CloudTrail in your source Region, with the processing Region recorded. Geographic profiles keep processing inside their geography.

Does the Bedrock console playground need an inference profile too?

Yes, for models that are profile-only in your Region. In the Chat or Text playground, choose the model, then under Inference choose Inference profiles and select one, such as the US profile, before sending a prompt.

This error looks like a wall the first time, and it is really a signpost: the model is there, it just has a different address now. Find the profile for your Region, decide where your data may travel, and give IAM the extra resources it now checks. Jake's support bot has been answering midnight questions on Claude Sonnet 5.5 all week, and the only thing Jake noticed was that its answers got better. Ethan now keeps model IDs in a config file, so the next upgrade is one line in one place.

📌 If you keep one line from this page

"On-demand throughput isn't supported" means "use the inference profile ID", and the right prefix depends on the model and the Region you call from.

Then give IAM the profile and the model in every destination Region, or the next error is AccessDenied.

Revision note. Written October 7, 2026, with the inference profile IDs for Claude Sonnet 5.5, GPT-6.1 Sol and Nova 2 Lite current that day. If a one-word model change took your app down, that happens to careful people every week; it is a one-line fix.

Related