Bedrock Mantle vs bedrock-runtime in AWS: Which Amazon Bedrock Endpoint to Use (Base URLs, IAM, Quotas, Errors)

Logeshwaran
—

Amazon Bedrock has two inference endpoints, and the choice between them decides which APIs, features and permissions you get. bedrock-runtime (bedrock-runtime.REGION.amazonaws.com) is the one AWS recommends for most new applications: it serves Converse, InvokeModel, the OpenAI-compatible Chat Completions and Responses APIs and the Anthropic Messages API, plus Guardrails and cross-Region inference. bedrock-mantle (bedrock-mantle.REGION.api.aws) serves only the OpenAI-compatible and Anthropic Messages APIs, and is the one to use when you need background inference, server-side tools such as web search, Projects or Workspaces, or one of the few models that live only there. Here is the surprise: Mantle is not a different engine. Both endpoints run on the same Mantle inference engine; "bedrock-mantle" is just the name of a second door. And the door has its own lock: a policy that allows every bedrock: action still does not let you call Mantle, because Mantle checks bedrock-mantle:CreateInference instead.

Jake learned about that lock on the busiest Saturday of his year. His phone repair shop has a small chat assistant on its website that answers "is my phone ready yet?" by looking up the repair ticket and replying in plain English. The freelancer who built it used the OpenAI SDK pointed at Amazon Bedrock, and it had worked quietly for months. Then a friend of Ethan's who does security work kindly tidied up Jake's AWS permissions, replacing a broad policy with a careful one that allowed bedrock:InvokeModel and nothing more. On Saturday morning the assistant started answering every customer with an apology. By lunchtime Jake had taken more than sixty phone calls asking the same question, while the repairs those customers were waiting for sat on his bench. "The policy says Bedrock is allowed," he told Ethan that evening. "Why is Bedrock saying no?" Ethan opened the code, saw a hostname ending in api.aws, and smiled. "Because your assistant isn't talking to the Bedrock you allowed. It's talking to Mantle." This page is that evening, in order: what each Bedrock endpoint is, how Mantle and runtime differ line by line, which one to choose, the base URLs and code for both, the IAM permissions, the quotas, the models that live on only one side, and every error that comes from picking the wrong door.

⚡ Quick Answer

• Bedrock Mantle vs bedrock-runtime → same inference engine, two endpoints. Use bedrock-runtime by default; use bedrock-mantle for background inference, server-side tools, Projects, Workspaces or Mantle-only models. Full comparison.

• "bedrock" vs "bedrock-runtime" → bedrock is the control plane (list models, create guardrails); bedrock-runtime runs the models. In boto3, invoke_model exists only on the runtime client. All endpoints.

• Base URLs → OpenAI SDK: https://bedrock-runtime.REGION.amazonaws.com/openai/v1 or https://bedrock-mantle.REGION.api.aws/v1. OpenAI's GPT-5.x and GPT-6 models and Gemma 4 use /openai/v1 on Mantle too. Every URL.

• IAM → runtime needs bedrock:InvokeModel; Mantle needs bedrock-mantle:CreateInference. API keys need CallWithBearerToken under the matching prefix. Policies.

Per-token prices are the same on both endpoints for most models, but quotas are separate pools, so a busy app on one endpoint does not borrow capacity from the other.

If Bedrock itself is new to you, our plain-English guide to Amazon Bedrock covers the idea in ten minutes. This page is for the moment after that, when you are staring at a code sample or an error message and wondering which of several similar hostnames you are supposed to be using.

🧭 NEW HERE? READ THESE FIRST

If the AWS side of this is still fuzzy, these five pages make the rest easy:

📌 Bookmark this; the URL table and the IAM policy below are what people come back for.

The Bedrock endpoint family: bedrock, bedrock-runtime, bedrock-mantle and the rest

"Bedrock" is not one address. It is a small family of endpoints, each with its own job, its own SDK client and sometimes its own IAM prefix. Most confusion comes from mixing them up, so here is the whole family in one table.

EndpointHostname patternWhat it doesboto3 client
bedrockbedrock.REGION.amazonaws.comControl plane: list and describe models, manage guardrails, inference profiles, custom and imported models, Provisioned Throughputboto3.client("bedrock")
bedrock-runtimebedrock-runtime.REGION.amazonaws.comInference: Converse, InvokeModel, plus the OpenAI-compatible and Anthropic Messages APIs on their own pathsboto3.client("bedrock-runtime")
bedrock-mantlebedrock-mantle.REGION.api.awsInference through the OpenAI-compatible APIs and the Anthropic Messages API only, with Mantle-only featuresNone; use the OpenAI or Anthropic SDK, or HTTP
bedrock-agentbedrock-agent.REGION.amazonaws.comBuild time for Bedrock Agents and Knowledge Bases: create agents, knowledge bases and data sourcesboto3.client("bedrock-agent")
bedrock-agent-runtimebedrock-agent-runtime.REGION.amazonaws.comRun time for Agents and Knowledge Bases: invoke an agent, retrieve from a knowledge baseboto3.client("bedrock-agent-runtime")
bedrock-data-automation (and -runtime)bedrock-data-automation.REGION.amazonaws.comBedrock Data Automation: set up projects and blueprints, then process documents, images, audio and videobedrock-data-automation, bedrock-data-automation-runtime
AgentCoreIts own endpointsBedrock AgentCore, the service that runs agents you write yourselfbedrock-agentcore, bedrock-agentcore-control

The pattern behind the names is old and useful: a name without "runtime" is where you set things up, and a name with "runtime" is where you use them. That is why the most common Bedrock mistake in Python looks like this:

import boto3
client = boto3.client("bedrock", region_name="us-east-1")
client.invoke_model(...)   # AttributeError: 'Bedrock' object has no attribute 'invoke_model'

The control-plane client has no invoke_model, because the control plane does not run models. Change the client to boto3.client("bedrock-runtime") and the call works. The reverse mistake, calling list_foundation_models on the runtime client, fails the same way for the same reason.

Mantle is the odd one out in two ways. Its hostname ends in api.aws instead of amazonaws.com, which is the quickest way to spot it in someone else's code, and there is no boto3 client for it, because it speaks the OpenAI and Anthropic API formats rather than AWS's own. You reach it with the OpenAI SDK, the Anthropic SDK or plain HTTP. FIPS variants exist for several of the AWS-style endpoints, such as bedrock-runtime-fips, in a handful of US, Canadian and GovCloud Regions; if you are not required to use FIPS endpoints, you can ignore them.

Bedrock Mantle vs bedrock-runtime: the full comparison

Both endpoints run the same inference engine, the one AWS calls Mantle, built around what it describes as a zero operator access design. The difference is the door: which APIs each endpoint accepts, which Bedrock features sit in front of it, and how it is authorized and metered.

bedrock-runtimebedrock-mantle
AWS's recommendationRecommended for most new applicationsUse when you need something only it has
Converse and InvokeModelYesNo
OpenAI Chat Completions and ResponsesYes, on /openai/v1Yes, usually on /v1
Anthropic Messages APIYes, on /anthropicYes, on /anthropic, but no structured outputs
Cross-Region inference profiles (us., global.)YesNo
Background (asynchronous) inferenceNo; background=true returns a 400 errorYes
Server-side tools and pre-configured tools, including web searchNoYes
Client-side tool useYesYes
Projects (OpenAI) and Workspaces (Anthropic)Default project only; no WorkspacesCreate your own
Guardrails and intelligent prompt routingYesNo
Prompt cachingYesYes, depending on the model
AuthenticationSigV4 or Bedrock API keySigV4 or Bedrock API key (the OpenAI SDK needs an API key)
IAM action for inferencebedrock:InvokeModel (and InvokeModelWithResponseStream for streaming)bedrock-mantle:CreateInference
QuotasPer model: requests per minute and one combined tokens-per-minute quotaPer model: separate input and output tokens per minute; no request limit
Usage attributionIAM principal, request metadata tags, application inference profilesProjects and Workspaces
Model invocation logging to CloudWatch or S3Yes, including the OpenAI-compatible APIs on this endpointNo; Mantle calls are not captured
Anthropic first-time-use formRequired once before invoking ClaudeNot required
Per-token priceStandardThe same for most models; check the model card
RegionsEvery Bedrock Region14 Regions, listed below
Private connectivity (PrivateLink)com.amazonaws.REGION.bedrock-runtimecom.amazonaws.REGION.bedrock-mantle

Read the table as two personalities. bedrock-runtime is the AWS-shaped endpoint: every Bedrock feature you configure in the console, such as Guardrails, cross-Region profiles and cost tags, works there, and it speaks the popular third-party API formats as a convenience. bedrock-mantle is the API-shaped endpoint: it speaks the OpenAI and Anthropic formats natively, including their newer stateful features, but it sits outside most of the Bedrock toolbox.

"So which one was my assistant using?" Jake asked. "Mantle," Ethan said. "Plenty of samples online point the OpenAI SDK at Mantle, and there is nothing wrong with that. It still works fine. It just needs its own permission."

How to tell which Bedrock endpoint your code is using

Before choosing, it helps to know where you are. When someone else wrote the code, as with Jake's assistant, these four checks answer the question in a couple of minutes.

  1. Search the code for the hostname. Anything ending in api.aws, such as bedrock-mantle.us-east-1.api.aws, is Mantle. Anything ending in amazonaws.com is one of the AWS-style endpoints, and the first word tells you which.
  2. Look at how the client is created. boto3.client("bedrock-runtime") or any AWS SDK runtime client means bedrock-runtime, because Mantle has no AWS SDK client. An OpenAI or Anthropic client with a base_url can be either, so read the URL.
  3. Check the environment variables. OPENAI_BASE_URL or ANTHROPIC_BASE_URL often holds the endpoint instead of the code. AWS_BEARER_TOKEN_BEDROCK means a Bedrock API key is in use, which works on both endpoints, so it does not settle the question by itself.
  4. Read the model ID. A us., eu. or global. prefix can only be bedrock-runtime, because Mantle has no cross-Region profiles. A bare ID such as anthropic.claude-sonnet-5 fits either.

Write the answer down in the code itself, as a comment next to the client, and in whatever notes your team keeps. Jake's freelancer now does exactly that, and it is the cheapest fix on this page.

Which Bedrock endpoint should you use? A decision guide

Most people should start with bedrock-runtime and never think about it again. Move to bedrock-mantle only when one of its specific features is the reason. Here is the decision, one question at a time.

  1. Do you use Converse, InvokeModel or boto3? Then you are on bedrock-runtime already, because Mantle does not offer those APIs.
  2. Do you need Guardrails, intelligent prompt routing, or a cross-Region inference profile such as global. or us.? bedrock-runtime. Mantle has none of them.
  3. Is your model available only through cross-Region inference? Newer models such as GLM 5.3, Grok 4.7 and Kimi K3 are, and they are bedrock-runtime only.
  4. Do you need background (long-running) requests, server-side tools, web search, or your own Projects or Workspaces for cost tracking? bedrock-mantle.
  5. Is your model available only on Mantle? A handful are, listed further down. bedrock-mantle.
  6. None of the above, and you just want the OpenAI or Anthropic SDK to work? Either works. bedrock-runtime is the safer default because every Bedrock feature you might add later lives there.

Both endpoints can be used side by side in the same application. A common split is Converse on bedrock-runtime for the main chat, with Guardrails in front, and the Responses API on Mantle for one background job that uses web search. Nothing stops you, as long as the IAM policy covers both and you plan quota for each separately.

Base URLs and model IDs for every API, on both endpoints

This is the table most people need. Replace REGION with your Region code, such as us-east-1.

APIbedrock-runtimebedrock-mantle
OpenAI SDK (Chat Completions, Responses)https://bedrock-runtime.REGION.amazonaws.com/openai/v1https://bedrock-mantle.REGION.api.aws/v1 (GPT-5.x, GPT-6 and Gemma 4 use /openai/v1)
Anthropic SDK (Messages)https://bedrock-runtime.REGION.amazonaws.com/anthropichttps://bedrock-mantle.REGION.api.aws/anthropic
Converse, InvokeModelboto3.client("bedrock-runtime"); AWS builds the URLNot available
List models (GET /models)Not implemented; use the Bedrock control plane insteadSupported
Count tokens for ClaudeSupported for models with Region-specific accessThe count_tokens path; the only option for Claude models that launch with cross-Region inference only

Two traps hide in that table.

The Mantle path is not always /v1. The default for OpenAI-compatible calls on Mantle is /v1, and open-weight models such as gpt-oss use it. But OpenAI's own GPT-5.x and GPT-6 models are served at /openai/v1 on Mantle, and so is Google's Gemma 4, whose model card says it is served at /openai/v1/responses, "not the default /v1/responses." The GPT-6 model card goes further and warns not to use /v1 at all; our GPT-6 on Bedrock guide covers that case. These paths have also moved during 2026, and some tools still ship the older one, so the habit to build is simple: whenever you use a model on Mantle, read the "Programmatic access" section of its model card, because that table is the source of truth for the exact URL and ID. One more path to remember: the list of models you can call is always at https://bedrock-mantle.REGION.api.aws/v1/models, even for models whose inference path is /openai/v1.

Model IDs differ between the doors. On bedrock-runtime, many newer models must be called through a cross-Region inference profile, such as global.anthropic.claude-sonnet-5, and the bare model ID is refused. On Mantle, there are no inference profiles at all, so you send the bare ID, such as anthropic.claude-sonnet-5, and the request is served in the Region you called. Copying an ID from one endpoint's sample to the other is the single most common reason a working request suddenly fails.

The same request on both endpoints: working code

Here is one request, a one-sentence answer from Claude Sonnet 5, sent four ways. Sonnet 5 is available on both endpoints, so the comparison is fair.

Anthropic SDK on bedrock-runtime, with a short-term key

pip install -U anthropic aws-bedrock-token-generator
from anthropic import Anthropic
from aws_bedrock_token_generator import provide_token

client = Anthropic(
    base_url="https://bedrock-runtime.us-east-1.amazonaws.com/anthropic",
    api_key=provide_token(region="us-east-1"),
)
response = client.messages.create(
    model="global.anthropic.claude-sonnet-5",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Explain an API endpoint in one sentence."}],
)
print(response.content[0].text)

The token generator turns your normal AWS credentials into a short-lived Bedrock API key, so no long-term secret sits in a file. Note the model ID: a cross-Region profile, because that is what bedrock-runtime expects for this model.

Anthropic SDK on bedrock-mantle, with an API key

import os
from anthropic import Anthropic

client = Anthropic(
    base_url="https://bedrock-mantle.us-east-1.api.aws/anthropic",
    api_key=os.environ["BEDROCK_API_KEY"],
)
response = client.messages.create(
    model="anthropic.claude-sonnet-5",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Explain an API endpoint in one sentence."}],
)
print(response.content[0].text)

Two lines changed: the base URL and the model ID. Everything else, including the Anthropic SDK itself, is identical. That is the whole promise of Mantle: code written for Anthropic's or OpenAI's own API runs with two edits.

Converse on bedrock-runtime with boto3

import boto3

client = boto3.client("bedrock-runtime", region_name="us-east-1")
response = client.converse(
    modelId="global.anthropic.claude-sonnet-5",
    messages=[{"role": "user", "content": [{"text": "Explain an API endpoint in one sentence."}]}],
    inferenceConfig={"maxTokens": 1024},
)
print(response["output"]["message"]["content"][0]["text"])

Converse uses your normal AWS credentials and the same request shape for every model on Bedrock, which makes it the easiest way to swap models later. It exists only on bedrock-runtime.

OpenAI SDK on either endpoint

import os
from openai import OpenAI

# bedrock-runtime
client = OpenAI(
    base_url="https://bedrock-runtime.us-east-1.amazonaws.com/openai/v1",
    api_key=os.environ["BEDROCK_API_KEY"],
)
# bedrock-mantle: swap in the line below, and use the model ID from the model card
# client = OpenAI(base_url="https://bedrock-mantle.us-east-1.api.aws/v1", api_key=os.environ["BEDROCK_API_KEY"])

resp = client.chat.completions.create(
    model="MODEL_ID_FROM_THE_MODEL_CARD",
    messages=[{"role": "user", "content": "Explain an API endpoint in one sentence."}],
)
print(resp.choices[0].message.content)

Claude models on Bedrock use the Messages API rather than Chat Completions, so for the OpenAI SDK pick a model whose model card lists Chat Completions for your endpoint, such as one of the OpenAI models, and copy its exact ID from that card. Never use your OpenAI API key or OpenAI's own base URL here; those send the request to OpenAI, not to Bedrock.

The same thing with curl, for quick tests

When an SDK hides too much, a raw request shows exactly what each endpoint expects. The OpenAI-compatible paths take the key as a Bearer token; the Anthropic paths take it in an x-api-key header with an anthropic-version header.

curl -X POST "https://bedrock-runtime.us-east-1.amazonaws.com/openai/v1/chat/completions" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $BEDROCK_API_KEY" \
  -d '{"model": "MODEL_ID_FROM_THE_MODEL_CARD", "messages": [{"role": "user", "content": "Hello"}]}'

curl -X POST "https://bedrock-mantle.us-east-1.api.aws/anthropic/v1/messages" \
  -H "x-api-key: $BEDROCK_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{"model": "anthropic.claude-sonnet-5", "max_tokens": 1024, "messages": [{"role": "user", "content": "Hello"}]}'

If curl works and your application does not, the problem is in the application's configuration, usually a stale base URL or a key in the wrong environment variable, not in Bedrock.

"Two lines," Jake said, reading over Ethan's shoulder. "That's all that separates the version that works from the version that cost me Saturday?" "Two lines in the code," Ethan said. "And one line in the IAM policy. That's the next part."

Using Mantle or bedrock-runtime from OpenCode, LiteLLM and other tools

Most coding assistants and gateways pick an endpoint for you, and knowing which one explains a lot of odd behavior.

  • Built-in "Amazon Bedrock" providers in tools such as OpenCode and LiteLLM use bedrock-runtime with your AWS credentials, through the AWS-native APIs. That is why they can use cross-Region profiles and Guardrails, and why Mantle-only models do not appear in them.
  • To use Mantle from those tools, add it as a generic OpenAI-compatible provider instead: base URL https://bedrock-mantle.REGION.api.aws/v1, a Bedrock API key, and the bare model ID. In LiteLLM that means an openai/ model name with api_base set to the Mantle URL; in OpenCode, a custom provider in opencode.json.

A minimal OpenCode custom provider for Mantle looks like this, with the key read from an environment variable:

{
  "$schema": "https://opencode.ai/config.json",
  "provider": {
    "bedrock-mantle": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "Bedrock Mantle",
      "options": {
        "baseURL": "https://bedrock-mantle.us-east-1.api.aws/v1",
        "apiKey": "{env:BEDROCK_API_KEY}"
      },
      "models": {
        "MODEL_ID_FROM_THE_MODEL_CARD": { "name": "My Mantle model" }
      }
    }
  }
}

Our OpenCode guide covers the built-in Bedrock provider and the install. The general rule for any tool: if it asks for AWS credentials and a Region, it is using bedrock-runtime; if it asks for a base URL and an API key, you choose the endpoint by the URL you give it.

Bedrock Mantle IAM: the permission that catches everyone

Each endpoint checks a different IAM action, under a different service prefix. This is the detail that broke Jake's assistant, and it is easy to miss because both prefixes start with the word "bedrock".

What you are doingbedrock-runtimebedrock-mantle
Run inferencebedrock:InvokeModelbedrock-mantle:CreateInference
Stream responsesbedrock:InvokeModelWithResponseStreamNo separate streaming action is listed; CreateInference is the inference permission
Use a Bedrock API keybedrock:CallWithBearerTokenbedrock-mantle:CallWithBearerToken
Responses API stored responsesbedrock:GetInvoke, CancelInvoke, DeleteInvoke on the default project, plus InvokeModel on the projectManaged through Projects
Restrict key typesCondition key bedrock:bearerTokenTypeCondition key bedrock-mantle:bearerTokenType

Because the prefixes differ, a policy with "Action": "bedrock:*" covers every bedrock-runtime and control-plane action and still says nothing about Mantle. If your application calls both endpoints with an API key, the policy needs both sets. A minimal example, before you narrow the resources for production:

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "RuntimeInference",
      "Effect": "Allow",
      "Action": ["bedrock:InvokeModel", "bedrock:InvokeModelWithResponseStream"],
      "Resource": "*"
    },
    {
      "Sid": "MantleInference",
      "Effect": "Allow",
      "Action": ["bedrock-mantle:CreateInference"],
      "Resource": "*"
    },
    {
      "Sid": "ApiKeysOnBothEndpoints",
      "Effect": "Allow",
      "Action": ["bedrock:CallWithBearerToken", "bedrock-mantle:CallWithBearerToken"],
      "Resource": "*"
    }
  ]
}

If you would rather start from an AWS managed policy, AmazonBedrockMantleInferenceAccess grants the published Mantle inference actions; listing and viewing models on Mantle use the matching bedrock-mantle list and get actions. For production, replace the wildcards with the inference profile and foundation model ARNs you actually use on bedrock-runtime, and with your project ARNs on Mantle, which look like arn:aws:bedrock-mantle:REGION:ACCOUNT-ID:project/proj_.... Our explainer on IAM policy JSON walks through each field if this is new.

Bedrock API keys come in two kinds, and the difference matters more than people expect. A short-term key lasts up to 12 hours, or as long as the session that created it, and inherits the permissions of the IAM identity that generated it; AWS recommends these for production. A long-term key lasts until the expiry you set and quietly creates an IAM user with attached policies behind the scenes; AWS recommends these only for exploring. Keys are sent as authorization headers and are not written to CloudTrail, although the API calls themselves are.

Jake's fix took Ethan's friend one minute: add bedrock-mantle:CreateInference and bedrock-mantle:CallWithBearerToken to the careful policy. The assistant was answering customers again before Jake had finished his tea. "I wasn't wrong to tighten it," the friend said, a little embarrassed. "You weren't," Ethan said. "The names just aren't on your side."

Quotas and throttling: two separate pools

The two endpoints are metered differently, and a request that fits easily on one can be throttled on the other.

bedrock-runtime uses classic per-account quotas: for each model, a requests-per-minute limit and a tokens-per-minute limit that counts input and output tokens together. You see and request increases for them in the Service Quotas console.

bedrock-mantle uses two per-model quotas instead: input tokens per minute and output tokens per minute, with no requests-per-minute limit at all. The way the input quota is checked surprises people. When a request arrives, Mantle counts the input tokens plus the value of max_tokens, or the model's maximum if you did not set one, against your input-tokens-per-minute quota. If admitting the request would exceed the quota, you get an HTTP 429 immediately. When the response finishes, the unused part of that reservation is given back.

The practical lesson: on Mantle, set max_tokens to what you actually need. An application that asks for 1,000 input tokens with no max_tokens reserves the model's full output maximum on every request, and can hit 429 errors at a small fraction of its real traffic. Two more details help on Mantle: cached input tokens read through prompt caching do not count against the input quota, and if the output quota runs out mid-response, generation stops and the response comes back with a finish reason that says so.

Quota increases also work differently. bedrock-runtime increases go through Service Quotas as usual. Mantle quotas are visible in Service Quotas, under Amazon Bedrock if you search for "bedrock-mantle endpoint", but increase requests are not processed there; you file them through the AWS Support limit-increase form, naming the endpoint, Region, model, quota and value, with recent usage to justify it. And because the pools are separate, a busy app on one endpoint does not borrow capacity from the other, so plan each one on its own.

When throttling happens, read the error type rather than just the status code. A 429 with a throttling error means a quota was exceeded. A 503 means the service is short of capacity right now, not that you did anything wrong. Some model APIs return 529 for the same kind of temporary overload. For all three, retry with exponential backoff and random jitter, honor any Retry-After header, and cap the attempts. One small trap when you set this up: botocore's total_max_attempts includes the first request, while the OpenAI and Anthropic SDKs' max_retries counts only the retries.

What only Bedrock Mantle can do: background requests, web search, Projects and Workspaces

These are the reasons to choose Mantle on purpose.

Background (long-running) inference

With the Responses API on Mantle, you can send a request with background=true. The call returns at once with a response ID, the model works in the background, and you check back for the result. That suits jobs that take minutes, such as long reports or multi-step agent runs, where holding an HTTP connection open is fragile. On bedrock-runtime, every request is synchronous, and background=true is rejected with a 400 error.

In Python with the OpenAI SDK, the pattern looks like this:

import time
from openai import OpenAI

client = OpenAI(base_url="https://bedrock-mantle.us-east-1.api.aws/v1", api_key="YOUR_BEDROCK_API_KEY")
resp = client.responses.create(model="MODEL_ID_FROM_THE_MODEL_CARD",
                               input="Summarize these 40 support tickets into five themes: ...",
                               background=True)
while resp.status in ("queued", "in_progress"):
    time.sleep(5)
    resp = client.responses.retrieve(resp.id)
print(resp.status, resp.output_text)

Poll gently, every few seconds rather than in a tight loop, and remember that the stored response lives in the Region that served it, so retrieve it from the same Regional URL.

Server-side tools and web search

On Mantle, some tools run inside Bedrock instead of in your code. The best-known is Web Search: when it is enabled, the model decides whether it needs current information, searches an index that Amazon builds and maintains, and can fetch page content, with citations in the answer. It is available through the Responses API for OpenAI's GPT models on Mantle; in the commercial US Regions that means the GPT-5.6 family and the earlier GPT-5.4 and GPT-5.5.

Web Search has one setting worth knowing before your first call. external_web_access defaults to true, which lets Fetch go to the live web when a page is not in Bedrock's cache, but only if the caller also holds the bedrock-websearch:ExternalWebAccess permission. The AWS managed policy AmazonBedrockFullAccess does not include it. So with that policy and the default setting, every Fetch quietly fails its authorization check, and answers rely on search snippets only. Set external_web_access to false to keep retrieval inside AWS and avoid the silent failure, or grant the permission if you really want live fetches.

Projects and Workspaces

A Project is a named boundary for one application, team or environment, with its own IAM access rules and its own cost tracking through tags in Cost Explorer. Projects are for the OpenAI-compatible APIs; Workspaces are the same underlying resource used with the Anthropic Messages API. Every account has a default project. On Mantle you can create more, with IDs that start with proj_, and attach a request to one with the OpenAI-Project header (or OpenAI(project="proj_...") in the SDK) or, for Workspaces, the anthropic-workspace-id header. On bedrock-runtime only the default project exists; the equivalent tool there is the application inference profile.

Stored conversations

Both endpoints store Responses API conversations by default, keeping the input and output for 30 days so you can continue a conversation with previous_response_id. On both, a stored response belongs to the Region that served it, so follow-ups must go to that same Region.

What only bedrock-runtime can do: Guardrails, cross-Region inference and structured outputs

And these are the reasons to stay on bedrock-runtime.

  • Guardrails. Content filters, denied topics, PII masking and the rest of Bedrock Guardrails apply only on bedrock-runtime. If your application must filter what users send or what the model says, that decides it.
  • Cross-Region inference. The us., eu. and global. inference profiles that route a request to whichever Region has capacity exist only on bedrock-runtime. On Mantle, a request is served in the Region you call.
  • Converse and InvokeModel. The AWS-native APIs, which work with your normal AWS credentials through boto3 and every other AWS SDK.
  • Structured outputs for Claude. The Messages API on Mantle rejects output_config.format with a 400 error. For JSON that must match a schema with Claude models, use Converse or InvokeModel on bedrock-runtime.
  • Intelligent prompt routing, which sends each prompt to a cheaper or stronger model in the same family, and application inference profiles and request metadata tags for cost tracking, with one exception: the Responses API on bedrock-runtime attributes usage by IAM principal only, and rejects a request that names an application inference profile with a 400 error.

If you are not sure whether you will need any of these later, that uncertainty is itself the argument for bedrock-runtime. Moving from runtime to Mantle when you need a Mantle feature is easy; discovering you need Guardrails after building on Mantle means moving the other way.

Which models are on which endpoint

AWS publishes an endpoint availability table for every model. As of October 6, 2026 it lists 134 models: 54 on both endpoints, 71 on bedrock-runtime only, and 9 on bedrock-mantle only.

The nine Mantle-only models are Google's Gemma 4 31B, Gemma 4 26B-A4B and Gemma 4 E2B; OpenAI's GPT-5.4 and GPT-5.5 and two Daybreak variants of GPT-5.6; xAI's Grok 4.3; and Anthropic's Claude Mythos 5. If you want any of those, Mantle is the only door.

The runtime-only group includes every Amazon Nova and Titan model, older Claude versions such as Claude 3.5 Haiku and the Claude Sonnet 4 line, and the newest cross-Region-only models: GLM 5.3, Grok 4.7 and Kimi K3. That last group makes sense once you know the rule from the comparison table: Mantle has no cross-Region profiles, so a model offered only through cross-Region inference cannot be on Mantle. Our guides to GLM 5.3 on Bedrock and Grok 4.7 on Bedrock show what that looks like in practice.

The models on both endpoints include the current Claude models (Sonnet 5.5, Opus 5.5, Fable 5.1, Sonnet 5, Opus 5, Haiku 4.5), OpenAI's GPT-6 and GPT-5.6 families and gpt-oss, plus Qwen, Mistral, DeepSeek, MiniMax, NVIDIA Nemotron and GLM 5.

Here is the catch that the table cannot show: "available on Mantle" does not mean available on Mantle in your Region. Claude Sonnet 5.5 is marked as supported on Mantle, and its model card shows exactly one Mantle Region: AWS GovCloud (US-West). OpenAI's GPT-6 Sol is on Mantle in US East (N. Virginia) only. On bedrock-runtime, the same models reach dozens of Regions through cross-Region profiles. Always check the model card's Regional availability table for the endpoint you plan to use.

Mantle itself runs in 14 Regions as of this writing:

Geographybedrock-mantle Regions
United Statesus-east-1 (N. Virginia), us-east-2 (Ohio), us-west-2 (Oregon), us-gov-west-1 (GovCloud US-West)
Europeeu-central-1 (Frankfurt), eu-west-1 (Ireland), eu-west-2 (London), eu-south-1 (Milan), eu-north-1 (Stockholm)
Asia Pacificap-northeast-1 (Tokyo), ap-south-1 (Mumbai), ap-southeast-2 (Sydney), ap-southeast-3 (Jakarta)
South Americasa-east-1 (SΓ£o Paulo)

The hostname is always bedrock-mantle.REGION.api.aws, for example bedrock-mantle.eu-west-2.api.aws for London. If you call Mantle in a Region outside that list, the hostname does not exist and you get a connection error rather than an AWS error.

Bedrock Mantle pricing: the same per token, with model exceptions

For the same model, the per-token price is the same on bedrock-runtime and bedrock-mantle; AWS's own guidance is to choose the endpoint by the APIs and features you need, not by cost. The full picture of what you pay for is in our Bedrock pricing guide.

There are exceptions, and they live in the model cards. OpenAI's GPT-6 models are the clearest one: their pricing table lists a "Mantle in-Region" row with a 10% premium, the same premium as US-only cross-Region routing, while global cross-Region routing on bedrock-runtime uses the base rate with no premium. For GPT-6 Sol that is $2.00 in and $10.00 out per million tokens globally, against $2.20 and $11.00 on Mantle. If a model's card lists separate Mantle prices, those win over the general rule.

Indirect costs can matter more than the token price. Projects and Workspaces on Mantle give you cost tracking by application; on bedrock-runtime you get the same with application inference profiles and tags. If you call either endpoint from inside a VPC through a NAT gateway, the data processing charge on the NAT gateway can be a surprise line on the bill, which is what the next section fixes.

Private networking: VPC endpoints for Mantle and runtime

Both endpoints support AWS PrivateLink, so instances in private subnets can reach Bedrock without an internet gateway or NAT device. Each endpoint has its own interface endpoint service name:

  • com.amazonaws.REGION.bedrock-runtime for bedrock-runtime
  • com.amazonaws.REGION.bedrock-mantle for bedrock-mantle
  • com.amazonaws.REGION.bedrock for the control plane
  • com.amazonaws.REGION.bedrock-agent and com.amazonaws.REGION.bedrock-agent-runtime for Agents and Knowledge Bases
  • bedrock-fips and bedrock-runtime-fips variants in a few US, Canadian and GovCloud Regions

Turn on private DNS when you create each endpoint, and your existing code keeps using the normal public hostnames, such as bedrock-mantle.REGION.api.aws, while the traffic stays inside the AWS network. If your application uses both inference endpoints, create both interface endpoints; one does not cover the other. Keeping this traffic off a NAT gateway also removes its per-gigabyte processing charge, which is the problem our NAT gateway bill guide digs into.

Bedrock Mantle and bedrock-runtime errors, and the fix

If you are here with an error on screen, take a breath: nearly every one of these comes from a mismatch between the endpoint, the API, the model ID and the permission. Find yours below.

AttributeError: 'Bedrock' object has no attribute 'invoke_model'

You created the control-plane client. Use boto3.client("bedrock-runtime") for inference.

AccessDeniedException on Mantle, although the policy allows bedrock:InvokeModel

Mantle checks bedrock-mantle:CreateInference, and bedrock-mantle:CallWithBearerToken if you use an API key. Add them; bedrock:* does not cover them. Our Bedrock AccessDeniedException guide covers the runtime-side causes.

AccessDeniedException on bedrock-runtime with an API key

The identity needs bedrock:CallWithBearerToken, plus bedrock:InvokeModel on both the inference profile and the foundation model it routes to.

400 error: background=true on bedrock-runtime

Background requests exist only on Mantle. Move that call to https://bedrock-mantle.REGION.api.aws, or make the request synchronous.

400 error: output_config.format on Mantle

Structured outputs are not supported on Mantle's Messages API. Use Converse or InvokeModel on bedrock-runtime for schema-constrained JSON from Claude.

400 error: application inference profile with the Responses API on bedrock-runtime

That API on runtime attributes usage by IAM principal only. Send a model ID or a system-defined profile such as global., or use Projects on Mantle for per-application tracking.

"Invocation with on-demand throughput isn't supported" on bedrock-runtime

The model needs a cross-Region inference profile. Use its us. or global. ID instead of the bare model ID.

Model not found or not supported on Mantle

Either the model is runtime-only, such as GLM 5.3 or Kimi K3, or it is on Mantle only in another Region, such as Sonnet 5.5 in GovCloud. Check the model card's endpoint and Region tables, and send the bare model ID, never a us. or global. profile, to Mantle.

client.models.list() fails on bedrock-runtime

bedrock-runtime does not implement the OpenAI GET /models operation. List models with the control-plane API, or call /v1/models on Mantle, where it is supported.

404 or path errors on Mantle

Check the base path for your model. Open-weight models such as gpt-oss use /v1; OpenAI's GPT-5.x and GPT-6 models and Gemma 4 use /openai/v1. The model card's Programmatic access table is the source of truth.

Could not connect to the endpoint URL

The hostname is wrong for the Region: either Mantle is not offered there, or the Region code has a typo. Mantle hostnames end in .api.aws, runtime hostnames in .amazonaws.com.

HTTP 429 on Mantle at low traffic

Each request reserves its input tokens plus max_tokens against the input quota. Set max_tokens to a realistic value, use prompt caching, and request an increase through the AWS Support form if you still need more.

Requests go to OpenAI instead of Bedrock

The SDK is using OpenAI's own base URL or key. Set OPENAI_BASE_URL to a Bedrock URL and OPENAI_API_KEY to a Bedrock API key, and remove any OpenAI key from the environment.

Real Bedrock Mantle errors developers hit, with the exact messages

The list above covers the mistakes the documentation predicts. This one covers what developers have actually run into on Mantle in 2026, word for word, so you can match the message on your screen. Where a fix comes from other developers' experience rather than from AWS, it says so.

"is not available for this account. You can explore other available models on Amazon Bedrock. For additional access options, contact AWS Sales."

This is the most common one, and the most frustrating, because everything else checks out. On Mantle it arrives as an HTTP 401 with the code access_denied; on bedrock-runtime the same sentence comes as a 403 AccessDeniedException. It names one model, such as openai.gpt-5.6-terra or anthropic.claude-sonnet-5, while other models work with the same key, Region and code.

It is an account-level access decision, not an IAM problem. For some of the newest models, AWS opens access gradually; its own launch post for one recent Claude model said access was expanding to all accounts "depending on your Bedrock usage," with AWS Support as the fast route. To confirm you are in this situation rather than a configuration mistake:

  1. Run a control test. Call a model you know works, such as an open-weight gpt-oss model on Mantle, with the same key and Region. If it answers, your key, endpoint and path are fine.
  2. Check availability from the control plane. aws bedrock get-foundation-model-availability --model-id MODEL_ID reports agreement, authorization, entitlement and Region status. Developers have reported every field showing as available while invocation still fails, so treat an all-clear here as "not your setup," not as proof of access.
  3. For Claude on bedrock-runtime, make sure Anthropic's first-time-use form has been submitted for the account or organization. That form is not required on Mantle.
  4. Open an AWS Support case naming the model, Region, endpoint and a request ID from the error. Several developers report waiting days for a reply on basic support, so start the case early and keep a working model as a fallback meanwhile.

"Too many tokens per day, please wait before trying again"

Developers on brand-new accounts have reported this on the very first call of the day, for every model. The Service Quotas console tells the story: the applied account-level value is 0 while the AWS default shows millions of tokens. Only AWS Support can raise a quota that is set to zero, so open a case with a short description of your use case and expected volume. A common explanation in those threads is that new accounts start with conservative limits until they have some billing history; a valid payment method on the account does not hurt either. One widely shared answer on that thread claimed that bedrock-mantle "represents cross-region inference profiles." It does not: Mantle has no cross-Region profiles at all, and its quotas are simply a separate pool from bedrock-runtime's.

404 "The model 'openai.gpt-5.6-terra' does not exist"

Mantle returns this when the model exists but is not served in the Region you called. One developer reported exactly this split: the same request returned 404 from eu-west-1 and 200 from us-east-1, for several GPT-5.x models and Grok 4.3, while gpt-oss worked in both. Tools that build the Mantle URL from your AWS_REGION make this easy to hit without noticing. Call GET https://bedrock-mantle.REGION.api.aws/v1/models in that Region to see what it actually serves, and pin the Region in your configuration.

404 "The model 'anthropic-claude-opus-4-7' does not exist"

Look closely at the name: the dots have become dashes. Some tools "normalize" model names for Anthropic-style endpoints, which works on Anthropic's own API but breaks on Bedrock, where IDs such as anthropic.claude-opus-4-7 keep their dots. One agent framework had exactly this bug for Mantle URLs and fixed it in an update. Update the tool, or find its setting that preserves model names as written.

"Missing 'authorization' or 'x-api-key' header"

The full error looks like this:

{"error":{"code":"invalid_api_key","message":"Missing 'authorization' or 'x-api-key' header","type":"permission_denied_error"}}

Mantle received a request with no credentials at all. It is typical of proxies and gateways running on an EC2 or container IAM role that sign bedrock-runtime calls automatically but send Mantle calls as plain HTTP. Either give the tool a Bedrock API key, ideally a short-term one generated from the role with aws-bedrock-token-generator, or use a version of the tool that signs Mantle requests with SigV4. Remember the key also needs bedrock-mantle:CallWithBearerToken.

"Only S3 URLs are supported for file_url."

Reported with code validation_error when a Responses API request on Mantle passed a pre-signed HTTPS link in an input_file item. In that report, Mantle wanted an s3:// URI for file_url. Use the S3 URI of an object the caller can read, and check the model card for which input types and file sizes the model accepts.

A background request stays "in_progress" forever

Developers have reported background requests on Mantle that were accepted, returned an ID and never finished, notably with Gemma 4 models and video or file inputs. A request the model cannot process may sit in in_progress instead of failing cleanly. Before blaming the endpoint, confirm the model accepts that input type and size on its model card, and that you are using the API path the card lists, since Gemma 4 moved to the Responses API on /openai/v1. Always give background jobs a client-side deadline, then cancel and retry, so a stuck job cannot stall your application.

OpenAI .NET SDK: System.FormatException when reading a Mantle response

A developer using the official OpenAI .NET SDK with the Responses API on Mantle reported that successful responses could not be read, because Mantle returned created_at as a floating-point number in scientific notation, such as 1.772195752E9, where the SDK expects an integer. The same code worked with Chat Completions on Mantle. Until that changes, use Chat Completions from .NET, handle the field yourself, or call Mantle from a language whose SDK accepts the format.

Nothing from Mantle appears in model invocation logs

This one is by design. Bedrock model invocation logging captures calls made through bedrock-runtime only, including the OpenAI-compatible APIs on that endpoint; calls through Mantle are not captured. Mantle does publish CloudWatch metrics under the AWS/BedrockMantle namespace, with token totals by account, project and model, but developers have reported those metrics missing for some models outside the OpenAI and Anthropic families, such as Grok 4.3 and Gemma 4. If you need a full audit trail of prompts and responses, log them in your own application or use bedrock-runtime. For cost alarms, AWS Budgets works regardless of endpoint.

"So half of these aren't even mistakes," Jake said. "Half of them are the platform being new," Ethan said. "That's why you keep a second model that works, and why you write down which door you're using. When something odd happens, you'll know in a minute whether it's you or them."

What Jake's shop chose, and why

Once the assistant was working again, Ethan asked the question that matters more than any table: should it stay on Mantle at all? They went through it the way this page suggests.

The assistant used the OpenAI SDK and one model that is available on both endpoints, so either door would work. It did not need background requests, web search or Projects; it answers a short question and looks up a ticket. On the other side, Jake had one thing he wanted that only bedrock-runtime offers: Guardrails, so the assistant would refuse politely when people asked it for legal or medical advice instead of repair updates, which had happened twice that month.

So they moved it. The freelancer changed the base URL to the runtime's /openai/v1 path, swapped the model ID for its cross-Region profile, and added a guardrail. The IAM policy got simpler too: back to bedrock: actions only, with the Mantle lines kept but marked as unused, in case a future feature needs them. The whole change took an afternoon, most of it testing.

The lesson Ethan drew was not "Mantle is worse." It was that the endpoint is a decision, and it is worth making on purpose. A small business with one assistant needs the endpoint that carries the features it will actually use. For Jake, that was Guardrails. For a team running long research jobs with web search, it would be Mantle.

Moving between Mantle and bedrock-runtime

Whichever direction you go, the change is small if you work through it in this order.

  1. Confirm the model is on the target endpoint in your Region, using the model card.
  2. Change the base URL: /openai/v1 or /anthropic on runtime, /v1 or /openai/v1 (per the model card) or /anthropic on Mantle.
  3. Change the model ID: a us. or global. profile where runtime requires one; the bare ID on Mantle.
  4. Update IAM: bedrock: actions for runtime, bedrock-mantle: actions for Mantle, and the matching CallWithBearerToken for API keys.
  5. Check features: background requests, server-side tools and Projects do not exist on runtime; Guardrails, structured outputs for Claude and cross-Region profiles do not exist on Mantle.
  6. Plan quotas separately, because the two endpoints do not share capacity, and test under realistic load before you switch traffic.

Existing Mantle applications are fully supported and do not have to move. Move only when a feature on the other side is worth it.

Bedrock endpoint terms, in plain English

If some of the words on this page were new, here they are in one place, each in a sentence or two.

  • Endpoint: the web address your code sends requests to. Bedrock has several, one per job.
  • Control plane: the part of a service where you set things up, such as listing models or creating a guardrail. For Bedrock, that is the bedrock endpoint.
  • Runtime (data plane): the part that does the work, here running the model on your prompt. For Bedrock, that is bedrock-runtime, and in a different shape, bedrock-mantle.
  • Inference: the act of sending a prompt to a model and getting an answer back. Every call on this page is an inference call.
  • Inference profile: an ID such as global.anthropic.claude-sonnet-5 that lets Bedrock route your request to another Region with spare capacity. Only bedrock-runtime understands them.
  • SigV4: the way AWS SDKs sign each request with your AWS credentials. You rarely see it; boto3 does it for you.
  • Bedrock API key: a single string you put in a header instead of signing requests, so the OpenAI and Anthropic SDKs can talk to Bedrock. Short-term keys expire within 12 hours; long-term keys last until the date you set.
  • Project and Workspace: named boundaries on Mantle for one application or team, used for access control and cost tracking. Projects are the OpenAI-side name; Workspaces are the Anthropic-side name for the same thing.
  • PrivateLink (interface endpoint): a private connection from your own network to an AWS service, so traffic never crosses the public internet.

Frequently asked questions

What is Bedrock Mantle?

Bedrock Mantle is one of Amazon Bedrock's two inference endpoints, at bedrock-mantle.REGION.api.aws. It serves the OpenAI-compatible Responses and Chat Completions APIs and the Anthropic Messages API, with extras such as background requests, server-side tools, Projects and Workspaces. Mantle is also the name of the inference engine behind both endpoints.

What is the difference between Bedrock Mantle and bedrock-runtime?

They run the same engine but offer different doors. bedrock-runtime adds Converse, InvokeModel, Guardrails and cross-Region inference, and is recommended for most new apps. Mantle offers only the OpenAI and Anthropic API formats, plus background inference, server-side tools, web search, Projects and Workspaces.

What is the difference between bedrock and bedrock-runtime?

bedrock is the control plane, used to list models and manage resources such as guardrails and inference profiles. bedrock-runtime runs the models. In boto3, invoke_model and converse exist only on the bedrock-runtime client.

What is the Bedrock Mantle endpoint URL?

https://bedrock-mantle.REGION.api.aws. For the OpenAI SDK, the base URL is https://bedrock-mantle.REGION.api.aws/v1 for open-weight models such as gpt-oss and https://bedrock-mantle.REGION.api.aws/openai/v1 for OpenAI's GPT-5.x and GPT-6 models and Gemma 4; for the Anthropic SDK, https://bedrock-mantle.REGION.api.aws/anthropic.

Is Bedrock Mantle more expensive than bedrock-runtime?

Not usually. Per-token pricing for the same model is the same on both endpoints. Some model cards list exceptions; GPT-6 models, for example, carry a 10% premium for Mantle in-Region use, the same as US-only cross-Region routing.

What IAM permission does Bedrock Mantle need?

bedrock-mantle:CreateInference for inference, and bedrock-mantle:CallWithBearerToken when you use a Bedrock API key. These are separate from the bedrock: actions, so a bedrock:* policy does not cover Mantle.

Does Bedrock Mantle support Guardrails?

No. Guardrails and intelligent prompt routing are available only on bedrock-runtime. If your application needs content filtering through Bedrock Guardrails, use bedrock-runtime.

Does Bedrock Mantle support cross-Region inference?

No. Mantle serves each request in the Region you call, with bare model IDs. Cross-Region profiles such as us. and global. work only on bedrock-runtime, which is why cross-Region-only models such as GLM 5.3 are not on Mantle.

Which models are only on Bedrock Mantle?

As of October 6, 2026: Gemma 4 31B, Gemma 4 26B-A4B and Gemma 4 E2B, GPT-5.4, GPT-5.5, two Daybreak variants of GPT-5.6, Grok 4.3 and Claude Mythos 5. Check AWS's endpoint availability table for changes.

Can I use the OpenAI SDK with Amazon Bedrock?

Yes, on either endpoint. Set the base URL to https://bedrock-runtime.REGION.amazonaws.com/openai/v1 or the Mantle URL, use a Bedrock API key as the key, and use a Bedrock model ID that supports the OpenAI-compatible APIs.

Can I use the Anthropic SDK with Amazon Bedrock?

Yes. Point it at https://bedrock-runtime.REGION.amazonaws.com/anthropic or https://bedrock-mantle.REGION.api.aws/anthropic and use a Bedrock API key. Use a cross-Region profile ID on runtime and the bare model ID on Mantle.

Why does client.models.list() fail on Bedrock?

bedrock-runtime does not implement the OpenAI list-models operation. Use the Bedrock control-plane API to list models, or call /v1/models on the Mantle endpoint, which supports it.

How do I increase Bedrock Mantle quotas?

Mantle quotas appear in Service Quotas under Amazon Bedrock, but increases are requested through the AWS Support limit-increase form. Name the endpoint, Region, model, quota (input or output tokens per minute) and the value, with recent usage.

What is bedrock-agent-runtime?

The run-time endpoint for Bedrock Agents and Knowledge Bases: it invokes agents and retrieves from knowledge bases. Its build-time partner, bedrock-agent, creates them. Both are separate from model inference on bedrock-runtime.

Is Bedrock AgentCore the same as bedrock-agent-runtime?

No. AgentCore is a separate service for running agents you write yourself, with its own clients, bedrock-agentcore and bedrock-agentcore-control. bedrock-agent-runtime serves Knowledge Bases and the older console-built Bedrock Agents, now called Bedrock Agents Classic.

Does bedrock-runtime support the Responses API?

Yes, on the /openai/v1 path, with differences: requests are always synchronous, server-side tools are unavailable, only the default project exists, and application inference profiles are rejected. Stored conversations work normally.

Can I use Bedrock Mantle from a private VPC?

Yes. Create an interface VPC endpoint with the service name com.amazonaws.REGION.bedrock-mantle and enable private DNS. bedrock-runtime needs its own endpoint, com.amazonaws.REGION.bedrock-runtime.

Why does Bedrock say a model "is not available for this account"?

Your account does not yet have access to that specific model, even if IAM is correct. AWS opens some new models gradually. Confirm with a control test on a model that works, submit Anthropic's first-time-use form if you are calling Claude on bedrock-runtime, then open an AWS Support case.

Why are Bedrock Mantle calls missing from CloudWatch logs?

Model invocation logging captures bedrock-runtime calls only. Mantle calls are not logged there. Use the AWS/BedrockMantle CloudWatch metrics, your own application logging and AWS Budgets alerts, or move the calls to bedrock-runtime if you need full prompt and response logs.

Do I need the Anthropic first-time-use form on Bedrock Mantle?

No. The first-time-use form is required before invoking Anthropic models through bedrock-runtime, once per account or organization, but it does not apply to Anthropic models accessed through the bedrock-mantle endpoint.

Two doors, one engine, and a lock on each door with a slightly different name. That is all bedrock-mantle and bedrock-runtime really are, and once you see it that way, most of the confusing errors turn into a quick check of four things: the hostname, the path, the model ID and the IAM prefix. Jake's assistant now runs on bedrock-runtime with a guardrail in front of it, and his freelancer added a note to the code that says, in capital letters, which endpoint it talks to. On the next busy Saturday, the phone rang for repairs, not for status questions. If you are fixing one of these errors right now, you are one or two lines away from done.

📌 If you keep one line from this page

Same engine, two doors: bedrock-runtime checks bedrock: permissions, Mantle checks bedrock-mantle: permissions.

Start on bedrock-runtime, and move a call to Mantle only when you need something only Mantle has.

Revision note. Written October 6, 2026, from the two Bedrock endpoints as they stand today. If an AccessDenied brought you here, you did nothing wrong; the permission names really are that easy to mix up, and the fix is one line.

Related