GLM 5.3 on Amazon Bedrock: Model ID, Pricing, Who Can Use It, Regions and Code
Yes, GLM is on Amazon Bedrock, and the newest one arrived this week. Z.ai's GLM 5.3 joined Bedrock on October 5, 2026. You call it as global.zai.glm-5.3 or us.zai.glm-5.3, it costs $1.68 per million input tokens and $5.28 per million output tokens with global routing, and it keeps the full 1-million-token context with up to 128K tokens of output. Here is the part the launch headlines skip: "available on Bedrock" does not mean available to you. AWS limits GLM 5.3 to eligible enterprise customers, so a small-business or personal AWS account may not be able to call it at all. And if your account can, you pay about 20% more per token than Z.ai charges on its own API for the same model.
Jake runs a phone repair shop, and a freelancer built him a small booking site that runs on his own AWS account. Last December a hole in the booking form let someone fill forty repair slots with fake names, and Jake lost a Saturday of walk-in customers he could not fit in, about $600 he still thinks about. So when he read that GLM 5.3, the open model everyone was calling the strongest for coding, had landed on Bedrock, he had one plan for his evening: have it read the booking site's code and point out anything fragile before the holiday rush. Twenty minutes later he called Ethan. "The AWS page says it's available. My console says nothing. What am I doing wrong?" Ethan laughed, not unkindly. "Nothing. You're just not on the guest list yet. Let me show you what is." This page is that conversation, written down: who can use GLM 5.3 on Bedrock, the exact model IDs, the real prices next to Z.ai's, working code for every API, the Regions, the features that are missing, how to use it from a coding tool, and every error you are likely to meet.
New to Bedrock itself? The plain-English guide to Amazon Bedrock covers the idea in ten minutes: one AWS service that rents you models from many companies, billed on your AWS account and governed by your IAM rules. GLM 5.3 arriving on that shelf is the news. Most of the surprises below come from the shelf's rules, not from the model, and once you know them they stop being surprises.
What GLM 5.3 is, and what Bedrock adds
GLM 5.3 is the flagship model from Z.ai, the Beijing company also known as Zhipu AI. It is built for coding and long agentic work: refactoring a repository that spans hundreds of files, running a task that takes many tool calls over hours, or reading a large codebase in one go. It reads and writes text only. There is no image input on Bedrock and no image, audio or video output.
Under the hood it is a mixture-of-experts model with 744 billion parameters in total and roughly 40 billion active for each token. In plain terms, it is a huge library where only a small team of specialists works on each word, which is why it can be smart without being impossibly slow. GLM 5.3 uses the same base model as GLM 5.2. Every gain came from post-training, the stage where a model is taught to follow instructions, use tools and reason, rather than from a bigger or newer base.
If you read AWS's launch announcement you may have seen 753B instead of 744B, and wondered which is right. Both describe the same model. 753 billion is what Hugging Face reports when it counts every tensor stored in the weight files; Z.ai's model card and the Bedrock model card both use 744 billion. Nothing is different about the model you call.
The headline numbers for Bedrock:
- Context window: 1 million tokens, roughly 750,000 words of input, enough for a mid-sized codebase or a long contract archive in one request.
- Maximum output: 128K tokens per response.
- Reasoning: always on, with selectable effort levels, so you trade speed and output tokens against quality. The reasoning section below has the values.
- Tools: function calling, structured JSON output and streaming, including streamed reasoning content.
- APIs: the OpenAI-compatible Responses and Chat Completions APIs, plus Bedrock's own Converse and InvokeModel.
- Prompt caching: automatic (implicit) by default, with explicit cache points when you want control.
What changed from GLM 5.2, in the order most teams will feel it. Z.ai reports a 50% improvement on its in-house coding benchmark, and its public scores jumped most on long, hard engineering tasks: Terminal-Bench 3.0 went from 4.6 to 28.3, and DeepSWE from 46.2 to 66.9. Z.ai also reports the highest score in its comparison table on CyberGym, a benchmark for finding bugs in real software, which is part of why the model got so much attention this autumn.
GLM 5.3 is also an open-weight model. Anyone can download the weights from Hugging Face and run them, which is the subject of our GLM-5.3 local installation guide. The license is MIT-style, with one unusual condition: a company that sells model access as a service and earns more than $10 billion a year must pass a security review by Z.ai before using the model commercially. For almost every reader that clause will never apply. It matters only because it explains why the largest cloud providers take a little longer to offer a GLM model than the smaller hosts do.
Bedrock's GLM shelf now has four models. Knowing them saves you from searching for versions that are not there:
| Model on Bedrock | Model ID | Context | Status in October 2026 |
|---|---|---|---|
| GLM 5.3 | zai.glm-5.3 (call it through us. or global.) | 1M | Active since October 5, 2026; eligible customers only |
| GLM 5 | zai.glm-5 | 200K | Active; open to all accounts in 11 Regions |
| GLM 4.7 | zai.glm-4.7 | 203K | Legacy; end of life no sooner than December 22, 2026 |
| GLM 4.7 Flash | zai.glm-4.7-flash | 203K | Active; the budget option at $0.07 / $0.40 in US Regions |
Notice what is missing. GLM 5.1 and GLM 5.2 never came to Bedrock, and neither did GLM-5.3-Flash. If you searched for "bedrock glm 5.1" or "is glm 5.2 on aws bedrock" and found nothing, that is why: AWS went straight from GLM 5 to GLM 5.3.
Who can use GLM 5.3 on Bedrock: the eligibility catch
This is the section most pages leave out, and the one that cost Jake his first twenty minutes. The Bedrock model card for GLM 5.3 says access is limited to eligible customers, and AWS's launch post describes it as available to eligible enterprise customers. Every other detail on the card, the IDs, the Regions, the prices, applies only once your account is on that list.
"So what makes an account eligible?" Jake asked. Ethan shrugged. "AWS doesn't publish a checklist. The card just says to contact your AWS account team. If you've never heard of your account team, that's your answer for now." Big companies with an enterprise agreement have a named account manager at AWS. A phone shop paying a few dollars a month for a booking site does not, and that is not a mark against you. It is simply how AWS rolls out a model it is being careful with.
How to tell where you stand, without guessing:
- Open the Bedrock console in a US Region, such as US East (N. Virginia), and go to the model catalog. Filter by the provider Z AI.
- Open GLM 5.3's card and try the playground. In the console, choose Test and then Playground, pick GLM 5.3, and send one short prompt.
- If the playground answers, your account is in. Anything that fails later is your code or your IAM policy, and the errors section below covers both.
- If GLM 5.3 is missing or the playground refuses while other models such as GLM 5 answer normally, eligibility is the reason. No IAM change on your side will fix it. Ask your AWS account team, or use one of the three routes below.
No account team? You can still ask. AWS runs a Connect with the Amazon Bedrock team form, the same contact link AWS's GLM 5.3 launch post points readers to. Keep the request short and specific, because a clear request is easier to route:
- The model you want: GLM 5.3 (
zai.glm-5.3), and which profile, US or global. - Your AWS account ID and the Region you will call from.
- What you will use it for, in one or two sentences, for example "code review for our booking app."
- A rough monthly volume, even a guess such as "about 5 million tokens a month."
There is no published timeline and no guarantee of a yes, so do not pause your project waiting for the reply. Build on GLM 5 or Z.ai's API in the meantime; if access arrives later, switching is a one-line change of the model ID.
If you do not have access yet, you have three honest options, and none of them is a hack:
- Use GLM 5 on Bedrock. It is open to every account, cheaper, and runs in-Region in 11 Regions. It is an older model with a 200K context, but for everyday coding help and document work it is a solid choice that stays inside AWS billing and IAM. The comparison table below shows what you give up.
- Call GLM 5.3 through Z.ai or another host. Z.ai's own API charges $1.40 per million input tokens and $4.40 per million output tokens, cheaper than Bedrock. The trade-off is that your prompts go to a company outside AWS, under its terms rather than your AWS agreement.
- Run it yourself. The weights are free. The hardware is not: even a heavily quantized copy needs well over 200 GB of memory. Our local guide has the real numbers.
Jake chose the second route for his own site, and Ethan chose Bedrock for his clients' code, because his clients' contracts require AWS billing and AWS data handling. Both of them were right for their situation. The rest of this page assumes you have access, and it flags every place where GLM 5 is the fallback.
GLM 5.3 Bedrock model ID: the string that decides whether you get an answer
Every Bedrock model has a model ID, and the newer ones also have inference profile IDs. For GLM 5.3 the difference is not a detail. GLM 5.3 is offered only through cross-Region inference, so the plain model ID cannot be used on its own to send a request. You must name one of two inference profiles.
| ID | What it is | Where requests run | Use it when |
|---|---|---|---|
zai.glm-5.3 | The base model ID | Not callable on its own | Writing IAM policies and reading the model card. Never as the model value in a request. |
us.zai.glm-5.3 | US geographic inference profile | US Regions only: N. Virginia, Ohio, N. California and Oregon | You must be able to say the request is processed in the United States. |
global.zai.glm-5.3 | Global inference profile | Any commercial AWS Region the profile can route to | You have no residency rule, or you call from outside the US. It is also about 9% cheaper. |
The prefix before zai. is a routing promise, not a different model. us. and global. send your prompt to exactly the same GLM 5.3. The difference is which data centers AWS may use to answer it, and the price. If Regions are still fuzzy, our map of Regions and Availability Zones helps: the inference profile sits one layer above the Region and lets a call made in Ohio be served from Oregon when Ohio is busy.
The address you send requests to depends on the API:
- OpenAI-compatible APIs (Responses and Chat Completions): base URL
https://bedrock-runtime.REGION.amazonaws.com/openai/v1, where REGION is the Region your client connects to, for exampleus-west-2. - Bedrock's own APIs (Converse and InvokeModel): the normal boto3
bedrock-runtimeclient with aregion_name. AWS builds the URL,https://bedrock-runtime.REGION.amazonaws.com, for you.
The Region you connect to is your "source" Region. The Region that actually processes the request may be a different one inside the profile. That is the whole point of a cross-Region profile, and it is why a busy afternoon in one Region does not stop your app.
Full ARNs are needed for IAM policies and for a few tools, not for normal requests. The inference profile ARN looks like arn:aws:bedrock:us-east-1:111122223333:inference-profile/global.zai.glm-5.3, with your own Region and account ID, and the foundation model ARN looks like arn:aws:bedrock:us-east-1::foundation-model/zai.glm-5.3. If ARNs read like line noise, our explainer on what an AWS ARN is takes each field apart.
One more thing that trips people who used GLM 5 first: GLM 5 is also served on the newer bedrock-mantle endpoint, but GLM 5.3 is not. Only bedrock-runtime works. If you copy a Mantle base URL from a GLM 5 project, the GLM 5.3 request fails before it reaches the model.
How to turn on GLM 5.3 in Amazon Bedrock, step by step
Turning on a model in Bedrock is mostly about permissions now, not a sign-up form. This is the order that works, assuming your account is eligible.
- Pick a source Region where the profile is offered. For
us.zai.glm-5.3, your client must connect to one of the four US Regions. Forglobal.zai.glm-5.3, any of the 31 listed Regions works, including London, Frankfurt, Mumbai, Singapore, Sydney and Tokyo. The full list is in the Regions section. - Find GLM 5.3 in the model catalog. Filter by the provider Z AI and open the card. It shows the model ID, both profile IDs, the context window and the lifecycle dates. If your account has never used a model from outside AWS, accept the provider terms when the console asks.
- Send one prompt in the playground. Choose Test, then Playground, select GLM 5.3, and ask something short. If it answers, the account side is finished.
- Choose how your code will sign in. For the OpenAI SDK route you need a Bedrock API key. A long-term key from the console is quickest for a test; for anything real, generate short-term keys from your normal AWS credentials, which expire on their own. For boto3 with Converse or InvokeModel, your usual AWS credentials work as they are.
- Give the caller the right IAM permissions. The identity that calls GLM 5.3 needs
bedrock:InvokeModel, plusbedrock:InvokeModelWithResponseStreamif you stream, on both the inference profile and the foundation model. If it signs in with a Bedrock API key, it also needsbedrock:CallWithBearerToken. - Send a test request from code. Use one of the four snippets in the next section, and print the whole response the first time, not only the text, so you can see the token counts you are paying for.
A small IAM policy that allows GLM 5.3 through both profiles looks like this. Replace the account ID with yours. The wildcard Region on the foundation model is there because a global profile can route to any Region in its list; if your security team will not accept a wildcard, list the Regions one by one.
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "GlmProfilesAndModel",
"Effect": "Allow",
"Action": [
"bedrock:InvokeModel",
"bedrock:InvokeModelWithResponseStream"
],
"Resource": [
"arn:aws:bedrock:*:111122223333:inference-profile/us.zai.glm-5.3",
"arn:aws:bedrock:*:111122223333:inference-profile/global.zai.glm-5.3",
"arn:aws:bedrock:*::foundation-model/zai.glm-5.3"
]
},
{
"Sid": "ApiKeyAccess",
"Effect": "Allow",
"Action": "bedrock:CallWithBearerToken",
"Resource": "*"
}
]
}
If you have never written one of these by hand, our walk-through of IAM policy JSON, line by line explains Effect, Action and Resource with examples. The thing to remember here is that the policy names two kinds of resource, the profile and the model behind it. Leave either one out and you get the same AccessDenied as if there were no policy at all, which is exactly why that error feels so unfair the first time.
How to call GLM 5.3 on Bedrock: working code for every API
GLM 5.3 on Bedrock supports four APIs: Responses, Chat Completions, Converse and InvokeModel. The first two are OpenAI-compatible, and AWS recommends them for new projects because they carry the most complete set of features. That sounds odd, so to be clear: the OpenAI SDK here is only a client library. With the base URL pointed at Bedrock and a Bedrock key, the request goes to AWS and is answered by Z.ai's model running inside AWS. Nothing is sent to OpenAI.
1. Responses API with short-term keys (the recommended route)
This is the cleanest setup for real work. The aws-bedrock-token-generator library turns your normal AWS credentials into a short-lived Bedrock key, so there is no long-term secret sitting in a file.
pip install -U openai aws-bedrock-token-generator
from aws_bedrock_token_generator import provide_token
from openai import OpenAI
region = "us-west-2" # your source Region
client = OpenAI(
api_key=provide_token(region=region),
base_url=f"https://bedrock-runtime.{region}.amazonaws.com/openai/v1",
)
resp = client.responses.create(
model="global.zai.glm-5.3",
input="Rewrite this Python function so it is iterative instead of recursive: ...",
)
print(resp.output_text)
print(resp.usage)
Run it with python bedrock-request.py. The last line prints the token counts. Keep that habit for the first week; it is the fastest way to learn what your prompts really cost.
2. Chat Completions API with a Bedrock API key
If you already have code written for the Chat Completions API, this is the smallest change. Set two environment variables and the OpenAI SDK picks them up.
pip install boto3 openai
export OPENAI_API_KEY="your-bedrock-api-key"
export OPENAI_BASE_URL="https://bedrock-runtime.us-east-1.amazonaws.com/openai/v1"
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
model="us.zai.glm-5.3",
messages=[{"role": "user", "content": "Explain what an inference profile is in two sentences."}],
)
print(response.choices[0].message.content)
On Windows, set the variables with setx OPENAI_API_KEY "..." and open a new terminal, or set them in your IDE's run configuration. A long-term API key is fine for this kind of test. Delete it in the console when you are done, the same way you would change a lock after a contractor finishes.
3. Converse API with boto3
Converse is Bedrock's own API with one request shape for every model, which makes it the easy choice when one app talks to several providers.
import boto3
client = boto3.client("bedrock-runtime", region_name="us-east-1")
response = client.converse(
modelId="us.zai.glm-5.3",
messages=[{"role": "user", "content": [{"text": "List three risks of storing passwords in plain text."}]}],
)
print(response["output"]["message"]["content"][0]["text"])
print(response["usage"])
To sign in with a Bedrock API key instead of your AWS credentials, set AWS_BEARER_TOKEN_BEDROCK to the key before you run it. For streaming, call converse_stream with the same arguments and read the events as they arrive.
4. InvokeModel with an explicit reasoning effort
InvokeModel sends the model's native request body. It is the most direct route, and the place to set the reasoning effort by name.
import json
import boto3
client = boto3.client("bedrock-runtime", region_name="us-east-1")
response = client.invoke_model(
modelId="us.zai.glm-5.3",
body=json.dumps({
"messages": [{"role": "user", "content": "Summarize what this function does: ..."}],
"reasoning_effort": "high",
"max_tokens": 4096,
}),
)
print(json.loads(response["body"].read()))
"Four ways to say the same thing," Jake said. "Which one do I learn?" Ethan didn't hesitate. "The first one. Short-term keys, Responses API. The others exist because people already have code written for them. Don't collect APIs like phone cases."
GLM 5.3 reasoning effort: low, high and max
GLM 5.3 always thinks before it answers. You cannot switch reasoning off, but you can choose how hard it thinks. Z.ai's model card lists three levels: low, high and max. If you pass nothing, or pass a value it does not recognize, it uses max, the slowest and most token-hungry setting.
That default matters for your bill. Reasoning tokens are output tokens, and output is the expensive side at $5.28 per million with global routing. A model that thinks at maximum effort on every request writes far more tokens than the visible answer suggests. The fix is to choose the level on purpose:
- low: quick answers, short edits, classification, anything a person would answer without thinking hard.
- high: most coding and document work. A sensible default for an app.
- max: the hard problems: a bug nobody can find, a long refactor, a design question with real trade-offs. This is the setting Z.ai uses for its benchmark scores.
Note one trap. By the model's own rules, a value it does not recognize means max, so a typo can show up as a larger bill rather than a failed request. When your output token counts look higher than you expected, check the spelling of the effort value first.
GLM Bedrock pricing: the table, the tiers, and a worked month
Here are the on-demand prices for GLM 5.3 on Bedrock, Standard tier, per million tokens.
| Inference profile | Input | Output | Cache read | Cache write (30 min) |
|---|---|---|---|---|
Global (global.zai.glm-5.3) | $1.68 | $5.28 | $0.312 | $2.10 |
US (us.zai.glm-5.3) | $1.848 | $5.808 | $0.3432 | $2.31 |
Two service tiers change those numbers. Priority costs 75% more and gets the fastest responses, for latency-critical traffic. Flex costs half as much and accepts slower, less predictable service, for work that can wait, such as overnight reports or bulk code review. Reserved capacity is not offered for GLM 5.3. You pick the tier per request with a service_tier field: "default" (or leave it out) for Standard, "priority" or "flex".
Now the comparison most buyers want, per million tokens at standard rates:
| Where you call it | Input | Output | What you get for the difference |
|---|---|---|---|
| GLM 5.3, Z.ai's own API | $1.40 | $4.40 | The list price; prompts go to Z.ai under its terms |
| GLM 5.3, Bedrock global | $1.68 | $5.28 | AWS billing, IAM, Guardrails, CloudWatch; 20% above Z.ai |
| GLM 5.3, Bedrock US | $1.848 | $5.808 | The same, plus processing kept in US Regions |
| GLM 5, Bedrock US Regions | $1.00 | $3.20 | Older model, 200K context, open to every account |
| Kimi K3, Bedrock global | $3.00 | $15.00 | The other large open model on Bedrock; about 2.8 times GLM 5.3's output price |
| Grok 4.7, Bedrock global | $2.00 | $6.00 | A closed model with a 500K context |
Is the 20% worth it? For a hobby project, probably not. For a business, it often is, and not because of the model. The premium buys one bill, one set of IAM rules, logging you already know, Guardrails you already configured, and an answer to the question a client's lawyer will ask: who saw our data? On Bedrock, prompts and responses are not shared with the model provider, so Z.ai never sees them. For Ethan's clients that one sentence is worth more than the price gap.
GLM 5 prices vary by Region, which matters if you need GLM 5 for residency reasons: $1.00 / $3.20 in N. Virginia, Ohio and Oregon; $1.20 / $3.84 in Jakarta, Mumbai, Tokyo, SΓ£o Paulo and Stockholm; $1.03 / $3.30 in Sydney; and $1.55 / $4.96 in London, the most expensive place on the list.
A worked month
Ethan's team runs a coding assistant for one client. Each request sends about 50,000 input tokens, of which 40,000 are the same every time (the system instructions and a map of the repository), plus 10,000 new tokens of code and question. Each answer, reasoning included, runs about 4,000 output tokens. They make 3,000 requests a month, about 100 a working day. With global routing:
- Without caching: 150 million input tokens cost $252.00 and 12 million output tokens cost $63.36. Total: $315.36, about 10.5 cents a request.
- With explicit caching, assuming the 40,000-token prefix is written to the cache about 200 times a month (whenever it expires or changes) and read from it the other 2,800 times: cache writes $16.80, cache reads $34.94, fresh input $50.40, output $63.36. Total: about $165.50, close to half.
- Same workload on Flex, no caching: $157.68, if the work can wait.
- Same workload on the US profile, no caching: $346.90.
- Same workload on GLM 5 in a US Region (no prompt caching listed for it): $188.40.
The numbers are rounded on purpose so you can redo them with your own traffic. The lesson in them is bigger than any single figure: for an agent that resends the same context every turn, caching is a bigger lever than the choice between Bedrock and Z.ai's API.
GLM 5.3 prompt caching on Bedrock: when it pays
Prompt caching stores the start of a prompt so the next request that begins the same way does not pay full price for it. GLM 5.3 caches automatically by default. That implicit mode helps whenever repeated calls share the same opening text. Explicit mode lets you mark exactly where the reusable part ends, which raises the hit rate, and AWS recommends it.
The math is friendlier than it looks. Writing tokens to the cache costs $2.10 per million, 25% more than a normal input token. Reading them back costs $0.312, about 81% less. On a 40,000-token prefix, the write costs under two cents extra, and every later read saves about five and a half cents. Caching pays for itself on the first reuse.
The rules: each cache point must cover at least 1,024 tokens, and a cached prefix lives for at least 30 minutes. To use explicit mode with the Responses API, switch it on in the request and mark the end of each reusable block:
resp = client.responses.create(
model="global.zai.glm-5.3",
extra_body={"prompt_cache_options": {"mode": "explicit"}},
input=[
{
"type": "message",
"role": "system",
"content": [{
"type": "input_text",
"text": SYSTEM_PROMPT, # long and stable: the best thing to cache
"prompt_cache_breakpoint": {"mode": "explicit"},
}],
},
{
"type": "message",
"role": "user",
"content": [{"type": "input_text", "text": USER_INPUT}],
},
],
)
print("cached tokens:", resp.usage.input_tokens_details.cached_tokens)
Put the stable material first and the changing material last. A cache only matches from the beginning of the prompt, so one changed word near the top, a timestamp in the system prompt for example, turns every request into a cache miss. It is the most common reason people say caching "does nothing."
GLM 5.3 Regions: where you can call it, and where your data goes
GLM 5.3 has no in-Region option anywhere. Every request goes through one of the two profiles, so the question is not "which Region hosts it" but "which Regions can I call from, and where might the request be processed".
| Geography | Source Regions you can call from | Profiles available |
|---|---|---|
| United States | us-east-1 (N. Virginia), us-east-2 (Ohio), us-west-1 (N. California), us-west-2 (Oregon) | us. and global. |
| Canada and Mexico | ca-central-1 (Canada), ca-west-1 (Calgary), mx-central-1 (Mexico) | global. only |
| Europe | Frankfurt, Zurich, Stockholm, Milan, Spain, Ireland, London, Paris | global. only |
| Asia Pacific | Taipei, Tokyo, Seoul, Osaka, Mumbai, Hyderabad, Singapore, Sydney, Jakarta, Melbourne, Malaysia, New Zealand, Thailand | global. only |
| Middle East, Africa, South America | il-central-1 (Tel Aviv), af-south-1 (Cape Town), sa-east-1 (SΓ£o Paulo) | global. only |
That is 31 source Regions for the global profile and four for the US profile. The gap that matters is Europe and Asia Pacific: there is no EU or APAC geographic profile for GLM 5.3. A team in Frankfurt can call it, but only through global routing, which may process the request in any Region on the global list. If you have promised a customer that their data is processed in the EU, GLM 5.3 on Bedrock cannot keep that promise today.
What stays the same with every profile: the request travels on the AWS network, it is billed and logged in your source Region, and it never goes to Z.ai. What changes is the Region where the model runs. If that sentence makes a compliance person nervous, the honest options are the us. profile for US-only processing, or GLM 5 in-Region, which runs in 11 Regions: N. Virginia, Ohio, Oregon, Stockholm, London, Tokyo, Mumbai, Sydney, Jakarta, Melbourne and SΓ£o Paulo. In those, a GLM 5 request is processed in the Region you call.
What works with GLM 5.3 on Bedrock, and what does not
The Bedrock features you can use with GLM 5.3 decide whether it fits your existing setup. Here is the list as of launch.
| Works with GLM 5.3 | Does not work with GLM 5.3 |
|---|---|
| Response streaming | Knowledge Bases |
| Implicit and explicit prompt caching | Count Tokens API |
| Guardrails | Intelligent prompt routing |
| Bedrock Agents and Flows | Prompt optimization |
| Prompt management and model evaluation | Abuse detection |
| Structured outputs, function calling | Reserved service tier, the bedrock-mantle endpoint |
Two gaps catch people. The first is Knowledge Bases. If you built a document search with a Bedrock Knowledge Base, you cannot pick GLM 5.3 as its answering model. You can still use the two together: call the Knowledge Base's Retrieve API to fetch the relevant passages, then send them to GLM 5.3 yourself in the prompt. It is one extra step, and many teams prefer that control anyway. GLM 5 has the same gap, so switching models does not help here.
The second is the Count Tokens API. You cannot ask Bedrock in advance how many tokens a prompt will use with GLM 5.3. Read the usage block in each response instead, and after a day of traffic you will know your averages better than any estimate.
GLM 5 vs GLM 5.3 on Bedrock: which one should you use?
For many readers the real choice is not GLM 5.3 against another company's model. It is GLM 5.3 against GLM 5, the GLM that every account can already call.
| GLM 5 | GLM 5.3 | |
|---|---|---|
| Who can use it | Every account | Eligible customers only |
| Model ID | zai.glm-5, called directly | us.zai.glm-5.3 or global.zai.glm-5.3 |
| Where it runs | In-Region, 11 Regions | Cross-Region only, 31 source Regions |
| Context / max output | 200K / 128K | 1M / 128K |
| APIs | Chat Completions, Converse, InvokeModel | All of those plus Responses |
| Endpoints | bedrock-runtime and bedrock-mantle | bedrock-runtime only |
| Prompt caching | Not listed | Implicit and explicit |
| Price (US, per million) | $1.00 in / $3.20 out | $1.848 in / $5.808 out (US profile), $1.68 / $5.28 global |
| Retirement promise | Not before February 11, 2027, then at least 6 months of legacy | No date; at least 45 days' notice and a legacy period of at least 45 days |
Look at the last row twice. GLM 5's card promises it will not be retired before February 2027. GLM 5.3's card promises only 45 days' notice. That does not mean AWS plans to remove it; it means the commitment is shorter than for older models. If you are building something a client will rely on for a year, keep the model ID in configuration rather than in code, so a forced switch is a one-line change.
"So which do I use?" Jake asked. Ethan's rule was simple. "If you can call 5.3 and your prompts are long, use 5.3; the million-token context and the caching are worth the price. If you can't call it, use GLM 5 and stop worrying. It was a top open model six months ago and it still writes good code."
GLM 5.3 vs Kimi K3 and the closed models: benchmarks and price
Benchmark tables from a model's own maker always flatter the maker, so read these as Z.ai's claims, not as a verdict. They are still useful for one thing: seeing where GLM 5.3 is close to the expensive closed models and where it is not.
| Benchmark (Z.ai's table) | GLM-5.3 | GLM-5.2 | Kimi K3 | Opus 4.8 | GPT-5.6 Sol |
|---|---|---|---|---|---|
| Terminal-Bench 3.0 | 28.3 | 4.6 | 17.4 | 21.1 | 34.6 |
| DeepSWE (v1.1) | 66.9 | 46.2 | 67.5 | 58.0 | 72.7 |
| NL2Repo | 58.0 | 48.9 | 58.0 | 69.7 | not reported |
| Toolathlon Verified | 73.0 | 59.9 | 76.5 | 76.2 | 74.9 |
| HLE with tools | 62.5 | 54.7 | 59.8 | 57.9 | 64.5 |
| GDPval-AA v2 (Elo) | 1769 | 1508 | 1682 | 1588 | 1730 |
The honest reading: GLM 5.3 is a large step up from GLM 5.2, roughly level with Kimi K3 on coding, and trails the best closed model in most coding rows while leading on Z.ai's office-work measure, GDPval. Now put the price next to it. On Bedrock, Kimi K3 costs $15 per million output tokens with global routing and GLM 5.3 costs $5.28. If the two are close on your own tasks, that gap decides it.
The only benchmark that counts is your own. Take ten real tasks from last month, run them through GLM 5.3 and whatever you use today at the same reasoning effort, and compare both the answers and the usage blocks. An afternoon of that beats any table, including this one.
Using GLM 5.3 on Bedrock from OpenCode and other coding tools
Most people will not call GLM 5.3 from a script. They will use it inside a coding assistant. Here is how the common ones stand at launch.
OpenCode
OpenCode supports Amazon Bedrock as a provider, and our OpenCode guide covers the install. The catch on day one: OpenCode's built-in model list does not include GLM 5.3 on Bedrock yet, so it will not appear when you run /models. Add it yourself in opencode.json:
{
"$schema": "https://opencode.ai/config.json",
"provider": {
"amazon-bedrock": {
"models": {
"glm-5.3": {
"id": "global.zai.glm-5.3",
"name": "GLM 5.3 (Bedrock)",
"limit": { "context": 1000000, "output": 128000 }
}
}
}
}
}
Then start OpenCode with your AWS profile and Region set, for example AWS_PROFILE=my-profile AWS_REGION=us-east-1 opencode, or with a Bedrock API key in AWS_BEARER_TOKEN_BEDROCK, which takes priority over every other credential if both are set. Run /models and pick GLM 5.3. The limit line matters: without it, OpenCode does not know how much context you have left.
LiteLLM and tools built on it
Many agents and gateways use LiteLLM to talk to Bedrock. At launch, LiteLLM did not recognize the short name bedrock/global.zai.glm-5.3. It added GLM 5.3 to its model list on October 6, so a fresh upgrade may be all you need. If your version still refuses it, name the Converse route and the full inference profile ARN instead:
bedrock/converse/arn:aws:bedrock:us-east-1:111122223333:inference-profile/global.zai.glm-5.3
Replace the Region and account ID with yours. That string works anywhere a tool asks for a LiteLLM model name.
Tools that speak the OpenAI API
Any tool that lets you set an OpenAI base URL, an API key and a model name can use GLM 5.3 on Bedrock: base URL https://bedrock-runtime.REGION.amazonaws.com/openai/v1, a Bedrock API key, and model global.zai.glm-5.3. If the tool offers a choice between the Chat Completions and Responses APIs, either works with GLM 5.3; GLM 5 supports only Chat Completions.
GLM 5.3 not working on Bedrock? Every error, and the fix
If you are reading this section with an error on your screen, take a breath. Almost every GLM 5.3 failure in the first week comes from one of these, and none of them means you broke your account.
"Invocation of model ID zai.glm-5.3 with on-demand throughput isn't supported"
You sent the bare model ID. Bedrock answers with a ValidationException that asks you to retry with the ID or ARN of an inference profile. Change the model value to global.zai.glm-5.3 or us.zai.glm-5.3.
The US profile fails from a European or Asian Region
us.zai.glm-5.3 can only be called from the four US source Regions. If your client connects to eu-west-1 or ap-south-1, switch to global.zai.glm-5.3, or point the client at a US Region if you need US-only processing.
GLM 5.3 is missing from the console, or the playground refuses it
This is eligibility, not configuration. If GLM 5 and other models work and GLM 5.3 does not, no IAM change will fix it. Contact your AWS account team, or use GLM 5 or Z.ai's API in the meantime.
AccessDeniedException even though your role can invoke models
Check that the policy names both the inference profile and the foundation model, arn:aws:bedrock:*::foundation-model/zai.glm-5.3, and that it covers every Region the profile can route to. If you sign in with a Bedrock API key, the identity also needs bedrock:CallWithBearerToken. Our Bedrock AccessDeniedException guide walks through the other causes, such as a service control policy at the organization level.
The request fails on a Mantle URL
GLM 5.3 is served only on bedrock-runtime. Replace a bedrock-mantle base URL with https://bedrock-runtime.REGION.amazonaws.com/openai/v1.
LiteLLM says the model is not supported
Upgrade LiteLLM, or use the bedrock/converse/ route with the full inference profile ARN shown in the tools section.
An image upload is rejected
GLM 5.3 on Bedrock is text only. Describe the image in words, extract its text first, or send the image to a model that accepts images.
The answer stops halfway
Look at the stop reason in the response. If it says the length limit was reached, raise max_tokens, up to the 128K ceiling, or lower the reasoning effort so fewer tokens go to thinking.
ThrottlingException under load
Your account hit its quota for the model. Retry with exponential backoff, spread steady background work onto the Flex tier, and request a quota increase in the Service Quotas console if the load is permanent.
Where else GLM 5.3 runs: Z.ai, OpenRouter, Mistral and your own hardware
Bedrock is one of many homes for GLM 5.3. Knowing the others keeps your options open, especially if the eligibility list keeps you out for a while.
- Z.ai's API. The maker's own service: $1.40 per million input tokens and $4.40 per million output tokens, with the full 1M context. The cheapest first-party route.
- OpenRouter. One key and many hosts. Z.ai's own endpoint there is the same $1.40 / $4.40. Some hosts list far lower prices, but many of them run compressed versions of the weights, so check the listed precision before you trust the bargain.
- Mistral's API. Mistral has served GLM 5.3 as
zai-glm-5-3since September 28, 2026, at the same list price. It is a handy option for teams already billing through Mistral in Europe. - Your own hardware or a rented GPU. Free weights, expensive memory. The local installation guide has the hardware math for Windows and Kali.
The model is the same everywhere. What changes is who sees your prompts, who bills you, and how much control you have. For Ethan's clients that pointed to Bedrock. For Jake's own code, Z.ai's API was the right call that evening.
GLM 5.3 on Bedrock: the go-live checklist
Before you move real traffic, walk through this list once. It takes ten minutes and saves a week of surprises.
- The playground answers for GLM 5.3 in your source Region, so eligibility is confirmed.
- Your code sends
global.zai.glm-5.3orus.zai.glm-5.3, read from configuration, not hard-coded. - The IAM policy names both profiles and the foundation model, plus
bedrock:CallWithBearerTokenif you use API keys. - Production signs in with short-term keys or IAM credentials, and any long-term test key is deleted.
- The reasoning effort is set on purpose for each kind of request, not left to the default.
- Stable prompt text comes first and is cached explicitly, and you have seen a nonzero cached token count.
- Work that can wait is routed to Flex.
- You have written down, for your customer or your own records, which profile you use and what it means for where requests are processed.
- A billing alarm is set, and someone has looked at the usage numbers after the first day. Our guide to reading AWS Cost Explorer shows where model charges appear.
- You know your fallback: GLM 5 on Bedrock, or GLM 5.3 through Z.ai, if access or the model changes.
Frequently asked questions
Is GLM on Bedrock?
Yes. Amazon Bedrock offers four Z.ai models: GLM 5.3 (since October 5, 2026), GLM 5, GLM 4.7 and GLM 4.7 Flash. GLM 5.3 is limited to eligible customers; the other three are open to every account in the Regions where they are offered.
Is GLM 5.3 available on AWS Bedrock?
Yes, since October 5, 2026, through the inference profiles us.zai.glm-5.3 and global.zai.glm-5.3. Access is limited to eligible customers, so it may not appear in every account. If yours does not show it, contact your AWS account team.
Is GLM 5.2 on AWS Bedrock?
No. GLM 5.1 and GLM 5.2 never came to Bedrock. AWS went from GLM 5 straight to GLM 5.3, which is built on the same base model as GLM 5.2 with better post-training. If you were waiting for GLM 5.2 on Bedrock, GLM 5.3 is the model to use.
What is the GLM 5.3 model ID on Bedrock?
Send global.zai.glm-5.3 for global routing or us.zai.glm-5.3 for US-only routing. The base model ID zai.glm-5.3 is used in IAM policies, but requests that name it directly are refused because GLM 5.3 is offered only through cross-Region inference.
How much does GLM 5.3 cost on Bedrock?
With global routing, $1.68 per million input tokens, $5.28 per million output tokens, $0.312 per million cached tokens read and $2.10 per million written. The US profile costs 10% more. Priority adds 75% and Flex halves the price.
What is GLM 5 Bedrock pricing?
GLM 5 costs $1.00 per million input tokens and $3.20 per million output tokens in N. Virginia, Ohio and Oregon. It costs $1.20 and $3.84 in Jakarta, Mumbai, Tokyo, SΓ£o Paulo and Stockholm, $1.03 and $3.30 in Sydney, and $1.55 and $4.96 in London.
Is GLM 5.3 cheaper on Bedrock or on Z.ai's API?
Z.ai's API is cheaper: $1.40 in and $4.40 out per million tokens, against $1.68 and $5.28 on Bedrock with global routing. The Bedrock premium buys AWS billing, IAM, Guardrails and the assurance that prompts are not shared with the model provider.
Why can't I see GLM 5.3 in my Bedrock console?
Most likely your account is not on the eligibility list yet. Check that you are in a supported Region, then try GLM 5. If GLM 5 works and GLM 5.3 does not, contact your AWS account team or use AWS's Connect with the Amazon Bedrock team form; no IAM setting will add it.
Is the GLM 5.3 API free?
No. Bedrock has no free tier for GLM 5.3, and Z.ai's API charges per token. The weights are free to download under the GLM-5.3 license, so the only free route is running it yourself, which needs well over 200 GB of memory.
Is GLM-5.3-Flash on Bedrock?
No. Bedrock offers the full GLM 5.3 only. GLM-5.3-Flash is available from Z.ai and other hosts. On Bedrock, the budget GLM is GLM 4.7 Flash at $0.07 per million input tokens and $0.40 per million output tokens in US Regions.
Can I keep GLM 5.3 requests in the EU?
Not today. GLM 5.3 has only a US profile and a global profile. From European Regions it is reachable only through global routing, which may process requests in any Region on the global list. For EU processing, GLM 5 runs in-Region in Stockholm and London.
Does Z.ai see my prompts when I use GLM 5.3 on Bedrock?
No. Bedrock runs the model inside AWS and does not share prompts or responses with model providers. Your requests stay on the AWS network and are billed and logged in your AWS account.
Does GLM 5.3 on Bedrock work with Knowledge Bases?
Not as the Knowledge Base's answering model. You can still combine them: call the Knowledge Base's Retrieve API to get the relevant passages, then send those passages to GLM 5.3 in your own prompt.
Does GLM 5.3 support images?
Not on Bedrock. GLM 5.3 takes text in and returns text. Image, audio and video inputs are not supported, so extract the text or describe the image before sending it.
What is the GLM 5.3 context window on Bedrock?
One million tokens of context, with up to 128K tokens of output per response. GLM 5 on Bedrock has a 200K context, so long documents and large codebases are the clearest reason to use GLM 5.3.
How do I set the reasoning effort for GLM 5.3?
Pass reasoning_effort with low, high or max. If you pass nothing, the model uses max, which writes the most output tokens. For most app traffic, high is a sensible default and low suits quick, simple requests.
Can I use GLM 5.3 on Bedrock in OpenCode?
Yes. Add a model under the amazon-bedrock provider in opencode.json with the id global.zai.glm-5.3 and a 1M context limit, set your AWS profile and Region, then pick it with /models. It does not appear in the built-in list at launch.
Is GLM 5.4 or GLM 5.5 out?
Not as of October 6, 2026. GLM 5.3 is Z.ai's newest released model, and the newest GLM on Bedrock. Names like GLM 5.4 and 5.5 appear in searches and rumors, but there is no model card or release for either yet.
The model itself is the easy part: one profile ID, one base URL, one key. The evening goes on the things around it, the guest list nobody mentions, the US profile that refuses European callers, the reasoning default that bills like a novel. Now you know them in advance. Jake did get his code review that night, through Z.ai's API, and it found the same missing limit on the booking form that had cost him last December. His freelancer fixed it in an afternoon, and Jake spent the next Saturday doing what he likes best, fixing phones for people who actually showed up. Ethan's clients got GLM 5.3 on Bedrock, with the bill and the paper trail they asked for. Whichever door you walk through, you are not late, and you did not miss anything. Good luck with the first request.
📌 If you keep one line from this page
If GLM 5.3 will not show up on Bedrock, it is the guest list, not your setup.
When you are in, send global.zai.glm-5.3, choose the reasoning effort on purpose, and cache the part of the prompt that never changes.
Revision note. Written October 6, 2026, the day after GLM 5.3 reached Amazon Bedrock. If your console showed nothing and you wondered what you broke, you broke nothing. The door is narrow for now, and the other doors are open.