Boto3 Error Handling: ClientError, the Exceptions List, and the IAM EntityAlreadyExists Trap (Role With Name Already Exists, Fixed)
EntityAlreadyExists is IAM telling you that a role, user, group, managed policy or instance profile with that name is already in your AWS account. In Python it arrives as botocore.exceptions.ClientError: An error occurred (EntityAlreadyExists) when calling the CreateRole operation: Role with name shop-backup-role already exists. with HTTP status 409. The fix is not to delete the old role. It is to write the create step so a second run picks up what is already there instead of crashing, a create-or-get pattern that takes about a dozen lines. Here is the surprise that breaks half the handlers people copy: the error code and the exception class have different names. The code is EntityAlreadyExists; the class boto3 raises is EntityAlreadyExistsException. A check like err.response["Error"]["Code"] == "EntityAlreadyExistsException" is never true, so your except block runs, matches nothing, and the script dies anyway. The second surprise is stranger: your own script can trip over a role it made a second earlier, because boto3 quietly retries a call whose answer got lost on the way back.
Jake runs a phone repair shop. Last spring his friend Ethan, a developer, wrote him a small Python script that sets up nightly photo backups: an S3 bucket, an IAM role, a Lambda function. It worked once. When Jake ran it again for his second shop, it stopped at the role with the red line above. He found a forum answer, wrapped the call in try/except, compared the code to "EntityAlreadyExistsException", and it still crashed. This page is what Ethan showed him over two coffees: what the error means and its six causes, why that except block never fired, how boto3 error handling works from zero (ClientError, the exceptions list, the fields worth logging), the create-or-get fix for roles, policies and instance profiles, the retry settings behind the surprise duplicate, the S3 and DynamoDB errors that break the pattern, the Terraform and CloudFormation versions of the same error, and how to test all of it without an AWS account.
New to IAM? Our plain-English guide to AWS IAM explains roles, users and policies in about ten minutes. You can follow this page without it; every step starts from zero.
What "An error occurred (EntityAlreadyExists) when calling the CreateRole operation" means
Think of IAM as the key cabinet for your AWS account. Every hook has a name tag: a role called shop-backup-role, a user called jake, a policy called photo-bucket-write. You can't hang two keys on hooks with the same tag. When your code asks IAM to create something and the tag is already in use, IAM refuses with EntityAlreadyExists and leaves the existing key exactly where it was. Nothing is overwritten and nothing is damaged. The request simply didn't happen.
Three rules decide when a name counts as "taken," and each one surprises someone:
- Names are unique per account, not per project. A role made by a tutorial two years ago, by a teammate, or by another stack in the same account owns that name until someone deletes it.
- IAM is global. There is one IAM per account, shared by every Region. A role you create while your script points at
us-east-1already exists when the same script runs againsteu-west-1. - Names ignore case. You can't have both
Role1androle1in one account. IfShopBackupRoleexists, creatingshopbackuprolefails with EntityAlreadyExists even though the two strings look different to Python.
Roles get most of the attention, but the same error comes from every IAM call that creates something with a name or a unique address:
| IAM operation | boto3 method | What is already there |
|---|---|---|
| CreateRole | create_role | A role with that name (any case) |
| CreateUser | create_user | A user with that name |
| CreateGroup | create_group | A group with that name |
| CreatePolicy | create_policy | A customer managed policy with that name |
| CreateInstanceProfile | create_instance_profile | An instance profile with that name |
| CreateLoginProfile | create_login_profile | The user already has a console password |
| CreateOpenIDConnectProvider | create_open_id_connect_provider | A provider for that issuer URL |
| CreateSAMLProvider | create_saml_provider | A SAML provider with that name |
| UploadServerCertificate | upload_server_certificate | A server certificate with that name |
The message itself always follows one template, which is worth knowing because every ClientError you will ever see in Python uses it:
An error occurred (EntityAlreadyExists) when calling the CreateRole operation: Role with name shop-backup-role already exists.
^^^^^^^^^^^^^^^^^^^ ^^^^^^^^^^ ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
the error code the API call the message, written for humans
The part in parentheses is the error code, the one string your program should test. The operation name tells you which call failed, which matters in a script that makes twenty calls. The message is for people and can change wording; the code doesn't. When retries ran out before the error, boto3 inserts a note before the colon, such as (reached max retries: 4). And the HTTP status, 409 Conflict here, is in the response but not in the printed line.
Jake: "So IAM isn't broken. It's telling me the role I asked for is already hanging on the wall."
Ethan: "Right. Your script is a new employee who walks in every morning and tries to cut a fresh key for the back door. The locksmith says 'there's already one on hook four.' A sensible employee takes the key off hook four. Yours throws a tantrum and goes home."
Why it happens: the six real causes of "role already exists" in AWS
The error always means the same thing, but the reason the name was taken decides the right fix. These are the six patterns behind nearly every report, roughly in order of how often they bite.
| Cause | How you recognize it | Right fix |
|---|---|---|
| The script ran before | First run worked, every later run fails at the same line | Create-or-get (code) |
| Same name, second Region | Works in one Region, fails the first time in another | Treat IAM as account-wide; create once, or put the Region in the name |
| Case collision | Your name isn't in the list you expected, but a differently cased one is | Pick one spelling, lowercase everywhere |
| Leftover from old work | The role predates your code; a deleted stack or lost Terraform state left it behind | Import it into your tool, or rename yours |
| Automatic retry after a timeout | Fails once in a blue moon; afterwards the role exists and looks right | Create-or-get makes it harmless (why) |
| Two runs at once | Only in CI, or when two people deploy at the same time | Create-or-get, plus one deploy at a time |
The script ran before
This is Jake's case and the most common one by far. A setup script written as a straight line of create calls only works on an empty account. The first run builds everything; the second run hits the first thing that already exists and stops, often leaving the rest of the setup half-applied. The cure isn't "delete everything before each run." It's making each step idempotent, a long word for a simple idea: running it twice ends in the same place as running it once.
Same name, second Region
Most AWS services are Regional, so a script that creates a Lambda function or a DynamoDB table in each Region feels natural. IAM breaks that habit. If your code builds a role named backup-role in a loop over Regions, the first Region creates it and every other Region fails. The fix is either to create IAM resources once per account, outside the Region loop, or to include the Region in the name, such as backup-role-eu-west-1.
Case collision
Because IAM compares names without caring about case, a teammate's Backup-Role blocks your backup-role. Python sees two different strings, so a check like if name in existing_names says "not there" and the create call fails anyway. Settle on lowercase names with hyphens across the team, and compare with .lower() on both sides when you search.
Leftover from old work
Infrastructure tools keep a record of what they made. When that record is lost or a stack is deleted with a retain setting, the real role stays in the account while the tool believes it never existed. The next deploy tries to create it fresh and gets EntityAlreadyExists. The infrastructure section covers the clean way out for Terraform and CloudFormation.
Automatic retry after a timeout
This is the one that makes people doubt their sanity. Your create call reaches IAM and the role is made, but the answer gets lost on the way back: a dropped connection, a slow network, a read timeout. boto3 treats that as a temporary failure and sends the same request again. The retry finds the role your first attempt created and fails with EntityAlreadyExists. From your code's point of view, a single call to create_role raised "already exists" for a role that didn't exist when you called it.
Two runs at once
Two CI jobs, two developers, or a scheduler that starts a new run before the old one finishes can both decide the role is missing and both try to create it. One wins; the other gets EntityAlreadyExists. Checking first doesn't help here, because both runs check before either creates. Only a create step that tolerates "already exists" survives a race.
Why your try/except is not catching the exception: the code is not the class name
Here is the handler from the forum answer that Jake copied. It looks careful. It catches the right exception type, reads the code, and re-raises anything it doesn't expect:
from botocore.exceptions import ClientError
try:
iam.create_role(RoleName="shop-backup-role", AssumeRolePolicyDocument=trust_json)
except ClientError as err:
if err.response["Error"]["Code"] == "EntityAlreadyExistsException": # never true
print("Role already there, carrying on")
else:
raise
The except block does run. The if compares the code IAM sent, EntityAlreadyExists, with a string that has nine extra letters, gets False, falls to raise, and your script crashes with the same error as before. It feels like Python ignored your try/except. It didn't; your comparison quietly failed. Two fixes work, and you can use either:
# Fix 1: compare against the real code
except ClientError as err:
if err.response["Error"]["Code"] == "EntityAlreadyExists":
print("Role already there, carrying on")
else:
raise
# Fix 2: catch the specific class (no string to get wrong)
except iam.exceptions.EntityAlreadyExistsException:
print("Role already there, carrying on")
Why do the names differ at all? boto3 builds an exception class for every error a service describes, and names the class after the error's shape in the service model. IAM's shapes end in Exception; the codes IAM actually sends don't. Here are the IAM errors you are most likely to meet, with both names side by side:
| Class you catch (iam.exceptions.…) | Code in err.response["Error"]["Code"] | HTTP |
|---|---|---|
| EntityAlreadyExistsException | EntityAlreadyExists | 409 |
| NoSuchEntityException | NoSuchEntity | 404 |
| DeleteConflictException | DeleteConflict | 409 |
| LimitExceededException | LimitExceeded | 409 |
| MalformedPolicyDocumentException | MalformedPolicyDocument | 400 |
| ConcurrentModificationException | ConcurrentModification | 409 |
| InvalidInputException | InvalidInput | 400 |
| UnmodifiableEntityException | UnmodifiableEntity | 400 |
| ServiceFailureException | ServiceFailure | 500 |
| CredentialReportExpiredException | ReportExpired | 410 |
That last row shows the rule can't even be reduced to "drop the word Exception": the class for an expired credential report is called CredentialReportExpiredException and its code is ReportExpired. Other services follow other habits. DynamoDB sends codes that include the suffix, such as ConditionalCheckFailedException, so class and code match. S3 uses plain names on both sides, such as NoSuchKey. And a few newer IAM errors, such as OrganizationNotFoundException, keep the suffix in the code too. You can't guess. Either catch the class, or print the code once from a real error, or read the map straight out of boto3 (the exceptions list section shows the two-line snippet).
Jake: "So the forum answer was wrong by one word."
Ethan: "By nine letters. It's like your parts bin is labeled 'screen' and the work order says 'screen assembly.' Same part, but if your rule is 'only pull from the bin whose label matches the order exactly,' you'll stand there forever. Catching the class is reading the part number instead of the label."
Boto3 error handling from zero: the two families of exceptions
Every exception boto3 raises belongs to one of two families, and knowing which one you are holding tells you where the problem is.
ClientError means AWS received your request and said no. The call reached the service, the service looked at it, and it sent back an error code: EntityAlreadyExists, AccessDenied, NoSuchKey, ThrottlingException and hundreds more. Every service-specific class, such as iam.exceptions.EntityAlreadyExistsException, is a subclass of ClientError, so except ClientError catches all of them.
BotoCoreError means the request never got a proper answer from AWS, usually because something on your side stopped it: no credentials, no Region, a bad parameter caught before sending, a network that won't connect, a certificate that won't validate. These are raised by botocore itself, the library underneath boto3, and their messages are fixed strings you can search for:
| Exception (botocore.exceptions.…) | Message you see | What it means |
|---|---|---|
| NoCredentialsError | Unable to locate credentials | No keys, profile, SSO session or role was found on this machine |
| PartialCredentialsError | Partial credentials found in {provider}, missing: {cred_var} | Half a key pair, usually the secret is missing |
| ProfileNotFound | The config profile ({profile}) could not be found | AWS_PROFILE or profile_name names a profile that isn't in your config files |
| NoRegionError | You must specify a region. | A Regional service was called with no Region set |
| ParamValidationError | Parameter validation failed: … | A required parameter is missing or has the wrong type; nothing was sent |
| EndpointConnectionError | Could not connect to the endpoint URL: "…" | DNS, proxy, firewall, or a service name and Region that don't exist together |
| ConnectTimeoutError | Connect timeout on endpoint URL: "…" | The connection couldn't be opened in time |
| ReadTimeoutError | Read timeout on endpoint URL: "…" | Connected, sent, but the answer didn't arrive in time |
| SSLError | SSL validation failed for {endpoint_url} {error} | The certificate chain couldn't be validated, often a corporate proxy |
| UnauthorizedSSOTokenError | The SSO session associated with this profile has expired or is otherwise invalid… | Run aws sso login for that profile |
| WaiterError | Waiter {name} failed: {reason} | A waiter gave up or hit a failure state |
Here is the shape of a handler that respects both families, with the specific cases first and the broad ones last. Python checks except clauses from top to bottom and stops at the first match, so a broad except ClientError placed above a specific class would swallow it:
import boto3
from botocore.exceptions import BotoCoreError, ClientError, NoCredentialsError
iam = boto3.client("iam")
try:
iam.create_role(RoleName="shop-backup-role", AssumeRolePolicyDocument=trust_json)
except iam.exceptions.EntityAlreadyExistsException:
pass # expected on a second run: the role is there
except ClientError as err:
raise # AWS said no for another reason: let it surface
except NoCredentialsError:
raise SystemExit("No AWS credentials on this machine. Run aws configure or aws sso login.")
except BotoCoreError as err:
raise SystemExit(f"Could not talk to AWS: {err}")
There is also a third, much smaller group: boto3.exceptions, raised by boto3's own helper layers rather than by AWS or botocore. The ones you are likely to meet are S3UploadFailedError (the S3 upload helper wraps the underlying ClientError in it, with a message starting "Failed to upload"), S3TransferFailedError, and ResourceNotExistsError (you asked for a resource type that doesn't exist, such as boto3.resource("lambda")). So if you search for "import boto3 exceptions," this is the module: from boto3.exceptions import S3UploadFailedError. For AWS errors you still import ClientError from botocore.
Jake: "So ClientError is AWS saying no, and the other family is my laptop not even getting through the door."
Ethan: "Exactly. One is the supplier refusing your order. The other is your phone line being dead. You don't fix a dead phone line by arguing with the supplier."
Reading a ClientError: the fields worth logging
Every ClientError carries the parsed reply from AWS in err.response, a plain dictionary. For the IAM error on this page it looks like this:
{
"Error": {
"Type": "Sender",
"Code": "EntityAlreadyExists",
"Message": "Role with name shop-backup-role already exists."
},
"ResponseMetadata": {
"RequestId": "4d5c9a1e-...",
"HTTPStatusCode": 409,
"HTTPHeaders": {...},
"RetryAttempts": 0
}
}
Four fields earn a place in your logs. Code is what your program decides on. Message is what a person reads, and it often names the exact resource or limit. HTTPStatusCode tells you the category at a glance: 400s are your request, 500s are the service. RequestId identifies this exact call in AWS's logs and is what AWS Support asks for when you open a case. Add err.operation_name, which holds the API call name (CreateRole), and RetryAttempts, which tells you whether boto3 already retried before giving up. A non-zero count on an "already exists" error is the fingerprint of the retry-after-timeout cause.
A small helper turns all of that into one readable line. Use .get() everywhere, because some errors, such as the bare 404 from an S3 HEAD request, arrive with fewer fields:
import logging
log = logging.getLogger("shop-setup")
def describe(err):
e = err.response.get("Error", {})
meta = err.response.get("ResponseMetadata", {})
return (f"{err.operation_name} failed: {e.get('Code')} "
f"(HTTP {meta.get('HTTPStatusCode')}, request {meta.get('RequestId')}, "
f"retries {meta.get('RetryAttempts', 0)}): {e.get('Message')}")
try:
iam.create_role(RoleName="shop-backup-role", AssumeRolePolicyDocument=trust_json)
except ClientError as err:
log.error(describe(err))
raise
The output reads like CreateRole failed: EntityAlreadyExists (HTTP 409, request 4d5c9a1e-..., retries 0): Role with name shop-backup-role already exists., which is everything you need to fix the problem or hand it to someone else.
The same structure answers most "botocore.exceptions.ClientError: An error occurred (…)" searches. The word in parentheses is the code; search it plus the operation name. (AccessDenied) on CreateRole means your identity lacks permission to create roles at all, which our guide to "not authorized to perform iam:CreateRole" walks through. (ExpiredToken) and (InvalidClientTokenId) are credential problems, linked at the end of this page. (403) or (404) with nothing else is almost always S3 answering a HEAD request, covered below.
Catching by name: client.exceptions and the boto3 exceptions list
Every boto3 client carries an exceptions attribute holding one class per error that service describes. That is what makes except iam.exceptions.EntityAlreadyExistsException possible. Three details make it pleasant to use.
The classes are subclasses of ClientError. Catching one gives you the same err.response and err.operation_name as a plain ClientError, so the logging helper above works unchanged.
You can print the full list, with codes. This is the honest answer to "where is the boto3 exceptions list": it lives inside each client, built from the service model, so it is always current for your installed version. Two lines show every class and the code it matches:
iam = boto3.client("iam")
for shape in iam.meta.service_model.error_shapes:
print(f"{shape.name:50} code: {shape.error_code}")
With boto3 1.43, IAM lists 38 errors, including the mismatched pairs in the table above. Swap "iam" for "s3", "dynamodb" or any other service name to see its list.
You can go from a code to a class. iam.exceptions.from_code("EntityAlreadyExists") returns the EntityAlreadyExistsException class. For a code the service never described, it returns plain ClientError, which is a useful hint that the error you are seeing comes from somewhere else, such as an S3 HEAD request.
If you use the resource interface (boto3.resource("s3")) instead of clients, the exceptions live one level down, on the client the resource wraps:
s3 = boto3.resource("s3")
try:
s3.Object("shop-photos", "2026/10/front-window.jpg").get()
except s3.meta.client.exceptions.NoSuchKey:
print("No such photo")
Which style to choose? Catching classes is clearer when you handle one or two expected errors. Comparing codes inside a single except ClientError is better when you handle many, or when the error might come from an operation that returns codes outside the model, like S3 HEAD. Many teams use both: classes for the expected case, a ClientError fallback for logging everything else.
The fix: create-or-get, so a second run is boring
Here is the function Ethan put in Jake's script. It tries to create the role. If IAM says the name is taken, it fetches the existing role instead and returns its ARN, along with a flag saying whether this run created it:
import json
import boto3
iam = boto3.client("iam")
TRUST = {
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Principal": {"Service": "lambda.amazonaws.com"},
"Action": "sts:AssumeRole",
}],
}
def ensure_role(name, trust_policy):
"""Create the role, or return the existing one. Returns (arn, created)."""
try:
resp = iam.create_role(
RoleName=name,
AssumeRolePolicyDocument=json.dumps(trust_policy),
)
return resp["Role"]["Arn"], True
except iam.exceptions.EntityAlreadyExistsException:
role = iam.get_role(RoleName=name)["Role"]
if role["AssumeRolePolicyDocument"] != trust_policy:
iam.update_assume_role_policy(
RoleName=name,
PolicyDocument=json.dumps(trust_policy),
)
return role["Arn"], False
arn, created = ensure_role("shop-backup-role", TRUST)
print(arn, "created" if created else "already existed")
Three things about it are deliberate.
It creates first and asks second. The tempting version checks with get_role and only creates if the role is missing. That costs an extra call on every run, and worse, it leaves a gap between the check and the create where a parallel run or a retried request can sneak in. The create-first version has no gap: IAM itself decides, atomically, whether the name is free.
It makes the existing role match. A role that already exists might not be the role you want. Someone may have edited its trust policy in the console, or an older version of your script may have created it differently. The function compares the trust policy and puts yours back if it differs. A nice detail helps here: boto3 hands you AssumeRolePolicyDocument from get_role already decoded into a Python dictionary, so you compare dict to dict, with no JSON parsing of your own. Do the same for attached policies if your setup depends on them, using list_attached_role_policies and attach_role_policy, which is safe to repeat.
It reports whether it created anything. The created flag lets the caller wait only when a role is brand new (next section) and makes your logs honest about what a run changed.
The same shape works for the other IAM objects. A customer managed policy needs one extra step, because get_policy wants an ARN rather than a name. Build it from your own identity, which also gets the partition right in AWS GovCloud and China accounts:
sts = boto3.client("sts")
def ensure_managed_policy(name, document):
try:
resp = iam.create_policy(PolicyName=name, PolicyDocument=json.dumps(document))
return resp["Policy"]["Arn"], True
except iam.exceptions.EntityAlreadyExistsException:
ident = sts.get_caller_identity()
parts = ident["Arn"].split(":") # arn:PARTITION:sts::ACCOUNT:...
partition = parts[1]
arn = f"arn:{partition}:iam::{ident['Account']}:policy/{name}"
iam.get_policy(PolicyArn=arn) # raises NoSuchEntity if it lives under another path
return arn, False
Instance profiles, the wrapper EC2 needs around a role, have two steps that can each hit an existing object: the profile itself, and the role inside it. Handle the first with the same catch, then look before adding the role, since a profile holds a single role:
def ensure_instance_profile(name, role_name):
try:
iam.create_instance_profile(InstanceProfileName=name)
except iam.exceptions.EntityAlreadyExistsException:
pass
iam.get_waiter("instance_profile_exists").wait(InstanceProfileName=name)
profile = iam.get_instance_profile(InstanceProfileName=name)["InstanceProfile"]
if role_name not in [r["RoleName"] for r in profile["Roles"]]:
iam.add_role_to_instance_profile(InstanceProfileName=name, RoleName=role_name)
return profile["Arn"]
For users, the pattern is identical to roles with create_user and get_user. For console passwords, EntityAlreadyExists on create_login_profile means the user already has one; call update_login_profile if you meant to change it.
Jake: "And when should I actually delete the old role and start fresh?"
Ethan: "Almost never from a script. Deleting a role breaks everything that uses it, right now, and other things might be using it that you don't know about. If the existing role is wrong, fix it in place, like the function does. If it belongs to someone else, pick a different name. Deleting is a decision for a person looking at the console, not for a loop."
Right after creating a role: waiters and IAM's propagation delay
Creating the role is half the job. The other half is using it, and IAM has a quirk that catches nearly everyone who automates it: changes take a little while to reach every part of AWS. IAM is eventually consistent. A role you created a moment ago might not be visible to the next get_role call yet, and might not be usable by another service for a bit longer.
For the first part, boto3 ships waiters: small loops that poll until something is true. IAM has four of them:
| Waiter name | What it calls | Default polling |
|---|---|---|
role_exists | GetRole until it stops returning NoSuchEntity | Every 1 second, up to 20 tries |
user_exists | GetUser until it stops returning NoSuchEntity | Every 1 second, up to 20 tries |
policy_exists | GetPolicy until it stops returning NoSuchEntity | Every 1 second, up to 20 tries |
instance_profile_exists | GetInstanceProfile until it stops returning 404 | Every 1 second, up to 40 tries |
arn, created = ensure_role("shop-backup-role", TRUST)
if created:
iam.get_waiter("role_exists").wait(
RoleName="shop-backup-role",
WaiterConfig={"Delay": 2, "MaxAttempts": 30},
)
If the waiter runs out of tries, it raises WaiterError with a message like Waiter RoleExists failed: Max attempts exceeded. Previously accepted state: Matched expected service error code: NoSuchEntity. Catch it from botocore.exceptions if you want a friendlier message.
Now the second part, which is the real shock of this section: the waiter passing doesn't mean every service can use the role yet. The waiter proves IAM itself can return the role. A different service, handed the brand-new role a moment later, can still refuse it. Lambda is the famous case: create_function with a role created seconds earlier fails with InvalidParameterValueException and the message "The role defined for the function cannot be assumed by Lambda," even after role_exists succeeded. The same message also appears when the trust policy really is wrong, which sends people on a long hunt through a policy that is perfectly fine.
The reliable fix is a short retry loop around the call that consumes the role, only for that specific error, only for a role this run just created:
import time
from botocore.exceptions import ClientError
lam = boto3.client("lambda")
def create_function_patiently(tries=8, pause=5, **kwargs):
for attempt in range(1, tries + 1):
try:
return lam.create_function(**kwargs)
except ClientError as err:
e = err.response["Error"]
new_role_lag = (e["Code"] == "InvalidParameterValueException"
and "cannot be assumed" in e.get("Message", ""))
if not new_role_lag or attempt == tries:
raise
time.sleep(pause)
If the loop exhausts all its tries, the problem is no longer timing; check that the trust policy names lambda.amazonaws.com as the principal, as the TRUST document above does.
Boto3 retry config: legacy, standard and adaptive, and why retries cause "already exists"
boto3 retries failed calls on its own, before your code ever sees an error. That is usually a kindness, and occasionally the reason for a puzzling EntityAlreadyExists. Understanding the three retry modes takes five minutes and pays back every time a call fails under load.
| Mode | Total attempts by default | Backoff | Notes |
|---|---|---|---|
| legacy | 5 (the first try plus 4 retries) | Exponential: a random fraction of a second, doubling each retry | The default in boto3 when you set nothing |
| standard | 3 (the first try plus 2 retries) | Exponential with jitter, capped at 20 seconds | A wider, consistent list of throttling and transient errors; a retry budget that stops retry storms |
| adaptive | 3, like standard | Like standard | Adds client-side rate limiting that slows your own calls when AWS throttles you |
Yes, the default is still legacy. A fresh client with no configuration reports {'mode': 'legacy'} in client.meta.config.retries. Switching is one line, and standard is the better choice for most code:
import boto3
from botocore.config import Config
cfg = Config(retries={"mode": "standard", "total_max_attempts": 5})
iam = boto3.client("iam", config=cfg)
Watch the two attempt settings, because they count differently. total_max_attempts counts every attempt including the first, so 5 means one try and four retries. The older max_attempts key inside retries counts only retries, so 5 there means six attempts in total. If you pass both, total_max_attempts wins. The same settings work without code changes through environment variables (AWS_RETRY_MODE=standard and AWS_MAX_ATTEMPTS=5, where AWS_MAX_ATTEMPTS counts total attempts) or in ~/.aws/config under your profile as retry_mode = standard and max_attempts = 5.
What standard mode retries is a short list: throttling codes (Throttling, ThrottlingException, TooManyRequestsException, ProvisionedThroughputExceededException, RequestLimitExceeded, SlowDown and a handful more), transient codes (RequestTimeout, RequestTimeoutException, PriorRequestNotComplete), HTTP 500, 502, 503 and 504, connection or read timeouts, and any error a service's model marks as retryable. What never gets retried: errors that mean your request itself is wrong or conflicts with what exists, including EntityAlreadyExists, AccessDenied, NoSuchEntity and validation errors. Retrying those would only produce the same answer more slowly.
Now the double-create, in slow motion:
- Your
create_rolerequest reaches IAM, and IAM makes the role. - The response is on its way back when the connection stalls past the read timeout (60 seconds by default, set with
Config(read_timeout=…)). - boto3 sees a
ReadTimeoutError, which both legacy and standard modes treat as temporary, and sends the identical request again. - IAM, correctly, says the role already exists.
- Your code receives EntityAlreadyExists, with
RetryAttemptsof 1 or more in the metadata.
Nobody did anything wrong, and the create-or-get function turns the whole episode into a non-event: it catches the error, fetches the role your first attempt made, and returns it.
Throttling deserves one more line, because scripts that create many users or roles in a tight loop can run into it. Switching to standard or adaptive mode and raising total_max_attempts is the first fix; spacing the loop out is the second. Our guide to ThrottlingException on Bedrock walks through the same retry settings under heavy load.
Jake: "So boto3 tried to be helpful, sent the order twice, and then blamed me for the duplicate."
Ethan: "Like a courier who never got your signature, so he delivers the parcel again, and the warehouse says you already have one. The fix isn't firing the courier. It's a receiving desk that says 'oh, that's ours, thanks' instead of panicking."
S3 and DynamoDB error handling: the two errors that don't look like the rest
The create-or-get idea, try the action and handle the expected error, works across AWS. Two services have quirks that trip up the obvious version of it.
S3 head_object: a bare 404, not NoSuchKey
The most common way to ask "does this file exist in S3" is head_object. When the object is missing, you would expect s3.exceptions.NoSuchKey. You get a plain ClientError with the code "404" and the message "Not Found". The reason is mechanical: a HEAD response has no body, so S3 can't send the XML document that normally carries the error code. boto3 falls back to using the HTTP status number as the code. s3.exceptions.from_code("404") returns plain ClientError, so an except s3.exceptions.NoSuchKey around head_object never fires. The same goes for 403, 400, 412 and the rest: on HEAD, you only ever get the number.
import boto3
from botocore.exceptions import ClientError
s3 = boto3.client("s3")
def s3_key_exists(bucket, key):
try:
s3.head_object(Bucket=bucket, Key=key)
return True
except ClientError as err:
if err.response["Error"]["Code"] in ("404", "NoSuchKey"):
return False
raise # 403 and everything else: a real problem, not "missing"
Two more traps hide in that function. First, a missing key turns into 403, not 404, when your identity lacks the s3:ListBucket permission on the bucket; S3 won't confirm what isn't there to someone who can't list. That is why the function re-raises 403 instead of returning False: a permission problem shouldn't be reported as "file missing." Second, on HEAD a missing bucket also reads as 404, so if the bucket name could be wrong, check it once with head_bucket. get_object, which does get a body back, raises the proper s3.exceptions.NoSuchKey. Our page on S3 "The specified key does not exist" covers every cause of that one, and S3 also has a waiter, object_exists, that polls head_object every 5 seconds up to 20 times.
DynamoDB: "create if not exists" is a condition, and its error has the suffix
DynamoDB doesn't have names to collide on; it has items. The equivalent of create-or-get is a write with a condition that only succeeds when no item with that key exists yet. The error this time is ConditionalCheckFailedException, and since DynamoDB codes include the suffix, the class name and the code string match:
ddb = boto3.client("dynamodb", region_name="us-east-1")
def put_if_absent(table, item):
try:
ddb.put_item(
TableName=table,
Item=item,
ConditionExpression="attribute_not_exists(pk)",
)
return True
except ddb.exceptions.ConditionalCheckFailedException:
return False
put_if_absent("bookings", {"pk": {"S": "booking#1042"}, "name": {"S": "Jake"}})
Add ReturnValuesOnConditionCheckFailure="ALL_OLD" to the call and the exception carries the existing item in err.response["Item"], so you get the "get" half of create-or-get without a second request. Our full guide to DynamoDB ConditionalCheckFailedException covers optimistic locking, transactions and the other jobs a condition does.
EntityAlreadyExists in Terraform, CloudFormation and CDK
Infrastructure tools hit the same IAM error inside their own error formats. The causes are the same six; the fixes use each tool's own way of saying "this role already exists, adopt it."
Terraform
Terraform uses the Go SDK, so the message looks different but carries the same code:
Error: creating IAM Role (shop-backup-role): operation error IAM: CreateRole, https response error
StatusCode: 409, RequestID: …, EntityAlreadyExists: Role with name shop-backup-role already exists.
This almost always means the role exists in the account but not in Terraform's state: created by hand, by another workspace, or left behind when state was lost. If the role is meant to be managed by this configuration, import it instead of deleting it. On Terraform 1.5 and later, add an import block and run a plan; on older versions use the command:
import {
to = aws_iam_role.backup
id = "shop-backup-role"
}
# or, on any version:
terraform import aws_iam_role.backup shop-backup-role
The import ID for a role is its name. If the role belongs to something else and your configuration only needs a role of its own, stop hard-coding name: leave it out and Terraform assigns a random, unique name, or use name_prefix for a readable start with a unique tail. Either way, two workspaces or two Regions can never collide again.
CloudFormation and CDK
In CloudFormation, RoleName on AWS::IAM::Role is optional, and leaving it out is the strongest fix: CloudFormation then generates a unique name for the role, so the same template can be deployed any number of times. Setting a name has two costs. You must acknowledge CAPABILITY_NAMED_IAM on every deploy, and the same template deployed to a second Region fails, because the role name is already taken account-wide. If you need a fixed name, build the Region into it, for example with Fn::Join and AWS::Region. CDK synthesizes CloudFormation, so the same rule applies: leave the role name unset unless something outside the stack has to find the role by name. When the role already exists and you want the stack to own it, CloudFormation can import it; our guide to CloudFormation "already exists" errors covers the import routes step by step.
Before AWS can answer: no module named boto3, no Region, SSL errors
A good share of boto3 errors happen before any request leaves your computer. They aren't AWS errors at all, and they are fast to fix once you know which is which.
ModuleNotFoundError: No module named 'boto3'
boto3 isn't installed in the Python that is running your script. The usual reason is two Pythons on one machine: you installed boto3 with one and ran the script with another. Install it with the same interpreter you run, which python3 -m pip guarantees:
python3 -m pip install boto3
# inside a virtual environment, activate it first, then:
python -m pip install boto3
To check whether boto3 is installed, and which version, ask the interpreter you plan to use: python3 -c "import boto3; print(boto3.__version__)". python3 -m pip show boto3 prints the version and install location too. In VS Code, "Import "boto3" could not be resolved" is the editor's language server looking at a different interpreter than the one you installed into; pick the right one with Python: Select Interpreter from the command palette, and the warning clears.
NoRegionError: You must specify a region
Regional services need a Region before boto3 can build an address. IAM, STS and S3 have global defaults, so boto3.client("iam") works with no Region configured; DynamoDB, Lambda, EC2 and most others raise NoRegionError. Pass region_name="us-east-1" to the client, set AWS_DEFAULT_REGION, or put region = us-east-1 in your profile in ~/.aws/config. This is why a script can create the IAM role happily and then fail on the very next line, creating the Lambda function.
NoCredentialsError: Unable to locate credentials
boto3 searched every place it knows (environment variables, the shared credentials and config files, SSO, container and instance roles) and found nothing. It is common in cron jobs and Docker containers, which don't inherit your shell's setup. Our guide to boto3 NoCredentialsError in Docker and cron walks through the search order and the fix for each environment.
botocore.exceptions.SSLError: SSL validation failed
Your machine couldn't validate AWS's certificate chain. Behind a corporate proxy that inspects HTTPS traffic, the proxy presents its own certificate, which Python's certificate bundle doesn't trust. Point boto3 at your company's CA bundle with the AWS_CA_BUNDLE environment variable, or ca_bundle in your profile, or the verify="/path/to/company-ca.pem" argument on a client. Turning verification off with verify=False makes the error disappear and leaves your traffic open to interception, so keep it for a quick diagnosis at most.
Testing your boto3 error handling without an AWS account
Error paths are the code you run least and trust most, which is a bad combination. botocore ships a tool called Stubber that answers your client's calls with responses you script in advance, without sending anything to AWS. It is the easiest way to prove that your create-or-get function handles a second run before you find out in production:
import copy
from botocore.stub import Stubber
EXISTING = {
"Path": "/",
"RoleName": "shop-backup-role",
"RoleId": "AROAEXAMPLEEXAMPLE123",
"Arn": "arn:aws:iam::123456789012:role/shop-backup-role",
"CreateDate": "2026-10-11T00:00:00Z",
"AssumeRolePolicyDocument": json.dumps(TRUST),
}
def test_second_run_reuses_the_role():
stub = Stubber(iam)
stub.add_client_error(
"create_role",
service_error_code="EntityAlreadyExists",
service_message="Role with name shop-backup-role already exists.",
http_status_code=409,
)
stub.add_response("get_role", {"Role": copy.deepcopy(EXISTING)})
with stub:
arn, created = ensure_role("shop-backup-role", TRUST)
assert created is False
assert arn.endswith(":role/shop-backup-role")
Two lessons from writing tests like this. First, use the real code string in service_error_code, "EntityAlreadyExists"; the stubbed error then raises the same EntityAlreadyExistsException class a real one would, which makes the test a check on the trap from earlier. Second, give each stubbed response its own fresh dictionary (that is what copy.deepcopy is for). boto3 decodes the policy document in place, so a response dictionary reused for a second call arrives already decoded and the test fails with a confusing error about a dict having no split.
Jake: "So I can break AWS on purpose without touching AWS."
Ethan: "That's the idea. It's the practice phone you hand a new hire. They learn what to do when a customer shouts, and no real customer gets shouted at."
Still stuck? Troubleshooting by symptom
My except block never runs
Work through three causes, in this order:
- You are comparing the code to the class name (
"EntityAlreadyExistsException"instead of"EntityAlreadyExists"). - A broader except clause above yours catches the error first, so move specific classes to the top.
- The error is raised somewhere you didn't wrap, such as a waiter (WaiterError) or a later call that uses the role.
Print type(err) and err.response["Error"]["Code"] once and the cause is obvious.
IAM says the role exists, but I can't find it
Check which account you are in with aws sts get-caller-identity; a different profile means a different set of roles. Then search for the name ignoring case, because Shop-Backup-Role blocks shop-backup-role. Region doesn't matter for IAM; switching Regions in the console shows the same roles.
It fails only in CI, or only sometimes
That pattern points to two runs racing each other or to an automatic retry after a slow response. Look at RetryAttempts in the error metadata: anything above zero means boto3 resent the request. The create-or-get pattern handles both; limiting CI to one deploy at a time handles the race at its source.
EntityAlreadyExists on create_login_profile
The user already has a console password. If your goal is to set a new one, call update_login_profile instead; if your goal is to make sure a password exists, treat the error as success.
LimitExceeded instead of EntityAlreadyExists
You have reached an account quota, such as the number of roles, policies or policy versions allowed. The message names the quota. Delete what you no longer use or request a higher quota; retrying won't help, and boto3 doesn't retry it.
ConcurrentModification when creating or editing a role
Two changes reached the same IAM object at the same moment, often two pipelines or a script running in parallel threads. Wait a little and retry that call, and serialize IAM changes where you can.
MalformedPolicyDocument right after fixing EntityAlreadyExists
Once the name clash is out of the way, the next most common CreateRole failure is a trust policy IAM can't accept: a missing Version, a principal written wrong, or text that isn't valid JSON. (A Python dict passed without json.dumps never reaches IAM; boto3 stops it first with ParamValidationError, "Invalid type for parameter AssumeRolePolicyDocument.") Our guide to MalformedPolicyDocument is in the card below.
Boto3 error handling and IAM EntityAlreadyExists: frequently asked questions
What does EntityAlreadyExists mean in AWS?
IAM refused to create a role, user, group, managed policy, instance profile or similar object because one with that name already exists in the account. Role, user, group and policy names are unique per account across all Regions and ignore case. Nothing was changed. The HTTP status is 409 Conflict.
How do I fix "Role with name already exists" in AWS?
Make the create step tolerate an existing role: call create_role, and when it raises EntityAlreadyExistsException, fetch the role with get_role and continue. If the role belongs to something else, choose another name. Delete the existing role only after confirming nothing uses it.
How do I catch EntityAlreadyExists in boto3?
Catch the modeled class with except iam.exceptions.EntityAlreadyExistsException, where iam is your boto3 IAM client. Or catch botocore.exceptions.ClientError and check that err.response["Error"]["Code"] equals "EntityAlreadyExists", without the Exception suffix.
Why is my try except not catching the boto3 exception?
Usually the code comparison is wrong: IAM sends EntityAlreadyExists while the class is EntityAlreadyExistsException. Other causes are a broader except clause above yours, or the error coming from a call outside the try block, such as a waiter or a later API call.
What is botocore.exceptions.ClientError?
It is the exception boto3 raises when an AWS service receives a request and returns an error. It carries the parsed reply in err.response (Error.Code, Error.Message, ResponseMetadata) and the API call name in err.operation_name. Every service-specific exception class is a subclass of it.
How do I get the error code from a boto3 ClientError?
Read err.response["Error"]["Code"]. The message is in err.response["Error"]["Message"], the HTTP status in err.response["ResponseMetadata"]["HTTPStatusCode"], and the request ID in err.response["ResponseMetadata"]["RequestId"]. Use .get() for safety, since some errors have fewer fields.
Where can I find the list of boto3 exceptions?
Each client carries its own list. Loop over client.meta.service_model.error_shapes and print shape.name and shape.error_code to see every class and its code. botocore.exceptions holds the client-side errors, and boto3.exceptions holds a few helper errors such as S3UploadFailedError.
What is the default retry mode in boto3?
Legacy, which makes up to five attempts in total, the first try plus four retries, with exponential backoff. Standard mode makes three attempts by default with jittered backoff capped at 20 seconds. Adaptive mode adds client-side rate limiting on top of standard.
How do I configure retries in boto3?
Pass a botocore Config to the client: Config(retries={"mode": "standard", "total_max_attempts": 5}). total_max_attempts counts the first try; the older max_attempts key counts retries only. AWS_RETRY_MODE and AWS_MAX_ATTEMPTS, or retry_mode and max_attempts in ~/.aws/config, do the same without code.
Is EntityAlreadyExists a retryable error?
No. It means the request conflicts with something that exists, so retrying returns the same answer. boto3 doesn't retry it. It can appear because of a retry, though: when a create call times out after IAM already made the role, the automatic retry finds that role.
Why does head_object return 404 instead of NoSuchKey?
A HEAD response has no body, so S3 can't send its usual XML error code. boto3 uses the HTTP status as the code, so the error is ClientError with Code "404". Check for "404" in your handler. A missing key returns 403 instead when you lack s3:ListBucket.
How do I fix NoRegionError: You must specify a region in boto3?
Pass region_name to the client, such as boto3.client("dynamodb", region_name="us-east-1"), set the AWS_DEFAULT_REGION environment variable, or add region to your profile in ~/.aws/config. IAM, STS and S3 work without one; most other services don't.
How do I check if boto3 is installed?
Run python3 -c "import boto3; print(boto3.__version__)" with the same interpreter your script uses. It prints the version, or raises ModuleNotFoundError if boto3 is missing. python3 -m pip show boto3 shows the version and install location.
How do I fix No module named boto3?
Install boto3 into the Python that runs your script: python3 -m pip install boto3, or activate your virtual environment first and run python -m pip install boto3. The error usually means boto3 went into a different Python on the same machine.
What does botocore.exceptions.SSLError mean?
Python couldn't validate the certificate chain for the AWS endpoint, most often because a corporate proxy inspects HTTPS. Point boto3 at your company CA bundle with AWS_CA_BUNDLE, ca_bundle in your profile, or verify="path/to/ca.pem" on the client. Turning verification off removes protection.
How do I fix EntityAlreadyExists in Terraform?
The role exists but isn't in Terraform state. Import it with an import block (to = aws_iam_role.name, id = "role-name") on Terraform 1.5 or later, or terraform import aws_iam_role.name role-name. If you only need a role of your own, drop name and let Terraform generate one.
How do I fix EntityAlreadyExists in CloudFormation?
Remove RoleName from AWS::IAM::Role so CloudFormation generates a unique name, which also lets the template deploy to several Regions. If you need a fixed name, include the Region in it. To adopt an existing role into a stack, use CloudFormation resource import.
How do I wait until an IAM role exists in boto3?
Use iam.get_waiter("role_exists").wait(RoleName="name"), which polls GetRole every second up to 20 times by default. Other services may still reject a brand-new role briefly, so wrap the call that uses it, such as Lambda create_function, in a short retry loop.
Jake's script now has three small functions where a straight line of create calls used to be. It ran for the second shop without a word of complaint, printed "already existed" for the role and the policy, and set up the new shop's bucket and function in under a minute. He ran it a third time just to watch nothing happen, and grinned when nothing did. Ethan's parting advice was the line below, and a promise that the day Jake opens a third shop, the script won't need him at all.
📌 If you keep one line from this page
Catch the class or compare the code, never the class name as a string, and write every create so that running it twice is boring.
IAM names are global and ignore case; boto3 retries for you, so "already exists" is sometimes your own success arriving twice.
Revision note. Written October 11, 2026, for everyone whose setup script worked exactly once. "Measure twice, cut once," goes the carpenter's proverb; good automation is the rare craft where cutting twice is safe too.