EKS ImagePullBackOff: Read the Real Error, Then Fix It

Logeshwaran
—

An EKS pod stuck in ImagePullBackOff means the node tried to download the container image, failed, and is now waiting longer and longer between retries. The status itself tells you nothing about why. The reason is one line further down, in the pod's events, and on EKS it is almost always one of six things: the node's IAM role cannot pull from ECR, the image tag does not exist, the node sits in a private subnet with no path to the registry, you hit Docker Hub's pull limit, the image was built for the wrong CPU architecture, or a repository policy says no. Each one prints a different sentence in the events. This page is the decoder, then the fix for each, in the order that wastes the least time.

Ethan hit it the expensive way. He moved a node group to Graviton to cut the cluster bill, and the next deployment went straight to ImagePullBackOff with a message about "no match for platform". Jake hit it the cheap way the same week: his first ever deployment on Ethan's cluster pulled a public image from Docker Hub, and at eleven in the morning every new pod started failing with a sentence about a pull rate limit, even though Jake had pulled that image exactly once. The surprise in Jake's case is the one most people never work out: all of Ethan's nodes sit behind a single NAT gateway, so Docker Hub sees the whole cluster as one address, and the limit is per address.

⚡ Quick Answer

• Read the real error first → kubectl describe pod <pod> and look at the Failed to pull image line under Events. The status is a symptom; that line is the cause. Decoder table.

• 401 or "no basic auth credentials" → the node role lacks ECR pull permission. Attach AmazonEC2ContainerRegistryPullOnly to the node IAM role. Steps.

• "manifest unknown" or "not found" → the tag is wrong or was deleted. Check with aws ecr describe-images. Steps.

• "i/o timeout" or "dial tcp" → private subnet with no NAT or missing ECR and S3 VPC endpoints. Steps.

• "toomanyrequests" → Docker Hub limit; mirror the image into ECR or use a pull-through cache. Steps.

After any fix, restart the deployment with kubectl rollout restart deployment <name>; the pod will not retry faster on its own.

If the pod says CrashLoopBackOff instead, the image pulled fine and the container is dying after it starts; that is a different page. If it says ErrImagePull, you are reading the same problem a few seconds earlier, before the backoff kicked in. Everything below applies to both.

What ImagePullBackOff means, and why the status hides the cause

When the scheduler places a pod on a node, the kubelet on that node asks its container runtime to fetch the image named in the pod spec. If the fetch fails, the pod goes to ErrImagePull. The kubelet retries, and each failure doubles the wait: 10 seconds, 20, 40, up to a cap of five minutes. While it is waiting, the status reads ImagePullBackOff. That is the whole mechanism. The status means "I am between retries", nothing more.

The reason for the failure is recorded as a Warning event on the pod, and it is the first thing to read, every time:

kubectl describe pod my-pod -n my-namespace
# or, just the events, newest last:
kubectl get events -n my-namespace --field-selector involvedObject.name=my-pod --sort-by=.lastTimestamp

You are looking for a line that starts Failed to pull image "...". Everything after the image name is the registry's own explanation, passed through by the runtime. Two things to note before you decode it. First, the image name in that line is what the node actually asked for, so a typo in the deployment shows up here before anywhere else. Second, the Back-off pulling image event that follows is noise; it only repeats that a retry is pending.

The decoder: what the event text means on EKS

What the event saysWhat it meansGo to
401 Unauthorized, no basic auth credentials, authorization failedThe node could not get a valid ECR token. Its IAM role lacks the pull permissions, or the image is in another AWS account that has not allowed this role.IAM and the node role
403 ForbiddenAuthenticated, then refused. An ECR repository policy, a service control policy, or an explicit IAM deny is blocking the node role.IAM and the node role
manifest unknown, manifest for ...:tag not found, not foundThe registry answered, but no image has that tag in that repository. Typo, wrong account or Region in the URI, or a lifecycle policy deleted the tag.Tag and URI
pull access denied ... repository does not exist or may require 'docker login'Docker Hub's wording for either a private repository with no credentials or a misspelled name.Docker Hub
toomanyrequests: You have reached your pull rate limitDocker Hub's anonymous limit, counted per public IP. Your whole cluster shares the NAT gateway's IP.Docker Hub
dial tcp ...:443: i/o timeout, context deadline exceeded, no such hostThe node cannot reach the registry at all. Private subnet with no NAT gateway, VPC endpoints missing or not attached to the subnet, or a security group on the endpoint blocking 443.Network path
no match for platform in manifestThe image was built for one CPU architecture and the node is the other. Typically an amd64-only image on a Graviton (arm64) node.Architecture
x509: certificate signed by unknown authorityA private registry with a self-signed certificate, or a corporate proxy intercepting TLS. Not an ECR problem.Private registries
Only Back-off pulling image, no Failed line visibleThe Failed event has scrolled off. Delete the pod so it is recreated and read the fresh events.Read events

Cause 1: the node role cannot pull from ECR (401, 403, no basic auth)

On EKS, nodes do not log in to ECR with a username and password. The kubelet uses the EC2 instance's IAM role to request a short-lived registry token, and that request is an IAM call. If the role is missing the ECR actions, the token request fails, and the runtime reports it as 401 Unauthorized or no basic auth credentials, which sends people off to check their own docker login. Your laptop is not involved. The node is.

Step 1: find the node role

# which node is the pod on?
kubectl get pod my-pod -n my-namespace -o wide

# managed node group: the role EKS launched the nodes with
aws eks describe-nodegroup --cluster-name my-cluster --nodegroup-name my-ng \
  --query nodegroup.nodeRole --output text

# any node: ask EC2 which instance profile it carries
aws ec2 describe-instances --filters "Name=private-dns-name,Values=ip-10-0-1-23.ec2.internal" \
  --query "Reservations[].Instances[].IamInstanceProfile.Arn" --output text

Step 2: check what the role can do

aws iam list-attached-role-policies --role-name myAmazonEKSNodeRole

You want to see one of two managed policies. AmazonEC2ContainerRegistryPullOnly is the right one for nodes; it exists since October 4, 2024 and grants exactly four actions: ecr:GetAuthorizationToken, ecr:BatchGetImage, ecr:GetDownloadUrlForLayer and ecr:BatchImportUpstreamImage, the last of which is what pull-through caches need. AmazonEC2ContainerRegistryReadOnly is the older, broader one that also lets the role list and describe repositories; it works too, and most clusters created before late 2024 carry it. If neither is there, attach the narrow one:

aws iam attach-role-policy --role-name myAmazonEKSNodeRole \
  --policy-arn arn:aws:iam::aws:policy/AmazonEC2ContainerRegistryPullOnly

No node restart is needed; the kubelet requests a fresh token on the next pull. Restart the deployment so it retries now rather than in five minutes. If attaching the policy itself fails with an access-denied on iam:PassRole or iam:AttachRolePolicy, that is your identity's problem, not the node's, and the PassRole guide untangles it.

Step 3: if the policy is there and you still get 401 or 403

Three places can still say no:

  1. The image lives in a different AWS account. Pulling cross-account needs a repository policy on the ECR repository in the owning account that allows the node role (or the whole node account) to call ecr:BatchGetImage, ecr:GetDownloadUrlForLayer and ecr:BatchCheckLayerAvailability. The node role's own permissions are necessary but not sufficient. Check with aws ecr get-repository-policy --repository-name app in the owning account; if it returns "does not exist", there is no cross-account allow and the pull will be refused.
  2. A repository policy with an explicit Deny. This is the case the EKS workshop uses as its worked example: identity permissions fine, repository policy denying the node role, result 403 Forbidden. Fix the policy with aws ecr set-repository-policy, then restart the deployment.
  3. A service control policy or permissions boundary. If your account sits in an Organization, an SCP that restricts ECR to certain Regions or denies ecr:GetAuthorizationToken for roles outside a tag will produce 401 or 403 on the node while every console check looks fine. The AccessDenied decoder shows how to read the policy chain.

One more trap for Fargate: Fargate pods pull with the pod execution role, not a node role, because there is no node. Attach the ECR policy to that role instead. The Fargate primer explains where that role comes from.

Cause 2: the tag or URI does not exist (manifest unknown, not found)

The registry answered, authenticated you, and then said it has no such image. This is the cheapest cause and the one worth ruling out before anything else, because a single wrong character produces it. Read the image URI from the Failed event character by character against reality:

# does the repository exist in this account and Region?
aws ecr describe-repositories --region us-east-1 --query "repositories[].repositoryUri"

# does the tag exist?
aws ecr describe-images --region us-east-1 --repository-name app --image-ids imageTag=1.2.1

# what tags does it have?
aws ecr list-images --region us-east-1 --repository-name app --query "imageIds[].imageTag"

Things that bite here, in rough order of how often Ethan has seen them on other people's clusters:

  • The Region in the URI. 123456789012.dkr.ecr.us-east-1.amazonaws.com/app:1.2.1 and the same string with us-west-2 are two different registries. A build pipeline that pushes to one Region while the cluster runs in another produces "manifest unknown" forever. ECR replication can mirror the repository across Regions if you want the same URI to work everywhere.
  • The account ID in the URI. Copying a deployment manifest from a staging account into production keeps the staging account's URI. Production nodes then try to pull cross-account without a repository policy and get 401, or the repository does not exist there and they get "not found".
  • A lifecycle policy expired the tag. ECR lifecycle rules such as "keep the last 30 tagged images" quietly delete older tags. A rollback to a six-month-old tag then fails. aws ecr get-lifecycle-policy tells you the rule; the only fix is to rebuild or re-push that version.
  • latest that was never pushed. ECR does not create a latest tag for you. If the pipeline tags images with the commit hash only, a manifest that says :latest finds nothing.
  • Using a digest from another registry. A @sha256:... reference is only valid in the registry that computed it; copying the image through a tool that re-compresses layers changes the digest.

Fix the manifest, apply, and the new ReplicaSet pulls the right image; the old pod can be left to disappear on its own.

Cause 3: the node cannot reach the registry (i/o timeout, dial tcp, no such host)

A timeout means the TCP connection to the registry never completed. The node's IAM and the image name are irrelevant until packets flow. On EKS this is nearly always a private-subnet story, and it comes in two shapes.

Shape A: private subnet, no NAT gateway

Nodes in a private subnet have no public IP. To reach ECR's public endpoints they need a route to a NAT gateway in a public subnet. If the route table for the node subnet has no 0.0.0.0/0 entry pointing at a nat- target, nothing outside the VPC is reachable, and every pull times out. Check the route table of the subnet the node is in:

aws ec2 describe-route-tables --filters "Name=association.subnet-id,Values=subnet-0abc123" \
  --query "RouteTables[].Routes[]"

Adding a NAT gateway fixes it and adds a line to your bill that surprises people; the NAT gateway cost guide shows what image pulls through NAT actually cost, which is the reason for Shape B.

Shape B: private subnet with VPC endpoints, one of them missing

Fully private clusters replace NAT with VPC endpoints, and ECR needs three of them, not one. The interface endpoint com.amazonaws.<region>.ecr.api handles the token and manifest calls, com.amazonaws.<region>.ecr.dkr handles the Docker-protocol pulls, and a gateway endpoint for S3 carries the actual layer downloads, because ECR stores layers in S3. A cluster with the two ECR endpoints and no S3 gateway authenticates fine, fetches the manifest fine, and then hangs on the first layer. That half-configured state is the most common private-cluster pull failure, and it reads as a timeout. eksctl's fully private mode creates all of these plus endpoints for EC2 and STS (and CloudWatch Logs if logging is on) precisely because nodes do not function without them.

Endpoint serviceTypeWhat it carriesSymptom if missing
com.amazonaws.<region>.ecr.apiInterfaceThe authorization token request and ECR API callsTimeout before any pull starts
com.amazonaws.<region>.ecr.dkrInterfaceThe registry protocol: manifests and layer URLsdial tcp ...dkr.ecr...:443: i/o timeout
com.amazonaws.<region>.s3GatewayThe image layers themselves, which ECR stores in S3Manifest fetched, then the pull hangs on the first layer
com.amazonaws.<region>.ec2InterfaceNode bootstrap and the cloud-provider integrationNodes fail to join, so pods never schedule
com.amazonaws.<region>.stsInterfaceCredentials for Fargate and IAM roles for service accountsFargate pulls and IRSA pods fail to get credentials
aws ec2 describe-vpc-endpoints --filters "Name=vpc-id,Values=vpc-0abc123" \
  --query "VpcEndpoints[].{svc:ServiceName,type:VpcEndpointType,subnets:SubnetIds,state:State}"

Check four things in that output: both ECR services are present, the S3 entry is type Gateway, the interface endpoints list the node subnets (an endpoint created in a different subnet is invisible to the nodes), and private DNS is enabled on the interface endpoints so the normal ECR hostname resolves to the endpoint's private address. Then check the endpoint's security group allows inbound 443 from the node security group or the VPC CIDR; a default-deny security group on an endpoint produces exactly the same timeout as no endpoint at all.

A quick test from inside the cluster, without SSH to the node:

kubectl run nettest --rm -it --restart=Never --image=public.ecr.aws/amazonlinux/amazonlinux:2023 -- \
  curl -sS -m 10 -o /dev/null -w "%{http_code}\n" https://123456789012.dkr.ecr.us-east-1.amazonaws.com/v2/

A 401 there is good news: the registry is reachable and merely wants a token. A hang until the ten-second limit is the network. (If even that test pod cannot pull, use any image already cached on the node; kubectl get nodes -o json | grep -o '"names":\[[^]]*' | head lists what is there.)

Cause 4: Docker Hub limits and private Docker Hub images

Jake's morning. The image was public, the manifest was right, the node role was fine, and the event said toomanyrequests: You have reached your pull rate limit. Docker Hub allows anonymous pulls of 100 per six hours, counted per public IPv4 address. Every node in a private subnet leaves the VPC through the NAT gateway, so every node, and every pod on every node, counts against the same 100. A node group that scales from three nodes to twelve at the start of the day, each pulling five base images, is sixty pulls before anyone deploys anything. Jake's single pull was the one that crossed the line, not the one that caused it.

Three fixes, from quickest to best:

  1. Authenticate the pulls. A free Docker account doubles the limit to 200 per six hours and makes it per account rather than per IP. Create a Kubernetes secret and reference it from the pod spec:
    kubectl create secret docker-registry dockerhub -n my-namespace \
      --docker-server=https://index.docker.io/v1/ --docker-username=jake --docker-password='...'
    
    # in the pod spec:
    #   imagePullSecrets:
    #   - name: dockerhub
    The same secret is how you pull private Docker Hub images, which otherwise fail with "pull access denied".
  2. Set up an ECR pull-through cache. ECR can act as a caching proxy for Docker Hub, GitHub Container Registry, Quay and others. You create a cache rule once, reference images as 123456789012.dkr.ecr.us-east-1.amazonaws.com/docker-hub/library/nginx:1.27, and ECR fetches upstream on first use and serves from the cache after that. The node role needs ecr:BatchImportUpstreamImage for the first pull, which is why the PullOnly policy includes it. Docker Hub rules require stored upstream credentials in Secrets Manager.
  3. Mirror the images you depend on into your own ECR repository as part of the build, and never reference Docker Hub from a production manifest. Slower to set up, immune to anyone else's outage or policy change.

Whichever you pick, the pods in backoff will not notice; restart the deployment after the change.

Cause 5: the image was built for the wrong architecture (no match for platform)

Ethan's case. Graviton nodes are arm64. An image built on a developer's Intel laptop, or by a CI runner, with a plain docker build is amd64 only. The registry has the image, the node is allowed to pull it, and the runtime refuses because the manifest list contains no arm64 variant. The event says no match for platform in manifest; if an older runtime pulls it anyway, the container starts and dies instantly with exec format error, which is the same mismatch one step later.

# what architecture are the nodes?
kubectl get nodes -L kubernetes.io/arch,kubernetes.io/os

# what architectures does the image offer?
docker buildx imagetools inspect 123456789012.dkr.ecr.us-east-1.amazonaws.com/app:1.2.1

Two fixes. The right one is a multi-architecture build, which produces one tag that works on both:

docker buildx build --platform linux/amd64,linux/arm64 \
  -t 123456789012.dkr.ecr.us-east-1.amazonaws.com/app:1.2.1 --push .

The stopgap is to pin the deployment to nodes of the architecture you have an image for, with a nodeSelector of kubernetes.io/arch: amd64, which keeps the pods off the Graviton nodes until the build is fixed. Ethan did the stopgap on Monday and the multi-arch build on Tuesday, and the savings from the Graviton move survived. Third-party images are usually multi-arch already; it is your own images that need the flag.

Cause 6: private registries and certificates (x509)

If the image comes from a registry you run, or from a vendor's registry behind a corporate proxy, the error can be x509: certificate signed by unknown authority. The node's container runtime does not trust the certificate presenting itself. This is never ECR, which uses public certificates. Options: put a publicly trusted certificate on the registry; add the private CA to the node AMI's trust store (on EKS-optimized AMIs, through a launch template user-data step that drops the CA into /etc/pki/ca-trust/source/anchors/ and runs update-ca-trust); or, if a proxy is intercepting TLS, exclude the registry hostnames from interception. Credentials for the private registry go in a docker-registry secret exactly as for Docker Hub.

After the fix: make the pod retry now

The backoff timer does not know you fixed anything. A pod that has backed off to the five-minute cap will sit for up to five minutes before trying again, and people routinely conclude a correct fix did not work. Force it:

# for a Deployment, StatefulSet or DaemonSet
kubectl rollout restart deployment my-app -n my-namespace

# for a bare pod
kubectl delete pod my-pod -n my-namespace

Then watch the events on the new pod. If the Failed line has changed to a different sentence, you have moved one step down the decoder, which is progress: a timeout that becomes a 401 means the network is fixed and IAM is next.

🧭 NEW HERE? READ THESE FIRST

If the pieces under this error are new to you, these five explain them without the jargon:

📌 Bookmark this page; the next ImagePullBackOff will have a different sentence under it.

Stop it coming back

  • Give every node role AmazonEC2ContainerRegistryPullOnly in the node group definition, in Terraform or eksctl, so a new node group never launches without it.
  • Pin images by digest in production manifests, or at least by immutable tags, and turn on tag immutability on the ECR repository so a tag cannot be silently repointed or deleted under a running deployment.
  • Keep ECR in the same Region as the cluster, or use replication, so cross-Region URIs never appear in a manifest.
  • Never reference Docker Hub directly from a cluster. Pull-through cache or a mirrored repository, decided once.
  • Build multi-arch by default the day you add the first Graviton node group, not the day a deployment fails.
  • For private clusters, make the S3 gateway endpoint part of the definition of done. Two ECR endpoints without it is the trap.

EKS ImagePullBackOff: the questions people ask

What does ImagePullBackOff mean in EKS?

The node failed to download the pod's container image and is waiting between retries, with the wait doubling up to five minutes. The status does not say why. Run kubectl describe pod and read the "Failed to pull image" event; the text after the image name is the registry's actual reason.

How do I fix ImagePullBackOff on an EKS pod?

Read the Failed to pull image event first. A 401 or "no basic auth credentials" means the node IAM role needs AmazonEC2ContainerRegistryPullOnly. "Manifest unknown" means the tag or URI is wrong. A timeout means the node has no route to ECR. "toomanyrequests" is Docker Hub's limit. "No match for platform" is an architecture mismatch. Fix the matching cause, then run kubectl rollout restart on the deployment.

Why does my EKS pod get 401 Unauthorized pulling from ECR?

The node's IAM role could not obtain an ECR authorization token, or the image is in another account that has not granted this role access through a repository policy. Attach AmazonEC2ContainerRegistryPullOnly or AmazonEC2ContainerRegistryReadOnly to the node role, and for cross-account images add a repository policy in the owning account.

What is the difference between ErrImagePull and ImagePullBackOff?

ErrImagePull is the state immediately after a failed pull attempt. ImagePullBackOff is the waiting state between retries once the kubelet has started backing off. They are the same problem seen at different moments, and the fix is identical.

Which IAM policy does an EKS node need to pull from ECR?

AmazonEC2ContainerRegistryPullOnly, which grants ecr:GetAuthorizationToken, ecr:BatchGetImage, ecr:GetDownloadUrlForLayer and ecr:BatchImportUpstreamImage. The older AmazonEC2ContainerRegistryReadOnly also works and additionally allows listing and describing repositories. Fargate pods use the pod execution role instead of a node role.

How do I pull an ECR image from another AWS account in EKS?

Two permissions are needed: the node role in the cluster's account must have ECR pull permissions, and the repository in the owning account must have a repository policy allowing that role or account to call ecr:BatchGetImage, ecr:GetDownloadUrlForLayer and ecr:BatchCheckLayerAvailability. Without the repository policy the pull fails with 401 or 403.

Why do pods in a private subnet get i/o timeout pulling images?

The node has no path to the registry. Either the subnet's route table has no NAT gateway route, or the VPC endpoints are incomplete. ECR needs the ecr.api and ecr.dkr interface endpoints plus an S3 gateway endpoint for the image layers, attached to the node subnets, with a security group that allows 443.

Does an EKS private cluster need an S3 VPC endpoint to pull from ECR?

Yes. ECR stores image layers in S3, so the two ECR interface endpoints alone let the node authenticate and fetch the manifest and then hang on the first layer download. Add a gateway endpoint for com.amazonaws.region.s3 to the node subnets' route tables.

What does "toomanyrequests: You have reached your pull rate limit" mean on EKS?

Docker Hub's anonymous limit of 100 pulls per six hours per public IPv4 address was reached. Nodes in private subnets share the NAT gateway's single public IP, so the whole cluster counts as one address. Authenticate pulls with a docker-registry secret, set up an ECR pull-through cache, or mirror the images into ECR.

How do I use a Docker Hub private image in EKS?

Create a secret with kubectl create secret docker-registry using your Docker Hub username and an access token, then reference it under imagePullSecrets in the pod spec, or attach it to the service account so every pod in the namespace uses it. The same secret lifts the anonymous rate limit.

What does "no match for platform in manifest" mean?

The image has no variant for the node's CPU architecture, usually an amd64-only image scheduled onto an arm64 Graviton node. Build a multi-architecture image with docker buildx build --platform linux/amd64,linux/arm64, or temporarily pin the deployment to amd64 nodes with a nodeSelector of kubernetes.io/arch: amd64.

Why does my image pull fail on Graviton nodes but work on other nodes?

Graviton instances are arm64 and the image was built only for amd64. Inspect the image with docker buildx imagetools inspect to see which platforms it carries, then rebuild it as multi-arch. Most official third-party images already are; custom images built with plain docker build are not.

How do I check whether an image tag exists in ECR?

Run aws ecr describe-images --repository-name app --image-ids imageTag=1.2.1 in the Region in the image URI. An ImageNotFoundException means the tag is gone or never existed; aws ecr list-images shows which tags are present. Lifecycle policies can delete old tags without warning.

Why does my pod still show ImagePullBackOff after I fixed the permissions?

The kubelet's retry backoff can be up to five minutes, and it does not reset when you change IAM or a manifest. Run kubectl rollout restart deployment, or delete the pod, and the replacement pulls immediately. Then read the new pod's events to confirm the error text has changed or gone.

Does ImagePullBackOff happen on Fargate?

Yes, with the same causes, except that permissions come from the pod execution role rather than a node role, and network access depends on the subnets the Fargate profile uses. Private Fargate subnets need the same ECR and S3 endpoints or a NAT route.

How do I set up an ECR pull-through cache for Docker Hub?

Create a pull-through cache rule in ECR with an upstream of registry-1.docker.io and a prefix such as docker-hub, store Docker Hub credentials in Secrets Manager as the rule requires, then reference images as your-registry/docker-hub/library/image:tag. The node role needs ecr:BatchImportUpstreamImage for the first pull of each image, which AmazonEC2ContainerRegistryPullOnly includes.

Jake's fix was a pull-through cache rule, which Ethan set up in ten minutes and which has since absorbed every Docker Hub hiccup without a pod noticing. Ethan's fix was a build flag. Neither error was a Kubernetes problem in the end; one was a line in a build file and the other was an address-counting rule at a company neither of them works for. If your pod is in backoff right now, read the sentence under the status before you touch anything. It is telling you which of the six it is, and five of the six are a single command away.

📌 If you keep one line from this page

ImagePullBackOff is the waiting room. The reason is in the "Failed to pull image" event, and on EKS it is almost always the node role, the tag, or the route.

Fix the sentence you read, then rollout restart. The pod will not hurry on its own.

Revision note. Written October 2, 2026, with the Docker Hub limits and ECR policies as they stand this autumn. If your deployment is stuck right now, the decoder table is the fastest route out; the rest of the page can wait until it is running.

Related