CloudFormation UPDATE_ROLLBACK_FAILED: Fix It in 4 Routes
If your AWS CloudFormation stack is stuck in UPDATE_ROLLBACK_FAILED, the fix almost always starts with one command: aws cloudformation continue-update-rollback --stack-name your-stack-name, or the "Continue update rollback" button under Stack actions in the console. Most stacks come back from that alone within a few minutes. But here is the part almost nobody explains: the locked-down state you are staring at is not CloudFormation breaking down on you. It is CloudFormation deliberately refusing to touch anything else until you tell it what to do, because the alternative is letting it guess, and guessing is exactly how a recoverable stack turns into an unrecoverable one.
Jake had a point-of-sale terminal running on an EC2 instance that a Lambda function restarted every night at 2 a.m. He changed one environment variable in the stack, clicked deploy, and went to make coffee. When he came back, the console was showing him a status he had never seen before, in red, refusing to move no matter how many times he refreshed the page: UPDATE_ROLLBACK_FAILED. His usual "Update stack" button was gone. His deploy pipeline was red. The only thing CloudFormation would let him click was something called "Continue update rollback" — and to Jake, that looked exactly like the button that had already failed once.
Here is the sentence that would have saved Jake twenty minutes of nervously re-reading the same warning banner: a stack in UPDATE_ROLLBACK_FAILED is not corrupted, and it is not unrecoverable by default. CloudFormation tried to undo a failed update, could not finish undoing it, and stopped — on purpose — rather than leave your infrastructure in a state nobody described anywhere. From that state, the CloudFormation API only accepts two actions: ContinueUpdateRollback and DeleteStack. Every other action, including a normal update, is blocked until you use one of those two.
What UPDATE_ROLLBACK_FAILED actually means
Every stack update in CloudFormation follows the same shape. You submit a change, CloudFormation applies it resource by resource, and one of two things happens: every resource finishes and the stack lands on UPDATE_COMPLETE, or one resource fails and CloudFormation automatically tries to reverse everything it already changed — that reversal is the rollback, and while it is happening the stack shows UPDATE_ROLLBACK_IN_PROGRESS.
A clean rollback lands you on UPDATE_ROLLBACK_COMPLETE. Your stack is back to exactly how it looked before you touched it, and you can try the update again once you have fixed whatever broke it. That is the safety net doing its job.
UPDATE_ROLLBACK_FAILED is what happens when even the reversal cannot finish. CloudFormation goes to undo a change — recreate a resource it deleted, restore a property it altered, point a reference back at something it removed — and it hits an error doing that too. Now it cannot go forward (the update already failed) and it cannot go back (the rollback failed). It stops, mid-repair, and locks the stack so nothing else can make the mess worse.
That last part is the piece worth sitting with. CloudFormation is a service designed to be idempotent and predictable: given a template, it should always know exactly what the current state of every resource is, so it always knows what "undo" means. The moment reality and its internal record of reality disagree — a resource got deleted by someone in the console, a permission got revoked, a resource is still waiting on a signal that is never going to arrive — CloudFormation has no safe assumption left to make. Rather than pick one and possibly delete or duplicate something you needed, it stops and asks a human to resolve the disagreement. UPDATE_ROLLBACK_FAILED is CloudFormation being honest about not knowing what happened, not CloudFormation failing at its job.
🙋♂️ Jake's Reality Check
"So it's not actually broken? Then why won't it just fix itself?"
Because CloudFormation would rather stop and ask than guess wrong on your behalf. The two paths still open to you from this state — continue the rollback, or delete the stack — both require you to make a call CloudFormation is not willing to make on its own. That is a safety feature aimed squarely at a phone-shop owner's actual worst case: a script that silently deletes the wrong database because a service assumed it knew best.
Find out why it failed before you touch anything
Clicking "Continue update rollback" immediately, before you know why the first rollback failed, works often enough that it is the right first move — it is step one in the Quick Answer box above for a reason. But when it does not work, you need the reason, and CloudFormation already recorded it for you.
In the console, open the stack, go to the Events tab, and read from the top down (events are listed newest first, so scroll to where the rollback started). Look for the resource whose status reads UPDATE_FAILED during the rollback, not during the original forward update — those are two different failures and only the rollback one matters here. The Status reason column next to it is CloudFormation telling you, in plain language, exactly what it could not do.
From the CLI, the same information comes from:
aws cloudformation describe-stack-events \ --stack-name your-stack-name \ --max-items 30
Read the ResourceStatusReason field for the first UPDATE_FAILED entry you find that happened after the rollback started. That single line usually tells you which of the three common causes you are dealing with, covered in full a little further down: a resource was changed or deleted outside CloudFormation, the IAM identity CloudFormation is using does not have permission to make the change it needs, or a resource is waiting on a signal (from an Auto Scaling group, a WaitCondition, or a custom resource) that is never going to show up.
Why this step is worth the two extra minutes
Skipping straight to "continue update rollback, skip whatever fails" without reading the reason is how people end up skipping resources they did not need to skip. Every resource you skip becomes a permanent mismatch between your template and your live infrastructure until you manually reconcile it — that is not a minor housekeeping detail, it is the thing that turns future updates into new failures. Two minutes reading the status reason now is cheaper than untangling a stack where three resources silently drifted out of sync with the template three updates ago.
If nobody on the team admits to touching anything
Out-of-band changes rarely come with a confession attached. If the status reason points to a resource that "should" still exist and nobody remembers deleting it, AWS CloudTrail is the next stop, not a guess. CloudTrail keeps a record of the AWS API calls made in your account — who made the call, whether it came from the console or a script, and when — and by default you already have access to the last 90 days of that activity through CloudTrail Event history, with no separate trail to set up first. Search it for the resource's name around the time it went missing, and you will usually find the exact console action, and the exact identity behind it, that CloudFormation is now failing to reconcile.
Telling a forward-update failure apart from a rollback failure
The Events tab does not visually separate "the update that failed" from "the rollback of that update that also failed" — both show up as ordinary rows, and both can list the exact same resource. The way to tell them apart is the stack-level status column, read top to bottom in time order: the stack first shows UPDATE_IN_PROGRESS, then, if a resource fails, UPDATE_FAILED and immediately UPDATE_ROLLBACK_IN_PROGRESS. Everything with a resource status of UPDATE_FAILED logged before that UPDATE_ROLLBACK_IN_PROGRESS marker belongs to the original failed update — interesting for understanding what you broke, but not eligible to be skipped. Everything logged after it, while the stack-level status still reads UPDATE_ROLLBACK_IN_PROGRESS, is the one that matters for --resources-to-skip: only resources that failed during that second pass are eligible, which is exactly why CloudFormation rejects an attempt to skip a resource that only ever failed during the forward update.
Route 1: continue the rollback without skipping anything
This is the plain version of the operation CloudFormation built specifically for this state: ContinueUpdateRollback. It tells CloudFormation to try the rollback again, resuming from wherever it stopped, with no resources excluded. If the failure was something transient — a brief service hiccup, a rate limit that has since cleared, a permission you just fixed — this alone gets you all the way to UPDATE_ROLLBACK_COMPLETE.
Console steps
- Sign in to the AWS Management Console and open CloudFormation.
- Use the Region selector at the top of the screen to confirm you are in the same Region as the stuck stack — a stack stuck in one Region will not show up, or worse, look like it does not exist, if you are looking in another.
- On the Stacks page, select the stack showing
UPDATE_ROLLBACK_FAILED. - Choose Stack actions, then Continue update rollback.
- In the dialog that opens, leave Advanced troubleshooting closed and choose Continue update rollback.
CLI steps
aws cloudformation continue-update-rollback \ --stack-name your-stack-name
A successful call from the CLI produces no output at all — that silence is the expected result, not a sign that something went wrong. Poll the stack status afterward with aws cloudformation describe-stacks --stack-name your-stack-name and watch for it to move to UPDATE_ROLLBACK_IN_PROGRESS and then, hopefully, UPDATE_ROLLBACK_COMPLETE.
If your stack normally deploys through an IAM role rather than your own user credentials (common in CI/CD pipelines), you can hand CloudFormation that role explicitly with --role-arn. If you leave it out, CloudFormation reuses whichever role was already associated with the stack, and if none was ever associated, it falls back to a temporary session built from your own credentials.
| Route | Use it when | What it leaves behind |
|---|---|---|
| Continue rollback, no skips | First attempt, always. Also right after you fix the underlying cause (permission, deleted resource, missing signal). | Nothing — a clean UPDATE_ROLLBACK_COMPLETE with the template and live resources back in agreement. |
| Continue rollback, skip resource(s) | Route 1 failed again with the same reason, and the resource genuinely cannot be reconciled (it is gone, or its replacement was created some other way). | A mismatch between the skipped resource and the template that you must manually fix before the next update, or that update fails too. |
| Delete the stack | Both routes above have failed, or the stack is disposable (a test/dev environment you can recreate from the template). | The stack, and by default every resource in it, unless you catch a resulting DELETE_FAILED and retain specific resources. |
✅ Why this is the one to try first
Skipping resources is the technique everyone remembers because it is the one that "works" even when you do not understand the failure. But it trades a five-minute problem today for a template-versus-reality mismatch that resurfaces on your next update, at a worse time. Fix the actual cause and run the plain continue-update-rollback first; treat skipping as the fallback it was designed to be, not the default move.
Route 2: when it fails again, skip the resources that are actually stuck
Sometimes the resource CloudFormation is trying to roll back to genuinely cannot be restored — it was deleted outside CloudFormation and you have no intention of recreating it under that exact name, or it is in a state CloudFormation's rollback logic does not know how to reverse. For those cases, ContinueUpdateRollback accepts a list of resources to skip. CloudFormation marks them UPDATE_COMPLETE without touching them and carries on rolling back everything else.
Console steps
- Stacks > select the stuck stack > Stack actions > Continue update rollback.
- In the dialog, expand Advanced troubleshooting.
- Under Resources to skip, select the logical IDs of the resource(s) whose
Status reasonyou read earlier. Only resources currently inUPDATE_FAILEDbecause the rollback itself failed on them appear here — a resource that failed for an unrelated reason, such as a canceled update, will not be eligible. - Choose Continue update rollback.
CLI steps
aws cloudformation continue-update-rollback \ --stack-name your-stack-name \ --resources-to-skip LogicalResourceId1 LogicalResourceId2
Use the resource's logical ID — the name you gave it inside your template's Resources block — never its physical ID (the actual instance ID, bucket name, or ARN). You can find the logical ID on the stack's Resources tab in the console, in the "Logical ID" column, next to the resource's physical ID.
⚠️ What this actually breaks
A skipped resource is marked UPDATE_COMPLETE whether or not it matches your template anymore. CloudFormation's internal record now disagrees with reality on purpose, with your permission, and it stays that way until you fix it. If you run another update on that stack without first reconciling the skipped resource — updating the template to match what actually exists, or recreating the resource to match the template — that next update can fail for the exact same reason, and in the worst case the stack becomes unrecoverable through normal means, leaving deletion as the only path out. Skip the minimum number of resources needed and go fix them immediately afterward, not "eventually."
The three real causes, and how to fix each one properly
Reading a hundred versions of this problem across teams turns up the same three root causes, over and over. Fixing the actual cause means Route 1 (no skipping) succeeds, which is always the better outcome.
1. Something was changed outside CloudFormation
This is the single most common cause, and it is almost always a person, not a bug. Someone opens the console and manually deletes a security group rule, resizes an RDS instance, or deletes an S3 bucket "just to clean up," entirely outside CloudFormation. CloudFormation's internal template state has no idea any of that happened. When the rollback later tries to restore that resource to its pre-update configuration, it reaches for something that either does not exist anymore or does not match what it expects, and the rollback fails.
The fix is to make reality match the template again: manually recreate the deleted resource with the same name and the same properties it had before, or manually change the resource back to the configuration your template describes, then run Route 1 again. If recreating it exactly is not realistic (an S3 bucket name that got taken by someone else in the meantime, for example), your only remaining option for that specific resource is to skip it in Route 2 and then update your template afterward so it reflects what actually exists now.
2. The IAM identity does not have permission to finish the rollback
A rollback is, under the hood, a set of ordinary create, update, and delete calls to the underlying AWS services — EC2, IAM, RDS, whatever your template touches. If the IAM user, role, or service role CloudFormation is using does not have permission for one of those calls, the rollback fails on that resource exactly as if it were a normal update. This shows up most often in accounts that recently tightened permissions, or in a CI/CD pipeline role that was scoped down after the stack's resources were first created.
Check the Status reason on the failed resource for language like "not authorized to perform" — that is IAM telling you directly which action and which resource it blocked. Fix the underlying role or policy so it can perform that action, then run Route 1 again. If your organization keeps a dedicated deployment role separate from your own console user, remember that continue-update-rollback accepts --role-arn so you can hand CloudFormation the correct role explicitly rather than relying on whatever is already attached to the stack.
3. A resource is waiting on a signal that is never coming
Some resources do not just get created, they have to report back that they finished starting up correctly — instances in an Auto Scaling group signaling that a launch configuration script completed, a WaitCondition or custom resource that expects a manual or Lambda-driven confirmation. If that signal was supposed to arrive during the rollback and never did — because the script that was meant to send it crashed, or the instance never launched, or a custom resource's Lambda function itself failed — the rollback times out waiting.
The fix is to manually send the signal CloudFormation is waiting for using the signal-resource command, so CloudFormation sees the success it needed and moves on:
aws cloudformation signal-resource \ --stack-name your-stack-name \ --logical-resource-id YourResourceLogicalId \ --unique-id INSTANCE_OR_UNIQUE_ID \ --status SUCCESS
Once you have sent the number of successful signals the resource is configured to wait for, continue the rollback with Route 1 and it should proceed normally.
| Symptom in the status reason | Likely cause | Fix before you continue the rollback |
|---|---|---|
| "does not exist" / "no such resource" / resource-not-found style errors | Resource deleted or renamed outside CloudFormation | Recreate it to match the template, or accept you must skip it |
| "not authorized to perform" | Missing IAM permission on the deployment role | Add the missing action to the role, then retry |
| "Received FAILURE signal" / timeout waiting on resource signal | An Auto Scaling instance, WaitCondition, or custom resource never confirmed success | Send the missing signal manually with signal-resource |
Deploying through the CDK, SAM, or Terraform changes nothing underneath
None of these tools replace CloudFormation; they generate a CloudFormation template and deploy it as an ordinary stack, so a stack they created can land in UPDATE_ROLLBACK_FAILED exactly the way a hand-written template can, and everything above still applies to it. The difference is only in how you get there and how much of it your tool automates for you.
The AWS CDK's own CLI has a command built specifically for this: cdk rollback. By default, cdk deploy already rolls back a failed deployment automatically, the same way CloudFormation itself does. If you deployed with --no-rollback, or a deployment left your stack paused rather than cleanly rolled back, cdk rollback takes it back to its last stable state. And if some resources fail to roll back — the same underlying situation this whole post is about — the CDK CLI gives you --orphan LogicalId to force a specific resource past the rollback, which is the CDK's own wrapper around exactly the same idea as --resources-to-skip: it works stack by stack, and repeating --orphan for more than one resource is supported the same way repeating logical IDs is on the plain CLI.
SAM and Terraform's AWS provider do not add an equivalent command of their own; a SAM stack or a Terraform-managed stack that hits UPDATE_ROLLBACK_FAILED is recovered with the plain aws cloudformation continue-update-rollback command, exactly as described above. The one extra step worth taking afterward, for either tool, is checking that its own record of the world — SAM's packaged template, or Terraform's state file — still agrees with what CloudFormation actually did, since a rollback you triggered manually outside the tool's own workflow is exactly the kind of change that tool did not see coming.
This matters more than it sounds like it should, because it is easy to assume a "managed" deployment tool insulates you from ever seeing a raw CloudFormation status code. It doesn't, and the insulation runs the other way around: the tool is a thin layer on top of the same stack lifecycle, so anything that can put a hand-written stack into UPDATE_ROLLBACK_FAILED — a deleted security group rule, a tightened IAM role, a signal that never arrived — can do the same thing to a stack the CDK, SAM, or Terraform generated for you. Knowing the underlying CloudFormation mechanics, rather than only your tool's wrapper commands, is what lets you recover a stuck deployment even on the day your tool's own recovery command does not quite fit the situation.
Nested stacks make this worse, and the skip syntax changes
If your root stack is built from smaller nested stacks — a common pattern where, say, a WebInfra parent stack contains a WebInfra-Compute child and a WebInfra-Storage child, each with their own resources — a failed rollback can knock the entire hierarchy into UPDATE_ROLLBACK_FAILED at once. Rolling back the parent always attempts to roll back every child stack along with it, so a single resource stuck deep inside one child can stall the whole tree.
You continue the rollback from the root stack only — never target a nested child stack directly with continue-update-rollback. When you need to skip a resource that lives inside a nested stack rather than in the root template, the logical ID you give CloudFormation has to include the nested stack's own name, separated by a period:
aws cloudformation continue-update-rollback --stack-name WebInfra \ --resources-to-skip myCustom WebInfra-Compute-Asg.myAsg \ WebInfra-Compute-LB.myLoadBalancer WebInfra-Storage.DB
Resources belonging directly to the root stack (myCustom above) only need their own logical ID. Resources inside a child stack need the format NestedStackName.ResourceLogicalID. You can find a nested stack's own name in its ARN — it appears right after stack/, for example WebInfra-Storage-Z2VKC706XKXT in an ARN like arn:aws:cloudformation:us-east-1:123456789012:stack/WebInfra-Storage-Z2VKC706XKXT/ea9e7f90.... You can find its logical ID either in the parent template, where it is defined as an AWS::CloudFormation::Stack resource, or on the parent stack's Resources tab in the console under the "Logical ID" column.
One extra rule applies only if the thing you want to skip is the entire nested stack resource itself, rather than a resource inside it: CloudFormation will only let you skip it if that embedded child stack is currently in DELETE_IN_PROGRESS, DELETE_COMPLETE, or DELETE_FAILED. If it is in any other state, that particular resource simply is not eligible to be skipped yet, and you will need to resolve whatever is happening inside the child stack first.
🕐 What tends to trip people up here
- Trying to run
continue-update-rollbackagainst the child stack's own name instead of the root stack — the API rejects this; it must be the top-level stack. - Skipping a resource inside a nested stack using just its plain logical ID, without the
NestedStackName.prefix — CloudFormation will not recognize which resource you mean. - With several nested levels, you sometimes need to skip resources at more than one level in the same call to get the whole hierarchy unstuck in one pass.
If continue rollback keeps failing: delete the stack
Ethan's take on this one is blunt: "If you have already tried fixing the cause, and you have already tried skipping the stuck resource, and it still will not budge, stop treating deletion as a last resort you are ashamed of. It is the second officially supported action from this state for a reason. CloudFormation is telling you it has run out of ideas too." Jake pushed back the first time he heard that: a stack holding a customer's live payment terminal is not something you delete on a hunch just because a dialog box offers it. Ethan's answer was the same one that runs through this whole guide — the option being available does not mean it is the first one you reach for, only that it exists once the safer ones are genuinely exhausted.
Remember, from the state itself: the API only permits ContinueUpdateRollback and DeleteStack once a stack is in UPDATE_ROLLBACK_FAILED. Deleting is not a workaround you found by accident — it is the built-in second door. This matters most for stacks you can safely recreate from the template — a staging environment, a disposable test stack, or infrastructure with no state that lives inside it (a Lambda-only stack, for example, versus one holding an RDS database with data you need).
For a stack that does hold something you cannot lose, do not delete first. Instead, snapshot or export whatever matters — take an RDS snapshot, copy the S3 bucket contents, export the DynamoDB table — and only then delete the stack, recreate it from the template, and restore what you exported.
If the delete itself gets stuck (a resource that will not delete cleanly puts the stack into DELETE_FAILED rather than finishing), CloudFormation gives you one more lever: you can retry the delete with a list of resources to explicitly retain, so the stack record disappears from CloudFormation without CloudFormation trying to also destroy those specific resources. That is a separate, narrower situation from the resource-skipping covered above — it only applies once you are already in DELETE_FAILED, not while you are still in UPDATE_ROLLBACK_FAILED.
Three versions of this, start to finish
The three causes above read cleanly on paper. In practice, they show up wrapped in whatever a real environment was doing that week. Here is what each one actually looked like for Jake, worked through the way you would work through it yourself.
The environment variable update (a signal timeout)
Jake's point-of-sale stack had an Auto Scaling group behind it, configured to wait for a success signal from each instance before CloudFormation considered the resource finished. His update changed an environment variable baked into the launch template, which meant a fresh instance had to launch and confirm it came up healthy. That night, the instance launched fine, but the startup script that was supposed to call cfn-signal at the end hit a network blip pulling a dependency and exited before it ever sent the signal. CloudFormation waited out its timeout, gave up, and tried to roll back — which meant launching yet another instance and waiting on another signal that, for the same underlying reason, never came either. Two timeouts later, the stack landed in UPDATE_ROLLBACK_FAILED.
The status reason on the Events tab named the Auto Scaling group directly and described a signal timeout. Jake didn't need to touch the template at all: he manually sent a success signal for the instance with signal-resource, watched the rollback pick up where it left off, and ran plain continue-update-rollback straight after. Clean recovery, nothing skipped.
The security group someone "cleaned up" (drift)
A different week, a different stack: Jake's shop had grown enough that a part-time contractor was helping manage the network side. The contractor found an EC2 security group rule that looked unused, opened the console, and deleted it — entirely reasonably, from where the contractor was sitting, and entirely outside CloudFormation. Weeks later, an unrelated template change forced an update to that same security group, the update failed for its own reasons, and the rollback tried to restore the rule the contractor had removed. CloudFormation reached for a rule that no longer existed in the form it expected, and the rollback failed.
"This is the one that actually worries me," Ethan said, looking at the Events tab. "Not because it's hard to fix — it's the easiest of the three. It worries me because nothing about deleting a security group rule in the console feels dangerous in the moment. It only becomes dangerous the next time CloudFormation has to reason about that resource." The fix was to manually re-add the exact rule the template described, then run continue-update-rollback with nothing skipped. Because the template and reality agreed again, the rollback finished without complaint.
The tightened deployment role (a permission gap)
The third one came from Jake's own side project, not the shop: a CI/CD pipeline whose deployment role got scoped down as part of a general security pass, without checking every action every stack in the account might eventually need during a rollback. The next update that touched an IAM-related resource failed forward, tried to roll back, and the rollback itself needed to modify that same IAM resource — a permission the freshly tightened role no longer had. The status reason spelled it out almost word for word: not authorized to perform the action on that resource.
Jake added the missing action back to the deployment role's policy, scoped to just that resource rather than reopening everything the security pass had closed, and reran continue-update-rollback with the same role. It finished on the first try. The role stayed tighter everywhere else; it just stopped being tighter than the stack it was responsible for rolling back.
Catch drift before it causes the next one
Since the most common cause of a failed rollback is a resource that no longer matches what CloudFormation expects, it helps to have a way to check for that mismatch before it ambushes you mid-update. CloudFormation has a built-in feature for exactly this, called drift detection, and it is worth running periodically on any stack you did not build entirely from scratch this week.
aws cloudformation detect-stack-drift --stack-name your-stack-name
This kicks off an asynchronous check — for each resource that supports drift detection, CloudFormation compares its live, actual configuration against what the template says it should look like, but only for the specific properties your template explicitly sets. Detection can take a few minutes on a large stack, so poll its progress with describe-stack-drift-detection-status, and once it finishes, pull the results with describe-stack-resource-drifts to see exactly which resources, and which properties on them, no longer match.
Two limits worth knowing before you rely on it: not every resource type supports drift detection yet, so a clean drift report is not an absolute guarantee nothing has changed, only that nothing checkable has. And running drift detection on a parent stack does not check its nested child stacks automatically — you have to run it again directly against each nested stack if you want the same coverage there. Used regularly, though, this is the single best early-warning system for the exact failure mode that sends stacks into UPDATE_ROLLBACK_FAILED in the first place: it lets you find and fix the mismatch on a quiet Tuesday afternoon, instead of discovering it mid-rollback at 11 p.m.
Stopping this from happening on your next deploy
None of the three causes above are exotic. All three are avoidable with habits that cost almost nothing once they are in place.
Treat every stack resource as CloudFormation's, not yours
The single biggest predictor of this whole problem is someone reaching into the console and changing something CloudFormation manages. Jake's shop rule for his own team now is simple: if a resource shows up in a CloudFormation stack's Resources tab, nobody touches it directly, ever — not "just this once," not "just to fix it faster." Any change goes through a template update, even a trivial one, so CloudFormation's record and reality never get the chance to disagree.
Keep the deployment role's permissions ahead of the template, not behind it
Permission errors during rollback usually mean the deployment role was scoped down after the stack's resources were already created, or a new resource type was added to the template without updating the role that deploys it. Before you widen a template to touch a new AWS service, check that the IAM role doing the deploying already has permission for that service's create, update, and delete actions — not just create.
Give signal-dependent resources a realistic timeout
If a resource in your template is configured to wait for a signal (an Auto Scaling group's creation policy, a WaitCondition, a custom resource backed by Lambda), make sure the timeout window is long enough for the thing sending that signal to realistically finish, including on a slow day. A timeout set for the best-case scenario turns an ordinary slow morning into a failed rollback.
Preview the update with a change set before you run it
A change set is CloudFormation's built-in preview: you give it a proposed template or parameter change, and it tells you exactly what will be created, updated, or deleted — including, for property-level changes, the before-and-after values — before anything actually happens. CloudFormation makes no changes to the stack until you explicitly execute the change set, so reviewing one costs you nothing but a look. It will not catch every possible runtime failure, since some causes (a service quota hit mid-deploy, a Lambda-backed custom resource that fails at execution time) only show up once the update is actually running, but it does flag the two things that most often turn into a painful rollback later: an unexpected resource replacement, and a property change that quietly takes down something you meant to leave alone. In the console, open the stack and choose Create change set from the Stack actions menu; from the CLI, it is aws cloudformation create-change-set followed by describe-change-set to review it before you execute.
Frequently asked questions
What does UPDATE_ROLLBACK_FAILED actually mean?
It means CloudFormation tried to undo a failed stack update and could not finish undoing it — usually because one resource cannot be restored to its previous state. The stack is locked so it cannot be updated or fully rolled back until you either continue the rollback or delete the stack.
Can I still update my stack while it's stuck in this state?
No. A stack in UPDATE_ROLLBACK_FAILED cannot be updated. It has to reach UPDATE_ROLLBACK_COMPLETE first, by continuing the rollback, before a normal update will be accepted again.
Why did Continue update rollback fail the first time I tried it?
Almost always because the underlying cause is still present — the resource CloudFormation cannot roll back to still does not exist, the missing IAM permission still is not granted, or the signal it is waiting for still has not arrived. Running the same command again without fixing that will usually fail the same way.
Which resources am I allowed to put in Resources to skip?
Only resources currently showing UPDATE_FAILED specifically because the rollback failed on them. A resource that ended up UPDATE_FAILED for some other reason, like a canceled update, is not eligible to be listed there.
What happens to a resource after I skip it?
CloudFormation marks it UPDATE_COMPLETE without actually touching it, and continues rolling back the rest of the stack. The skipped resource's real-world state and the stack template no longer necessarily match, and that mismatch stays there until you manually reconcile it.
Can I skip a resource that's inside a nested stack?
Yes, but the format changes. You cannot use the plain logical ID alone; you need NestedStackName.ResourceLogicalID so CloudFormation knows which nested stack the resource belongs to.
What if I don't know which resource caused the failure?
Check the stack's Events tab in the console, or run describe-stack-events from the CLI, and look for the first resource marked UPDATE_FAILED after the rollback began. Its status reason names the resource and explains, in plain language, what CloudFormation could not do.
Is it safe to just delete the stack instead of troubleshooting?
It is one of only two actions the API allows from this state, so it is a legitimate option, not a hack. It is only "unsafe" in the sense that deleting the stack will delete its resources by default, so it is the right move for disposable environments and the wrong first move for anything holding data or state you have not backed up.
Can a stack in UPDATE_ROLLBACK_FAILED actually be deleted?
Yes. CloudFormation explicitly permits deletion from this state; it is the second of the only two actions available once a stack is stuck here.
What if my Auto Scaling group is stuck waiting for a signal?
Send the missing success signal manually with the signal-resource CLI command, specifying the logical resource ID and the unique ID it is waiting on. Once CloudFormation receives the number of successful signals it needs, continuing the rollback should proceed normally.
What if a resource was deleted outside CloudFormation?
Recreate it manually with the same name and the same properties it had in the stack template, then continue the rollback with no resources skipped. If you cannot recreate it exactly, your remaining option is to skip that resource and update the template afterward to reflect what actually exists.
Do I need special IAM permissions to run continue-update-rollback?
You need permission to call the ContinueUpdateRollback action itself, and CloudFormation also needs the underlying permissions to actually perform the rollback on each resource, via whichever role is associated with the stack (or the one you pass explicitly with --role-arn).
What's the difference between ContinueUpdateRollback and RollbackStack?
ContinueUpdateRollback is specifically for recovering a stack already stuck in UPDATE_ROLLBACK_FAILED. RollbackStack is a separate, more general action for rolling back stacks that are stuck mid-operation, and it supports an "express mode" for faster rollbacks; they solve related but different problems.
Will skipping resources break my stack permanently?
Not by itself, but it leaves your template and your live infrastructure disagreeing about that resource. If you reconcile them before the next update, you are fine. If you do not, the next update can fail for the same underlying reason, and repeated skipping without reconciling is how stacks become genuinely unrecoverable.
How do I stop this from happening on my next update?
Never change a CloudFormation-managed resource directly in the console; keep the deployment role's permissions ahead of what the template needs, not behind it; and give any resource that waits on a signal a timeout long enough for a realistically slow deploy, not just the best case.
I already skipped a resource once. Now a new update is failing because of it — now what?
This is the mismatch catching up with you. Reconcile the skipped resource with the template before doing anything else: either update the template so it accurately describes the resource as it actually exists now, or manually change the live resource to match what the template still says. Once the two agree again, the new update should proceed normally.
Revision note. Written September 2026, covering the current ContinueUpdateRollback and DeleteStack behavior, including nested-stack handling, the CDK's own rollback tooling, and the signal-resource recovery path. AWS occasionally extends this recovery flow (the RollbackStack API and the newer express deployment mode are examples), so if a step here stops matching what you see in the console, the flow has probably grown again. If you are reading this at 11 p.m. with a stack that will not budge: you have not broken anything CloudFormation cannot recover from, and the fix is almost always smaller than the red banner makes it feel.