Terraform "Error Acquiring the State Lock": Safe Unlock
Terraform Error acquiring the state lock on an S3 backend means another operation holds the lock, or Terraform cannot create or read the locking record. First identify the lock ID, backend and active pipeline; wait or stop the real owner before unlocking. Modern S3 backends use use_lockfile = true, while DynamoDB locking is deprecated. The counterintuitive detail: terraform force-unlock does not fix your infrastructure or repair state; it only removes a coordination guard, which makes an incorrect unlock more dangerous than the original error.
Jake has a phone shop with a modest AWS booking system. He keeps its infrastructure in Terraform so a website change does not become an afternoon of clicking. One pipeline deploys a database adjustment while another begins updating the booking application. The second run stops before it changes anything and prints a state-lock error.
Jake: "Can I just delete the DynamoDB item and run again?"
Ethan: "The lock is the sign on the door saying somebody is already inside. Find that somebody first."
Understand what Terraform locked and why your apply stopped
You are looking at a failed command, but the failure may be evidence that the backend did exactly what it should. Terraform state is the recorded map between your configuration and the real AWS objects it manages. Two writers editing the same map at once could leave it contradictory or outdated.
A backend is the system storing that state. In this case the state JSON lives at an S3 object key, such as prod/network/terraform.tfstate. Locking is a separate coordination step that prevents two potentially writing operations from treating that one state snapshot as theirs to update simultaneously.
Locking is automatic where the backend supports and enables it. A rejected lock attempt means Terraform did not proceed with that protected operation. The error is not evidence that Terraform already changed the infrastructure during the failed attempt; you still need to inspect the other active run.
Error: Error acquiring the state lock
Error message: operation error S3: PutObject, ...
Lock Info:
ID: 11111111-2222-3333-4444-555555555555
Path: example-state/prod/network/terraform.tfstate
Operation: OperationTypeApply
Who: runner@ci-host
Created: 2026-10-09 08:40:00 +0000 UTC
That output is illustrative. The real cloud error under the first line varies: it may refer to S3 object creation, DynamoDB conditional writes, permissions, missing resources, or a lock held by another process. The ID, Path, Operation, Who, and Created entries are clues. They are not authorization to unlock a running deployment.
Diagnose the lock before you change anything
When your run is blocked, the quickest safe fix often costs zero commands: identify the already running job and let it finish. A lock created thirty seconds ago by your current production pipeline is very different from one left by a crashed job yesterday.
- Copy the exact lock ID, state path, operation, holder, timestamp and underlying AWS error to your incident notes.
- Identify the S3 bucket, backend key, workspace, account and Region used by the failing working directory.
- Look for another
terraform apply,planor state command in terminals, pipeline runs or deployment workers. - If the owner is active, wait for it or coordinate an orderly cancellation; retry with
-lock-timeout. - If the original owner has stopped, take a state backup and confirm the abandoned lock still matches before issuing
terraform force-unlock. - If no lock holder exists but AWS reports missing resources or AccessDenied, repair the backend configuration or permissions instead of force-unlocking.
Run aws sts get-caller-identity to confirm which account your shell currently uses. Then inspect the Terraform backend settings in your repository and any backend arguments supplied by CI. A surprisingly common source of confusion is a developer running against one bucket while their pipeline uses another.
aws sts get-caller-identity
terraform version
terraform workspace show
terraform init -reconfigure
The final terraform init -reconfigure command reinitializes backend configuration without moving state. Only use it when you have reviewed the configured bucket, key and Region; if you are deliberately migrating the backend, that is a different procedure. The displayed workspace matters because non-default workspaces can use a different state object path.
Jake: "The ID looks old. Can I just unlock?"
Ethan: "Age is a clue, not proof. Check the job and its AWS activity. Some legitimate applies are slow."
Compare native S3 lock files with the DynamoDB table
If your team configured Terraform years ago, you may still have a DynamoDB table named something like terraform-locks. Newer Terraform supports native S3 lock files instead. The distinction matters because the lock lives in a different place, uses different IAM permissions, and produces different failure messages.
| Feature | S3 lock file | DynamoDB locking |
|---|---|---|
| Backend option | use_lockfile = true |
dynamodb_table = "terraform-locks" |
| Lock location | S3 object at state key plus .tflock |
DynamoDB item keyed by LockID |
| Typical permissions | GetObject, PutObject, DeleteObject on lock object | DescribeTable, GetItem, PutItem, DeleteItem on table |
| Current direction | Preferred for supported Terraform versions | Deprecated; retained for compatibility during transition |
| State storage | State remains an S3 object | State remains an S3 object |
| Most useful first clue | S3 request/error plus .tflock |
DynamoDB table, Region and conditional write result |
Native S3 locking became available with Terraform 1.10. The backend option is not automatically on: its default is false. Setting use_lockfile = true gives the S3 backend a way to coordinate ownership through a lock object alongside the state object.
terraform {
required_version = ">= 1.10"
backend "s3" {
bucket = "example-terraform-state"
key = "prod/network/terraform.tfstate"
region = "us-east-1"
encrypt = true
use_lockfile = true
}
}
The bucket name above is a placeholder, not a bucket you can use as-is. Configure a bucket you control and protect it as an infrastructure-critical resource. State can contain secrets and identifiers even if you mark Terraform outputs sensitive.
DynamoDB-based locking remains important when an older Terraform release or shared team workflow still expects it. Its backend table requires a string partition key named LockID. Adding a table with a different partition-key name does not create an equivalent lock service.
terraform {
backend "s3" {
bucket = "example-terraform-state"
key = "prod/network/terraform.tfstate"
region = "us-east-1"
dynamodb_table = "terraform-locks"
}
}
Treat this second configuration as legacy, not a reason to build a new table for every project. The state and the lock still need coherent permissions, Region settings, and a clear owner across all writers.
Fix S3 native locking permission errors at the lock object
If your backend says use_lockfile = true but AWS returns AccessDenied, examine the lock object policy. Permission to read and write the state JSON does not automatically include permission to delete its lock file.
For the example key prod/network/terraform.tfstate, the lock object path is prod/network/terraform.tfstate.tflock. Terraform needs s3:GetObject, s3:PutObject and s3:DeleteObject on that lock object. Ordinary state writing needs its own GetObject and PutObject permissions, plus suitable bucket listing permissions.
aws s3api head-object \
--bucket example-terraform-state \
--key prod/network/terraform.tfstate.tflock \
--region us-east-1
A missing object is not automatically a problem when no lock is held; a lock file is temporary. An AccessDenied response is more important here. Review the pipeline role, bucket policy, SCPs and, where applicable, encryption permissions before issuing a force-unlock.
If S3 Object Lock, unusual retention controls, or restrictive bucket policies interfere with deleting temporary objects, you can create lock-release problems even when writes succeed. Keep state protection and temporary lock management deliberate rather than making broad public or administrator permissions the easy answer.
Repair legacy DynamoDB locks and the missing-table error
If your error names ResourceNotFoundException, the problem may not be somebody holding the lock. Terraform might be calling a DynamoDB table that does not exist in the Region or account selected by the backend. A typo, a deleted bootstrap resource, or the wrong credentials can all create that symptom.
aws dynamodb describe-table \
--table-name terraform-locks \
--region us-east-1 \
--query "Table.{Status:TableStatus,KeySchema:KeySchema}"
If this command fails with ResourceNotFoundException, inspect the configured table name, account and Region. If it succeeds, confirm the partition key is LockID with type String, then inspect IAM permissions and the complete Terraform error. An item-contention error differs from a missing-table error.
aws dynamodb get-item \
--table-name terraform-locks \
--key '{"LockID":{"S":"EXACT_LOCK_KEY_FROM_BACKEND"}}' \
--region us-east-1
The item-key example is intentionally a placeholder: it is not necessarily the lock UUID displayed to users. DynamoDB locking records use backend-specific path values. Use the exact key from your backend context if you inspect a legacy table. Reading is safer than deleting; avoid manually removing lock rows while an apply could still be running.
Another variant is AccessDeniedException. In that case the table may exist and still be unusable for the runner role. Check dynamodb:DescribeTable, GetItem, PutItem and DeleteItem on the correct table ARN.
Decide when terraform force-unlock is actually safe
You have found a lock that outlived its owner. Before clearing it, prove the work has stopped. Terraform force-unlock removes a lock for the current backend configuration. It does not roll back an incomplete apply, rewrite AWS resources, or repair a damaged state snapshot.
The ideal case is your own terminal crashed, there is no surviving Terraform process, the CI system has no active job for the state, and you can match the stale lock ID to your interrupted operation. That is the situation manual unlocking was designed to address.
terraform force-unlock 11111111-2222-3333-4444-555555555555
Terraform prompts for confirmation unless you provide -force. The -force option suppresses the confirmation; it does not make an unsafe unlock safer. Use the exact lock ID shown for this backend, not a copied ID from a different environment.
| Situation | Should you unlock? | Safer next action |
|---|---|---|
| Another CI apply is still running | No | Wait; coordinate or cancel cleanly |
| Your local Terraform process is still writing | No | Let it finish or stop it safely |
| Your failed job is terminated and lock remains | Usually, after owner verification | Back up state and force-unlock the matching ID |
| AWS says table not found | No | Fix backend table, account or Region |
| AWS says S3 AccessDenied | No | Fix permissions for state or lock object |
| Unknown lock owner or ambiguous timestamp | Not yet | Trace CI logs, host and backend identity |
| State seems corrupted after failed write | Unlock alone is insufficient | Preserve versions and inspect state recovery path |
What could actually go wrong
Unlocking while a second writer still operates can permit simultaneous state writes. A later clean-looking apply cannot guarantee that every remote resource and state record stayed consistent. Treat the lock as a safety mechanism, not an annoyance to bypass.
Jake: "This lock is the thing blocking Friday's release."
Ethan: "Or the thing stopping Friday's release from overwriting Thursday's production state. That is why verifying the owner is the fastest responsible step."
Handle Ctrl+C, terminated terminals and killed CI jobs
If you interrupted Terraform with Ctrl+C, the first interrupt is intended to allow a more graceful shutdown. A second interrupt or an external kill may prevent cleanup. A runner terminated by its CI platform can leave an abandoned lock if cleanup never completes.
A stale lock has no magical expiration you can always rely on. After your terminal closes or a runner disappears, confirm the owning process is actually gone. Check the CI job record, runner host, cloud execution logs and the lock timestamp. Then make the unlock decision for the exact state path.
Your laptop was closed halfway through apply
Start by checking whether a Terraform process is still running locally. Look for the same command in another terminal, background session or development environment. Then use the lock information to check whether a teammate or remote runner has since acquired the same state.
If the owner is genuinely dead, retrieve a backup first, unlock once, run a fresh plan and review any partial resource changes. A failed operation can make changes in AWS before crashing, so the next plan is an audit, not a ceremonial green checkmark.
Your CI runner was killed after its job timed out
Check whether the CI platform automatically retried the job or left a deployment child process working. Some organizations have both pipeline-level retries and external automation that invokes Terraform. The lock may belong to a successor run rather than the old one you remember.
Pause new applies for that state, identify the last completed run, and check active executions. Only after you have one clear owner should you clean up an orphaned lock and reopen the queue.
Give an active lock time with -lock-timeout
If you know another legitimate apply is finishing, waiting is better than unlocking. Terraform supports -lock-timeout for commands such as plan and apply, letting a runner retry acquiring the lock for a bounded period.
terraform plan -lock-timeout=5m
terraform apply -lock-timeout=10m
The values are durations, not requests to hold the lock for that long after the operation finishes. A ten-minute wait is a policy decision you can adjust to match the typical deployment time. If the owner never releases the lock, the command eventually fails and you investigate the stale owner.
For CI, a lock wait should fit within the job timeout while leaving room for the actual plan or apply. If every pipeline waits an hour for the same state, your deployment design may be generating more concurrency than it can safely execute.
Stop two CI pipelines fighting over one state
Your team may have separate workflows for a network stack and a web application but still accidentally point both at prod/terraform.tfstate. Terraform correctly serializes writes to that one state, yet the pipeline scheduler may start competing runs faster than they can finish.
Use an explicit serialization key that corresponds to the backend state identity: account, bucket, workspace and key. Two jobs that can write the same state must enter one deployment queue, whether their Git branches or repository names differ.
name: terraform-prod
on: [workflow_dispatch]
concurrency:
group: terraform-prod-network-state
cancel-in-progress: false
jobs:
apply:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
# Configure AWS credentials and Terraform securely
- run: terraform init
- run: terraform apply -lock-timeout=10m -auto-approve
This GitHub Actions example is illustrative, not a complete security-ready production workflow. The concurrency group serializes runs within the provider’s scope, but cross-repository or cross-system writers still need a coordinated queue. Keep approvals and reviewed plans in place for sensitive deployments.
If your organization uses GitLab, Jenkins, CodePipeline, or another runner, implement the equivalent resource lock or concurrency group. The backend lock remains essential as the last safety layer even when the CI scheduler is careful.
Keep workspaces, state keys and pipeline identities aligned
You may think two runs target different environments because one branch is named staging and another production. That separation is only real if the backend state keys or workspaces actually differ.
The S3 backend uses the configured key for the default workspace. Non-default workspaces use a workspace prefix and name in the object path. A CI job that silently selects the default workspace can collide with production or update the wrong state.
terraform workspace list
terraform workspace show
terraform init -reconfigure
Treat workspace selection as an explicit deployment input. Check the exact S3 state object path in logs and review the configured workspace_key_prefix. Different AWS provider profiles do not automatically mean separate Terraform state files, so check the backend block too.
Migrate safely from DynamoDB to S3 lock files
If your CI fleet uses Terraform 1.10 or later and the S3 backend, native locking is the forward-looking choice. Migration is not simply editing one file in one branch while half your runners continue relying on DynamoDB alone.
During a transition, the S3 backend can accept both dynamodb_table and use_lockfile = true. That compatibility can help coordinate a fleet of old and new clients, subject to the versions in use.
terraform {
backend "s3" {
bucket = "example-terraform-state"
key = "prod/network/terraform.tfstate"
region = "us-east-1"
dynamodb_table = "terraform-locks"
use_lockfile = true
}
}
Plan the cutover: identify all writers, update their Terraform versions, enable lock-object permissions, deploy compatible backend configuration everywhere, and verify no legacy-only runner remains. Only then remove the deprecated DynamoDB option and retire the table through a controlled infrastructure change.
Terraform backend settings are initialized locally and can be cached in runner workspaces. After changing the backend configuration, have every pipeline run terraform init -reconfigure in the correct working directory so an old cached setup does not silently persist.
Protect state with S3 Versioning and recover a bad write
A lock keeps cooperating writers apart. It does not protect you from every mistaken apply, accidental file deletion or broken update. Enable S3 Versioning on the state bucket so earlier object versions remain available for recovery.
aws s3api put-bucket-versioning \
--bucket example-terraform-state \
--versioning-configuration Status=Enabled
Before manually removing a stale lock after a suspicious failure, you can save a copy of the currently readable Terraform state. Treat this file like a secret: store it in a controlled location and do not commit it to Git or attach it to a public ticket.
terraform state pull > terraform-state-backup-2026-10-09.json
aws s3api list-object-versions \
--bucket example-terraform-state \
--prefix prod/network/terraform.tfstate
The first command writes sensitive state data to your local disk. Prefer an encrypted workstation or controlled CI artifact with strict retention. The second command lists object versions; it does not replace the current state or restore one by itself.
If your state is damaged, investigate its lineage, serial number, last successful apply and corresponding AWS resources before restoring an older S3 version. Restoring stale state without reconciliation can cause Terraform to plan incorrect recreations or deletions.
Recover when state was saved locally but not uploaded
An uncommon but important failure happens when Terraform cannot persist the new state back to the remote backend. Terraform may write a local recovery file to avoid losing the changes it already knows about. That file deserves special care.
You might see instructions from Terraform explaining that the state was not saved remotely. Preserve that generated state and do not rerun an uncontrolled apply over it. Compare it with the backend snapshot and the real resources created before the error.
The terraform state push command can overwrite remote state, but it is a high-risk recovery step. Terraform checks lineage and serial in ordinary circumstances; the -force option bypasses protections. Have an experienced owner review backups and the exact target before considering a push.
Protect recovery data
State snapshots can include passwords, tokens and private resource attributes. Keep recovered JSON and S3 object versions restricted, encrypted and outside source control.
Decode the “resource not found” state-lock variant
If Terraform reports ResourceNotFoundException while acquiring a DynamoDB lock, the missing resource is often the lock table, not an AWS instance managed by your configuration. The error is about backend infrastructure and should be addressed before planning changes to application resources.
| Error wording | Likely layer | First safe action |
|---|---|---|
ResourceNotFoundException from DynamoDB |
Missing or incorrectly addressed table | describe-table in correct Region/account |
ConditionalCheckFailedException from DynamoDB |
Lock contention or condition failure | Identify existing holder and wait |
AccessDenied from S3 |
State or lock-object permission | Inspect role, policy and key path |
NoSuchBucket |
Wrong bucket or missing bootstrap | Confirm backend bucket name/account |
PreconditionFailed creating .tflock |
Possible S3 lock contention | Inspect lock holder; avoid deleting blindly |
ExpiredToken or credentials error |
Authentication/session failure | Refresh role/session and retry |
The underlying service text wins over a generic Terraform headline. When the error names a missing table, running force-unlock against it will not create the missing table. When it names the S3 bucket, creating a new empty bucket under a different account would make matters worse.
Choose the least disruptive fix for each error
You do not need the same repair for every lock failure. Start with evidence that distinguishes concurrency, authentication and missing backend resources. Applying one narrow fix is faster than simultaneously rotating credentials, changing backends and deleting locks.
- Read the first underlying AWS error and the lock metadata, including the current state path.
- Confirm AWS identity and backend Region; compare local settings with your CI job.
- If a valid writer is active, retry with
-lock-timeoutor let the deployment queue drain. - If AWS reports access or resource errors, repair the named backend dependency and reinitialize.
- If the lock is orphaned, preserve state and use
terraform force-unlockwith the matching ID. - Run
terraform plan, compare the changes against the previous deployment, and resolve drift before applying.
Jake: "I had an approved plan before the lock. Can I skip the plan after I unlock?"
Ethan: "If another apply happened while yours waited, that old plan may no longer describe the same world. Check again."
Understand terraform apply -lock=false before you use it
The flag -lock=false tells Terraform not to acquire the state lock for supported operations. It is not equivalent to clearing a stale lock, nor is it a recommended production troubleshooting step.
terraform apply -lock=false
That command is shown to identify the risky flag, not as the fix to run. If two processes write the same backend, disabling locking removes the guard that would otherwise coordinate them. A better choice is to resolve the owner, wait with -lock-timeout, or carefully unlock an orphaned lock.
Even when you believe you are the only writer, permissions or missing-table errors should be repaired directly. Suppressing locking could hide an infrastructure configuration problem while leaving future pipelines vulnerable.
Work through a CI collision with actual timing and state paths
Consider a worked example with two production jobs sharing s3://example-terraform-state/prod/network/terraform.tfstate. Job A starts at 09:00 and holds the lock for eight minutes. Job B starts at 09:02. Those timestamps and durations are illustrative, not a benchmark.
With -lock-timeout=3m, Job B can stop around 09:05 if Job A still owns the lock. With -lock-timeout=10m, Job B can wait long enough for Job A to finish at 09:08, then attempt acquisition. The second result still depends on the health of the backend and the job remaining within CI timeout limits.
The least expensive fix is to queue Job B behind Job A before it launches. That reduces wasted runner time and protects operators from deciding a healthy lock is stale simply because their own build is impatient.
After Job A finishes, Job B must review any newly calculated plan against the updated state. A successful lock acquisition means it has exclusive access for that operation; it does not guarantee your proposed infrastructure changes are still correct.
Confirm account, Region and IAM scope before blaming state
In cross-account deployments, your Terraform AWS provider may assume a role in one account while the S3 backend authenticates separately. Confusing the two identities leads to backend errors that look unrelated to the application resources you meant to manage.
Run aws sts get-caller-identity with the same profile or environment variables your backend uses. Check the S3 bucket Region and, for a legacy table, DynamoDB Region and table ARN. Review bucket policy, role policy, KMS permissions and applicable organization restrictions.
Changing the provider block does not necessarily change backend authentication. You can have permission to create EC2 instances while lacking s3:DeleteObject on a .tflock file. Keep those policy scopes separate during diagnosis.
Use observability to prove which writer held the lock
If nobody remembers starting the conflicting apply, start with the lock holder string and CI run history. Compare timestamps across your build runner and AWS activity logs. The goal is attribution, not collecting everything in the account.
S3 object metadata and CloudTrail data events, when configured, can provide supporting context for object operations. DynamoDB activity can likewise be investigated through suitable audit logging. Event availability depends on what your organization has enabled, so absence of a data event alone does not prove nothing happened.
Have each pipeline print its non-secret deployment identifier, backend key, workspace and Terraform version before it initializes. Then the next lock error points to a concrete job instead of a mysterious runner hostname.
Keep the new backend from failing in future releases
After you unblock the deployment, choose a durable operating model. For supported Terraform versions, use S3 native locking, enable bucket versioning, apply least-privilege IAM policies, and keep all writers on consistent backend configuration.
Queue applies per state key; separate unrelated infrastructure into distinct state files where that separation reflects real ownership, not just to evade a lock. Require review of production plans and define how an operator proves a CI job is dead before unlocking.
Finally, pin or control your Terraform version in development and CI so backend capabilities do not differ unexpectedly. An old runner that does not understand use_lockfile should not be assumed to coordinate with a newer runner using native locking.
Know when the next plan indicates more than a lock problem
You have removed a confirmed orphaned lock and the backend now responds. The next terraform plan may still show resources to replace, destroy, import or recreate. That is not something to wave away because the lock error disappeared.
Look for changes made by the interrupted apply before it died. Compare the last approved plan, AWS resource inventory and currently saved state. If the plan is surprising, investigate drift and state recovery before releasing a second apply.
Ethan: "Unlocking opens the door. It does not tell us what the last person changed inside."
That is the principle that keeps lock repair from turning into an outage.
Trace a real CI lock incident without deleting evidence
When your overnight deployment has failed and a morning production release is already queued, your instinct may be to click Retry. The better first move is to open the previous job's timeline. Find the exact point where the job entered Terraform, whether it received a cancellation signal, and whether another run was launched afterward.
A CI job can be marked canceled while its underlying executor still completes cleanup. Another system may also be running the same Terraform directory under a service account. The state lock makes these overlapping systems visible, even when the pipeline interface makes them look independent.
Write down the deployment identifier, state key, runner identity, start time and final status for each run. Compare those entries with the holder shown in Terraform's Lock Info. If the holder is a hostname instead of a recognizable job ID, examine which pool of runners uses that hostname. Reused runners can have similar names from one run to the next, so time and process identity matter.
Then distinguish three timelines. In the first, Job A is actively applying and Job B simply arrived early. In the second, Job A was canceled, but its process survives on a runner. In the third, Job A disappeared entirely and left a stale lock. Only the third timeline calls for manual unlock after you confirm the owner is gone.
For the first timeline, queue Job B or give it a sensible lock timeout. For the second, terminate the remaining owner through the runner's normal management interface, then look for orderly cleanup. For the third, preserve the current state and release the exact orphaned lock. The paths look similar in a red CI dashboard but demand different actions.
Jake: "Two branches, two independent deployments?"
Ethan: "Branches separate code. The bucket key tells us whether they share state."
That distinction is why a preview branch must not inherit a production backend key merely because its Terraform files were copied from the production project.
When the state is shared deliberately, choose one place to authorize writes. Some teams use deployment environments with manual approvals. Others use an application-level mutex or a runner queue. No single CI option protects jobs in an entirely different CI provider, so the Terraform backend lock remains the backstop.
See the S3 permissions that make lock acquisition work
If your Terraform state object is readable but acquiring the lock fails, narrow the investigation to the lock-file path. A backend with the key prod/network/terraform.tfstate needs access to prod/network/terraform.tfstate.tflock too. It is easy to grant permissions to exactly the state file and forget the adjacent object.
Begin with the role assumed by your CI job, not the privileges attached to your own AWS account. Confirm the bucket name and object key first. If the policy was generated from a Terraform variable, inspect the actual value rather than the intended value in documentation.
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Action": "s3:ListBucket",
"Resource": "arn:aws:s3:::example-terraform-state"
},
{
"Effect": "Allow",
"Action": ["s3:GetObject", "s3:PutObject"],
"Resource": "arn:aws:s3:::example-terraform-state/prod/network/terraform.tfstate"
},
{
"Effect": "Allow",
"Action": ["s3:GetObject", "s3:PutObject", "s3:DeleteObject"],
"Resource": "arn:aws:s3:::example-terraform-state/prod/network/terraform.tfstate.tflock"
}
]
}
This is a narrow illustrative identity policy for one backend path, not a complete access configuration. Real installations may require a more restrictive ListBucket condition, workspace prefixes, KMS permissions, cross-account bucket policies or organizational constraints. IAM allows a request only when applicable policies collectively permit it and no applicable explicit deny blocks it.
If the state file uses SSE-KMS encryption, check key policy and role permissions separately. Reading and writing encrypted S3 objects can require KMS actions in addition to the S3 actions. The error may name the KMS service or it may surface through the S3 request, so preserve the complete message when troubleshooting.
Be especially careful if an administrator can read the state while the CI runner cannot. That is a permissions mismatch, not proof the state lock is damaged. Repair the runner's access according to your security design rather than copying an administrator access key into the pipeline.
Verify the state after a canceled apply, not just the lock
If your original apply was interrupted after it created a load balancer but before it completed the rest of the change, releasing its abandoned lock is only one part of recovery. You also need to understand which resources were created, modified or deleted before the failure.
Start with the most recent successful CI job and its approved plan. Record the time of the interrupted apply, then examine the AWS resources it was supposed to change. Look at current Terraform state and compare it with what the cloud actually contains. It is entirely possible for an API call to succeed while later steps in the same apply fail.
When the backend object has multiple S3 versions, compare their modification times and version IDs. Versioning gives you an opportunity to recover an earlier snapshot; it does not tell you which snapshot accurately reflects all resources. Restoration must be treated as a controlled reconciliation exercise, not a reflexive rollback button.
Suppose the state file before the interrupted apply did not include a newly created security group, but AWS now contains it. Restoring the older state without handling that group might make Terraform propose a duplicate or fail due to conflicting names. In other cases, the new resource exists and the current state already tracks it. The safe next step depends on the actual state and remote resources.
Use terraform plan to see Terraform's proposed reconciliation, then inspect unexpected replacements and deletions. Where necessary, Terraform provides state-management and import workflows, but they require the actual resource addresses and correct current configuration. Blindly importing or deleting resources to make a plan green can make the final state less trustworthy.
For a production incident, preserve the original backup file and relevant object versions while investigating. Keep a record of who approved each state change. When a change involves customer data, shared networking or encryption keys, coordinate with the service owner before applying a corrective plan.
Jake: "So every canceled deployment means hours of detective work?"
Ethan: "Only the risky ones need deep recovery. A plan showing no surprising changes is useful evidence. A plan proposing to delete your database is a reason to stop."
Check Terraform version compatibility before enabling use_lockfile
You may be reading a modern article while your company runs an older pinned Terraform binary. Native S3 locking is available beginning with Terraform 1.10. Older versions may reject the use_lockfile argument because they do not recognize it.
Check the version in the exact environment that runs the deployment. Your laptop version is not proof that a Dockerized runner or managed build agent uses the same release. Include terraform version in job logs and consider a compatible required_version constraint in the configuration.
During migration, test a nonproduction state with the new version and access policies before changing production. Review provider and Terraform compatibility separately. Changing backend lock mechanics should not require changing the application resource definitions in the same deployment.
If you have multiple repositories or several accounts sharing one backend infrastructure stack, build an inventory of each writer and its version. The awkward migration is when one forgotten nightly job continues using the DynamoDB table after everybody else moves to S3 lock files. Removing the table too early can cause those old jobs to fail, and mixing uncoordinated lock strategies risks concurrent writes.
Keep both backends' lock mechanisms configured during the supported transition until every writer has migrated. Then schedule retirement of the DynamoDB dependency and remove obsolete permissions. A controlled change avoids waking up to a table-not-found error from an old pipeline weeks later.
Version pinning also helps you reproduce lock behavior in development. If a feature behaves differently between environments, record the Terraform CLI version, backend configuration and AWS credentials used rather than assuming the provider version alone determines state locking.
Before the next planned release, run a short operational review: identify who can approve production writes, which repository owns each backend key, which CI runners may still be active, and where the last known-good state object version is stored. A written recovery checklist pays for itself the first time an interrupted pipeline appears at the end of a long deployment day.
FAQ: Terraform S3 state lock failures
These answers are the fast follow-ups you may need when a specific lock message, backend option or CI symptom sends you here.
What does terraform error acquiring the state lock mean?
Terraform could not acquire the coordination lock required for the configured backend operation. Another writer may hold it, or the backend may have an access, configuration or resource error. Read the underlying AWS message and lock metadata first.
How do I clear a Terraform state lock safely?
Identify the holder and confirm its process has finished. Back up the current state if a failure may have changed resources, then run terraform force-unlock with the exact lock ID for your current backend configuration.
Is terraform force-unlock safe during an active apply?
No. Releasing a lock while another writer is active can allow concurrent state changes. Use force-unlock only for your own abandoned lock after proving the original operation has stopped.
What is the difference between S3 use_lockfile and DynamoDB locking?
S3 native locking creates a temporary .tflock object alongside the state, while legacy DynamoDB locking uses a table item. DynamoDB locking is deprecated; compatible modern Terraform supports use_lockfile = true.
Why is my terraform state lock DynamoDB table not found?
A ResourceNotFoundException typically means the named table is missing or addressed in the wrong AWS account or Region. Inspect the backend table name and run aws dynamodb describe-table using the correct identity.
What does terraform apply -lock=false do?
It bypasses locking for the operation. It does not repair the backend or clear a stale lock, and it can permit concurrent writers. Resolve the real lock issue instead.
How do I set terraform -lock-timeout?
Use a duration such as terraform apply -lock-timeout=10m or terraform plan -lock-timeout=5m. Terraform waits up to that duration for the lock to become available.
Can Ctrl+C leave a stale Terraform lock?
Yes, especially if cleanup is interrupted or the process is forcibly killed. Confirm the process is gone and no other job owns the state before manually unlocking.
Why are two Terraform CI pipelines fighting over one state?
Both pipeline jobs are addressing the same backend state object, so their write-capable operations compete for one lock. Serialize them using a CI concurrency group keyed to the state identity.
Does Terraform S3 state locking require a DynamoDB table?
No. Terraform 1.10 and later can use native S3 lock files with use_lockfile = true. DynamoDB is a deprecated compatibility option, not a required dependency of native locking.
What IAM permissions does an S3 .tflock object need?
Terraform requires s3:GetObject, s3:PutObject and s3:DeleteObject on the lock object, alongside the proper state-object and bucket permissions.
How do I back up Terraform state before force-unlock?
Enable S3 Versioning and securely save the current state using terraform state pull if it is readable. Protect downloaded state because it may contain sensitive information.
Can I delete the .tflock file directly from S3?
Deleting an S3 lock object can bypass normal lock ownership checks. Prefer terraform force-unlock after verifying the original owner is gone; investigate permissions or backend errors first.
Why does Terraform state lock error say resource not found?
The missing resource can be the DynamoDB locking table or another backend resource, not necessarily an infrastructure object in your Terraform configuration. Inspect the named AWS service error and account/Region.
How do I migrate Terraform state locking from DynamoDB to S3?
Upgrade all writers to compatible Terraform, grant S3 lock-object permissions, temporarily configure both mechanisms where supported, update every pipeline, and remove the DynamoDB option only after legacy writers are gone.
Will terraform force-unlock fix a corrupted state file?
No. It releases lock ownership only. Recovering damaged or outdated state may require S3 object versions, comparison with actual resources and a carefully reviewed recovery procedure.
Finish with a safe state instead of just a cleared error
Your immediate goal is not to make the red lock message disappear at any cost. It is to establish one responsible writer, restore the integrity of the backend workflow, and understand whether an interrupted apply made partial changes before it stopped.
If the lock belonged to a healthy job, the answer may simply be waiting and queueing. If it was abandoned by your own stopped operation, a careful unlock can be right. If AWS reports a missing table or object permission, the right fix is a backend repair. I hope your next production run finishes with one clear owner and no mystery about the state it changed.
📌 If you keep one line from this page
Find the lock owner before removing the lock; a blocked Terraform apply is safer than two writers corrupting one state.
Revision note. Written October 9, 2026, with S3 native locking preferred and the DynamoDB table deprecated. Look before you leap: find the lock's owner before you force-unlock anything.