ECS Exit Code 137: OutOfMemoryError Fix, Every Cause

Logeshwaran
—

ECS exit code 137 means the Linux kernel killed your container's main process with SIGKILL, signal 9, and 128 plus 9 is 137. When the stopped task also says "OutOfMemoryError: Container killed due to memory usage", the container used more memory than its hard limit allowed and ECS pulled the plug. The fix is to find out which limit it hit (the container's own memory value, the Fargate task size, or the EC2 host underneath), then either give that one layer more room or stop the application eating it. Here is the part most guides skip: 137 on its own does not prove an out-of-memory kill. A container that ignores the polite stop signal during a deployment gets exactly the same number, and doubling its memory fixes nothing.

Jake runs a phone repair shop, and his online booking page talks to a small API that Ethan put on ECS Fargate for him last spring. One Tuesday the API restarted twice during the lunch rush, the console showed a stopped task with exit code 137, and Jake's first message to Ethan was "just give it more memory, I'll pay the difference". Ethan did not touch the memory for an hour. He read the stopped task, matched the time against the service events, and found that one restart was a real memory kill and the other was a deployment where the old container never answered the stop signal. Two different problems, one number. This page is the order he worked through, from the ten-minute checks that tell the two apart, through each fix, to what it costs when more memory really is the answer.

⚡ Quick answer

• What 137 means: SIGKILL. Read the stopped task's reason before you decide why.

• Console: ECS → Clusters → your cluster → Tasks → filter Stopped → open the task. You have one hour before it disappears.

• CLI: aws ecs describe-tasks --cluster my-cluster --tasks TASK_ARN and look at stopCode, stoppedReason and each container's reason and exitCode.

• Real OOM: raise the limit that was hit (container memory, Fargate task size, or EC2 capacity), or cut the application's peak. Raising memoryReservation does nothing.

• Java: cap the heap with -XX:MaxRAMPercentage and leave room for everything that is not heap. Node.js: --max-old-space-size limits one heap, not the process.

• 137 during a deployment with normal memory: the container ignored SIGTERM. Fix signal handling and stopTimeout, not memory.

Where the number 137 comes from, and why it is a clue rather than a verdict

Every process on Linux ends with an exit status. When a process dies because the kernel sent it a signal, the shell convention is to report 128 plus the signal number. SIGKILL is signal 9, so a SIGKILLed process reports 137. That is the whole story of the number. It tells you that something killed the process without warning. It tells you nothing about who or why.

SIGKILL is the one signal a program cannot catch, ignore, or delay. The process simply stops, mid-sentence if it has to. That is why an OOM-killed container so often leaves no useful stack trace: it was halfway through writing a log line, or holding an unsaved request, when it vanished. If you have been hunting through application logs for an error that explains the crash, this is why there is nothing there. The application did not crash. It was removed.

Two things send SIGKILL to an ECS container, and they could not be more different. The first is the kernel's memory enforcement: the container crossed a hard limit, and the kernel killed it to protect everything else. The second is ECS itself: it asked the container to stop with SIGTERM, waited for the stop timeout, got no answer, and sent SIGKILL. Same number, opposite fixes. Everything on this page is about telling those two apart before you spend money on the wrong one.

Three exit codes keep each other company in ECS stopped tasks, and it helps to know the other two by sight.

Exit codeSignalWhat it usually means in ECSWhere to look
137SIGKILL (9)Memory limit enforced, host under pressure, or a forced stop after SIGTERM was ignoredStopped reason, container reason, service events at that minute
139SIGSEGV (11)Segmentation fault: a native crash, a bad pointer, a broken native moduleApplication logs, native dependencies, runtime version
143SIGTERM (15)A polite stop that the container obeyed: deployments, scale-in, a manual stopService events; usually nothing is wrong

People search for this in a dozen ways: ecs container exit code 137, exited with code 137, stopped exit code: 137, process finished with exit code 137. Outside ECS, in Docker on your laptop or in Kubernetes, a plain Linux exit code 137 means the same thing: the process was killed, and the platform around it decides what the kill means. A 139 deserves its own investigation, because a segmentation fault is a bug, not a budget. A 143 during a deployment is usually just a container doing what it was told.

Jake: "So 137 is like a knock on the door without a name on the parcel."

Ethan: "Exactly that. The parcel has a label, though. ECS writes the reason on the stopped task, and that label is the first thing we read."

The first ten minutes: read the stopped task before you touch a number

You have a short window. Stopped tasks only appear in the ECS console for one hour, and after that the evidence is gone unless you captured it. If your incidents are usually looked at the next morning, this is the single reason they so often end with a shrug. Turn on event capture and log retention now, before the next one, and the rest of this section stops being a race.

Here is the console route, which is fine for one task on one bad afternoon.

  1. Open the Amazon ECS console and choose Clusters, then the cluster the service runs in.
  2. Open the Tasks tab and change Filter desired status to Stopped. Running tasks hide the one you want.
  3. Pick the task whose stop time matches the restart you saw. In the Last status column, choose Stopped; the pop-up shows the stopped reason.
  4. Open the task itself and read every container's reason and exit code, not just the first one in the list.
  5. Note the task definition revision and the exact stop time. You will match that time against service events and metrics in a moment.

Here is the CLI route, which is the one to use when you want the answer in a file you can keep. First list the stopped tasks, because you need the full ARN.

aws ecs list-tasks --cluster my-cluster --desired-status STOPPED --region us-east-1

Then describe the one you want, and ask only for the fields that matter.

aws ecs describe-tasks --cluster my-cluster --tasks TASK_ARN --region us-east-1 \
  --query 'tasks[].{stopCode:stopCode,stoppedReason:stoppedReason,stoppedAt:stoppedAt,
            taskDef:taskDefinitionArn,containers:containers[].{name:name,exitCode:exitCode,reason:reason}}'

Three fields do the work. stopCode is the category: EssentialContainerExited when the main container died on its own, ServiceSchedulerInitiated when ECS stopped it for a deployment or scale-in, UserInitiated when a person did. stoppedReason is the sentence ECS writes about the task. And each container's reason is where OutOfMemoryError: Container killed due to memory usage appears when the kernel's memory limit was the cause. If that phrase is on the container, you have a real memory kill and you can skip ahead to the limits. If the container shows 137 with no OutOfMemoryError and the stop code says the scheduler initiated it, you are looking at a shutdown problem, and the section on SIGTERM is the one you need.

With a task that runs more than one container, read all of them. The application container is often the bystander: a log router or a sidecar grew, the task hit its ceiling, and everything in it was stopped together. The first container in the list is rarely the one that caused it.

Now pull the service events for the same minute.

aws ecs describe-services --cluster my-cluster --services orders-service --region us-east-1 \
  --query 'services[].events[:15]'

You are looking for a deployment, a scale-in, or a health-check failure at the stop time. A failed load balancer health check makes ECS replace the task; if that container then ignores SIGTERM, it dies with 137, and the memory was never the problem. Ethan's second lunchtime restart was exactly this: a deployment event two seconds before the stop, no OutOfMemoryError on the container, and a Node process that had no handler for SIGTERM.

AWS's own troubleshooting pages now mention one more route: the Amazon ECS MCP server lets an AI assistant read task failures and container logs for you. It is a reasonable way to speed up the reading, but it reads the same fields you just saw, so it cannot find what you did not retain. Capture first.

Nine causes, one table: match the clue you already have

You do not need to read all nine rows. Find the clue you already hold in the middle column and let the row tell you where to go first. Several causes can be true at once, which is why the table sends you to an investigation rather than a verdict.

Likely causeThe clue you will seeFirst move
Container hard limit hitOutOfMemoryError on the container; usage near its memory valueRaise that container's limit or cut its peak
Fargate task size exhaustedAll containers together approach the task's memoryNext supported task size, or trim a sidecar
EC2 host under pressureHost memory low; several tasks stopped around the same timeRealistic reservations, spread tasks, add capacity
Java heap plus native overheadJVM sized at or near the container limitMaxRAMPercentage and non-heap headroom
Node.js buffers and external memoryRSS rises while the JavaScript heap looks fineMeasure RSS, cap concurrency, stream large work
Short spike between samplesGraph looks calm; task still diedFiner telemetry, Maximum statistic, bound the work
Sidecar is the eaterTask high, application container normalContainer Insights per-container metrics
Forced stop after SIGTERM137 at a deployment or scale-in, no OutOfMemoryErrorSignal handling and stopTimeout
Unbounded growthMemory climbs with uptime and never comes backCaches, retained objects, queue backlogs

Hard limit versus soft reservation: the two settings that share a word

This is the section that trips up the most people, so take it slowly. ECS gives each container two memory settings whose names look like the same knob at two sizes. They are not. One of them can kill your container. The other cannot, and raising it when you meant the first is the most common "fix" that fixes nothing.

memory is the hard limit, in MiB. Cross it and the kernel kills the container, and that is the OutOfMemoryError you saw. memoryReservation is the soft reservation. It is a promise to the scheduler about how much the container normally needs, used to decide which EC2 host has room for it. The container is free to use more than its reservation when the host has memory to spare, right up to the hard limit if one is set, or to whatever the host has left if there is not.

Look at what this definition actually promises.

{
  "name": "orders-api",
  "memoryReservation": 512,
  "memory": 1024
}

The scheduler counts 512 MiB against the host when it places this container. The kernel kills it at 1,024 MiB. The two numbers are not added; the container does not have 1,536 MiB. If you raise the reservation to 768 and leave the hard limit alone, you have changed where the scheduler puts the container and nothing else. It still dies at 1,024. When you set both, AWS requires memory to be greater than memoryReservation, and if you set neither at the container level, the task must carry a task-level memory value instead.

Jake: "I raised the reservation twice last month and it kept dying. I thought the console was broken."

Ethan: "The console was telling you the truth, just in a word that sounds like the other word. You moved the promise, not the ceiling."

So why have a soft reservation at all? Because on the EC2 launch type it is how you share a host honestly. Picture one container instance running several small services. Each one idles at 300 MiB and occasionally bursts to 900. If you reserve 300 for each, the scheduler happily packs ten of them onto a host that cannot survive all ten bursting at once. If you reserve 900 for each, the host sits half empty most of the day. There is no right answer, only a choice about how much overcommitment you can live with, and a payment service deserves a more honest reservation than a nightly report generator.

⚠ The memory fix that fixes nothing

Raising memoryReservation feels like giving the container more room, and it does not move the hard limit by a single byte. Check which boundary was actually hit first. Otherwise you change your scheduling math and keep the same kill.

On Fargate, the memory lives at the task level, and every container shares it

If you run on Fargate, the picture shifts one level up. Fargate makes you declare CPU and memory for the whole task, and that task size is the real ceiling. Containers inside the task share it. You can still give individual containers their own hard limits with memory, and those limits apply on top, but leaving them out does not make anything unlimited. The task size is always there.

Here is Jake's booking API as a Fargate task with 2 GiB, an application container, and a log router beside it.

{
  "family": "orders-example",
  "requiresCompatibilities": ["FARGATE"],
  "networkMode": "awsvpc",
  "cpu": "1024",
  "memory": "2048",
  "containerDefinitions": [
    { "name": "orders-api",      "image": "example/orders-api:1", "essential": true,  "memory": 1536 },
    { "name": "logging-sidecar", "image": "example/log-router:1", "essential": false, "memory": 256 }
  ]
}

Read the three ceilings. The task has 2,048 MiB. The application may use up to 1,536 of it before its own limit kills it, even if the task as a whole still has room. The sidecar may use 256. If the sidecar were left without a limit and quietly grew to 700 MiB, the application could be killed at well under its own 1,536 because the task ran out. That is the sidecar row in the decision table, and it is the one that makes people swear the application was innocent. It was.

Which number do you raise? If the task has spare memory but the application keeps hitting its own container limit, raise the container limit and keep the task total in mind. If the containers together fill the task, move to the next supported task size. If a sidecar is the one growing, fix the sidecar; giving the application a bigger limit just lets the two fight over the same ceiling.

The Fargate sizes you are allowed to pick, including the new 32-vCPU row

You will need this table the moment you decide to resize, because Fargate only accepts certain pairs. Type a memory value it does not offer for that CPU and the task definition fails to register, with Invalid 'cpu' setting for task from the CLI or No Fargate configuration exists for given values from Terraform. The values are in MiB in the JSON, so 1 GB is 1024.

Task CPUMemory you can chooseNote
0.25 vCPU (256)512 MiB, 1 GB, 2 GBLinux only
0.5 vCPU (512)1, 2, 3 or 4 GBLinux only
1 vCPU (1024)2 to 8 GB in 1 GB stepsLinux and Windows
2 vCPU (2048)4 to 16 GB in 1 GB stepsLinux and Windows
4 vCPU (4096)8 to 30 GB in 1 GB stepsLinux and Windows
8 vCPU (8192)16 to 60 GB in 4 GB stepsLinux platform 1.4.0 or later
16 vCPU (16384)32 to 120 GB in 8 GB stepsLinux platform 1.4.0 or later
32 vCPU (32768)60, 120 or 244 GBAdded June 5, 2026; Linux, x86 and ARM

The 32-vCPU row is new. AWS added it on June 5, 2026, for both x86 and ARM, on Fargate and Fargate Spot, in every commercial Region, and existing Compute Savings Plans apply to it. It exists for data processing and inference jobs, and it is not where a memory problem in a web API should take you. If your CPU is already enough, pick the smallest pair that gives you the memory headroom you measured. Buying vCPUs to get memory is paying twice.

One more thing the table hides: on Windows tasks, the task-level CPU and memory are not enforced at runtime, so if you run Windows containers, the container-level memory is the only ceiling that actually bites.

Why Java containers die at 137 even when the heap looks right

If your service is Java, there is a trap that catches almost everyone once. The heap is the only memory number most people set, and the heap is only part of what a JVM needs. Thread stacks, class metadata, the code cache, garbage collector structures, direct buffers, and the native memory of every library all live outside the heap. Give a JVM a 2 GiB heap inside a 2 GiB container and you have left it nothing for the rest. The heap never fills. The container dies anyway.

Older JVMs added a second trap: they sized themselves from the host's memory instead of the container's, because they could not read cgroup limits. A JVM that believes it has 64 GB to play with inside a 2 GB container is going to find out otherwise. Any current JDK is container-aware, but "any current JDK" is a claim worth checking on the image you actually ship.

The clean fix is to size the heap as a share of what the container has, and leave the rest alone.

java -XX:MaxRAMPercentage=70.0 -jar app.jar

In a 2,048 MiB container, 70 percent gives the heap about 1,434 MiB and leaves roughly 614 MiB for everything else. Whether that is enough depends on your application: a service with hundreds of threads or heavy direct-buffer use may need a lower percentage, a simple API can often take more. The point is that the split is now a decision you made, not an accident the kernel resolved for you. You can pass the flag without touching the command line, which keeps it consistent across environments.

{ "name": "JAVA_TOOL_OPTIONS", "value": "-XX:MaxRAMPercentage=70.0" }

To see what the JVM thinks it has, run java -XshowSettings:system -version inside the container. If the memory it reports is the host's and not the container's, no heap percentage will save you; fix the JDK first.

Jake: "Mine is a Java app. I gave it a 2 GB heap inside a 2 GB container and it still died."

Ethan: "That is the classic one. The heap is only part of what Java needs. Leave it room to breathe, and the number you set for the heap stops being the number the container sees."

One last distinction, because the words collide. A Java OutOfMemoryError: Java heap space is an exception the JVM throws when the heap fills; you get a stack trace and often a heap dump. The ECS OutOfMemoryError is the kernel killing the process from outside; you get nothing. If you have the stack trace, your heap is too small. If you have only a 137, your container is.

Node.js: the heap flag you already set is doing less than you think

Node.js runs on V8, and V8's managed JavaScript heap is only one part of the process. Buffers, native modules, the runtime itself, and anything allocated outside the heap all count toward the resident memory the kernel measures. The flag everyone reaches for, --max-old-space-size, caps V8's old-generation heap in MiB. It does not cap the process. Set it to 1,280 inside a 2 GiB container and you have a sensible starting point, not a guarantee.

node --max-old-space-size=1280 server.js
# or, in the task definition:
{ "name": "NODE_OPTIONS", "value": "--max-old-space-size=1280" }

What you actually want to know is where the memory is going, and Node will tell you if you ask. Log this at a quiet moment and again during the operation you suspect.

const m = process.memoryUsage();
console.log({
  rssMiB: Math.round(m.rss / 1048576),
  heapUsedMiB: Math.round(m.heapUsed / 1048576),
  heapTotalMiB: Math.round(m.heapTotal / 1048576),
  externalMiB: Math.round(m.external / 1048576),
  arrayBuffersMiB: Math.round(m.arrayBuffers / 1048576)
});

If the heap stays flat and RSS keeps climbing, the problem is outside the heap and no heap flag will touch it: look at buffers, native modules, and anything that parses whole files into memory. If heap and RSS rise together, something is holding on to objects: caches without a cap, sessions that never expire, global arrays, work that was started and never finished. Jake's API turned out to be the first kind. A photo upload for the repair form was read whole into a Buffer, six customers uploaded at once, and the task died at exactly the moment the shop was busiest. Ethan streamed the uploads and capped concurrency at three, and the memory graph flattened without a single change to the task size.

Why the CloudWatch graph said 60 percent and the task died anyway

If you have stared at a flat memory graph and thought "but it never went above 60 percent", you are not imagining things and you are not wrong about what you saw. Both things are true. A CloudWatch graph is a series of samples, not a recording of every byte at every instant. Suppose your container idles at 620 MiB, picks up one large document, climbs to 1,100 MiB for four seconds, and is killed because its hard limit is 1,024. A one-minute average will show you something comfortably under the limit, and the task will still be gone.

Three habits make the graph honest. Use the finest period available for the incident window. Compare the Maximum statistic against the Average; a spike that disappears in the average often survives in the maximum. And make sure you are looking at the task or container that died, not the service-wide line, because the service average is the other tasks' calm diluting one task's crisis. Even then, a maximum is the highest reported sample, not the true instantaneous peak.

🕐 June 2026: faster scaling, same blind spot

ECS service auto scaling now accepts 20-second CPU and memory metrics for target tracking, alongside the 60-second default, using the ECSServiceAverageMemoryUtilizationHighResolution metric type; AWS's own test cut the time to trigger a scale-out from 363 seconds to 86. It helps a service grow before it is swamped. It does not make a four-second spike visible, and scaling out never stops one container from crossing its own hard limit. The high-resolution metrics are billed under CloudWatch pricing.

For short spikes, the honest instrument is the application. Log memory before and after the expensive operation, then try it with one input, then six. If one file is safe and six at once are not, a concurrency limit is a more direct fix than a bigger task, and it is free.

Container Insights: the tool that names the hungry container

When a task has several containers and the service graph will not tell you which one is eating, Container Insights with enhanced observability will. It adds per-task and per-container memory metrics, and in a multi-container task that is the difference between knowing and guessing. It has a cost, like all CloudWatch metrics, and in an incident it pays for itself in the first hour.

Turn it on for one cluster, or for every new cluster in the account.

# one existing cluster
aws ecs update-cluster-settings --cluster my-cluster --settings name=containerInsights,value=enhanced

# every new cluster, account-wide (the account root makes it apply to all users and roles)
aws ecs put-account-setting --name containerInsights --value enhanced --principal-arn arn:aws:iam::123456789012:root

Two details catch people. Without --principal-arn, the account setting applies only to the user or role that ran the command. And the account setting covers clusters created from then on; an existing cluster needs the first command.

Once it is on, the metrics you want are MemoryUtilized and TaskMemoryUtilization at the task level, and ContainerMemoryUtilized and ContainerMemoryUtilization per container, all in the ECS/ContainerInsights namespace. Read the dimensions: a line grouped by service tells a different story from one keyed by TaskId and ContainerName. The raw performance records land in a log group named /aws/ecs/containerinsights/CLUSTER_NAME/performance, and one Logs Insights query pulls the peak each container reached in the window.

filter Type = "Container"
| filter TaskId = "YOUR_TASK_ID"
| stats max(MemoryUtilized) as peakObservedMiB by ContainerName
| sort peakObservedMiB desc

That query answered Jake's sidecar question in one line: the log router had been holding 400 MiB of buffered output because its destination was slow, and the application was the innocent party. The stopped reason tells you the task died of memory, Container Insights tells you which container, and the application log tells you what it was doing. Together they are the diagnosis. Any one of them alone is a guess.

The other 137: a container that would not stop

Now the second cause, the one that looks identical and costs people the most money. When ECS stops a container on purpose, for a deployment, a scale-in, or a failed health check, it sends SIGTERM first. That is the polite request: finish what you are doing, close your connections, exit. ECS then waits for the container's stop timeout. The default is 30 seconds and the largest value you can set is 120; on Fargate the parameter needs Linux platform version 1.3.0 or later. If the process is still alive when the timeout ends, ECS sends SIGKILL, and the container exits with 137.

Jake: "So the 137 I saw during the deployment was not about memory at all?"

Ethan: "Quite possibly not. ECS asked the old container to stop, it did not answer, and ECS pulled the plug. That is a shutdown bug, and more memory would not have changed a thing."

The tell is simple. A 137 at a deployment or scale-in, with normal memory on the graph and no OutOfMemoryError on the container, is a shutdown problem. Work through it in this order.

  1. Confirm from the service events that ECS stopped the task on purpose at that time.
  2. Read the container's last log lines. Did it log anything when SIGTERM arrived? Silence usually means the signal never reached the process.
  3. Check the entrypoint. A shell script that launches the app with node server.js on its own line receives the signal itself and never passes it on. Use exec node server.js, or run the process directly, so the application is PID 1 and gets the signal.
  4. Make sure the application actually handles SIGTERM: stop accepting new work, finish in-flight requests, close the database pool, exit.
  5. Look at how long cleanup really takes. If it is honestly 60 seconds, set stopTimeout to 90 in the container definition. If the handler waits forever, no timeout is long enough.
{ "name": "orders-api", "essential": true, "stopTimeout": 90 }

Raising the timeout buys time for an application that is genuinely tidying up. It does nothing for one that never got the message. Jake's API fell into the entrypoint trap: a wrapper script started Node, the script got the SIGTERM, Node never did, and every deployment ended with a 137 that had nothing to do with memory. One exec fixed a month of mystery restarts.

Changing the right number, from the console or the CLI

Only now, with the evidence in hand, do you change a number. Two rules keep you honest. If a container with a 1,024 MiB limit keeps needing 1,200, raising the task size alone will not help while that container limit stands. And raising a container limit cannot create memory beyond the Fargate task size. Change the ceiling that was hit, keep the total in mind, and keep the old revision so you can go back.

From the console: open Task definitions, choose the family, create a new revision, adjust the task-level CPU and memory and each container's hard limit and reservation, register it, then update the service to the new revision and watch the deployment. From the CLI, pull the current definition first.

aws ecs describe-task-definition --task-definition orders-api:12 --region us-east-1

You cannot feed that whole output straight back in; it carries read-only fields that registration rejects. Trim it to the definition itself, edit the memory values, and register the new revision, then point the service at it.

aws ecs register-task-definition --cli-input-json file://task-definition.json --region us-east-1
aws ecs update-service --cluster my-cluster --service orders-service --task-definition orders-api:13 --region us-east-1

The new allocation applies to tasks created from the new revision. Running tasks do not grow because you registered something; they are replaced as the deployment rolls. Before you judge the change, confirm the replacement tasks are actually on revision 13 and not an older one still finishing its deployment.

If you want a starting number for a service you have not measured yet, here are honest ones. Treat them as a first guess you will measure against, not as a recommendation from AWS, because it is not one.

WorkloadA Fargate starting pointWhat to measure before trusting it
Small, low-traffic API0.25 vCPU / 1 GiBPeak memory per request
Node.js web service0.5 vCPU / 2 GiBRSS, heap and buffers together
Java API1 vCPU / 3 GiBHeap plus non-heap peak
Python background worker0.5 vCPU / 2 GiBBatch size and object growth
Document or image processing2 vCPU / 8 GiBDecoded size times concurrency
In-memory data batch4 vCPU / 16 GiBPeak working set

Size from the peak you measured plus headroom, not from the steady state. A service that sits at 500 MiB and bursts to 1.5 GiB needs the 1.5 figure. A service that climbs from 500 MiB to 2 GiB over eight hours and never comes down needs a leak hunt, not a bigger box.

What doubling the memory actually costs

Before you double anything, it helps to know what doubling costs. For one small service the answer is reassuring. For a fleet of ten it is worth a second look. Fargate bills the CPU and memory you request, per second with a one-minute minimum, whether the application uses it or not. In US East (N. Virginia), Linux on x86 is $0.000011244 per vCPU-second and $0.000001235 per GB-second; the first 20 GB of ephemeral storage is included. ARM is cheaper, at $0.0000089944 and $0.0000009889.

Take one task running all month, 30 days, which is 2,592,000 seconds.

Task sizeCPU for the monthMemory for the monthTotal
1 vCPU, 2 GB$29.14$6.40$35.54
1 vCPU, 4 GB$29.14$12.80$41.94
Ten tasks, 2 GB to 4 GBunchanged+$64.00about $64 more a month

So doubling Jake's memory would have cost him $6.40 a month, which is less than one screen protector. That is the reassuring part, and it is why "just give it more memory" is not a silly instinct for a single small service. The catch is the other 137. Paying $6.40 a month, or $64 for a fleet, to fix a container that ignores SIGTERM buys nothing at all, and the restarts keep coming. Logs, data transfer, and any load balancer are on top of these numbers, and prices differ by Region.

Jake: "Six dollars? I spent an afternoon worrying about six dollars?"

Ethan: "You spent an afternoon worrying about the wrong six dollars. The memory kill was worth the six. The deployment one was free to fix and would have cost you six a month forever."

When it still dies after you raised the memory

If you have raised the memory and the task died again, stop doubling. One of the five things below is true, and in my experience it is usually the first or the fourth.

  1. The new revision is not the one running. A deployment can leave old tasks alive for a while, or the service may still point at the previous revision. Check the running task's definition ARN before anything else.
  2. The container limit did not move. You raised the task size, and the container's own memory value is still the old ceiling. The task has room; the container does not.
  3. The peak grew to meet the new limit. If usage climbs to whatever you give it, you are feeding a cache without a cap, a leak, or unbounded concurrency. More memory delays the kill; it does not prevent it.
  4. It is the other 137. The kill happens at deployments or scale-in, memory is normal, and there is no OutOfMemoryError on the container. Go back to the SIGTERM section.
  5. You are reading the wrong line. The service average looks fine because nine tasks are calm and one is dying. Look at the task that stopped, by its ID.

If the stopped task is already gone from the console, you are working from retained logs and captured events, and that is still enough to answer the first four questions. Resist the urge to back-fill a stop reason from the number alone; 137 by itself is a kill, not a cause. For a support case, gather the cluster, service, task ARN, task definition revision, Region, timestamps, stopped reason, exit code, the Container Insights peak, and sanitized application logs. Support can read the service side with you. It cannot reconstruct a spike that was never recorded.

Not meeting it again at 2 a.m.

You have fixed it once. This last section is about not meeting it again at 2 a.m., and most of it is one alarm and one habit. The alarm: a CloudWatch alarm on sustained task memory utilization, at the task or container dimension rather than the service average, set below the level where the kill happens. You want to hear about 85 percent for ten minutes, not about 100 percent for one second. The habit: ECS task state-change events captured to a log group or an EventBridge rule, so the stopped reason survives the one-hour console window and the next incident starts with evidence instead of a search.

Then the small things that each remove one of the rows from the decision table. Give Java explicit non-heap headroom, and measure Node's RSS rather than its heap. Cap concurrency and input sizes for anything that processes files. On the EC2 launch type, keep reservations honest about what containers really use at their busiest. Make the application handle SIGTERM and start it with exec, so the second kind of 137 cannot happen. And after any release that adds a cache, a dependency, a background job, or bigger uploads, look at the peak once; that is the moment the number changes, and it is a quiet moment to catch it.

The thing worth holding on to is the distinction between a bounded workload that genuinely needs more room and an unbounded one that will take whatever you give it. The first is a line in a task definition and six dollars. The second is a bug, and it is a kindness to your future self to call it one.

Exit code 137 questions people actually ask

What does exit code 137 mean in ECS?

The container's process was killed with SIGKILL, signal 9, and 128 plus 9 is 137. It was either killed by the kernel for crossing a memory limit, or killed by ECS because it ignored SIGTERM during a stop. The stopped task's reason tells you which.

Why is my ECS container exiting with code 137?

The four usual causes are a container hard limit that was hit, a Fargate task size that ran out, an EC2 host under memory pressure, or a forced stop after the container ignored the stop signal. Read the container's reason field: OutOfMemoryError means memory; no OutOfMemoryError at a deployment means shutdown.

Does ECS exit code 137 always mean out of memory?

No. A container that does not exit within its stop timeout after SIGTERM is killed with SIGKILL and reports the same 137. If it happens at deployments and the memory graph is normal, fix signal handling, not memory.

How do I fix "OutOfMemoryError: Container killed due to memory usage"?

Find which ceiling was hit: the container's memory value, the Fargate task size, or the EC2 host. Raise that one, or reduce the application's peak by capping concurrency, streaming large inputs, or sizing the Java heap with headroom. Then register a new task definition revision and update the service.

What is the difference between memory and memoryReservation in ECS?

memory is the hard limit; cross it and the container is killed. memoryReservation is a soft reservation the scheduler uses to place the container on an EC2 host; the container may use more when the host has room. They are not added together, and raising the reservation never raises the hard limit.

Can a Fargate task run out of memory when the main container looks fine?

Yes. Every container in the task shares the task's memory. A log router or sidecar that grows can exhaust the task and take the healthy application down with it. Container Insights with enhanced observability shows each container's usage separately.

Why does CloudWatch show low memory right before an OOM kill?

The graph is sampled, and a spike of a few seconds can fall between samples or vanish in a one-minute average. Use the finest period, compare Maximum with Average, and look at the task that died rather than the service-wide line. For short spikes, measure inside the application.

How do I confirm an ECS memory kill with Container Insights?

Enable enhanced observability on the cluster, then read ContainerMemoryUtilized and ContainerMemoryUtilization for the task ID that stopped. A Logs Insights query on the cluster's performance log group returns each container's peak in the window.

How do I fix Java ECS exit code 137?

Size the heap as a share of the container with -XX:MaxRAMPercentage, leave room for thread stacks, metaspace, code cache and native buffers, and confirm the JVM reads the container's memory rather than the host's with java -XshowSettings:system -version.

Does --max-old-space-size prevent Node.js OOM kills?

Not on its own. It caps V8's old-generation heap, not the process. Buffers, native modules and the runtime count toward the memory the kernel measures. Log process.memoryUsage() and watch RSS, not just heapUsed.

What does ECS exit code 139 mean?

SIGSEGV, a segmentation fault. It is a crash in native code or a bad memory access, not a memory budget problem, so a bigger task will not fix it. Look at native dependencies and the runtime version.

What does ECS exit code 143 mean?

SIGTERM, the polite stop signal, obeyed. It is normal during deployments, scale-in and manual stops. Check the service events to confirm the stop was intended.

Why does ECS show stopped exit code 137 during a deployment?

The old container did not exit within its stop timeout after SIGTERM, so ECS sent SIGKILL. The usual causes are an entrypoint script that swallows the signal, or an application without a SIGTERM handler. Fix those before raising stopTimeout, which has a maximum of 120 seconds.

How much memory should I give a Fargate task?

Measure the real peak, add headroom, then pick the nearest supported CPU and memory pair; Fargate only accepts certain combinations, from 0.25 vCPU with 512 MiB up to 32 vCPU with 244 GB. Size from the peak, not the steady state.

Why does my ECS task still exit with code 137 after increasing memory?

Either the running task is not on the new revision, the container's own hard limit was left at the old value, the usage grew to meet the new limit, or it was the shutdown kind of 137 all along. Check them in that order.

What does "process finished with exit code 137" mean?

The process was killed with SIGKILL. On a laptop it is usually the operating system or Docker enforcing a memory limit; in ECS, read the stopped task's reason to learn whether memory or a forced stop was the cause.

How long do stopped tasks stay visible in the ECS console?

One hour. After that, the stopped reason is gone unless you captured task state-change events to a log group or EventBridge. Set that up before the next incident.

What does doubling Fargate memory cost?

In US East (N. Virginia) on Linux x86, memory is $0.000001235 per GB-second, so going from 2 GB to 4 GB on one task running all month adds about $6.40. Ten tasks add about $64. CPU is the larger part of the bill and does not change.

If your task is restarting right now, give the next stopped task one honest look before you change anything, because that one line of reason is the difference between a six-dollar fix and a free one. Jake's booking API has not restarted since the week Ethan read the stop reasons instead of the memory graph: the upload is streamed, the entrypoint says exec, and the task is exactly the size it always was. The number 137 did not change. What changed is that it stopped being a verdict and became a question with a short answer.

📌 If you keep one line from this page

Exit code 137 tells you a process was killed. The stopped task's reason tells you whether memory is the problem you need to fix.

Read the reason before you raise the number.

Revision note. Written October 8, 2026, with the 32-vCPU Fargate sizes and the 20-second scaling metrics in place. If a task of yours died during real traffic, keep its stop reason before you touch the allocation; that one line saves the most expensive wrong turn.

Related