S3 503 Slow Down Error: Why It Happens and Every Fix
An Amazon S3 503 Slow Down error (you will also see it as SlowDown: Please reduce your request rate) means requests hit one prefix, the front part of an object's key name like backups/2026-09-29/, faster than S3 had scaled that prefix to take them. The published rate is at least 3,500 write requests (PUT/COPY/POST/DELETE) or 5,500 read requests (GET/HEAD) per second per partitioned prefix. The twist: that figure is a floor S3 scales toward, not a wall you hit at a fixed number, so a sudden jump in traffic can draw 503s while S3 is still catching up. The fix is boring and reliable: retry with backoff, ramp traffic up gradually, and spread keys across several prefixes.
Jake's Saturday started with a log file full of red. His backup script was copying 1,200,000 receipt scans into a single path called backups/, and every few seconds S3 answered Slow Down. His first guess was that S3 was down. His second was that he'd been banned. Neither was true, and the real cause was much less dramatic.
"Nothing is broken," Ethan told him. "S3 adds capacity for a prefix as your traffic shows up, and it says Slow Down while it catches up. Your job is to arrive more gradually and spread out." This post is the hour of explaining that followed: the real numbers, the myths, the fixes from cheapest to most drastic, and what to do when none of it works. Jake asks the questions you were about to type.
503 Slow Down: the exact error text and what S3 is telling you
The same error shows up under different names depending on which SDK (the software library your code uses to talk to AWS) or tool printed it. Here is the whole family side by side, so you can match what is sitting in your log.
| What you see | What it means | Where to go |
|---|---|---|
AmazonS3Exception: Slow Down (Service: Amazon S3; Status Code: 503; Error Code: 503 Slow Down; Request ID: A4DBBEXAMPLE2C4D) |
Too many requests against a prefix right now. Retryable. | Fix 1 |
Please reduce your request rate. with Error Code: SlowDown |
The same error in different words. | Fix 1 |
ServiceUnavailable, HTTP 503, "Service is unable to handle request." |
S3 could not take the request at that moment. Retry. | Fix 1 |
500 Internal Error |
S3 could not manage the request at that time. Retry. | Fix 1 |
HTTP 200 with InternalError or SlowDown in the body of a copy |
An error tucked inside a "success" response. | Edge cases |
ThrottlingException with HTTP 400 from KMS |
Not S3. Your encryption key service hit its quota. | Edge cases |
GlacierExpeditedRetrievalNotAvailable |
No Expedited retrieval capacity at that moment. | Edge cases |
Two details are worth pinning down. First, the Request ID in that message and, when present, the S3 Extended Request ID are your receipts. Save several pairs; they are the first thing Support asks for. Second, SlowDown is a 5xx status, meaning server-side, even though the cause is usually your traffic pattern. That is why S3 does not bill it, which the cost section covers.
Jake: "So is S3 down?"
Ethan: "No. S3 is a distributed service, so a very small percentage of 5xx errors is expected in normal use. Picture a checkout counter that opens more registers as the line grows. While the new register warms up, the cashier says 'one moment.' That is a 503. It only becomes your problem when the retries run out, or when the rate of 503s stays high."
What S3 rate limits actually are (as of September 29, 2026)
Here are the numbers with the scope on each, because scope is where most bad advice goes wrong. A prefix is the front part of an object's key, which is the object's full name in the bucket. In backups/2026-09-29/receipt-0001.jpg, both backups/ and backups/2026-09-29/ are prefixes. The console draws them as folders, but S3 keeps one flat list of keys.
| Limit | Number | Scope | What to remember |
|---|---|---|---|
| Writes: PUT, COPY, POST, DELETE | At least 3,500 per second | Per partitioned prefix, general purpose bucket | "At least" and "partitioned" both matter |
| Reads: GET, HEAD | At least 5,500 per second | Per partitioned prefix, general purpose bucket | Ten prefixes reading in parallel can reach 55,000 per second |
| Number of prefixes | No limit | Per bucket | Each prefix scales on its own |
| Hot objects | Sustained rates over 5,000 requests per second to a small number of objects can draw 503s | Per object | Cache them instead of hammering S3 |
| S3 Express One Zone directory bucket | Default 200,000 reads and 100,000 writes per second; up to 2 million reads and 200,000 writes | Per directory bucket | Higher than the default through AWS Support |
| Expedited Glacier retrieval | At least 3 retrievals every 5 minutes and up to 150 MB/s per provisioned capacity unit | Per unit | Without provisioned capacity there is no Expedited SLA |
Read the first row twice: at least, and partitioned. S3 partitions a busy prefix as your request pattern demands, and the 3,500 and 5,500 apply to each partitioned prefix. Nothing in the published figures caps how many prefixes you can use, and a bucket's overall rate is a multiple of how many prefixes carry traffic.
The figures date from July 17, 2018, when S3 raised its per-prefix rates to at least 3,500 writes and 5,500 reads per second at no extra charge, and they have not moved since.
🙋♂️ Jake's Reality Check
"My script sends 6,000 writes a second and works most days. Are the numbers wrong?"
The straight answer. No. They are a floor for a prefix that has scaled to your traffic, not a promise about any single second. Traffic that lands on a prefix S3 has not scaled yet, or that jumps suddenly, is what draws 503s. Once S3 finishes scaling, requests are generally served without retries.
Myth 1: the limit is 5,500 requests per second per bucket
This is the most common misreading, and it makes people design around the wrong number. The figures are per partitioned prefix. A bucket that keeps images under images/ and videos under videos/ can handle double the rate of one that keeps everything under a single prefix, because S3 scales each prefix separately as its request rate rises. Ten prefixes reading in parallel can reach 55,000 read requests per second.
Ethan: "Stop asking how many requests your bucket can take. Ask how many each of your prefixes can take, and how your keys are spread across them."
Two caveats keep this honest. Scaling is gradual, not instant, and actual performance varies with your workload. So 55,000 is what ten scaled prefixes can reach, not what ten brand-new ones will swallow in the first second.
Myth 2: you must randomize key names to avoid 503s
Old advice said to put random characters at the front of every key. That guidance was retired: the July 17, 2018 announcement removed the earlier guidance to randomize prefixes and said logical or sequential names carry no performance implications. So 2026-09-29/receipt-0001.jpg is a perfectly fine key.
Here is where the myth still bites. All of tonight's traffic lands on tonight's prefix. A job that writes a million files under one date prefix is a million requests aimed at one prefix, however tidy the names look. The current performance guidance shows the fan-out idea with a short leading segment, like a1b2/log-2024-01-01.txt instead of a bare log-2024-01-01.txt. Randomness is not required to get the base rate. Spreading traffic across prefixes is how you go past it.
Jake: "So naming order does not matter, but where the traffic lands does?"
Ethan: "Exactly that."
Myth 3: creating more prefixes gives you more capacity right away
When you create a prefix, S3 does not automatically assign extra resources for the supported request rate. S3 scales based on request patterns. A prefix is not something you reserve. It is a piece of a key name, and a key that does not exist yet has no traffic history to scale from.
Jake: "So I can't pre-create 100 folders and get 100 times the speed on day one?"
Ethan: "No. Each prefix ramps on its own. What 100 prefixes buy you is a much higher ceiling once each has ramped, with every prefix absorbing its own share. If you know a launch-day spike is coming that one prefix cannot absorb, there is a human route: ask AWS Support to provision partitions for your prefixes. I cover that in Fix 5."
Myth 4: if my average rate is under the limit, I cannot get a 503
This is the answer to the search that starts S3 SlowDown when not close to limit. Averages hide the moments that matter. CloudWatch request metrics arrive in one-minute buckets, and S3's limits are per second. A job that sends 120,000 requests in the first ten seconds of a minute and then goes quiet averages 2,000 per second on the minute graph and looks harmless, while 12,000 per second actually arrived in that burst.
A 503 without a scary average usually comes from one of these:
- A sudden jump. A high request rate, or a sudden increase in the rate for an object or prefix, can return Slow Down while S3 scales.
- Hot objects. Sustained rates over 5,000 requests per second to a small number of objects can draw 503s, however many prefixes you have.
- Rapid concurrent requests to the same key. Rapid concurrent requests to one key can return 503, and retrying is the right response.
- Concurrent PUTs to the same object. Storage Lens has a dedicated metric for this, covered in the measuring section.
- Cross-Region copy bandwidth. S3 can return Slow Down when copies exceed the bandwidth available to you across Regions.
Jake's Lambda that texts order confirmations behaves perfectly all month and misbehaves only on the last day, when every invoice run fires at once. Averages love hiding that kind of day.
Myth 5: 503s mean S3 is down, and every one costs you money
Neither. Bucket owners are billed for requests that return HTTP 200 OK and for HTTP 4xx client errors, with some listed exceptions. They are not billed for HTTP 5xx server errors, including 503 Slow Down. The failed attempts are free. What you pay for is the retry that finally succeeds, once, plus the compute time your code spends waiting between attempts. The cost section works through the arithmetic.
And a very small percentage of 5xx responses is normal for a distributed service. What deserves attention is a sustained rise in the rate, or an application that gives up after the SDK's retries run out.
Which kind of SlowDown do you have? A five-minute triage
Match the symptom, then jump to the fix. If more than one row fits, work in the order the fixes are numbered, because they run from cheapest to most drastic.
| Symptom | Likely cause | Go to |
|---|---|---|
| Errors start when a job starts and fade after a while | Traffic jumped faster than S3 scaled the prefix | Fix 2: ramp up |
| One prefix throws errors while others are fine | That prefix carries more than its scaled rate | Fix 3: spread out |
| Errors cluster on the same handful of objects | Hot objects | Fix 4: fewer requests |
| 503 on PUT to one key from several writers | Concurrent PUTs to the same object | Edge cases |
| Only copies or replication across Regions fail | Cross-Region bandwidth | Edge cases |
HTTP 400 ThrottlingException |
KMS quota, not S3 | Edge cases |
| Glacier restore returns 503 | No Expedited capacity | Edge cases |
| A Spark or EMR job fails with Slow Down | Too many tasks writing at once | Edge cases |
| Slow Down shows up mostly on LIST calls | A big bucket is being scanned with LIST | Fix 4: fewer requests |
| Several jobs, or several accounts, share one prefix | Their combined load lands on one prefix | Edge cases |
| Nothing above fits, or errors persist under steady load | Needs a closer look | Support checklist |
| Fix | Effort | Extra cost | Use it when |
|---|---|---|---|
| 1. Retries | A config change | None | Always. It is the baseline. |
| 2. Ramp up | A scheduler or a lower concurrency setting | None, but the job runs longer | A job starts at full speed |
| 3. More prefixes | A key layout change | None for new writes; $0.005 per 1,000 to COPY old objects | One prefix carries too much |
| 4. Fewer requests | Caching or batching | Often lowers request charges; CloudFront has its own pricing | Hot objects and repeated reads |
| 5. Support or Express One Zone | A case, or a migration | Express has its own pricing | Predictable spikes or latency-critical work |
Ethan: "Whatever the row says, turn retries on first. It costs nothing and it hides a lot of small trouble."
Fix 1 (cheapest): make sure retries are on and sized for the job
The cheapest fix is the one already inside your SDK. Whether the error surfaces from boto3, the Java SDK, the JavaScript SDK or the AWS CLI, the AWS SDKs retry S3 503 responses automatically with exponential backoff, which means waiting longer after each failed attempt. If you skip the SDK and call the REST API yourself, retry logic on 503 is your job.
SlowDown sits on the SDK's throttling list, so it gets the long base delay: 1,000 ms, doubling on each retry, with full jitter (a random fraction of the wait, so a thousand clients do not all retry at the same instant), capped at 20 seconds. The formula is delay = random(0,1) x min(20,000 ms, 1,000 ms x 2^retry).
| Retry number | Longest possible wait | Worst-case total so far |
|---|---|---|
| 1 | 1,000 ms | 1 second |
| 2 | 2,000 ms | 3 seconds |
| 3 | 4,000 ms | 7 seconds |
| 4 | 8,000 ms | 15 seconds |
| 5 | 16,000 ms | 31 seconds |
| 6 | 20,000 ms (the cap) | 51 seconds |
The average wait is about half of each figure. With the default of 3 attempts, one try plus two retries, the worst-case added wait is 3 seconds. For a batch job that is often too stingy. For a customer waiting on a web page it is about right.
🕐 What changed between versions
- Before: Java, Python, Ruby, PHP, C++ and the AWS CLI defaulted to a legacy retry mode with no shared retry quota, and backoff numbers differed by language.
- Now: retry updates were announced on May 20, 2026. You opt in with
AWS_NEW_RETRIES_2026=true, and the new behavior becomes the default in November 2026. Standard mode adds a retry quota of 500 tokens, and each throttling retry costs 5. - What that means here: the numbers above describe that standardized behavior. An older SDK version without the opt-in may use different delays and defaults, and any value you set explicitly always wins. Check your SDK's tracking issue to see whether your version has shipped the flag.
There is no console switch for any of this. Retries live in your code or config. Set them like this:
- Set the mode explicitly.
AWS_RETRY_MODE=standardin the environment, orretry_mode = standardin~/.aws/config. Standard is the default in the current SDK reference, but the SDKs that still default to legacy today (see the box above) will not use it until the change lands, so do not leave it to chance. - Raise max attempts for batch work with
AWS_MAX_ATTEMPTS=8ormax_attemptsin the config file. The count includes the first try, and 1 turns retries off. - Leave user-facing calls at the default, so a customer's page does not hang for half a minute.
- If you wrote your own retry loop, add jitter. Without it, every worker retries in lockstep.
- Log the request ID when the final attempt fails, so a Support case has something to hold.
# Shell: read by the AWS CLI and by SDKs that honor these variables
export AWS_RETRY_MODE=standard
export AWS_MAX_ATTEMPTS=8
# Or write them to ~/.aws/config for the default profile
aws configure set retry_mode standard
aws configure set max_attempts 8
When the same setting exists in several places, code beats the environment variable, which beats the config file, which beats the SDK default. The way to set retries in code differs per language, so use your SDK's developer guide for that part.
Adaptive mode adds a client-side rate limiter that can delay even the first request when it sees throttling. It suits a client that hammers one resource, like a batch processor writing to one bucket. It does not suit a client serving many tenants or prefixes, because throttling on one slows every request from that client.
Jake: "Why not set max attempts to 100 and be done?"
Ethan: "Because retries are more traffic, and a hundred patient retries turn a two-minute job into a two-hour mystery. The SDK protects you a little: each throttling retry spends 5 tokens from a 500-token budget, so a client that never succeeds stops retrying after about 100 of them and returns the error. Retries are a seatbelt, not a steering wheel. Turn them on, size them, then fix the traffic shape."
One more retry idea applies to slow requests rather than errors. For small requests under 512 KB, retry after 2 seconds and again after 4 more, on a new connection with a fresh DNS lookup. That is for latency-sensitive apps, not a cure for Slow Down.
Fix 2: ramp traffic up instead of slamming the door
A sudden jump in request rate is the classic trigger, so the remedy is to start below where you need to be and climb. Here is the routine:
- Start at a rate you know is fine, for example 500 writes per second.
- Raise it in small steps on a timer, say every 10 minutes. AWS publishes no step size, so start conservative and let your error graph set the pace.
- Watch
5xxErrorson the prefix. If they rise, hold the current step until they drain. Do not climb during errors. - Stop at the rate you need. After S3 finishes scaling, requests are generally served without retries.
AWS publishes no figure for how long S3 takes to scale a prefix, so the error graph is your clock: hold the current rate while 5xxErrors drain, then climb again. For aws s3 sync and aws s3 cp, the lever is concurrency. The default is 10 concurrent requests, and lowering it eases the pressure:
aws configure set default.s3.max_concurrent_requests 5
Concurrent requests are not requests per second, since the rate also depends on object size, but fewer in flight means a gentler start. Now the worked example. Jake's job writes 1,200,000 small receipt files. Prices are US East (N. Virginia), where writes cost $0.005 per 1,000 requests, so 1,200,000 writes cost 1,200 x $0.005 = $6.00 every time, because 503 responses are not billed.
| Strategy | Rate | Time to write 1,200,000 files | Write cost |
|---|---|---|---|
| One prefix at the published floor, best case | 3,500 per second | 343 seconds (5.7 minutes) | $6.00 |
| One prefix, start at 500 per second and add 500 every 10 minutes | 500, then 1,000, then 1,500 | 600 + 600 + 200 = 1,400 seconds (23.3 minutes) | $6.00 |
| Eight prefixes at 500 per second each | 4,000 per second in total | 300 seconds (5 minutes) | $6.00 |
Read the last two rows together. Ramping one prefix is safe but slow, while eight prefixes at a modest 500 each finish in a fifth of the time and never ask any single prefix for much. Ramping protects you from the jump. Spreading raises the ceiling. Do both when the job is big.
Fix 3: spread the load across more prefixes
Each prefix scales separately as its request rate rises, so more prefixes means a higher total ceiling. The cheap way to get there is a shard segment in the key, chosen by your code. Here is Jake's key before and after, with eight shards picked by the file number modulo 8:
before: backups/2026-09-29/receipt-000001.jpg
after: backups/2026-09-29/s01/receipt-000001.jpg
backups/2026-09-29/s02/receipt-000002.jpg
...
backups/2026-09-29/s00/receipt-000008.jpg
There are trade-offs, and I would rather you hear them now. Whatever reads the data must know how to compute the shard. Listing becomes one call per shard. Lifecycle rules and metrics filters that select by prefix have to cover every shard, and one bucket can hold up to 1,000 metrics configurations, which is plenty. Each shard still ramps on its own, so sharding does not skip Fix 2.
Jake: "Can't I just rename my existing 1,200,000 files into shards?"
Ethan: "Not in a general purpose bucket. RenameObject exists only in S3 Express One Zone. Renaming here means COPY to the new key and then DELETE the old one. The copies cost $0.005 per 1,000 in US East, so 1,200,000 of them come to $6.00, and DELETE is free. But those copies are themselves write requests, aimed at the same 3,500-per-second floor. My advice: shard new writes and leave old data where it is, unless the old data is your hot read path."
Fix 4: send fewer requests in the first place
The best request is the one you never send. These options are in rough order of payoff:
- Cache reads. CloudFront, a content delivery network that keeps copies of your objects at locations near your readers, or ElastiCache, an in-memory cache, means your repeated GETs never reach S3. That also lowers your request charges. A hot object is the textbook case.
- Combine small files. If 1,200,000 receipts were packed into 1,200 archives of 1,000 each, the upload would be 1,200 PUTs, about $0.006, instead of 1,200,000. The cost is that reading one receipt now means fetching an archive, so it fits backups better than a website.
- Stop scanning with LIST. LIST requests are charged at the PUT rate, $0.005 per 1,000, which is 12.5 times the $0.0004 per 1,000 for a GET. S3 Inventory can hand you a list of your objects without a scan.
- Be careful with parallel range reads. Downloading a big object in byte ranges is good for throughput, but every range is its own request. Start with one request at a time, measure bandwidth and CPU, then add concurrency. If one request uses 25 percent of your CPU, up to four in parallel is a sensible ceiling.
- Reuse connections and spread DNS. A pool of HTTP connections avoids setup cost on every request. Code that pins a single IP address misses the load balancing that comes from S3's wide pool of addresses.
Note that browsing a bucket in the S3 console also generates GET and LIST requests, billed at the same rates as API calls. It is a small thing, and an easy one to forget when you are watching a request graph.
Fix 5 (most drastic): ask Support to pre-scale, or change bucket type
Provisioned partitions. If you expect traffic spikes that exceed what one prefix supports, contact AWS Support and ask them to provision partitions for your prefixes. Send them the prefix names, your expected peak rates, and the dates, because that is what turns a vague worry into a doable request. Console route: Support Center, then Create case.
S3 Express One Zone. This storage class lives in directory buckets. Each directory bucket defaults to 200,000 reads and 100,000 writes per second, can support up to 2 million reads and 200,000 writes, and does not depend on key names or access pattern. Higher throughput than the default goes through AWS Support. Directory buckets organize keys into directories instead of prefixes, authenticate with session-based credentials from CreateSession, and store data in one Availability Zone, an isolated data center location inside a Region, designed for 99.95 percent availability within that zone. Express also has its own request and storage prices, so read the S3 pricing page before you migrate.
✅ Why this is the one to use
Try Fix 1 through Fix 4 first. They cost a few lines of config and a key layout, and they fix most Slow Down cases. Express One Zone is a real answer to a real problem, like latency-sensitive analytics or AI training data, and a poor answer to a nightly backup.
When the fix does not work: what each failure looks like
A fix that half works is the most confusing state, so here is how to read what you see.
- 5xxErrors is not zero, but nothing fails in your app. This is normal. The SDK retried and the retry worked. A small percentage of 5xx responses is expected.
- 503s at the start of every run, fading after a while. You are still climbing too fast. Lower the starting rate or shrink the step size in Fix 2.
- 503s that never fade at a steady rate. The prefix is carrying more than it supports, or a few hot objects are taking the traffic. Spread the keys, cache the hot objects, then ask Support.
- Errors on one key only. Look for concurrent writers to that key.
- The job dies with the error after the final attempt. Retries ran out. Raise max attempts for batch work and add the ramp. Also note that when the SDK's retry budget is empty it returns the error immediately instead of retrying, which is what a sustained storm looks like from the outside.
- Sharded the prefixes and saw no change in the first minutes. Each new prefix ramps on its own, so give it a gradual climb.
- HTTP 400 instead of 503. That is KMS. Prefix changes will not help.
See it happening: console, CLI and logs
You cannot fix what you cannot see, and the tools that count 503s have to be switched on first. Before you switch anything on, find out which operation is failing. If it is LIST, the fix may be an S3 Inventory report instead of a scan, and no prefix redesign is needed. You have three tools, and they answer different questions.
CloudWatch request metrics give you one-minute counts, including 5xxErrors and AllRequests. They are off until you create a metrics configuration, which can cover the whole bucket or a filter by prefix, object tag or access point. They are billed at the standard CloudWatch rate, and delivery is best-effort, so treat them as a near-real-time picture rather than a complete ledger.
- Console: open S3, choose the bucket, open the Metrics tab, go to Request metrics, and choose Create filter. Scope it to the busy prefix. Labels shift a little between console updates. Plan on about 15 minutes before the charts fill in.
- CLI: create the same configuration with the two commands below. The
Idyou pick becomes theFilterIddimension in CloudWatch. - Read the metric with
get-metric-statisticsand look for minutes where5xxErrorsis not near zero.
aws s3api put-bucket-metrics-configuration \
--bucket amzn-s3-demo-bucket --id busy-prefix \
--metrics-configuration '{"Id":"busy-prefix","Filter":{"Prefix":"backups/"}}'
aws cloudwatch get-metric-statistics --namespace AWS/S3 \
--metric-name 5xxErrors \
--dimensions Name=BucketName,Value=amzn-s3-demo-bucket Name=FilterId,Value=busy-prefix \
--start-time 2026-09-29T08:00:00Z --end-time 2026-09-29T09:00:00Z \
--period 60 --statistics Sum
S3 Storage Lens answers "which prefix?" If you upgrade a dashboard to the Advanced tier and enable detailed status code metrics, you get 503 counts, and prefix-level metrics show which prefixes are receiving throttling during ingestion. Since December 2, 2025 it also has daily performance metrics, including a concurrent PUT 503 count, which tracks 503s caused by simultaneous PUTs to the same object. Its suggested fixes: with a single writer, adjust retry behavior or use S3 Express One Zone; with several writers, use a consensus mechanism, meaning a way for the writers to agree who writes next, or use Express.
🕐 What changed between versions
- Before December 2, 2025: Storage Lens offered status code metrics, including 503 counts by prefix on the Advanced tier.
- Now: daily performance metrics at organization, account, bucket and prefix level, including a concurrent PUT 503 count and its percentage.
- What that means for the steps above: if 503s cluster on a few write-heavy keys, check this metric before you touch your prefix layout.
Server access logging plus Athena gives you every request, so you can filter for 500 and 503 responses and see exact keys and timestamps. It is the heavy tool, and the one that finds a hot object.
The cases most explanations skip
HTTP 200 OK with SlowDown or InternalError in the body
CopyObject, UploadPartCopy and CompleteMultipartUpload can return 200 OK first and then embed an error in the body if something fails during the copy. The message looks like We encountered an internal error. Please try again. (Service: Amazon S3; Status Code: 200; Error Code: InternalError). The SDKs detect the embedded error and retry using your configured settings. If you call the REST API directly, you must read the whole response body and parse it. A small percentage of these is normal, and retrying is the fix. Cross-Region copies can throttle this way.
Cross-Region bandwidth
S3 can return Slow Down when your copy requests exceed the bandwidth available to you across Regions, even if the request counts look modest. The steps are the same: retry, raise attempts, and lower how many copies run at once.
Concurrent PUTs to the same key
Rapid concurrent requests to one key can return 503, and the Storage Lens concurrent PUT 503 metric will show it. If several writers update one key, decide who writes, or write to different keys.
KMS ThrottlingException looks like the same problem, and is not
If your bucket uses SSE-KMS, meaning encryption with a key in AWS KMS, each request can count against KMS quotas for the account and Region. KMS answers with You have exceeded the rate at which you may call KMS. Reduce the frequency of your calls and HTTP 400. That is not an S3 503, and adding S3 prefixes will not fix it. An S3 Bucket Key cuts the request traffic from S3 to KMS, and it applies to new objects. Replication counts too: replicated objects use your KMS quota. Athena queries over a large number of small KMS-encrypted objects can hit the same throttling, and Athena's own backoff does not always prevent it. Bucket Keys or a higher KMS quota are the two levers.
Glacier Expedited retrievals
During sustained high demand, S3 can deny Expedited requests and return a 503 with GlacierExpeditedRetrievalNotAvailable. Each provisioned capacity unit provides at least three Expedited retrievals every five minutes and up to 150 MB/s. Otherwise switch to Standard or Bulk. There is no Expedited SLA without provisioned capacity.
EMR, Spark and Glue-style jobs
Big jobs fail with Slow Down when hundreds of tasks write at once. The same advice applies when a Glue job, which runs Spark, ends with Please reduce your request rate while writing lots of small files. EMRFS retries with exponential backoff and defaults to 15 retries, adjustable through fs.s3.maxRetries, though a very high value lengthens the job. Cut output partitions with .coalesce() or .repartition() before writing, or reduce cores per executor or the number of executors.
aws s3 sync and aws s3 cp
The CLI's transfer commands are multithreaded, with 10 concurrent requests by default, and they start at full speed. If every file sits under one prefix, every thread hits that prefix. Lower max_concurrent_requests, raise AWS_MAX_ATTEMPTS for the run, and where you can, sync sub-directories to different destination prefixes.
Lots of Lambda functions writing at once
Jake's booking page starts one Lambda function per booking. On a busy Saturday hundreds run at the same moment, each writing a small confirmation file under one prefix. That is a burst by construction. Cap how many workers can run at once in whatever launches them, and write to more than one prefix.
S3DistCp jobs on EMR
S3DistCp copies can fail with Error Code: 503 Slow Down when many reducers write under one prefix. All the paths under year=2019/ share the prefix year=2019/, however deep the folders below it go. Two levers help: fewer reducers, for example -Dmapreduce.job.reduces=10, and more retries, for example -Dfs.s3.maxRetries=20. Both numbers are placeholders to tune, not targets.
Several jobs, or several accounts, on one prefix
If Spark, Hive and s3-dist-cp jobs all read and write the same prefix, their loads add up, so lower the concurrency of each. If you have set up cross-account access, other AWS accounts may be submitting jobs to that same prefix too. That is the classic reason for a Slow Down when your own graph looks calm.
CloudFront in front of S3
If CloudFront uses S3 as its origin, an origin 503 is relayed to the visitor. The same per-prefix numbers apply at the origin, and the fixes are the same: spread the keys and keep the cache hit rate high.
Scope traps: Region, account, bucket type
Here is where each limit lives, since scope is where advice goes wrong.
- Prefix, not account. The S3 request figures are per partitioned prefix in a general purpose bucket, and the S3 quotas page lists no account-wide request cap. The account-and-Region limit that actually bites is KMS.
- Multi-Region Access Points. Each request behaves the way the bucket it lands in behaves, so that bucket's prefixes carry the load.
- Bucket type. This post covers general purpose buckets and, briefly, directory buckets. Table buckets and vector buckets have their own quota pages, so check those before you assume any number here applies.
- Cross-Region copies. Copy bandwidth between Regions is its own constraint, separate from the per-prefix request rates.
Jake's backup bucket is named whiskers-backups, after his cat. The cat has never once been throttled, which is more than Jake can say.
What a SlowDown costs you, and what it does not
Prices here are US East (N. Virginia). GET requests cost $0.0004 per 1,000 and PUT, COPY, POST and LIST cost $0.005 per 1,000. DELETE and CANCEL are free. The 503 responses themselves are not billed.
The bill that surprises people is compute time. Say Jake's Lambda function, allocated 512 MB, waits an extra 15 seconds per invocation on retries, and it runs 1,000,000 times. Lambda charges $0.0000167 per GB-second, so the extra wait costs 1,000,000 x 0.5 GB x 15 seconds x $0.0000167 = $125.25. The S3 requests behind it may cost pennies. The waiting is what you are paying for.
⚠️ What this actually costs
A retry storm inside paid compute, such as Lambda, EC2 or a Spark cluster, bills for every second spent waiting even though S3 bills nothing for the 503s. Cap max attempts in workloads where the clock is the expense.
When nothing works: the support-case checklist
If errors stay high under steady load after Fix 1 through Fix 4, open a case. Gather these first, because the case moves fast when they are in the first message:
- Several
Request IDandS3 Extended Request IDpairs from failing requests. - Timestamps with time zone, the bucket name and the Region.
- The busy prefixes and your rough request rates per prefix.
- Your SDK or CLI version and your retry settings.
- The
5xxErrorsgraph or the Storage Lens status code metrics. - What changed: a new job, a bigger batch, a new writer, a new Region.
- For a planned spike: the prefixes, expected peak rates and dates.
What Support can do: provision partitions for your prefixes ahead of a spike, and raise directory-bucket throughput above the default. What no one can do: make a client retry politely, or lift a KMS quota from inside an S3 case. Those live in your code and in KMS.
Automation: get paged before your customers notice
Build the alarm once. It goes off when 5xxErrors on the busy prefix stay high for a while, not for one blip. The SNS topic ARN, meaning the unique address of the notification channel, is a placeholder here:
aws cloudwatch put-metric-alarm --alarm-name s3-503-backups \
--namespace AWS/S3 --metric-name 5xxErrors \
--dimensions Name=BucketName,Value=amzn-s3-demo-bucket Name=FilterId,Value=busy-prefix \
--statistic Sum --period 60 --evaluation-periods 5 --datapoints-to-alarm 3 \
--threshold 50 --comparison-operator GreaterThanThreshold \
--treat-missing-data notBreaching \
--alarm-actions arn:aws:sns:us-east-1:111122223333:ops-alerts
The threshold of 50 is an arbitrary starting point. Set it from a normal week of your own traffic. Pair it with Storage Lens for the daily view of which prefix and which status code.
The 16 questions people ask about S3 503 Slow Down
What does Amazon S3 503 Slow Down mean?
It means S3 is asking you to reduce your request rate. The request itself was valid. S3 could not take it at that moment, usually because traffic to a prefix jumped faster than S3 had scaled it. It is a retryable server-side error, it is not billed, and the SDKs retry it automatically with exponential backoff.
How many requests per second can an S3 bucket handle?
The published figure is at least 3,500 PUT, COPY, POST or DELETE requests and at least 5,500 GET or HEAD requests per second per partitioned prefix, not per bucket. There is no limit on the number of prefixes, so a bucket's overall rate is a multiple of how many prefixes carry traffic. Ten prefixes reading in parallel can reach 55,000 reads per second once scaled.
Is the S3 rate limit per bucket or per prefix?
Per partitioned prefix. Two prefixes in one bucket scale separately, so a bucket that uses images/ and videos/ can handle double the rate of one that uses a single prefix. Directory buckets in S3 Express One Zone work differently, with request limits per bucket regardless of key names.
How do I fix Please reduce your request rate from S3?
Work through the fixes from cheapest to most drastic. Confirm SDK retries are on and sized for the job, ramp traffic up gradually instead of jumping to full speed, spread keys across more prefixes, cache or combine files to send fewer requests, and only then ask AWS Support about provisioning partitions or move latency-sensitive data to S3 Express One Zone.
Why do I get S3 SlowDown when I am under 5,500 requests per second?
Because 5,500 is a floor for a scaled prefix, not a guarantee for every second. Sudden jumps, bursts hidden inside one-minute averages, a few very hot objects, concurrent writes to the same key and cross-Region copy bandwidth can all draw 503 responses below that number.
Do I have to randomize S3 key names to avoid 503 errors?
No. The July 17, 2018 announcement removed the guidance to randomize prefixes and said logical or sequential names carry no performance implications. What matters is whether all your traffic lands on one prefix. Adding a short shard segment to spread writes across several prefixes raises your ceiling.
Are S3 503 Slow Down errors billed?
No. Bucket owners are billed for HTTP 200 OK and 4xx responses, with some listed exceptions, but not for 5xx errors such as 503 Slow Down. You pay for the retry that succeeds, and for any compute time your code spends waiting between retries.
Does creating more prefixes give me more S3 request capacity right away?
No. Creating a prefix does not assign extra resources. S3 scales each prefix based on request patterns as traffic rises, so each new prefix ramps on its own. More prefixes raise the ceiling you can reach, but they do not remove the need to climb gradually.
How long does S3 take to scale up after a 503 Slow Down?
AWS publishes no duration. Scaling is gradual, the 503 responses dissipate once it completes, and after that requests are generally served without retries. Watch the 5xxErrors metric, hold your rate while errors drain, and ask AWS Support if high error rates persist under a steady load.
Can I request an S3 request rate limit increase?
For general purpose buckets, contact AWS Support before an expected spike so they can provision partitions for your prefixes. For S3 Express One Zone directory buckets, which default to 200,000 reads and 100,000 writes per second, you can request higher throughput through AWS Support.
Why does aws s3 sync keep failing with SlowDown?
The CLI runs many transfers at once, 10 concurrent requests by default, which can push your rate up quickly. Lower it with aws configure set default.s3.max_concurrent_requests 5, raise AWS_MAX_ATTEMPTS for the run, and if every file shares one prefix, sync into several prefixes instead.
How many times should I retry an S3 503 error?
The SDK default is 3 attempts in total, one try plus two retries. For batch jobs I raise it to about 8, where each wait is jittered and capped at 20 seconds. Keep user-facing calls at the default, and always log the request ID when the last attempt fails.
Why did CopyObject return 200 OK but the copy failed with SlowDown?
CopyObject, UploadPartCopy and CompleteMultipartUpload can send 200 OK first and embed an error in the body if something fails mid-copy. The SDKs detect it and retry using your configured settings. If you call the REST API directly, you must parse the response body. Cross-Region copies can hit throttling this way.
Is a KMS ThrottlingException the same as S3 Slow Down?
No. KMS returns ThrottlingException with HTTP 400 when requests exceed its quota for an account and Region, and S3 requests on SSE-KMS objects can count against it. Turning on an S3 Bucket Key for new objects cuts the request traffic from S3 to KMS.
Does S3 Express One Zone have the same request limits?
No. Directory buckets default to 200,000 reads and 100,000 writes per second per bucket, and can support up to 2 million reads and 200,000 writes, regardless of key names. The trade-offs are a single Availability Zone, a different bucket type and session-based authentication.
What should I send AWS Support when S3 503 errors will not go away?
Send several request ID pairs, meaning the Request ID and S3 Extended Request ID, from failing requests, with timestamps and time zone, the bucket name and Region, the busy prefixes, rough request rates, your SDK version and retry settings, and the 5xxErrors graph or Storage Lens status code metrics.
Thanks for reading this far. If a nightly job or a busy Saturday ever made S3 tell you to slow down, that was the service asking for a gentler start, not passing judgment on you. Jake's backup, sharded into eight prefixes and ramped from a modest start, now finishes before his coffee cools, and the only thing S3 says to him is 200. If you hit a case this post does not cover, send me the error text and I will add it.
📌 If you keep one line from this page
A 503 Slow Down is S3 asking you to arrive more gradually and spread across more prefixes, not a wall and not a bill.
Retry with backoff, ramp up, shard the keys.
Revision note. I wrote this on September 29, 2026, with 3,500 writes and 5,500 reads per prefix, a default of 3 attempts and the November 2026 retry-default rollout in mind. Those are the first three things I would re-check next year. If your Saturday job just turned a calm log into a wall of red, that is a fair thing to be rattled by. Ramp it, spread it, and let the retries do their quiet work.