AWS CLI ThrottlingException and Rate exceeded in CI
A ThrottlingException in GitHub Actions is AWS telling your job it asked for too much too quickly, and it reaches your log only after the SDK inside the CLI has already spent its own retry budget. The fix is not another retry wrapped around the step: it is raising the budget the failing CLI actually reads, and then making fewer calls per run.

What this error means
A step that was working yesterday ends with a single line naming an operation you did not think was expensive, usually a describe or a parameter read, and the CLI exits 254. Three shapes carry the same meaning. The CLI formats the error itself, in the form "An error occurred (ThrottlingException) when calling the GetParameter operation", sometimes with a parenthetical count of the retries it already made before giving up. A script that calls boto3 directly raises botocore.exceptions.ClientError with the same text inside a traceback. And the exception name varies by service: Throttling, ThrottlingException, RequestLimitExceeded and TooManyRequestsException are the same answer from different APIs, and the AWS CLI documentation lists all of them among the errors its standard retry mode already handles. This page has no recorded run, because reproducing it honestly needs credentials for an account whose API limits you are willing to exhaust, and neither belongs on a public runner. The block below is the sample text the Latchkey pattern library matches on, labeled as such rather than dressed up as a run.
An error occurred (Throttling) when calling the DescribeStacks operation: Rate exceeded
An error occurred (ThrottlingException) when calling the GetParameter operation: Rate exceeded
botocore.exceptions.ClientError: An error occurred (ThrottlingException) when calling the InvokeFunction operationWhich retry budget your CLI was actually on
The AWS CLI retries throttling responses by itself, and how many times depends on a mode you have probably never set. Version 2 defaults to standard mode, which AWS documents as "A default value of 2 for maximum retry attempts, making a total of 3 call attempts". Three attempts against a limit that lasts longer than a second is the reason so many of these fail.
| Retry mode | Documented attempts | What it retries |
|---|---|---|
standard, the version 2 default | 2 retries, 3 calls | Transient errors, the throttling exception names, and HTTP 500, 502, 503, 504 |
legacy, the version 1 default | 4 retries, 5 calls | A narrower error list, plus HTTP 429, 500, 502, 503, 504 and 509 |
adaptive | As standard | Standard plus client-side rate limiting, documented as experimental |
Common causes
A matrix fanned out and every leg called the same API at once
The most common shape in CI. Ten jobs start within a second of each other, each reads four parameters or describes the same stack, and the account's per-second budget for that operation is spent before any of them finish. Nothing in the workflow looks expensive, because no single job is.
The limit is shared with everything else in the account
API limits are per account and per region, not per workflow. A deployment pipeline, a cost exporter running on a schedule and an engineer running a describe loop locally all draw on the same bucket. In our experience the throttled build is rarely the heaviest caller; it is just the one that was unlucky about timing.
A polling loop is spending calls while it waits
Waiters and hand-written while loops around a describe call make one request per iteration. A stack that takes eight minutes to settle, polled every two seconds by four jobs, is close to a thousand calls that produce no information until the last one.
The default budget is three calls and the burst outlasts it
Standard mode gives up after three attempts and a backoff measured in seconds. When the throttle comes from a genuine spike rather than a momentary blip, three attempts inside twenty seconds is simply not long enough, and the step fails with a message that makes it sound like the API refused you outright.
How to fix it
Raise the attempt budget in the job environment
- Put
AWS_RETRY_MODEandAWS_MAX_ATTEMPTSin the jobenv:block so every AWS call in the job inherits them. - Start at six attempts. The backoff is exponential with a documented ceiling of 20 seconds per wait, so six attempts is still under a couple of minutes.
- Remember the first call counts toward the number.
env:
AWS_RETRY_MODE: standard
AWS_MAX_ATTEMPTS: '6'Make one call where you were making many
Most throttled steps are a loop that could have been a single request. Read parameters in a batch rather than one at a time, filter server side with --query instead of describing everything, and let a waiter do the polling that your shell loop was doing by hand.
# instead of four get-parameter calls
aws ssm get-parameters --names /app/db-url /app/api-key /app/region /app/bucket \
--query "Parameters[].{n:Name,v:Value}"Stop the fan-out from arriving all at once
Cap how many matrix legs run together, or put the AWS-touching jobs in a concurrency group. Ten legs that finish in eleven minutes instead of ten are cheaper than ten legs that fail and get re-run by hand.
jobs:
deploy:
strategy:
max-parallel: 3
matrix:
region: [us-east-1, us-west-2, eu-west-1, ap-south-1]Try adaptive mode when the load is genuinely yours
Adaptive mode adds client-side rate limiting on top of standard mode, so the CLI slows itself down as it sees throttling responses rather than retrying at full speed. AWS documents it as experimental and subject to change, which is a fair reason to keep it out of a deploy path, and a poor reason to avoid it in a nightly job that sweeps a hundred resources.
env:
AWS_RETRY_MODE: adaptive
AWS_MAX_ATTEMPTS: '8'Set the budget where the failing command reads it
Both settings have environment variables, AWS_RETRY_MODE and AWS_MAX_ATTEMPTS, and in a workflow that is the right place for them: a job level env: block covers every AWS call in the job, including the ones inside composite actions and scripts you did not write. The config file works too, and is the better home when the same values have to apply inside a container image.
The count includes the first call. AWS_MAX_ATTEMPTS: 6 means six calls in total, not one call and six retries, which matters when you are reasoning about how much load you are adding to a limit you have already hit.
jobs:
deploy:
runs-on: ubuntu-latest
env:
AWS_RETRY_MODE: standard
AWS_MAX_ATTEMPTS: '6'
steps:
- uses: aws-actions/configure-aws-credentials@v6
with:
role-to-assume: ${{ secrets.AWS_ROLE_ARN }}
aws-region: us-east-1
- run: aws ssm get-parameters --names /app/db-url /app/api-keyA step-level retry is usually the wrong layer
Wrapping the step in a retry action is the first thing most workflows reach for, and it is the layer that helps least. The CLI has already made three calls by the time the step fails; a step retry makes three more, from the same account, against a bucket that is refilling at a fixed rate. If the limit is account wide and several jobs are in it together, that is a way to stay throttled for longer.
Raising the attempt budget inside the CLI is better because the backoff between attempts is the SDK's, and because adaptive mode, if you opt into it, slows the client down instead of hammering. Best of all is the fix that does not retry anything: ask for less.
What a runner does about it
Latchkey runs GitHub Actions jobs on managed runners that watch a failing step's output and act on it, and throttling is one of the classes they are built for, because the failure is real, transient and nothing to do with the code under test. This page claims no repair for this specific error, because that claim belongs with a recorded run and this page does not have one. See how self-healing works for what the runner does with a failing step, and what it deliberately does not touch.
How to prevent it
- Set the retry mode and attempt budget once, in the job environment, rather than per command.
- Prefer batch operations and server-side filters over loops that call an API per item.
- Cap matrix parallelism for jobs that all touch the same account and region.
- Keep scheduled sweeps and deploys off the same hour, so two pipelines do not share one bucket.
Frequently asked questions
What does Rate exceeded mean in an AWS CLI error?
Does the AWS CLI retry ThrottlingException automatically?
What does reached max retries mean in an AWS CLI error?
Should I add a retry action around my AWS CLI step?
AWS_MAX_ATTEMPTS first, reduce the number of calls second, and keep the step retry for failures that are not throttling at all.