Skip to content
Latchkey LogoLatchkey home

Docker unexpected status code from registry in CI

A docker unexpected status code from registry message tells you which client gave up, and each of the three clients in play words it differently and fires on a different range of status codes. Read the wording first: it decides whether you are looking at a registry outage worth retrying or an authorization decision that will be identical on every attempt.

Diagram of which client emits which status wording and which codes each one covers
The distribution client only reaches its "received unexpected HTTP status" wording outside the 200 to 499 range. A 401 or a 403 becomes a typed error instead.

What this error means

A pull, a push or a buildx imagetools call fails after the login step reported success, and the line names a status code rather than a reason. Three wordings circulate and they are not interchangeable. The distribution client says "received unexpected HTTP status" and a status. The containerd client names the method and the URL as well. A failure inside the token exchange says neither, and instead repeats whatever the authorization server put in its challenge. No run is recorded for this page, so the block below sets the three wordings side by side as their source formats them.

Three wordings as their source files format them, not a recorded run
--- distribution client, a registry that fell over
received unexpected HTTP status: 502 Bad Gateway
--- containerd client, same command, a token without the scope
unexpected status from HEAD request to https://registry.example.com/v2/acme/api/manifests/1.4.2: 403 Forbidden
--- the token exchange refusing before either of those
server message: insufficient_scope: authorization failed

A 401 and a 403 are different answers, not two spellings of one

The registry API defines an error code per situation, each pinned to an HTTP status. UNAUTHORIZED is 401 and its message is "authentication required": the registry could not work out who you are. DENIED is 403 and its message is "requested access to the resource is denied": it knows who you are and the answer is still no. A third, TOOMANYREQUESTS, is 429 and means the identity is fine and the pace is not.

That distinction survives into the client. When a registry answers a bearer challenge with an OAuth error, the distribution client maps invalid_token onto UNAUTHORIZED and insufficient_scope onto DENIED, which is the difference between a credential that is wrong and a credential that is real but narrow. A read-only token used for a push lands on the second one every time.

StatusRegistry error code and messageWhat it rules out
401UNAUTHORIZED, "authentication required"Nothing about permissions: the registry never identified you
403DENIED, "requested access to the resource is denied"The credential itself, which was accepted and then found insufficient
429TOOMANYREQUESTS, "too many requests"Both of the above: the identity and the scope were fine
5xxNo code, the client reports the status aloneEverything about your credentials, because the registry did not judge them

Common causes

The token is real and its scope is too narrow

The commonest case in CI, because a token minted for pulls is the one already in the secrets store when someone adds a publish job. The registry authenticates it, then refuses the write, and the challenge comes back with insufficient_scope. The giveaway is that the same job pulls from the same repository seconds earlier without complaint.

The registry is having a bad minute

A 502, 503 or 504 from a registry edge produces the distribution client wording with the status attached, and it is the one case here that a retry genuinely fixes. In our experience this shows up on multi-architecture publishes more than on single pushes, because they make many more requests and each one is another chance to catch the bad minute.

No credential was ever attached to this host

Logging in to one registry does not help with another, and a reference with no host in it goes to Docker Hub rather than to the registry you meant. The containerd path reports this as an authorization failure with "no basic auth credentials" appended, which reads like a rejection but is really the client saying it had nothing to send.

The token service the registry points at is not answering

A self-hosted registry names its token service in the challenge it sends back, and the client goes wherever that points. If the realm is wrong, unreachable from the runner, or behind a proxy the job does not use, the failure lands during the exchange rather than on the registry itself. The wording quotes the token service, not the registry.

How to fix it

Mint the token with the scope the job actually performs

  1. List what the job does to the repository: pull only, push and pull, or delete as well.
  2. Issue one token per job shape rather than one shared token widened until everything passes.
  3. Log in inside the job that needs it, so the failure names the step that lacks the scope.
.github/workflows/publish.yml
- uses: docker/login-action@v3
  with:
    registry: registry.example.com
    username: ${{ secrets.REGISTRY_USER }}
    password: ${{ secrets.REGISTRY_PUSH_TOKEN }}

Retry the 5xx range and only the 5xx range

Wrap the push in a bounded retry that inspects the wording before it tries again. A 502 deserves another attempt; a 403 deserves a different token. Retrying both costs you the queue wait and the layer upload twice for an answer that was already final.

Terminal
for i in 1 2 3; do
  if out=$(docker push registry.example.com/acme/api:1.4.2 2>&1); then exit 0; fi
  echo "$out"
  case "$out" in
    *"50"[0-9]*) sleep $((i * 10)) ;;
    *) echo "not a server error, not retrying"; exit 1 ;;
  esac
done
exit 1

Print the challenge when the exchange is what failed

The registry tells an unauthenticated client where to get a token and under what service name. When the failure quotes a token service rather than the registry, that header is the thing to read, because it names the host the client actually went to.

Terminal
curl -sS -D - -o /dev/null https://registry.example.com/v2/ | grep -i www-authenticate

Check the host in the reference before blaming the credential

A reference with a single name in it resolves to Docker Hub, so a job that logged in to a private registry and then pulled acme/api never contacted that registry at all. Print the fully qualified reference next to the login target once and this whole class of confusion disappears.

Terminal
echo "logged in to: registry.example.com"
echo "pulling: $IMAGE"
case "$IMAGE" in */*.*/*|*:*/*) ;; *) echo "::warning::no registry host in the reference" ;; esac

Which client is talking, and why it matters

The "received unexpected HTTP status" wording has a narrow job. In the distribution client, a response between 400 and 499 is parsed into a typed error, and the unexpected-status type is returned only when the code falls outside 200 to 499. In practice that means the 5xx range. So if you find that wording next to a 401 in a document, the document is wrong: that path does not produce it.

The containerd client, which is what Docker uses with the containerd image store and what buildx uses through BuildKit, does report a 403 that way, in its own wording with the method and URL included. It also does something helpful on a HEAD: because a HEAD response carries no body, a 403 on one is repeated as a GET purely to collect the error details the registry would have sent, and only the body is borrowed. That is why the same failure can read as a bare status one day and carry a registry sentence the next.

Terminal
docker info --format "{{json .DriverStatus}}"
# a pair ["driver-type","io.containerd.snapshotter.v1"] means the containerd client
# is the one resolving your references, so expect the containerd wording

Tell the two apart in one step

The registry will answer an unauthenticated ping and tell you where its token service is, which separates "the host is reachable and speaking the registry API" from everything downstream. Then ask that token service for the exact scope your job needs and read what comes back.

A 401 from the ping means the host is fine and you have no credential yet, which is normal. A 403 from the token request with the push scope, while the pull scope succeeds, is the read-only token case and no amount of logging in again will change it.

.github/workflows/publish.yml
- name: Separate authentication from authorization
  run: |
    curl -sS -i https://registry.example.com/v2/ | head -3
    curl -sS -o /dev/null -w "pull scope: %{http_code}\n" -u "$USER:$TOKEN" \
      "https://registry.example.com/token?service=registry&scope=repository:acme/api:pull"
    curl -sS -o /dev/null -w "push scope: %{http_code}\n" -u "$USER:$TOKEN" \
      "https://registry.example.com/token?service=registry&scope=repository:acme/api:push,pull"

Why no recorded run backs this page

The three wordings on this page are produced by three different status ranges, and only one of them can be triggered on demand without breaking something. Recording the 5xx wording would mean making a registry return 502 on cue, which is fault injection rather than reproduction, and it would prove that our fault injector works. Recording the 403 wording would mean publishing a credential that is real enough to be accepted and then refused.

What can be checked without a run is exactly what the page rests on: which range each client formats each wording for, and which code the registry attaches to each status. Those were read in the source rather than inferred from a log, which is the stronger evidence for a claim about which of two things failed.

How to prevent it

  • Keep the pull token and the push token separate, so the narrow one cannot be used widely.
  • Retry on the status range, never on the word "failed".
  • Log in to every registry the job touches, including the one the base image comes from.
  • Record which image store the runner uses, because it decides which wording you will get.

Frequently asked questions

Can received unexpected HTTP status ever mean 401?
Not from the distribution client. That client parses every response between 400 and 499 into a typed error and reaches the unexpected-status wording only for codes outside 200 to 499. A 401 becomes an UNAUTHORIZED error with the message "authentication required" instead. If your log pairs that wording with a 401, something other than that client wrote the line.
Why does the same failure look different on two runners?
Because two different clients resolve references depending on the image store. With the classic store the daemon uses the distribution client and you get its wording. With the containerd image store the containerd remotes package does the work and names the method and URL as well. Neither is more correct; they are separate implementations of the same protocol.
What does insufficient_scope actually mean here?
It is an OAuth error the authorization server puts in its challenge, and the client maps it onto the registry DENIED code, which is a 403. It means your credential was accepted and does not carry the permission the request needed, most often push on a repository you can already pull. A new login with the same token will produce it again.
Should I add a global retry around every registry step?
Only if it reads the error first. An unconditional retry turns a deterministic authorization failure into the same failure three times, and it hides the pattern from whoever reads the log next week. Retry the server-error range, fail fast on 401, 403 and 400, and let 429 back off longer than the rest.

Related guides

References

A 502 from a registry is weather. Latchkey runners handle the transient class and report the rest. Start free → 30-day trial · No credit card