Skip to content
Latchkey LogoLatchkey home

GitHub Actions concurrency job stuck pending is not a deadlock

A GitHub Actions concurrency job stuck pending is almost never a job waiting on itself, because the documented behavior makes that particular knot impossible to tie. A concurrency group holds at most one running item, and the default handling of a second waiter is to cancel it rather than to stack it, so an unbounded wait has to be coming from somewhere outside the block you are staring at.

Four mechanisms that leave a job pending, with the scope each one covers and how to confirm it
A concurrency group is scoped to the repository and its name is case insensitive, so the holder can be a run of a different workflow.

What this error means

A job or a whole run shows as pending. It is not queued for a runner, which is a different state with a different cause, and the distinction matters because a pending item is being held by a rule while a queued one is waiting for a machine. Nothing is red, nothing is annotated, and the run page offers no reason. The most misleading part is that the workflow you are looking at often contains a concurrency block, which makes the block look guilty when the item holding the group is frequently in another file entirely.

The query that separates pending from queued, not a log line
gh api repos/OWNER/REPO/actions/runs --jq '.workflow_runs[] | select(.status=="pending") | {name, status, html_url}'

What the group actually holds, and for how long

The documentation is specific about the shape of a concurrency group, and reading it closely removes the deadlock theory. There can be at most one running job or workflow in a group at any time. When a second one is queued while another item in the same group in the repository is in progress, the second becomes pending. By default any existing pending item in that group is then cancelled and the new one takes its place.

Read that last sentence again, because it is the one that matters. The default is not a queue. It is a slot for one waiter, and a new arrival evicts the previous waiter. A design that produced a permanent stand off would need two items each waiting on the other, and the eviction rule means the second waiter never survives long enough to wait on anything.

Two properties of the group name do most of the real damage instead. The group is scoped to the repository rather than to the workflow, so two different workflow files that both use a constant group name are in one group. And the name is compared case insensitively, so prod and Prod are the same group, which is not what somebody writing two blocks in two files usually intends.

What is holding itScope of the holdHow to confirm
A run in the same groupthe whole repositorylist runs and filter on status pending
queue: max with a full queuethe group, up to one hundredread the group block for a queue property
An environment protection rulethe named environmentthe run page shows a waiting deployment
A job waiting for a runnerthe runner label or groupthe state is queued, not pending

Common causes

Another workflow in the repository uses the same group name

The commonest cause by a distance, because the group is repository scoped and a constant name is the natural thing to write. Two files that both say group: deploy are one group, and neither file gives any hint of the other.

The group name differs only by case

Group names are compared case insensitively, so a block that says Prod and one that says prod are in the same group. This looks deliberate in review, because the two files clearly meant different things.

The block sets `queue: max` and the queue is doing its job

Up to one hundred items may wait, in order. A long wait here is the configured behavior, not a fault, and it is invisible unless you read the block rather than the run page.

The run is waiting on an environment, not on concurrency

A job with an environment that carries a required reviewer or a wait timer is held by the deployment rules. In our experience this is the one most often reported as a concurrency problem, because both end in a run that sits there.

The item is queued rather than pending

Waiting for a runner with matching labels is a different state with a different fix. If the API says queued, nothing on this page applies and the runner labels are where to look.

How to fix it

Establish the state before you change anything

  1. Ask the API for the runs and read the status field rather than trusting the color on the page.
  2. If the status is queued, stop here and go and look at runner availability and labels.
  3. If it is pending, list every run in the repository with that status and look for the one that is in progress in the same group.
shell
gh api repos/OWNER/REPO/actions/runs \
  --jq '.workflow_runs[] | {name, status, conclusion, html_url}' | head -30

Grep the repository for every concurrency block

The holder is frequently in a file nobody opened. One search across the workflows directory lists every group name in the repository, and duplicates jump out immediately once they are side by side.

shell
grep -rn -A2 'concurrency:' .github/workflows/

Make group names unique per workflow and ref

Build the name from github.workflow and github.ref so that a workflow only ever contends with other runs of itself on the same branch. Where a deliberate shared lock is wanted, give it a name that says so.

.github/workflows/ci.yml
concurrency:
  group: ${{ github.workflow }}-${{ github.ref }}
  cancel-in-progress: true

Check the environment before blaming the group

Open the run and look for a deployment waiting for review or for a timer. An environment hold shows on the run page as a deployment that needs approval, which is a different control from anything in the concurrency block.

Put the lock on the job that needs it

A workflow level block holds the whole run. If only the deploy step must be serialized, move the block to that job so the build half of the pipeline is free to run while another deployment finishes.

.github/workflows/deploy.yml
jobs:
  build:
    runs-on: ubuntu-latest
  deploy:
    needs: build
    concurrency:
      group: deploy-production-${{ github.ref }}
    runs-on: ubuntu-latest

The queue property, and the one combination that is refused

Where the newer behavior is available, the block takes a queue property that decides how many items may wait. The default is single, which is the one slot described above: a new arrival cancels the existing waiter. Setting max allows up to one hundred items to wait in the group, processed in the order they started waiting, and any arrival beyond that is cancelled.

This is the setting that turns an apparent stall into an intended one. A deployment group with queue: max in front of it is supposed to hold things, and a long wait there is the feature working. If somebody added it to serialize deploys and did not tell the rest of the team, every later deploy looks stuck.

One combination is not allowed. queue: max together with cancel-in-progress: true is refused with a validation error, because the two describe opposite handling of an item that is already running. If you have both, the file does not compile, which means you are not looking at a pending run at all, you are looking at a run that was never created.

Ordering within the queue is first in, first out by the time each item started waiting on the group, not by the time it was dispatched. The documentation adds that because actual start times vary, the ordering is not guaranteed. Do not build anything that depends on strict sequencing.

.github/workflows/deploy.yml (illustrative)
concurrency:
  group: production-deploy
  queue: max

# refused: the two options describe conflicting behavior
concurrency:
  group: production-deploy
  queue: max
  cancel-in-progress: true

Scoping a group so it holds what you meant

The advice that follows from the repository scope is to stop using constant group names. A group named deploy is a repository wide lock on the word deploy, and every workflow that ever borrows the name joins it. Building the name from the workflow and the ref is the documented way to keep one workflow from cancelling another, because github.workflow differs per file and github.ref differs per branch.

Where a shared lock genuinely is the point, say so in the name. A group called deploy-production used deliberately by three workflows is fine and readable. The problem is never the sharing, it is the sharing nobody knew about.

When you build a name from a property that only exists for some events, give it a fallback. github.head_ref is defined on pull request events and not on others, and a workflow that also responds to pushes needs something to fall back to. The run id is both unique and always defined, which makes it the safe right hand side.

.github/workflows/ci.yml (illustrative)
concurrency:
  group: ${{ github.workflow }}-${{ github.ref }}
  cancel-in-progress: true

# a name that is defined for every event this workflow handles
concurrency:
  group: ${{ github.head_ref || github.run_id }}

Why there is no recorded run on this page

The subject here is a run that is doing nothing, and a recording of a run doing nothing for an interval we chose would be a recording of our own patience rather than of a mechanism. Worse, the interesting variable is what else exists in the repository at that moment, and a fixture repository containing exactly the two runs needed to make a point is a demonstration rather than evidence.

The behavior is instead quoted from the documentation, which states the one running item rule, the eviction of a pending item, the repository scope, the case insensitivity and the queue limits directly. Those statements are the thing worth citing, because they are what GitHub commits to.

Nothing on this page is a failure a runner could repair. A held group is a lock behaving as configured, and a runner that decided to release somebody deployment lock would be a bug.

How to prevent it

  • Never use a constant word as a concurrency group name; build it from github.workflow and a ref.
  • Review every concurrency block in the repository together, because they share one namespace.
  • Write down which groups are deliberately shared, next to the workflows that share them.
  • Give any group name built from an event specific property a fallback that is always defined.

Frequently asked questions

Can a GitHub Actions job deadlock on its own concurrency group?
Not in the way it is usually described. A group holds one running item, and by default a second waiter is cancelled rather than queued behind the first, so a permanent mutual wait cannot form. A long pending state means something else in the repository holds the group.
What is the difference between a pending and a queued run?
Pending means a concurrency group or a deployment rule is holding the item. Queued means it is waiting for a runner with matching labels to become available. Ask the API for the status field rather than judging from the run page, because the two look similar.
Are concurrency groups scoped to one workflow?
No. They are scoped to the repository, and names are compared case insensitively. Two different workflow files using the same group name, or names differing only by case, are in the same group and will hold each other up.
How many runs can wait in a concurrency group?
One by default, and that one is cancelled when a newer arrival takes its place. With queue: max up to one hundred can wait in the order they started waiting, and anything beyond that is cancelled.
Can I combine queue: max with cancel-in-progress?
No. The combination is refused with a validation error because the two options describe conflicting handling of a run that is already in progress. A workflow that has both does not compile, so there is no run to be pending.

Related guides

References

Held by a group is not the same as waiting for a machine. Latchkey supplies the machine. Start free → 30-day trial · No credit card