GitHub Actions continue-on-error on a matrix, and the green run
GitHub Actions continue-on-error on a matrix is written once on the job and applied to every combination the matrix produces, because the matrix expands the job and the key goes with it. One leg you wanted to tolerate therefore buys tolerance for all of them, and the run reports success while a test suite is failing.

What this error means
The run is green and a leg is red. In the job list one combination carries a failure marker, its log ends in a real error, and the run conclusion at the top says the whole thing succeeded. Required checks are satisfied, the pull request merges, and nobody is notified. The usual way this is discovered is that a bug reaches production and somebody scrolls far enough down a run summary to find the leg that had been failing for weeks.
test:
continue-on-error: true # written once, applied to every leg
strategy:
matrix:
node: [18, 20, 22]One key, every combination
A matrix does not create a group of related jobs with shared settings. It expands one job definition into a set of configurations, and every key on that definition, including this one, belongs to each configuration. So writing the key at job level is writing it on each leg, which is why the tolerance you meant for one combination applies to all of them.
The documentation states the outcome from the side that matters most: if the key is true on a job, other jobs in the matrix will continue running even if that job fails, and the run is prevented from failing because of it. Both halves are load bearing. The first is about whether the rest of the matrix is canceled, and the second is about what the run concludes.
It is also worth separating this key from fail-fast, because the two are constantly confused. fail-fast controls whether one leg failing cancels the others that are still running, and it has no effect on the conclusion. continue-on-error controls whether a failure counts, and it has no effect on what gets canceled. Turning fail-fast off leaves a failing leg failing the run; turning continue-on-error on makes the leg advisory.
| Setting | Other legs keep running | A failing leg fails the run |
|---|---|---|
| neither key set | no, the rest are canceled | yes |
fail-fast: false | yes | yes |
continue-on-error: true on the job | yes | no, for every leg |
continue-on-error as a matrix expression | yes for the tolerated leg | only for the legs where it is false |
Common causes
The key was added for one experimental combination
A new version was added to the matrix, it failed, somebody made it tolerable, and the key went on the job because that is where the matrix is. Every existing combination became advisory in the same commit, and nothing about the run looked different until one of them broke.
The key was confused with fail-fast
They solve different problems. fail-fast decides whether a failure cancels the legs still running, and continue-on-error decides whether a failure counts at all. Reaching for the second when you wanted the first converts a cancellation question into a reporting one.
The key was inherited from a template
Workflow templates and generated pipelines sometimes carry the key from a context where the whole job really was advisory. Copied into a repository where the job gates merges, it silently removes the gate.
The value came from an expression that is true everywhere
An expression over a matrix key that is absent from most combinations should be false for them, but an expression written against the wrong key, or a flag added through the matrix rather than through an include entry, can be true on every leg. In our experience this is the version that survives review, because the file looks like the scoped fix.
How to fix it
Scope the key to the leg with a matrix flag
Add the flag through an include entry so only the intended combination carries it, and make the job level key an expression over that flag. Legs without the key evaluate as false and keep failing the run.
continue-on-error: ${{ matrix.experimental == true }}Decide whether you wanted fail-fast instead
- If the complaint was that one failure canceled the other legs before they reported, set
fail-fast: falseand leave continue-on-error alone. - If the complaint was that one known-broken combination failed the run, scope continue-on-error to that combination.
- If the complaint was both, set both, because they do different things.
- Never set continue-on-error at job level to solve a cancellation problem.
Audit what the run conclusion is hiding
List the recent runs of the workflow and look at each job conclusion rather than each run conclusion. A leg that has been failing under a tolerant key looks identical to one that has never failed if you only read the top line.
gh run list --workflow ci.yml --limit 20 --json databaseId \
--jq '.[].databaseId' | xargs -I{} gh run view {} --json jobs \
--jq '.jobs[] | select(.conclusion == "failure") | .name'Put tolerance on the step when only one step is unreliable
If the thing you are forgiving is a flaky upload or an optional lint pass, put the key on that step instead of on the job. A step level tolerance leaves the rest of the leg able to fail the run, which keeps the gate where you meant it.
Scoping it to one leg with the matrix itself
The schema types the job level key as a boolean that may read the strategy and matrix contexts among others, which is the mechanism that makes a per-leg answer possible. Add a flag to the matrix through an include entry, default it to false everywhere else, and let the key read it.
The include entry does the scoping. Written against a value that already exists in a dimension, it augments that combination with the flag, and every other combination is left without it. An expression reading a key that is absent evaluates as false, which gives the behavior you want without listing the negative cases.
This is worth doing even when the tolerated leg is the only one you care about today, because it documents the intent. A reviewer reading continue-on-error: true cannot tell whether the author meant one leg or all of them. A reviewer reading an expression over a matrix flag can see the answer in the same screen.
jobs:
test:
runs-on: ubuntu-latest
continue-on-error: ${{ matrix.experimental == true }}
strategy:
fail-fast: false
matrix:
node: [18, 20]
include:
- node: 22
experimental: true
steps:
- uses: actions/checkout@v7
- run: npm testWhat a required check sees
The reason this is worth more than a tidiness argument is what it does to branch protection. A matrix job produces one check per leg, named with the matrix values, and a leg that is tolerated reports a conclusion that does not block. If the required context is the job name without the matrix suffix, or if a gate job reads the matrix job's result, both see something that passed.
That is the whole failure mode: the signal is intact in the run page, where a person has to go looking, and gone from every surface that automates the decision. A test suite that cannot block a merge is documentation, not a gate, which may be exactly what you want for an experimental combination and is rarely what you want for the rest.
If the intent really is advisory, make it visible. Name the leg so the run page reads as advisory, and keep the required check on the legs that are not.
Why there is no recorded run on this page
The symptom here is a green run, which is the one outcome a recording cannot make an argument out of. There is no error line to capture, no annotation, and nothing in any log that says a result was discounted: the leg's own log ends in an ordinary failure and the run summary says success. The evidence is the pair, seen on two different surfaces at once, which is a screenshot of a user interface rather than a log.
The part of this that is checkable without a run is the part worth being precise about: which contexts the job level key is allowed to read, which is in the published schema, and what the documentation states the key does to other jobs in the matrix and to the run. Both are quoted above.
How to prevent it
- Treat a job level key as written on every leg, because that is what the expansion does.
- Express tolerance as a matrix flag, so the file says which combination is advisory.
- Review job conclusions, not run conclusions, when auditing a workflow.
- Keep
fail-fastandcontinue-on-errordecisions separate and commented.
Frequently asked questions
Does continue-on-error on a matrix job apply to every combination?
true makes all the legs advisory. Making the value an expression over a matrix key is what scopes it to one combination.What is the difference between continue-on-error and fail-fast?
How do I let one matrix combination fail without hiding the others?
Will a tolerated matrix leg still block a required status check?
true is dangerous on a job that gates merges.