# GitHub Actions continue-on-error on a matrix, and the green run

> GitHub Actions continue-on-error on a matrix applies to every leg, so one experimental combination turns the whole matrix into advisory checks.

Source: https://latchkey.dev/learn/github-actions/github-actions-continue-on-error-with-matrix-masks-real-failure  
Updated: 2026-09-20

GitHub Actions continue-on-error on a matrix is written once on the job and applied to every combination the matrix produces, because the matrix expands the job and the key goes with it. One leg you wanted to tolerate therefore buys tolerance for all of them, and the run reports success while a test suite is failing.

## What this error means

The run is green and a leg is red. In the job list one combination carries a failure marker, its log ends in a real error, and the run conclusion at the top says the whole thing succeeded. Required checks are satisfied, the pull request merges, and nobody is notified. The usual way this is discovered is that a bug reaches production and somebody scrolls far enough down a run summary to find the leg that had been failing for weeks.

```Illustrative job fragment, not a log line
test:
    continue-on-error: true   # written once, applied to every leg
    strategy:
      matrix:
        node: [18, 20, 22]
```

## Common causes

### The key was added for one experimental combination

A new version was added to the matrix, it failed, somebody made it tolerable, and the key went on the job because that is where the matrix is. Every existing combination became advisory in the same commit, and nothing about the run looked different until one of them broke.

### The key was confused with fail-fast

They solve different problems. fail-fast decides whether a failure cancels the legs still running, and continue-on-error decides whether a failure counts at all. Reaching for the second when you wanted the first converts a cancellation question into a reporting one.

### The key was inherited from a template

Workflow templates and generated pipelines sometimes carry the key from a context where the whole job really was advisory. Copied into a repository where the job gates merges, it silently removes the gate.

### The value came from an expression that is true everywhere

An expression over a matrix key that is absent from most combinations should be false for them, but an expression written against the wrong key, or a flag added through the matrix rather than through an include entry, can be true on every leg. In our experience this is the version that survives review, because the file looks like the scoped fix.

## How to fix it

### Scope the key to the leg with a matrix flag

Add the flag through an include entry so only the intended combination carries it, and make the job level key an expression over that flag. Legs without the key evaluate as false and keep failing the run.

```.github/workflows/ci.yml (illustrative)
continue-on-error: ${{ matrix.experimental == true }}
```

### Decide whether you wanted fail-fast instead

1. If the complaint was that one failure canceled the other legs before they reported, set `fail-fast: false` and leave continue-on-error alone.
2. If the complaint was that one known-broken combination failed the run, scope continue-on-error to that combination.
3. If the complaint was both, set both, because they do different things.
4. Never set continue-on-error at job level to solve a cancellation problem.

### Audit what the run conclusion is hiding

List the recent runs of the workflow and look at each job conclusion rather than each run conclusion. A leg that has been failing under a tolerant key looks identical to one that has never failed if you only read the top line.

```Terminal
gh run list --workflow ci.yml --limit 20 --json databaseId \
  --jq '.[].databaseId' | xargs -I{} gh run view {} --json jobs \
  --jq '.jobs[] | select(.conclusion == "failure") | .name'
```

### Put tolerance on the step when only one step is unreliable

If the thing you are forgiving is a flaky upload or an optional lint pass, put the key on that step instead of on the job. A step level tolerance leaves the rest of the leg able to fail the run, which keeps the gate where you meant it.

## How to prevent it

- Treat a job level key as written on every leg, because that is what the expansion does.
- Express tolerance as a matrix flag, so the file says which combination is advisory.
- Review job conclusions, not run conclusions, when auditing a workflow.
- Keep `fail-fast` and `continue-on-error` decisions separate and commented.

## One key, every combination

A matrix does not create a group of related jobs with shared settings. It expands one job definition into a set of configurations, and every key on that definition, including this one, belongs to each configuration. So writing the key at job level is writing it on each leg, which is why the tolerance you meant for one combination applies to all of them.

The documentation states the outcome from the side that matters most: if the key is true on a job, other jobs in the matrix will continue running even if that job fails, and the run is prevented from failing because of it. Both halves are load bearing. The first is about whether the rest of the matrix is canceled, and the second is about what the run concludes.

It is also worth separating this key from `fail-fast`, because the two are constantly confused. fail-fast controls whether one leg failing cancels the others that are still running, and it has no effect on the conclusion. continue-on-error controls whether a failure counts, and it has no effect on what gets canceled. Turning fail-fast off leaves a failing leg failing the run; turning continue-on-error on makes the leg advisory.

| Setting | Other legs keep running | A failing leg fails the run |
| --- | --- | --- |
| neither key set | no, the rest are canceled | yes |
| `fail-fast: false` | yes | yes |
| `continue-on-error: true` on the job | yes | no, for every leg |
| `continue-on-error` as a matrix expression | yes for the tolerated leg | only for the legs where it is false |

> The last row is what most people wanted when they reached for the key. The job level `continue-on-error` accepts the strategy and matrix contexts, so its value can be read from the matrix itself rather than fixed for the whole expansion.

## Scoping it to one leg with the matrix itself

The schema types the job level key as a boolean that may read the strategy and matrix contexts among others, which is the mechanism that makes a per-leg answer possible. Add a flag to the matrix through an include entry, default it to false everywhere else, and let the key read it.

The include entry does the scoping. Written against a value that already exists in a dimension, it augments that combination with the flag, and every other combination is left without it. An expression reading a key that is absent evaluates as false, which gives the behavior you want without listing the negative cases.

This is worth doing even when the tolerated leg is the only one you care about today, because it documents the intent. A reviewer reading `continue-on-error: true` cannot tell whether the author meant one leg or all of them. A reviewer reading an expression over a matrix flag can see the answer in the same screen.

```.github/workflows/ci.yml (illustrative)
jobs:
  test:
    runs-on: ubuntu-latest
    continue-on-error: ${{ matrix.experimental == true }}
    strategy:
      fail-fast: false
      matrix:
        node: [18, 20]
        include:
          - node: 22
            experimental: true
    steps:
      - uses: actions/checkout@v7
      - run: npm test
```

## What a required check sees

The reason this is worth more than a tidiness argument is what it does to branch protection. A matrix job produces one check per leg, named with the matrix values, and a leg that is tolerated reports a conclusion that does not block. If the required context is the job name without the matrix suffix, or if a gate job reads the matrix job's result, both see something that passed.

That is the whole failure mode: the signal is intact in the run page, where a person has to go looking, and gone from every surface that automates the decision. A test suite that cannot block a merge is documentation, not a gate, which may be exactly what you want for an experimental combination and is rarely what you want for the rest.

If the intent really is advisory, make it visible. Name the leg so the run page reads as advisory, and keep the required check on the legs that are not.

## Why there is no recorded run on this page

The symptom here is a green run, which is the one outcome a recording cannot make an argument out of. There is no error line to capture, no annotation, and nothing in any log that says a result was discounted: the leg's own log ends in an ordinary failure and the run summary says success. The evidence is the pair, seen on two different surfaces at once, which is a screenshot of a user interface rather than a log.

The part of this that is checkable without a run is the part worth being precise about: which contexts the job level key is allowed to read, which is in the published schema, and what the documentation states the key does to other jobs in the matrix and to the run. Both are quoted above.

## FAQ

### Does continue-on-error on a matrix job apply to every combination?

Yes. A matrix expands one job definition into several configurations and every key on that definition belongs to each of them, so a literal `true` makes all the legs advisory. Making the value an expression over a matrix key is what scopes it to one combination.

### What is the difference between continue-on-error and fail-fast?

fail-fast decides whether a failing leg cancels the legs still running, and does not change what the run concludes. continue-on-error decides whether a failing job fails the run, and does not change what is canceled. Wanting the other legs to finish is a fail-fast question.

### How do I let one matrix combination fail without hiding the others?

Add a flag to that combination through an include entry and write the job key as an expression over it. The schema allows the job level key to read the matrix and strategy contexts, so the value can differ per leg, and combinations without the flag evaluate as false.

### Will a tolerated matrix leg still block a required status check?

No. A tolerated leg reports a conclusion that does not block, so branch protection sees a passing check and a gate job reading the result sees success. That is the point of the key and the reason a blanket `true` is dangerous on a job that gates merges.

## References

- [GitHub Actions: workflow syntax, jobs.<job_id>.continue-on-error](https://docs.github.com/en/actions/reference/workflows-and-actions/workflow-syntax)
- [actions/runner: the published workflow schema, job-factory and strategy](https://github.com/actions/runner/blob/main/src/Sdk/WorkflowParser/workflow-v1.0.json)
- [GitHub Actions: run variations of a job with a matrix](https://docs.github.com/en/actions/how-tos/write-workflows/choose-what-workflows-do/run-job-variations)
- [GitHub: about protected branches and required status checks](https://docs.github.com/en/repositories/configuring-branches-and-merges-in-your-repository/managing-protected-branches/about-protected-branches)

---

Latchkey runs CI/CD that repairs its own failures. Agent entry points: https://latchkey.dev/agent.txt, https://latchkey.dev/openapi.json, https://latchkey.dev/llms.txt
