# Docker mkdir /var/lib/docker: read-only file system in CI

> Docker mkdir /var/lib/docker: read-only file system in CI means the daemon data root is not writable. Learn to find which mount went read-only.

Source: https://latchkey.dev/learn/docker/docker-mkdir-var-lib-docker-read-only-in-ci  
Updated: 2026-09-21

Docker mkdir /var/lib/docker: read-only file system means the daemon tried to create a directory under its own data root and the kernel refused because the filesystem holding it is mounted read-only. This is a host condition rather than anything about your image or your workflow, and the daemon is reporting it accurately from one level below where a job can reach.

## What this error means

Pulls, builds and runs all fail together, with paths under `/var/lib/docker` in the message and an operation word at the front that changes depending on what the daemon was doing. One public report of a single broken host carries `mkdir`, `unlinkat` and `open` forms of the same failure within a few lines of each other, which is the tell: when the data root goes read-only, everything that writes there fails, and each failure names whichever call happened to hit it.

```The line abiosoft/colima#883 reports (quoted, not a recorded run)
API error (500): mkdir /var/lib/docker/overlay2/a784aef8fbf3af0f2a38904f816e31b88e268056b829a6b35028d77864d4a4c2-init: read-only file system
```

## Common causes

### The kernel remounted the filesystem read-only after I/O errors

The most common cause on a machine that was working earlier. A failing virtual disk produces read or write errors, and the kernel protects the filesystem by refusing further writes. Every Docker operation that touches the data root fails from that moment, which is why the errors arrive in a burst rather than one at a time.

### The host root filesystem is read-only by design

Immutable and hardened host images ship a read-only root, and a daemon whose data root still points at the default path under it has nowhere to write. In our experience this is the cause when the failure is present from the very first container on a freshly provisioned host rather than appearing partway through its life.

### The data root was pointed at a mount that is not writable

A `data-root` set in daemon.json to a network share, a snapshot or a volume mounted with the ro option produces this immediately. The setting looks right in the config and the mount options are where the problem actually is, so checking the config alone will not find it.

### A container was given the data root as a read-only bind

Tooling that mounts `/var/lib/docker` into a container to inspect it sometimes mounts it read-only, and a nested daemon handed that mount reports exactly this. Here the host is fine and the container is the one that cannot write, which a mount listing inside the container settles.

## How to fix it

### Find out whether the mount or the config is the problem

1. Ask where the daemon thinks its root is, rather than assuming the default path.
2. Read the mount options for that path with `findmnt`, which prints `ro` or `rw` plainly.
3. Check the kernel log, because a mount that went read-only will say why it did.

```Terminal
docker info --format '{{.DockerRootDir}}'
findmnt -no OPTIONS -T "$(docker info --format '{{.DockerRootDir}}')"
dmesg | grep -i -E 'remount|read-only|I/O error' | tail
```

### Move the data root onto a volume you control

On a self-hosted runner, give the daemon a dedicated writable volume rather than leaving it on whatever the host image provided. This survives host images that harden the root filesystem, and it gives you something you can size and monitor separately from everything else on the machine.

```/etc/docker/daemon.json
{ "data-root": "/mnt/docker-data" }
```

### Remount read-write only when the kernel has not objected

If the mount is read-only because of how it was mounted rather than because the kernel demoted it, remounting is the fix. If the kernel demoted it after I/O errors, remounting without repairing the filesystem first is how a bad disk becomes a corrupt one, so read the kernel log before you touch it.

```Terminal
mount | grep /var/lib/docker
sudo mount -o remount,rw /var/lib/docker
```

### Fail the job on an unwritable data root instead of debugging Docker

Add a write probe as the first Docker step so the job reports the host condition in one line rather than producing a pile of unrelated-looking container errors. On hosted runners this converts a confusing failure into a clean rerun.

```Terminal
root=$(docker info --format '{{.DockerRootDir}}')
touch "$root/.ci-write-probe" && rm -f "$root/.ci-write-probe"
```

## How to prevent it

- Keep the daemon data root on a volume provisioned and monitored for it.
- Alert on kernel remount-read-only events on any long-lived runner host.
- Probe that the data root is writable at the start of jobs that depend on Docker.
- Never mount the data root read-only into tooling that will run a daemon against it.

## The operation word is Go, not Docker

The whole tail of this message comes from the Go standard library, which is why it looks identical across completely different failures. Go renders a filesystem error as the operation, a space, the path, a colon and the error text, in `io/fs`, and the final words come from the `syscall` errno table where entry 30 is spelled `read-only file system`. Docker contributes the path and nothing else.

So `mkdir /var/lib/docker/...: read-only file system` exists as a literal in no source file. Searching for it will not find the code, because there is no code to find: it is three pieces glued at run time. That also explains the shifting first word. `mkdir` is a directory being created, `open` is a file being written, `unlinkat` is something being removed, and all three mean the same thing about the mount.

| Operation in the message | What the daemon was doing when it failed |
| --- | --- |
| `mkdir` | Creating a layer directory, usually while pulling or starting a container |
| `open` | Writing metadata, for example the image repositories file |
| `unlinkat` | Removing a container root filesystem during cleanup |
| `read-only file system` | Errno 30 from the kernel, identical in all three cases |

> The number in an `API error (500)` prefix is the HTTP status the daemon returned, so it tells you the daemon answered and failed, not that anything about the request was malformed.

## Why a data root goes read-only mid-life

A mount does not usually start read-only on a machine that has been running containers. The kernel remounts a filesystem read-only when the block device underneath it reports errors, which is its way of refusing to make a corrupt filesystem worse. On a CI host that means a failing or full virtual disk, and the container errors are the symptom rather than the disease.

The other route is deliberate. Hardened and immutable host images keep the root filesystem read-only on purpose, and a daemon whose data root was never moved off it cannot write anywhere. That version fails from the first container rather than partway through a working day, which is a useful way to tell the two apart.

```Terminal
# is it the mount, and has the kernel said why
findmnt -no TARGET,SOURCE,OPTIONS -T /var/lib/docker
dmesg | tail -40
docker info --format '{{.DockerRootDir}}'
```

## On a hosted runner this is not yours to fix

A GitHub-hosted runner is a fresh virtual machine that the job does not administer, so if its data root is read-only the machine is faulty and the right move is to fail fast and rerun rather than to remount anything. Nothing in a workflow step can repair a filesystem the kernel has already given up on, and attempts to do so mostly produce a second confusing error on top of the first.

On a self-hosted runner you do own the host, and then this is worth fixing properly rather than rebooting around. Move the data root onto a volume you provision for it, and monitor that volume, because a data root sharing a disk with everything else is a data root that eventually fills or fails.

```.github/workflows/ci.yml
- name: Fail fast if the daemon cannot write
  run: |
    docker info --format '{{.DockerRootDir}}'
    touch "$(docker info --format '{{.DockerRootDir}}')/.ci-write-probe" 2>/dev/null \
      || { echo "daemon data root is not writable on this host"; exit 1; }
```

## What the runner does about it

No repair, and no recorded run, and here the reason is that the failure lives below the job. To record this we would have to put a runner daemon data root on a read-only mount before the job started, which means shipping a deliberately broken runner and photographing it. A capture of a machine we sabotaged is not evidence about machines you would use, and the mechanism is fully described by the Go formatting and the kernel behavior above without one.

It is also the wrong shape for automatic repair even in principle. Remounting a filesystem the kernel took read-only, underneath a running job, is an action that can turn a failed build into lost data. The honest offer on this failure is a fast, legible verdict rather than a retry.

## FAQ

### Why does the operation word change between mkdir, open and unlinkat?

Because Go names whichever call failed. The message is assembled by the Go standard library from the operation, the path and the errno text, so the first word is simply what the daemon was doing when it hit the read-only mount. All three forms mean the same thing about the filesystem.

### Is this the same as running out of disk space?

No, and the errno says so. A full disk gives you errno 28 and the words `no space left on device`; a read-only mount gives you errno 30 and these words. They often share a root cause on a failing host, but the checks are different and so is the fix.

### Can I fix this from inside a GitHub Actions job?

Essentially no. The daemon and its data root belong to the host, and a workflow step has neither the privileges nor the standing to remount a filesystem the kernel took read-only. The useful response on a hosted runner is to detect it, fail quickly and rerun on another machine.

### Why can I not find this error string in the Docker source?

Because it is not in it. The path is Docker, and everything else comes from Go: the shape from the `io/fs` error formatting and the closing words from the `syscall` errno table. The line only exists once a real path and a real errno have been slotted into that shape at run time.

## References

- [Go: PathError formatting, which produces the operation, path and errno shape](https://github.com/golang/go/blob/master/src/io/fs/fs.go)
- [Go: the syscall errno table where entry 30 is "read-only file system"](https://github.com/golang/go/blob/master/src/syscall/zerrors_linux_amd64.go)
- [abiosoft/colima#883: mkdir, open and unlinkat forms of this failure on one host](https://github.com/abiosoft/colima/issues/883)
- [Docker docs: dockerd and the data-root option](https://docs.docker.com/reference/cli/dockerd/)

---

Latchkey runs CI/CD that repairs its own failures. Agent entry points: https://latchkey.dev/agent.txt, https://latchkey.dev/openapi.json, https://latchkey.dev/llms.txt
