Docker device or resource busy, in CI cleanup
Docker device or resource busy is the kernel refusing to unlink or unmount something that is still in use, and in CI it almost always arrives in a cleanup step rather than in the work the job was doing. Our recorded run reproduces it on a runner and records the part nobody mentions: the failed removal deleted the files under the mount before it failed.


What this error means
A step that removes a directory, unmounts a path, or removes a container fails with a line ending in "device or resource busy", and the job is usually already finished with the real work. The same words reach you from three places: rm reports it as "cannot remove", umount as "target is busy", and the daemon as a refusal to remove a container filesystem, which is the shape moby/moby#31195 collects: "Unable to remove filesystem for <id>: remove <path>/shm: device or resource busy". Other runtimes spell the same errno unlinkat, as the kubelet does in kubernetes/kubernetes#131308. Our recorded run produced the first two on a runner by leaving a bind mount in place, which is what a leaked container mount looks like to a cleanup step. We made that mount with mount --bind rather than leaking one out of a container, because whether a container mount survives into the host depends on propagation settings that differ between runner images.
umount: /home/runner/busy-demo/mnt: target is busy.
--- files under src before the cleanup step: 1
--- the cleanup step: rm -rf over the mount
rm: cannot remove '/home/runner/busy-demo/mnt': Device or resource busy
rm exited 1
--- files under src after it: 0Reproduced on a Latchkey runner
Docker version 29.7.2, build a7dcaa6
docker.io/library/alpine:3
--- control: a container writes through a bind mount, then exits
2
--- control: a directory with nothing mounted on it removes cleanly
rm -rf on an unmounted directory: ok
--- the leftover mount, made with mount --bind to model one a job left behind
/home/runner/busy-demo/mnt /dev/nvme0n1p1[/home/runner/busy-demo/src]
--- umount while a process has its working directory inside
umount: /home/runner/busy-demo/mnt: target is busy.
--- files under src before the cleanup step: 1
--- the cleanup step: rm -rf over the mount
rm: cannot remove '/home/runner/busy-demo/mnt': Device or resource busy
rm exited 1
--- files under src after it: 0
--- the same cleanup step again, this time as the step's own exit
rm: cannot remove '/home/runner/busy-demo/mnt': Device or resource busy
[latchkey-bash-wrapper] BEGIN sidecar POST (boot_wait=30s max_time=320s url=http://localhost/diagnose socket=/run/latchkey-self-heal/sock)
[latchkey-bash-wrapper] END sidecar POST ok (attempts=1 http=200)The runner diagnosed the failure and did not retry it; this failure needs the fix below.
This is the kernel talking, not Docker
Device or resource busy is an errno, returned when something asks to remove or unmount a path that the kernel still has references to. Docker appears in the story only as the thing that created the mount, or as the process reporting the errno back to you. That is why the fix is never a Docker flag: you have to find the reference and release it.
Read your log against the table first, because the three wordings send you to three different places, and only one of them is about containers at all.
| Who printed it | What it looks like | What is holding the path |
|---|---|---|
rm | "cannot remove ...: Device or resource busy" | The directory is a mountpoint |
umount | target is busy. | A process has a file or a working directory inside |
| The daemon | "Unable to remove filesystem for ..." | A container mount the daemon has not released |
| A cleanup wrapper | Any of the above, exit 1 | Whatever the step before it left behind |
Common causes
The path is a mountpoint
The case our run records. A bind mount, a volume, or a tmpfs left in place makes the directory unremovable while everything inside it is reachable and deletable. A cleanup step that recurses into it does damage on the way to the error.
A process still has the file or directory open
An exec session that was never closed, a background process a test started, or a shell whose working directory is inside the path all keep it busy. Our run reproduces this half with a process whose working directory sits in the mount, which is enough for the unmount to be refused.
The daemon has not released a container mount
Reported over a path under the daemon storage directory, and moby/moby#31195 is the long-running report of it: removing a container takes at least two attempts, and the daemon answers that it is unable to remove the filesystem for that container. It is the one case where a second attempt sometimes does work, because the daemon may finish releasing the mount in between.
A previous job left state on a reused runner
On a self-hosted or long-lived runner the cleanup that did not happen yesterday is the failure today. In our experience this is the cause whenever the failing path has a name from a job that is not the one running.
How to fix it
Never recurse into a path that might be a mountpoint
- Check with
mountpoint -qbefore any recursive removal in a cleanup step. - Unmount first, then remove, so a failure stops before it deletes anything.
- Treat a refused unmount as a signal to look for the holder, not to force the removal.
mountpoint -q "$DIR" && sudo umount "$DIR"
rm -rf "$DIR"Ask what is holding it before forcing anything
Both questions have a one-line answer. List the mounts under the path you are cleaning, and list the processes with something open inside it. On a runner these run in milliseconds and they turn a mystery into a name.
findmnt -R "$DIR" || echo "nothing mounted under $DIR"
fuser -vm "$DIR" 2>&1 || echo "no process holding $DIR"Let containers clean themselves up
A container started with the remove-on-exit flag releases its mounts when it exits, which removes most of the leftovers that a later step trips over. Where a container has to outlive a step, stop it explicitly in a cleanup step that always runs, rather than relying on the job ending.
docker run --rm -v "$PWD":/src app:ci make test
# and, for the ones that must persist:
docker rm -f app-under-test 2>/dev/null || trueRetry only the daemon case, and only once
The daemon refusing to remove a container filesystem is the one shape where a short wait genuinely helps, because it may still be releasing the mount. One retry after a pause is reasonable; a loop is a way of hiding a leaked mount that will fail the next job too.
docker rm -f "$C" || { sleep 2; docker rm -f "$C"; }A failed cleanup is not only a failed step
The recorded run prints the number of files under the mount source before and after the cleanup: one, then zero. The removal walked into the mount, deleted what it could reach through it, and only then hit the directory it could not remove. The step failed, and the data that was underneath is gone anyway.
This is the practical reason to fix the ordering rather than retry the step. On a job that mounts a cache directory or a workspace into a container, a cleanup step that starts with a recursive delete can empty the real directory while reporting a failure that reads like a permissions problem.
# check before you delete
if mountpoint -q "$DIR"; then
echo "$DIR is a mountpoint, unmount first"
exit 1
fi
rm -rf "$DIR"Find the holder, then release it in order
The order that works is the reverse of the order that created the state: stop the containers, unmount what is mounted, then remove the paths. Each of those steps has a way to ask what is in the way, and all three are cheap enough to run unconditionally in a cleanup step.
On a hosted runner the whole question is usually moot, because the machine is discarded after the job. It matters on a self-hosted or long-lived runner, where the next job inherits whatever this one left: our own recorded runs start from a runner with no leftovers, which is why the script had to create the leftover mount itself.
docker ps -q | xargs -r docker rm -f
findmnt -R "$WORKSPACE" || true
fuser -vm "$DIR" 2>&1 || true
sudo umount "$DIR" && rm -rf "$DIR"What the runner does about it
No repair, and none is claimed. On the recorded run the wrapper posted the failure to the sidecar and the sidecar answered, and the script's own second attempt at the removal failed exactly as the first had. Latchkey carries no pattern for this, and unmounting a path on a job's behalf is not something a runner should decide to do: the mount may be the point of the job.
The one thing the run adds is the file count on either side of the failure, which is why this page argues for ordering rather than for a retry.
How to prevent it
- Order cleanup as stop containers, unmount, then remove paths.
- Guard every recursive delete in a cleanup step with a mountpoint check.
- Start throwaway containers with the remove-on-exit flag.
- On long-lived runners, make cleanup a step that always runs rather than a best effort.
Frequently asked questions
What does device or resource busy mean in Docker?
How do I fix rm cannot remove device or resource busy?
Why does docker rm fail with device or resource busy?
Does rm -rf delete files through a mount?
Related guides
References
- Docker docs: storage drivers and the daemon storage directory
- Docker docs: bind mounts
- moby/moby#31195: device or resource busy when removing containers
- kubernetes/kubernetes#131308: the unlinkat spelling of the same errno, from a kubelet
- Docker documentation
- Docker build cache
- GitHub Actions documentation