Skip to content
Latchkey LogoLatchkey home

Docker device or resource busy, in CI cleanup

Docker device or resource busy is the kernel refusing to unlink or unmount something that is still in use, and in CI it almost always arrives in a cleanup step rather than in the work the job was doing. Our recorded run reproduces it on a runner and records the part nobody mentions: the failed removal deleted the files under the mount before it failed.

Runner log: umount refused, then rm refused, with the file count dropping to zero
The recorded run: a busy unmount, a refused removal, and the file count either side of it. The mount survived the cleanup; the file under it did not.
Diagram of what holds a path busy and the order a cleanup step should use
Three things hold a path: a mount, an open file, and a running container. A cleanup step that removes before it unmounts finds out in the wrong order.

What this error means

A step that removes a directory, unmounts a path, or removes a container fails with a line ending in "device or resource busy", and the job is usually already finished with the real work. The same words reach you from three places: rm reports it as "cannot remove", umount as "target is busy", and the daemon as a refusal to remove a container filesystem, which is the shape moby/moby#31195 collects: "Unable to remove filesystem for <id>: remove <path>/shm: device or resource busy". Other runtimes spell the same errno unlinkat, as the kubelet does in kubernetes/kubernetes#131308. Our recorded run produced the first two on a runner by leaving a bind mount in place, which is what a leaked container mount looks like to a cleanup step. We made that mount with mount --bind rather than leaking one out of a container, because whether a container mount survives into the host depends on propagation settings that differ between runner images.

Actions log, cleanup step, Docker 29.7.2
umount: /home/runner/busy-demo/mnt: target is busy.
--- files under src before the cleanup step: 1
--- the cleanup step: rm -rf over the mount
rm: cannot remove '/home/runner/busy-demo/mnt': Device or resource busy
rm exited 1
--- files under src after it: 0

Reproduced on a Latchkey runner

Run 2026-09-20·Runner latchkey-small·Exit code 1

Docker version 29.7.2, build a7dcaa6
docker.io/library/alpine:3
--- control: a container writes through a bind mount, then exits
2
--- control: a directory with nothing mounted on it removes cleanly
rm -rf on an unmounted directory: ok
--- the leftover mount, made with mount --bind to model one a job left behind
/home/runner/busy-demo/mnt /dev/nvme0n1p1[/home/runner/busy-demo/src]
--- umount while a process has its working directory inside
umount: /home/runner/busy-demo/mnt: target is busy.
--- files under src before the cleanup step: 1
--- the cleanup step: rm -rf over the mount
rm: cannot remove '/home/runner/busy-demo/mnt': Device or resource busy
rm exited 1
--- files under src after it: 0
--- the same cleanup step again, this time as the step's own exit
rm: cannot remove '/home/runner/busy-demo/mnt': Device or resource busy
[latchkey-bash-wrapper] BEGIN sidecar POST (boot_wait=30s max_time=320s url=http://localhost/diagnose socket=/run/latchkey-self-heal/sock)
[latchkey-bash-wrapper] END sidecar POST ok (attempts=1 http=200)

The runner diagnosed the failure and did not retry it; this failure needs the fix below.

This is the kernel talking, not Docker

Device or resource busy is an errno, returned when something asks to remove or unmount a path that the kernel still has references to. Docker appears in the story only as the thing that created the mount, or as the process reporting the errno back to you. That is why the fix is never a Docker flag: you have to find the reference and release it.

Read your log against the table first, because the three wordings send you to three different places, and only one of them is about containers at all.

Who printed itWhat it looks likeWhat is holding the path
rm"cannot remove ...: Device or resource busy"The directory is a mountpoint
umounttarget is busy.A process has a file or a working directory inside
The daemon"Unable to remove filesystem for ..."A container mount the daemon has not released
A cleanup wrapperAny of the above, exit 1Whatever the step before it left behind

Common causes

The path is a mountpoint

The case our run records. A bind mount, a volume, or a tmpfs left in place makes the directory unremovable while everything inside it is reachable and deletable. A cleanup step that recurses into it does damage on the way to the error.

A process still has the file or directory open

An exec session that was never closed, a background process a test started, or a shell whose working directory is inside the path all keep it busy. Our run reproduces this half with a process whose working directory sits in the mount, which is enough for the unmount to be refused.

The daemon has not released a container mount

Reported over a path under the daemon storage directory, and moby/moby#31195 is the long-running report of it: removing a container takes at least two attempts, and the daemon answers that it is unable to remove the filesystem for that container. It is the one case where a second attempt sometimes does work, because the daemon may finish releasing the mount in between.

A previous job left state on a reused runner

On a self-hosted or long-lived runner the cleanup that did not happen yesterday is the failure today. In our experience this is the cause whenever the failing path has a name from a job that is not the one running.

How to fix it

Never recurse into a path that might be a mountpoint

  1. Check with mountpoint -q before any recursive removal in a cleanup step.
  2. Unmount first, then remove, so a failure stops before it deletes anything.
  3. Treat a refused unmount as a signal to look for the holder, not to force the removal.
Terminal
mountpoint -q "$DIR" && sudo umount "$DIR"
rm -rf "$DIR"

Ask what is holding it before forcing anything

Both questions have a one-line answer. List the mounts under the path you are cleaning, and list the processes with something open inside it. On a runner these run in milliseconds and they turn a mystery into a name.

Terminal
findmnt -R "$DIR" || echo "nothing mounted under $DIR"
fuser -vm "$DIR" 2>&1 || echo "no process holding $DIR"

Let containers clean themselves up

A container started with the remove-on-exit flag releases its mounts when it exits, which removes most of the leftovers that a later step trips over. Where a container has to outlive a step, stop it explicitly in a cleanup step that always runs, rather than relying on the job ending.

.github/workflows/ci.yml
docker run --rm -v "$PWD":/src app:ci make test
# and, for the ones that must persist:
docker rm -f app-under-test 2>/dev/null || true

Retry only the daemon case, and only once

The daemon refusing to remove a container filesystem is the one shape where a short wait genuinely helps, because it may still be releasing the mount. One retry after a pause is reasonable; a loop is a way of hiding a leaked mount that will fail the next job too.

Terminal
docker rm -f "$C" || { sleep 2; docker rm -f "$C"; }

A failed cleanup is not only a failed step

The recorded run prints the number of files under the mount source before and after the cleanup: one, then zero. The removal walked into the mount, deleted what it could reach through it, and only then hit the directory it could not remove. The step failed, and the data that was underneath is gone anyway.

This is the practical reason to fix the ordering rather than retry the step. On a job that mounts a cache directory or a workspace into a container, a cleanup step that starts with a recursive delete can empty the real directory while reporting a failure that reads like a permissions problem.

Terminal
# check before you delete
if mountpoint -q "$DIR"; then
  echo "$DIR is a mountpoint, unmount first"
  exit 1
fi
rm -rf "$DIR"

Find the holder, then release it in order

The order that works is the reverse of the order that created the state: stop the containers, unmount what is mounted, then remove the paths. Each of those steps has a way to ask what is in the way, and all three are cheap enough to run unconditionally in a cleanup step.

On a hosted runner the whole question is usually moot, because the machine is discarded after the job. It matters on a self-hosted or long-lived runner, where the next job inherits whatever this one left: our own recorded runs start from a runner with no leftovers, which is why the script had to create the leftover mount itself.

Terminal
docker ps -q | xargs -r docker rm -f
findmnt -R "$WORKSPACE" || true
fuser -vm "$DIR" 2>&1 || true
sudo umount "$DIR" && rm -rf "$DIR"

What the runner does about it

No repair, and none is claimed. On the recorded run the wrapper posted the failure to the sidecar and the sidecar answered, and the script's own second attempt at the removal failed exactly as the first had. Latchkey carries no pattern for this, and unmounting a path on a job's behalf is not something a runner should decide to do: the mount may be the point of the job.

The one thing the run adds is the file count on either side of the failure, which is why this page argues for ordering rather than for a retry.

How to prevent it

  • Order cleanup as stop containers, unmount, then remove paths.
  • Guard every recursive delete in a cleanup step with a mountpoint check.
  • Start throwaway containers with the remove-on-exit flag.
  • On long-lived runners, make cleanup a step that always runs rather than a best effort.

Frequently asked questions

What does device or resource busy mean in Docker?
It is the kernel refusing an unlink or an unmount because something still references the path. Docker is usually the thing that created the reference, not the thing refusing you. The fix is to find the reference, a mount or an open file or a running container, and release it before you remove anything.
How do I fix rm cannot remove device or resource busy?
Check whether the directory is a mountpoint and unmount it first. If the unmount is also refused, find the process holding it with a tool such as fuser, stop that process, then unmount and remove. Forcing the removal is not an option: the kernel will refuse it every time while the reference exists.
Why does docker rm fail with device or resource busy?
Because the daemon cannot remove the container filesystem while a mount under its storage directory is still referenced, which is what moby/moby#31195 reports as being unable to remove the filesystem for a container. Unlike the other cases on this page, one retry after a short pause sometimes succeeds, because the daemon may finish releasing the mount in the meantime.
Does rm -rf delete files through a mount?
Yes, and our recorded run measures it. The file count under the mount source was one before the cleanup step and zero after it, because the recursive delete removed everything it could reach through the mount before it reached the directory it could not remove. The step failed and the data was gone regardless.

Related guides

References

A fresh runner every job carries nothing forward. Latchkey runs them at $0.0025/min at 2 vCPU. Start free → 30-day trial · No credit card