Docker failed to create shim task runtime error in CI
Docker failed to create shim task runtime error is a relay, not a diagnosis: four or five programs each wrapped the failure below them before it reached your log, and every clause except the last one is a component saying that the component beneath it failed. The fix is always decided by the final clause, and the length of everything in front of it is what makes this error feel like an infrastructure fault when it usually is not.

What this error means
A container refuses to start and the message runs for two or three lines, naming containerd, an OCI runtime and runc in turn. The shape is stable and the ending is not, which is the whole difficulty: the same prefix precedes a missing binary, a refused mount, a denied capability and a genuinely broken runtime, and the prefix is the part people search for. Jobs that use service containers, Docker-in-Docker or a Kubernetes-based runner see it most, because those stacks add still more wrappers in front of the same tail.
failed to create containerd task: failed to create shim task: OCI runtime create failed: runc create failed: unable to start container process: error during container init: error mounting "sysfs" to rootfs at "/sys": mount src=sysfs, dst=/sys, dstFd=/proc/thread-self/fd/11, flags=MS_NOSUID|MS_NODEV|MS_NOEXEC: operation not permittedEvery clause has an author, and only one has news
You can attribute this line piece by piece, and doing so is the fastest way to stop reading it as one thing. failed to create shim task is containerd, from core/runtime/v2/task_manager.go, where the task manager wraps whatever the shim returned. Above it, when the caller is Docker rather than Kubernetes, the daemon adds failed to create task for container in moby daemon/internal/libcontainerd/remote/client.go, and the CLI puts Error response from daemon: in front of that. Below it, runc contributes unable to start container process from libcontainer/container_linux.go and error during container init from libcontainer/process_linux.go.
None of those five clauses is a cause. Each one is a component reporting that the layer beneath it failed, which is correct behavior and terrible search material. Because the chain is assembled by five programs at run time, the whole line is a literal in no source file, and searching for all of it finds other people pasting it rather than anything that explains it. Search the last clause.
| Clause in the chain | Which program wrote it | What it tells you |
|---|---|---|
Error response from daemon: | The Docker CLI | The daemon answered. Absent on Kubernetes paths |
failed to create task for container | The Docker daemon (moby) | The daemon asked containerd and containerd said no |
failed to create shim task | containerd | The shim was reached and the task creation failed |
OCI runtime create failed: runc create failed | containerd, then runc | runc was invoked and got as far as creating the container |
error during container init: and what follows | runc | The only part worth searching, and the only part with a fix |
Common causes
A mount or capability the container was refused
The tail in the line above, and the common case in nested and sandboxed environments. runc tries to set up the container rootfs, a mount is denied, and the entire chain reports it. Nothing about the image is wrong; the environment would not let the runtime build the container it was asked for.
The OCI runtime is missing, broken or mismatched
A daemon with no usable runc, or a runc too old for the containerd it is paired with, fails every container start this way. In our experience this is the cause on custom runner images and almost never on hosted ones, because hosted images install the three components together.
The host could not spawn another process
Shim creation needs a new process, and a host at its process or memory limit cannot provide one. This is the genuinely transient member of the family, it correlates with heavy container churn in the same job, and it is the only cause where retrying is a real answer rather than a delay.
The image cannot run the command it was given
When the final clause names an executable rather than a mount or a permission, the chain is reporting your image and not your infrastructure. That version has its own page, because the diagnosis and the fix are both entirely different from everything above.
How to fix it
Cut the message at the last colon and search that
- Take everything after the final
causedorinit:clause and treat it as the error. - Search that fragment, not the prefix, because the prefix belongs to five other programs.
- Fix at the layer the fragment names: the image, the host, or the runtime install.
docker run --rm "$IMAGE" true 2>&1 | sed 's/.*: //'Verify the runtime stack on any runner image you build
Print the daemon, containerd and runc versions in the job, and keep them installed together from the same source rather than pinned independently. A runtime triple that was never tested as a set is the most reliable way to produce this failure for every container at once.
docker info --format '{{.ServerVersion}}'
containerd --version
runc --version
ls -l "$(command -v runc)"Give nested runtimes the environment they need, explicitly
When the tail names a mount or a permission and you are running a daemon inside a container, the fix is in how that inner daemon was started rather than in the image it was asked to run. Set the privileges and namespaces deliberately, and scope them to the job that needs them.
docker run --privileged --cgroupns=host \
-v /sys/fs/cgroup:/sys/fs/cgroup:rw \
docker:28-dindRetry once, deliberately, and log that you did
A single scripted retry classifies the failure and costs one container start. Make it one retry rather than a loop, and print that the retry happened, so a log that ends green still records that the first attempt did not. A silent retry loop around a structural failure is how a five second error becomes a five minute one.
for attempt in 1 2; do
docker run --rm "$IMAGE" true && break
echo "attempt $attempt failed"
[ "$attempt" = 2 ] && exit 1
doneTransient and structural look identical in one log line
The advice that circulates about this error is to retry, and retrying is a reasonable first move exactly once. What makes it dangerous as a habit is that a structural failure and a transient one produce the same prefix, so a retry loop around a broken runtime turns a clear failure into a slow one and hides the log that would have explained it.
Decide from the tail rather than from the prefix. A tail naming a permission, a capability, a mount or a missing file is structural and will fail identically forever. A tail naming a resource, a timeout or a closed connection is worth one retry. If you cannot tell, run the same container twice in the same job and let the second attempt answer it, which is cheaper than reasoning about it.
- name: Start it twice before blaming infrastructure
run: |
docker run --rm "$IMAGE" true && exit 0
echo "first attempt failed, retrying once to classify"
docker run --rm "$IMAGE" trueCheck that a runtime is actually there
The one structural cause worth ruling out before anything else is a runtime that is missing or mismatched, because it fails every single container rather than some of them, and because it is answered in one command. A custom runner image that installs the Docker daemon without a working runc, or pins a runc too old for the containerd beside it, produces this chain for everything it is asked to start.
Print the versions of all three components once per job on any runner you build yourself. It costs a line and it converts the most confusing version of this failure into an obvious one.
docker info --format '{{.ServerVersion}} {{.DefaultRuntime}}'
containerd --version
runc --versionWhat the runner does about it
No repair, and none is claimed: this slug has no heal-evidence record, so nothing here may say a Latchkey runner retries or fixes it. There is also no recorded run, and the reason is the page itself. The subject of this page is that the tail varies and the prefix does not. A recorded run would have to pick one tail, and whichever we picked would be shown next to the general chain as though it were the representative case, which is the exact mistake the page exists to prevent. One capture would make the page worse.
That also settles the retry question honestly. A runner cannot know from the prefix whether the failure is transient, and neither can you, so a page that shows one retried capture would be teaching a rule that only holds for the tail in the picture. Classify on the tail and retry deliberately.
How to prevent it
- Install the daemon, containerd and runc as a set, and print their versions per job.
- Keep service containers on the runner daemon rather than a nested one where you can.
- Write log checks against the final clause, never against the wrapper prefix.
- Cap retries at one, and make the retry visible in the log.
Frequently asked questions
Which component writes the failed to create shim task clause?
core/runtime/v2/task_manager.go, where the task manager wraps whatever its shim returned. It is containerd reporting the layer below it, so on its own the clause carries no information about the cause. The clause at the end of the chain is the one that does.Should I just retry when I see this?
Why does the message look different in Kubernetes than in Docker?
Error response from daemon: and failed to create task for container; a CRI path adds failed to create containerd task and a sandbox identifier instead. The containerd and runc clauses in the middle are the same in both.Is this the same error as the one about an executable not being found?
Related guides
References
- containerd: the "failed to create shim task" wrap, core/runtime/v2/task_manager.go
- moby: the daemon wrap that precedes the containerd clause, in libcontainerd
- opencontainers/runc#5309: the full chain quoted above, ending in a refused mount
- nestybox/sysbox#1025: the same chain with the older runc wording in the middle
- Docker documentation
- Docker build cache
- GitHub Actions documentation