Skip to content
Latchkey LogoLatchkey home

Docker failed to create shim task runtime error in CI

Docker failed to create shim task runtime error is a relay, not a diagnosis: four or five programs each wrapped the failure below them before it reached your log, and every clause except the last one is a component saying that the component beneath it failed. The fix is always decided by the final clause, and the length of everything in front of it is what makes this error feel like an infrastructure fault when it usually is not.

Diagram of the wrapper chain from client to runc and which clause each layer adds
Each layer adds its own clause and passes the failure up. Only the last clause was produced by something that actually failed.

What this error means

A container refuses to start and the message runs for two or three lines, naming containerd, an OCI runtime and runc in turn. The shape is stable and the ending is not, which is the whole difficulty: the same prefix precedes a missing binary, a refused mount, a denied capability and a genuinely broken runtime, and the prefix is the part people search for. Jobs that use service containers, Docker-in-Docker or a Kubernetes-based runner see it most, because those stacks add still more wrappers in front of the same tail.

The line opencontainers/runc#5309 reports from a Kubernetes sandbox (quoted, not a recorded run)
failed to create containerd task: failed to create shim task: OCI runtime create failed: runc create failed: unable to start container process: error during container init: error mounting "sysfs" to rootfs at "/sys": mount src=sysfs, dst=/sys, dstFd=/proc/thread-self/fd/11, flags=MS_NOSUID|MS_NODEV|MS_NOEXEC: operation not permitted

Every clause has an author, and only one has news

You can attribute this line piece by piece, and doing so is the fastest way to stop reading it as one thing. failed to create shim task is containerd, from core/runtime/v2/task_manager.go, where the task manager wraps whatever the shim returned. Above it, when the caller is Docker rather than Kubernetes, the daemon adds failed to create task for container in moby daemon/internal/libcontainerd/remote/client.go, and the CLI puts Error response from daemon: in front of that. Below it, runc contributes unable to start container process from libcontainer/container_linux.go and error during container init from libcontainer/process_linux.go.

None of those five clauses is a cause. Each one is a component reporting that the layer beneath it failed, which is correct behavior and terrible search material. Because the chain is assembled by five programs at run time, the whole line is a literal in no source file, and searching for all of it finds other people pasting it rather than anything that explains it. Search the last clause.

Clause in the chainWhich program wrote itWhat it tells you
Error response from daemon:The Docker CLIThe daemon answered. Absent on Kubernetes paths
failed to create task for containerThe Docker daemon (moby)The daemon asked containerd and containerd said no
failed to create shim taskcontainerdThe shim was reached and the task creation failed
OCI runtime create failed: runc create failedcontainerd, then runcrunc was invoked and got as far as creating the container
error during container init: and what followsruncThe only part worth searching, and the only part with a fix

Common causes

A mount or capability the container was refused

The tail in the line above, and the common case in nested and sandboxed environments. runc tries to set up the container rootfs, a mount is denied, and the entire chain reports it. Nothing about the image is wrong; the environment would not let the runtime build the container it was asked for.

The OCI runtime is missing, broken or mismatched

A daemon with no usable runc, or a runc too old for the containerd it is paired with, fails every container start this way. In our experience this is the cause on custom runner images and almost never on hosted ones, because hosted images install the three components together.

The host could not spawn another process

Shim creation needs a new process, and a host at its process or memory limit cannot provide one. This is the genuinely transient member of the family, it correlates with heavy container churn in the same job, and it is the only cause where retrying is a real answer rather than a delay.

The image cannot run the command it was given

When the final clause names an executable rather than a mount or a permission, the chain is reporting your image and not your infrastructure. That version has its own page, because the diagnosis and the fix are both entirely different from everything above.

How to fix it

Cut the message at the last colon and search that

  1. Take everything after the final caused or init: clause and treat it as the error.
  2. Search that fragment, not the prefix, because the prefix belongs to five other programs.
  3. Fix at the layer the fragment names: the image, the host, or the runtime install.
Terminal
docker run --rm "$IMAGE" true 2>&1 | sed 's/.*: //'

Verify the runtime stack on any runner image you build

Print the daemon, containerd and runc versions in the job, and keep them installed together from the same source rather than pinned independently. A runtime triple that was never tested as a set is the most reliable way to produce this failure for every container at once.

Terminal
docker info --format '{{.ServerVersion}}'
containerd --version
runc --version
ls -l "$(command -v runc)"

Give nested runtimes the environment they need, explicitly

When the tail names a mount or a permission and you are running a daemon inside a container, the fix is in how that inner daemon was started rather than in the image it was asked to run. Set the privileges and namespaces deliberately, and scope them to the job that needs them.

Terminal
docker run --privileged --cgroupns=host \
  -v /sys/fs/cgroup:/sys/fs/cgroup:rw \
  docker:28-dind

Retry once, deliberately, and log that you did

A single scripted retry classifies the failure and costs one container start. Make it one retry rather than a loop, and print that the retry happened, so a log that ends green still records that the first attempt did not. A silent retry loop around a structural failure is how a five second error becomes a five minute one.

Terminal
for attempt in 1 2; do
  docker run --rm "$IMAGE" true && break
  echo "attempt $attempt failed"
  [ "$attempt" = 2 ] && exit 1
done

Transient and structural look identical in one log line

The advice that circulates about this error is to retry, and retrying is a reasonable first move exactly once. What makes it dangerous as a habit is that a structural failure and a transient one produce the same prefix, so a retry loop around a broken runtime turns a clear failure into a slow one and hides the log that would have explained it.

Decide from the tail rather than from the prefix. A tail naming a permission, a capability, a mount or a missing file is structural and will fail identically forever. A tail naming a resource, a timeout or a closed connection is worth one retry. If you cannot tell, run the same container twice in the same job and let the second attempt answer it, which is cheaper than reasoning about it.

.github/workflows/ci.yml
- name: Start it twice before blaming infrastructure
  run: |
    docker run --rm "$IMAGE" true && exit 0
    echo "first attempt failed, retrying once to classify"
    docker run --rm "$IMAGE" true

Check that a runtime is actually there

The one structural cause worth ruling out before anything else is a runtime that is missing or mismatched, because it fails every single container rather than some of them, and because it is answered in one command. A custom runner image that installs the Docker daemon without a working runc, or pins a runc too old for the containerd beside it, produces this chain for everything it is asked to start.

Print the versions of all three components once per job on any runner you build yourself. It costs a line and it converts the most confusing version of this failure into an obvious one.

Terminal
docker info --format '{{.ServerVersion}} {{.DefaultRuntime}}'
containerd --version
runc --version

What the runner does about it

No repair, and none is claimed: this slug has no heal-evidence record, so nothing here may say a Latchkey runner retries or fixes it. There is also no recorded run, and the reason is the page itself. The subject of this page is that the tail varies and the prefix does not. A recorded run would have to pick one tail, and whichever we picked would be shown next to the general chain as though it were the representative case, which is the exact mistake the page exists to prevent. One capture would make the page worse.

That also settles the retry question honestly. A runner cannot know from the prefix whether the failure is transient, and neither can you, so a page that shows one retried capture would be teaching a rule that only holds for the tail in the picture. Classify on the tail and retry deliberately.

How to prevent it

  • Install the daemon, containerd and runc as a set, and print their versions per job.
  • Keep service containers on the runner daemon rather than a nested one where you can.
  • Write log checks against the final clause, never against the wrapper prefix.
  • Cap retries at one, and make the retry visible in the log.

Frequently asked questions

Which component writes the failed to create shim task clause?
containerd does, in core/runtime/v2/task_manager.go, where the task manager wraps whatever its shim returned. It is containerd reporting the layer below it, so on its own the clause carries no information about the cause. The clause at the end of the chain is the one that does.
Should I just retry when I see this?
Once, and only to classify it. A missing mount, a denied capability or a broken runtime will produce the identical prefix on every attempt, so a retry loop hides a permanent failure behind a slow one. If the second attempt fails the same way, stop retrying and read the last clause.
Why does the message look different in Kubernetes than in Docker?
Because the outer wrappers belong to whoever called containerd. A Docker path adds Error response from daemon: and failed to create task for container; a CRI path adds failed to create containerd task and a sandbox identifier instead. The containerd and runc clauses in the middle are the same in both.
Is this the same error as the one about an executable not being found?
It is the same chain with a different ending, and the ending is what matters. When the final clause names an executable that is not in the image, the fix is in the image and that case has its own page. When it names a mount, a permission or a resource, the fix is in the host or the runtime and this is the page for it.

Related guides

References

Nested daemons are what you do when the runner is not yours. Latchkey runners are $0.0025/min at 2 vCPU. Start free → 30-day trial · No credit card