Skip to content
Latchkey

GitHub Actions "The self-hosted runner lost connection"

The self-hosted runner stopped responding to GitHub during a job, so the service marked it disconnected and the job failed; the host itself is usually the problem, not your workflow.

What this error means

A running job fails when the self-hosted runner drops offline, and the Runners page shows the runner as Offline. The job log often ends mid-step with no command error.

github-actions
The self-hosted runner: my-runner-3 lost connection with the server.

Diagnose it: is the job queued, or is the runner gone?

A job that never starts and a job whose runner disappeared mid-run look similar in the UI and have opposite causes. The first is a labelling or capacity problem, the second is the runner being killed, usually by memory pressure or a spot reclaim.

.github/workflows/ci.yml
- name: Runner facts
  run: |
    echo "runner name: $RUNNER_NAME"
    echo "os/arch:     $RUNNER_OS/$RUNNER_ARCH"
    nproc; free -h; df -h /
    echo "labels this job asked for: ${{ toJSON(job) }}"

Common causes

Host network instability

Flaky egress, VPN drops, or DNS failures on the host severed the runner's persistent connection to GitHub.

Host resource exhaustion

The runner process was starved by an out-of-memory condition or full disk and could no longer maintain the connection.

How to fix it

Stabilize the host

  1. Check host network egress and DNS to github.com and the Actions endpoints.
  2. Confirm the host had free memory and disk during the job.
  3. Restart the runner service and verify it reconnects as Idle.
  4. Re-run the job once the runner is healthy.

Offload to auto-retrying managed runners

Latchkey managed runners auto-retry transient infrastructure failures and run on monitored, right-sized hosts, so a flaky self-hosted box does not fail your builds, at $0.0025/min for 2 vCPU against $0.006 GitHub-hosted.

The failures that are not your workflow

  • Exit 137 is the kernel out-of-memory killer, not an application error. Check free -h above against your peak usage.
  • Disk exhaustion presents as unrelated write errors deep in a build. GitHub-hosted runners ship roughly 14 GB of free space, which a Docker-heavy job can exhaust.
  • A lost connection to the server on a self-hosted runner is usually the host being reclaimed or rebooted, not a network fault in your job.
  • A job that starts and immediately fails with no step output normally failed during runner setup, before your workflow ran at all.

How to prevent it

  • Monitor self-hosted host network and resource health.
  • Keep runners off memory and disk limits.
  • Alert when a runner flips Offline unexpectedly.

Frequently asked questions

What causes GitHub Actions "The self-hosted runner lost connection"?
There are 2 common causes: host network instability and host resource exhaustion. Flaky egress, VPN drops, or DNS failures on the host severed the runner's persistent connection to GitHub.
How do I fix GitHub Actions "The self-hosted runner lost connection"?
There are 2 fixes depending on which cause you have: stabilize the host and offload to auto-retrying managed runners. Work through them in order, since the first is the most common.
What does GitHub Actions "The self-hosted runner lost connection" actually mean?
A running job fails when the self-hosted runner drops offline, and the Runners page shows the runner as Offline.
How do I stop GitHub Actions "The self-hosted runner lost connection" happening again?
Monitor self-hosted host network and resource health. The prevention section lists 3 changes that keep it from recurring.

Related guides

References

This is a transient network failure, not a bug in your code. Latchkey detects, repairs, and retries it for you. Start free → 30-day trial · No credit card