# Self-Healing CI: Recovering Metaspace / Native-Thread OOM

> JVM metaspace and "unable to create native thread" OOMs are resource ceilings, not heap leaks. See the manual fix and how self-healing CI retries safely.

Source: https://latchkey.dev/learn/self-healing-ci/self-heal-metaspace-thread-oom  
Updated: 2026-06-25

Not every JVM OOM is a heap problem - metaspace and native-thread OOMs come from other ceilings, and the fix is more headroom, not a heap dump.

## What makes a failure safely retryable

Automatic retry is only correct for failures that are genuinely transient. Retrying a deterministic failure wastes minutes and hides a real defect, so the classification matters more than the retry mechanism.

- Safe to retry: network timeouts, registry 5xx, transient DNS failures, a service container that was not ready, a spot instance reclaimed mid-run.
- Not safe to retry: assertion failures, compile errors, lint violations, anything that fails identically on every attempt.
- Ambiguous, and worth investigating rather than retrying: out-of-memory kills, disk exhaustion, and flaky tests. These repeat under load and a retry only hides the trend.
- Always record that a retry happened. A pipeline that silently retries is a pipeline whose real failure rate you do not know.

## FAQ

### I already raised `-Xmx` and still get an OOM - why?

`-Xmx` only sizes the heap. A Metaspace or "unable to create native thread" OOM comes from a different region, so a bigger heap does nothing - it can even make native-thread OOMs worse by leaving less native memory. The fix is to size the region that actually ran out, or add overall headroom.

---

Latchkey runs CI/CD that repairs its own failures. Agent entry points: https://latchkey.dev/agent.txt, https://latchkey.dev/openapi.json, https://latchkey.dev/llms.txt
