dockerd goroutine leak: repro watcher + upstream contribution plan #181

Open
opened 2026-08-27 11:56:20 -04:00 by mysticalsoap · 0 comments
Owner

The 2026-08-06 leak (64k goroutines / 4.3GB RSS at ~51h uptime, fleet-wide healthcheck failure) is instrumented and waiting on a recurrence. Tracking the watcher and the contribution plan here; mitigations are #117.

Watcher (lives in dotfiles, not this repo): bin/docker-goroutine-watch.py + docker-goroutine-watch.{service,timer} user units. Polls dockerd NGoroutines every 5min; above 5k it captures a full pprof goroutine dump + docker system df-style info + last 15min of docker journal at checkpoints (5k/10k/20k/35k/50k/65k/80k), re-arming when the count returns below 5k. dockerd's /debug/pprof/* is always mounted on the socket — no debug mode needed.

Contribution plan, in order, once a repro dump lands:

  1. Use the dump to diagnose the suspected root cause for real: golang/go#69316 (x/net/http2/hpack eviction race, open since 2020, never reproduced by maintainers). Our fleet may be the best-instrumented reproduction environment anyone has pointed at it.
  2. Contribute that fix upstream, and alongside it revive moby/moby#52450 (bounded waits in daemon/health.go — reviewed positively, stalled, fork deleted; diff still fetchable from the PR). Decision 2026-08-15: both go together, grounded in the lived incident — no speculative fixes before a repro.

Confound: the leak happened under docker 29.7.1; the daemon has run 29.7.2 since 2026-08-10 (pacman upgraded 08-09, binary picked up at next reboot) and has since been clean through windows of 9+ days — over 4x the original time-to-failure, with more containers. This may never recur. If months pass clean, close this as fixed-upstream-somewhere-in-29.7.2.

The 2026-08-06 leak (64k goroutines / 4.3GB RSS at ~51h uptime, fleet-wide healthcheck failure) is instrumented and waiting on a recurrence. Tracking the watcher and the contribution plan here; mitigations are #117. **Watcher** (lives in dotfiles, not this repo): `bin/docker-goroutine-watch.py` + `docker-goroutine-watch.{service,timer}` user units. Polls dockerd `NGoroutines` every 5min; above 5k it captures a full pprof goroutine dump + `docker system df`-style info + last 15min of docker journal at checkpoints (5k/10k/20k/35k/50k/65k/80k), re-arming when the count returns below 5k. dockerd's `/debug/pprof/*` is always mounted on the socket — no debug mode needed. **Contribution plan, in order, once a repro dump lands:** 1. Use the dump to diagnose the suspected root cause for real: golang/go#69316 (`x/net/http2/hpack` eviction race, open since 2020, never reproduced by maintainers). Our fleet may be the best-instrumented reproduction environment anyone has pointed at it. 2. Contribute that fix upstream, and alongside it revive moby/moby#52450 (bounded waits in `daemon/health.go` — reviewed positively, stalled, fork deleted; diff still fetchable from the PR). Decision 2026-08-15: both go together, grounded in the lived incident — no speculative fixes before a repro. **Confound**: the leak happened under docker 29.7.1; the daemon has run 29.7.2 since 2026-08-10 (pacman upgraded 08-09, binary picked up at next reboot) and has since been clean through windows of 9+ days — over 4x the original time-to-failure, with more containers. This may never recur. If months pass clean, close this as fixed-upstream-somewhere-in-29.7.2.
Sign in to join this conversation.
No milestone
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
mysticalsoap/docker#181
No description provided.