dockerd cold-boot pile-up: overlapping image-usage walks saturate the daemon #182
Labels
No labels
audit-work
bug
docs
general-admin
major-upgrade
needs-vps-sync
new-service
on-hold
outside-work
post-podman
renovate
upstream
vps
No milestone
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
mysticalsoap/docker#182
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Second, separate bug found 2026-08-10 while chasing the goroutine leak (#181) — don't conflate them.
After a cold boot, dockerd climbed to ~31k goroutines within 20min and the whole fleet went unhealthy for ~50min. Captured dump showed zero leak signature — instead 14k+ goroutines in
containerd/core/images.Walkdoing per-image disk-usage calculation. Root cause from source:ImageService.Images()bounds each call internally (NumCPU*2workers) but nothing guards against overlapping top-level calls — when one call takes longer than the interval between callers (cold caches + 108GB of images at the time), capped batches stack ~300 deep. cAdvisor's 30s housekeeping is the suspected repeat caller, unconfirmed (logs were empty). During the pile-up evendocker imagestimes out; lightweight socket endpoints (/info,/_ping) stay fast — diagnose with those, don't add heavy CLI calls to the backlog.Mitigated since 2026-08-10 by the daily unused-image prune (dotfiles
docker-image-prune.timer,-a -f --filter until=24h): 213→84 images, 108→37GB, and no recurrence on subsequent reboots. Remaining work:ImageService.Images()is arguably a moby bug (no request coalescing/backpressure on an expensive endpoint)