fix(container): reap ghost containers so an app can't be locked out of itself
Demo images / Build & push demo images (push) Successful in 3m24s
Demo images / Build & push demo images (push) Successful in 3m24s
A ghost is a container whose process tree is still running while podman
has no record of it: the exit-command's `cleanup --rm` deletes the record,
conmon and the payload survive. It keeps owning exactly what the app needs
— the published host port and the file locks in its data dir — so the
replacement container either fails to bind ("address already in use") or
starts and dies on the lock, and Restart=always loops it there forever.
Nothing in the stack could see it: every podman-level stop/rm/recreate
misses a container podman lost.
Seen twice now: 752 restarts on a fleet node (2026-08-10) and again on the
dev box today, where Gitea flapped until it fell out of My Apps. Both were
cleared by hand; container-doctor.sh has the same logic but is an
out-of-band script the daemon never calls.
- New container::ghost_reaper: finds conmon processes whose 64-hex
container id is absent from `podman ps -a --no-trunc -q`, then kills the
payload's children and conmon (TERM, 5s grace, then KILL — the Gitea
ghost ignored TERM). Id-based, never name-based: killing by name would
hit the live managed container. A failed `podman ps` reaps nothing
rather than treating every container as a ghost.
- Hooked at repair_before_package_start (covers package.start,
package.restart and the orchestrator start path) and in the boot
reconciler's 30s tick, so ghosts are cleared before an app is asked to
start and swept for every app continuously.
Restart feedback: the lifecycle RPCs return {"status":"restarting"} in
milliseconds and work in the background, so "Restarting..." flashed for a
few frames and the buttons went idle while the app was still down — the
click read as a no-op. The hero buttons now show a spinner and hold it off
the node's own state (starting/stopping/restarting/updating, plus running
+ health=starting), and the just-clicked action is held until the backend
confirms it picked the work up, with a 12s cap so an unresponsive node
still releases the controls.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Fable 5
parent
b113fafee4
commit
9ccc325a4d
@@ -221,6 +221,12 @@ impl BootReconciler {
|
||||
}
|
||||
|
||||
async fn tick(&self) {
|
||||
// Sweep ghost containers first: a process tree podman has forgotten
|
||||
// still holds its app's ports and data locks, so reconcile would keep
|
||||
// restarting that app into the same wall (752 restarts on a fleet
|
||||
// node, 2026-08-10). Nothing else in the stack can see them —
|
||||
// every podman-level stop/rm misses a container podman lost.
|
||||
crate::container::ghost_reaper::reap_all().await;
|
||||
let report = self.orchestrator.reconcile_existing().await;
|
||||
Self::log_report(&report);
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user