fix(tests): bound the lnd getinfo probe + track the backend-recreate cascade gap

lnd.bats retry loop had no per-exec timeout — a wedged lnd RPC (chain-blind
after bitcoin-knots was recreated under it) hung one podman exec, and the
whole gate, indefinitely (.228, 30+ min at test 75). timeout 10 per attempt
keeps the intended ~4min bound. Tracker: new §C item for the root cause —
apps cache their backend container's IP; a backend recreate must cascade a
dependent restart/alert or lnd goes chain-blind silently.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
archipelago
2026-07-08 19:48:30 -04:00
co-authored by Claude Fable 5
parent e706d446e7
commit 81743874a1
2 changed files with 11 additions and 1 deletions
+4 -1
View File
@@ -53,8 +53,11 @@ teardown_file() {
# lnd's RPC readiness LAGS the container "running" state: after a (re)start the
# wallet must auto-unlock before lncli answers, so a single-shot getinfo races
# that window and false-fails. Retry until ready (~90s), like a health probe.
# `timeout 10` per attempt: a wedged lnd RPC (e.g. chain-blind after its
# bitcoin backend was recreated under it, .228 2026-07-08) otherwise hangs
# a single exec — and with it the whole suite — indefinitely.
run sh -lc 'for i in $(seq 1 80); do
podman exec lnd lncli \
timeout 10 podman exec lnd lncli \
--tlscertpath /root/.lnd/tls.cert \
--macaroonpath /root/.lnd/data/chain/bitcoin/mainnet/readonly.macaroon \
--rpcserver localhost:10009 getinfo >/dev/null 2>&1 && exit 0