feat(lnd): channel-peer watchdog — a dropped peer link heals itself
Demo images / Build & push demo images (push) Successful in 3m49s
Demo images / Build & push demo images (push) Successful in 3m49s
LND normally reconnects channel peers after a restart, but not reliably: after long or repeated downtime (an app update, a node reboot, reconciler churn) the peer link can stay down for hours while BOTH endpoints keep the channel flagged disabled in the routing graph. The node looks perfectly healthy, the wallet shows balance, and every payment in either direction fails "no route to the recipient" — observed live on framework-pt (2026-09-01): its only channel sat disabled on both policy sides for ~17 hours after the LND 0.21.2 update, while shorty had 583k spendable and the user was told, by a mis-mapped modal, that they had 'no payment channel'. The channel graph is desired state — every open channel should have a live peer connection. A daemon-side watchdog now enforces it: - every 2 minutes, list channels + peers over LND REST - for each channel whose remote peer is not connected, look the peer's advertised addresses up in the public graph and dial one - per-peer retries throttled to 10 minutes so an unreachable peer is not hammered; 'already connected' counts as done; a peer with no advertised address is logged once per pass (cannot be dialed) - no-ops quietly on nodes without LND (missing macaroon) and while a wallet is locked (503 body has no channels) Unit tests pin the selection against the live REST shapes (remote_pubkey in /v1/channels vs pub_key in /v1/peers). v1.8.10 CHANGELOG + What's New entries staged so the next release run is clean first time.
This commit is contained in:
@@ -841,6 +841,37 @@ impl Server {
|
||||
});
|
||||
}
|
||||
|
||||
// LND channel-peer watchdog — every 2 minutes, reconnect the peers
|
||||
// of open channels that LND has not re-established on its own. LND's
|
||||
// reconnect logic gives up with a long backoff after repeated or
|
||||
// extended downtime (an app update, a reboot, reconciler churn), and
|
||||
// while the peer link is down BOTH endpoints keep the channel flagged
|
||||
// `disabled` in the routing graph — payments fail "no route" in both
|
||||
// directions while the node itself looks perfectly healthy. The
|
||||
// channel graph is desired state; this keeps it (framework-pt,
|
||||
// 2026-09-01: only channel unroutable ~17h after the 0.21.2 update).
|
||||
// No-ops quietly on nodes without LND. Per-peer retries are throttled
|
||||
// to 10 minutes so an unreachable peer is not hammered every pass.
|
||||
{
|
||||
tokio::spawn(async move {
|
||||
let mut interval = tokio::time::interval(Duration::from_secs(120));
|
||||
let mut last_attempt: HashMap<String, Instant> = HashMap::new();
|
||||
loop {
|
||||
interval.tick().await;
|
||||
match crate::container::lnd::reconnect_disconnected_channel_peers(
|
||||
&mut last_attempt,
|
||||
Duration::from_secs(600),
|
||||
)
|
||||
.await
|
||||
{
|
||||
Ok(0) => {}
|
||||
Ok(n) => info!(n, "LND channel-peer watchdog reconnected channel peers"),
|
||||
Err(e) => debug!("LND channel-peer watchdog (non-fatal): {}", e),
|
||||
}
|
||||
}
|
||||
});
|
||||
}
|
||||
|
||||
// FIPS seed-anchor apply loop — every 5 minutes we re-push the
|
||||
// configured seed anchors into the running fips daemon via
|
||||
// `fipsctl connect`. This keeps the mesh bootstrap resilient:
|
||||
|
||||
Reference in New Issue
Block a user