feat(fips): resilience — connectivity watcher with immediate anchor re-apply, rebindable peer listener, cached service probe, warm-path union

Phase A3 of docs/FIPS-UPTIME-AND-UI-STATE-PLAN.md (RC5), measured against
the A2 dial_stats baseline:

- 25s connectivity watcher in the fips supervisor: re-applies seed anchors
  immediately on an anchor-link drop, on startup-disconnected, AND on
  silent data-path death (connect_fails growing with zero fips_ok — the
  live .198 failure where the daemon reported "connected" while every
  dial blackholed and the 300s tick never healed it). Bounded to one
  re-apply per 60s.
- anchors::apply is now concurrent with a 15s per-connect cap — the old
  serial loop waited unbounded on each `sudo fipsctl connect`, so one
  hung subprocess stalled the whole periodic tick.
- rebindable peer listener: the accept loop returns after persistent
  accept errors (was: continue forever = inbound-dead until restart) and
  peer_late_bind_loop rebinds — also on fips0 ULA change.
- is_service_active gets a 10s TTL cache (was up to 2 systemctl spawns
  per dial attempt and per warm-tick peer).
- the warm tick now warms the union of federation peers + configured
  seed anchors (direct anchor links used to go cold between 300s ticks),
  skipping the redundant per-peer service check.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
archipelago
2026-07-28 03:57:12 -04:00
co-authored by Claude Fable 5
parent c83bade022
commit a3f07d5ac6
5 changed files with 167 additions and 42 deletions
+30 -19
View File
@@ -232,26 +232,32 @@ pub async fn remove(data_dir: &Path, npub: &str) -> Result<Vec<SeedAnchor>> {
/// leaving `anchor_connected=false` and every peer dial falling back to
/// a slow Tor timeout.
pub async fn apply(anchors: &[SeedAnchor]) -> Vec<ApplyResult> {
let mut results = Vec::with_capacity(anchors.len());
for anchor in anchors {
let out = Command::new("sudo")
.args([
"-n",
"fipsctl",
"connect",
&anchor.npub,
&anchor.address,
&anchor.transport,
])
.output()
.await;
// Concurrent, each connect hard-capped: the old serial loop waited
// unbounded on every `sudo fipsctl connect`, so one hung subprocess
// stalled the whole apply — and the periodic anchor tick behind it,
// which is exactly when a wedged daemon most needs the re-apply.
let futs = anchors.iter().cloned().map(|anchor| async move {
let out = tokio::time::timeout(
std::time::Duration::from_secs(15),
Command::new("sudo")
.args([
"-n",
"fipsctl",
"connect",
&anchor.npub,
&anchor.address,
&anchor.transport,
])
.output(),
)
.await;
let result = match out {
Ok(o) if o.status.success() => ApplyResult {
Ok(Ok(o)) if o.status.success() => ApplyResult {
npub: anchor.npub.clone(),
ok: true,
message: String::from_utf8_lossy(&o.stdout).trim().to_string(),
},
Ok(o) => ApplyResult {
Ok(Ok(o)) => ApplyResult {
npub: anchor.npub.clone(),
ok: false,
message: format!(
@@ -260,11 +266,16 @@ pub async fn apply(anchors: &[SeedAnchor]) -> Vec<ApplyResult> {
String::from_utf8_lossy(&o.stderr).trim()
),
},
Err(e) => ApplyResult {
Ok(Err(e)) => ApplyResult {
npub: anchor.npub.clone(),
ok: false,
message: format!("sudo fipsctl launch failed: {}", e),
},
Err(_) => ApplyResult {
npub: anchor.npub.clone(),
ok: false,
message: "sudo fipsctl connect timed out after 15s".to_string(),
},
};
if result.ok {
tracing::debug!(npub = %result.npub, "Seed anchor applied");
@@ -275,9 +286,9 @@ pub async fn apply(anchors: &[SeedAnchor]) -> Vec<ApplyResult> {
"Seed anchor apply failed (non-fatal)"
);
}
results.push(result);
}
results
result
});
futures_util::future::join_all(futs).await
}
/// Outcome of a single `fipsctl connect` call.