fix(fmcd): cap CPU + watchdog-restart the iroh relay hot-loop

On NAT'd nodes that can reach the iroh federation neither directly nor
via iroh's public relays, fmcd's embedded iroh networking enters a
relay/hole-punch reconnect hot-loop that pegs its entire CPU allotment
indefinitely (observed ~1 core sustained for 4 days on a Tailscale node,
while LAN nodes that reach the guardian directly stay <3%). fmcd 0.8.0
exposes no iroh/relay knobs, so:

- fmcd-run now samples fmcd's own CPU and restarts it when it stays near
  its allotment for ~15 min (a restart demonstrably clears the stuck iroh
  state; real work is bursty and never flat-pegs a core for minutes).
- Lower cpu_limit 1 -> 0.25 core so a stuck instance can't starve the
  node (steady-state is <3% of a core; joins are brief).

Ships as fmcd:0.8.1 (launcher-only rebuild, same fmcd binary). Bumped the
image pin + cpu_limit in the manifest, image-versions.sh, the embedded
catalog manifest (releases/app-catalog.json), and the UI catalogs.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
archipelago
2026-06-28 12:19:27 -04:00
co-authored by Claude Opus 4.8
parent 4519dbf04f
commit 6734947c3e
6 changed files with 79 additions and 11 deletions
+3 -3
View File
@@ -1,6 +1,6 @@
{
"schema": 1,
"updated": "2026-06-24",
"updated": "2026-06-28",
"apps": {
"adguardhome": {
"version": "v0.107.55",
@@ -1219,7 +1219,7 @@
"version": "0.8.0",
"description": "Fedimint ecash client daemon (fmcd). Lets the node hold Fedimint ecash and join federations; the wallet talks to it over a local REST API.",
"container": {
"image": "146.59.87.168:3000/lfg2025/fmcd:0.8.0",
"image": "146.59.87.168:3000/lfg2025/fmcd:0.8.1",
"pull_policy": "if-not-present",
"network": "archy-net",
"generated_secrets": [
@@ -1242,7 +1242,7 @@
}
],
"resources": {
"cpu_limit": 1,
"cpu_limit": 0.25,
"memory_limit": "1Gi",
"disk_limit": "2Gi"
},