# Incident + follow-up tracker — 2026-09-01 (post-HTTPS-work, post-LND-0.21.2 breakage) Live incident spanning framework-pt and shorty-s after the HTTPS/launcher work and the LND 0.18.4→0.21.2 pin bump. Root causes found on real nodes; status updated as work lands. Each fix ships with a regression test so the same class cannot silently return. ## A. Root causes (all verified live) | # | Symptom | Root cause | |---|---------|-----------| | A1 | LND sends fail "Payment failed: Not Found" | LND 0.21 **removed** the deprecated `/v1/channels/transactions` REST route; backend still called it. Receive was fine; the "Failed to fetch" on framework-pt was A3 masking it. | | A2 | Shorty NPM restart-loops (counter 3176) | Manifest conversion (fc68c5b6) dropped (a) the `/etc/letsencrypt` mount NPM's s6 boot demands, and (b) `NET_BIND_SERVICE` — its internal nginx binds 80/443/81 and the orchestrator runs `--cap-drop=ALL`. | | A3 | framework-pt: every `/rpc/v1` fetch CORS-blocked, "Failed to fetch", dashboard "not responding", mempool/indeehub frames broken | nginx sent `Strict-Transport-Security: max-age=31536000; includeSubDomains` on **HTTPS**; browsers cached it, then silently upgraded the still-open **http** dashboard's fetches/frames to https → scheme change = cross-origin → CORS block. HTTP is a supported mode on purpose (self-signed cert, /ca.crt flow). | | A4 | Mempool/IndeeHub/bitcoin-UI frames stay `http://` on HTTPS pages (mixed content, "does not connect") | `portAuth()` looked the launch port up under the launch alias (`mempool-web`, `lnd`, `bitcoin-knots`…); the signed catalog declares those ports under the manifest id that owns them (`archy-mempool-web`, `lnd-ui`, `bitcoin-ui`) → miss → launcher fell back to http. Cache also only warmed in Store/Discover views. | | A5 | IndeeHub nostr sign-in dead over HTTPS | NIP-07 bridge compared `event.origin` for strict equality with the stored (http) app URL and replied to the **stored** URL as postMessage targetOrigin — both break when the frame was scheme-upgraded. | | A6 | Portainer "disappeared" after restart/update, then demands a setup token "see server logs" | Update to 2.45.0 recreated the container; on a fresh DB Portainer ≥2.21 mints a one-time setup token printed ONLY in container logs — hostile appliance UX. The "disappearance" was the recreate + this unknown-token first screen. | ## B. Fixes (code) | Fix | Files | Status | |-----|-------|--------| | B1 LND pay via `Router.SendPaymentV2` (`/v2/router/send`), pending-status + actionable failure reasons preserved | `core/archipelago/src/api/rpc/lnd/payments.rs` (+ unit tests) | ✅ code | | B2 Portainer setup token surfaced in the existing credentials interstitial (`package.credentials` → AppSidebar card with copy) | `core/archipelago/src/api/rpc/package/install.rs` (+ unit tests) | ✅ code | | B3 HSTS: none on :80, `max-age=0` on :443 (actively clears cached policy) | `image-recipe/configs/nginx-archipelago.conf` | ✅ code | | B4 NPM manifest: `/etc/letsencrypt` mount + `NET_BIND_SERVICE` | `apps/nginx-proxy-manager/manifest.yml` | ✅ code | | B5 `portAuth` alias resolution + unanimous port-wide fallback | `neode-ui/src/views/discover/curatedApps.ts` | ✅ code | | B6 Catalog cache warmed at dashboard bootstrap | `neode-ui/src/App.vue` | ✅ | | B7 NIP-07 bridge: host/port equality + reply to `event.origin` | `neode-ui/src/stores/appLauncher.ts` ✅ · `neode-ui/src/views/appSession/useNostrBridge.ts` ✅ | ✅ | | B8 Stale LND 0.18.4 refs in test expectations | `tests/lifecycle/remote-lifecycle.sh` | ✅ | ## C. Regression tests ("never again") | Test | Guards | Status | |------|-------|--------| | C1 Rust: router v2 response shape, nested errors, failure reasons | B1 | ✅ | | C2 Rust: setup-token log extraction (live-captured 2.45.0 line shape) | B2 | ✅ | | C3 bats: `lnd-api-compat` — POST `/v2/router/send` on the running LND must answer (never 404) | B1 vs image skew at gate time | ✅ (route probe verified live on shorty: HTTP 500 ≠ 404) | | C4 bats: nginx must NOT send HSTS on :80; :443 must send `max-age=0` | B3 | ✅ | | C5 neode-ui unit: portAuth alias + unanimous-scan (incl. bitcoin-knots→8334 https) | B5/B4-mixed-content | ✅ (6 tests) | | C6 neode-ui unit: bridge origin equality ignores scheme | B7 | ✅ (2 tests) | Backend suites: 34 targeted Rust tests green (payments v2 shape, setup-token extraction, lnd wallet/info regressions); middleware/dispatcher suite green; full neode-ui suite green (62 tests in the touched areas); production bundle built and verified to embed the alias fix. `cargo fmt` applied. ## D. Deploy & live verification | Step | Status | |------|--------| | D1 shorty NPM crash-loop stopped cleanly (user-stopped marker; public hosts keep serving via host nginx mirror) | ✅ 12:52Z | | D2 shorty live nginx HSTS patch + reload | ✅ verified: :80 and :443 both answer `max-age=0` | | D3 Regenerate catalog (releases/app-catalog.json + store copies) | ✅ semantic diff = exactly the two NPM fixes | | D4 **User runs `scripts/sign-catalog.sh`** (signer built at /tmp/archy-sign-bin) | ✅ catalog signed + committed + pushed | | D5 Commit + push (origin + gitea-vps2 OTA mirror) | ✅ 9 commits pushed | | D6 Release v1.8.9-alpha: `scripts/create-release.sh 1.8.9-alpha` (mnemonic) → `scripts/publish-release-assets.sh 1.8.9-alpha gitea-vps2` | ✅ PUBLISHED (tag v1.8.9-alpha, releases/manifest.json live, backend+frontend assets verified by the script) | | D7 OTA on shorty-s + framework-pt (Update button; shorty is on 1.8.8-alpha, daily check — hit Update now) | ⬜ user action | | D8 shorty: clear the NPM user-stopped marker + Start (or it starts via the fixed catalog) | ✅ NPM LIVE-HEALED via the signed catalog: unit regenerated with both fixes, container up, admin UI HTTP 200 on :8081 (verified 15:42Z) | | D9 framework-pt: Start Mempool — its containers are confirmed stopped (port 4080 refuses; gate answers on 7778/8334/50002/18083 so those apps will embed over https immediately) | ⬜ | | D10 Post-deploy live checks: LND send+receive; mempool/IndeeHub/bitcoin-UI frames over https; NPM healthy + admin :8081 ✅; portainer token card on fresh DB; zero CORS errors | ⬜ after nodes update | ## E. Follow-ups discovered during the incident (ride the NEXT release, v1.8.10+) - **LND channel-peer watchdog** (this release's headline platform fix): every 2 minutes the daemon reconnects peers of open channels that LND has not re-established on its own (per-peer retry throttled to 10 minutes), using the peer's advertised addresses from the public graph. Kills the whole class this incident exposed — a channel unroutable ~17h after an LND update while both nodes looked healthy. Unit tests pin the selection logic over the live REST shapes. - **Funding-modal honesty fix** (1464b1b2): the Lightning "no channel" modal now states the node's real state — pending channel confirming / balance on the far side / payment couldn't route / genuinely no channels. Note the stale-direction defect it fixes: the payment-failure mapper never set the direction, so a SEND failure showed the RECEIVE-branch copy ("Receiving needs inbound liquidity…") — the exact modal users saw while their node had a healthy 583k-outbound channel. Both fixes have their v1.8.10 CHANGELOG + What's New entries staged so the next `create-release.sh 1.8.10-alpha` runs clean first time. - Nodes poll for OTA updates on `daily_check` — after publishing, tell the user to hit Update rather than wait for the next check. - `origin` remote had a stale pushurl with a dead token (pushes failed); fixed to the canonical repo URL, stale `~/.git-credentials` entry with an encoded port removed. ## F. Post-v1.8.9 verification on shorty-s (2026-09-01 evening) - v1.8.9 applied; payment pipeline confirmed live: a 400,000 sat payment SUCCEEDED through the v2 router route; the 404s are gone. - App gate serves TLS on 4080/8334/18083/50002 (401 gate pages over https) — https app frames now answer. Mempool over https requires a hard refresh (PWA precaches the old bundle). - **"No route to the recipient" on sends is real**: the invoices being tested are from framework-pt, whose only channel (peer "Sandwich Farm", 0224c955…) is flagged `disabled` on BOTH policy sides in the routing graph after today's node churn — the peer connection never re-established (LND's reconnect backoff can stretch to hours). A disabled edge is unroutable in both directions, so payments to/from framework-pt fail regardless of shorty's 583k outbound. Fix: `lncli connect` the peer, wait for the channel_update to re-enable the edge (~minutes), then re-test. - The 577k attempt earlier failed for a different, correct reason: it exceeded the channel's spendable balance (583,542 − 9,850 reserve ≈ 573k max). framework-pt immediate workaround until its OTA lands: open the dashboard by IP (`http://192.168.x.x`) instead of `framework-pt.local`, and/or clear the cached policy once via `chrome://net-internals/#hsts` → Delete domain security policies → `framework-pt.local`.