21 KiB
phase, plan, subsystem, tags, requires, provides, affects, tech-stack, key-files, key-decisions, requirements-completed, coverage, duration, completed, status
| phase | plan | subsystem | tags | requires | provides | affects | tech-stack | key-files | key-decisions | requirements-completed | coverage | duration | completed | status | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 01-federation-mesh-hardening | 16 | security |
|
|
|
|
|
|
|
|
140min | 2026-08-01 | task-1-complete-checkpoint-FAILED-recreate-does-not-fire |
Phase 1 Plan 16: Rotate Existing Installs Off the Shipped Gateway Credential (FED-07) Summary
Task 1 rotates the credential correctly and was proven to do so on a real node. But the checkpoint DISPROVED the assumption it rests on: the rotated credential never reaches the running container, because the Quadlet path does not restart units and the gateway is classified restart-sensitive. FED-07 remains open.
Status
FED-07 is NOT closed, and this plan alone cannot close it. The checkpoint ran on archi-dev-box on 2026-08-01 and found that rotation does not propagate to the running gateway. A follow-up plan is required — see the checkpoint result below.
Accomplishments
rotate_compromised_gateway_credential(secrets_dir) -> Result<bool>incontainer::secrets: reads the canonical hash file, returnsOk(false)for absent/unreadable/unique/unrecognised, and only on an exact denylist match writes a fresh pair and returnsOk(true).write_bcrypt_pair(dir, name)factored out ofensure_one'sBcryptarm so there is exactly one bcrypt-generation implementation, called by both generation and rotation.- Wired into
resolve_dynamic_envbesideensure_generated_secrets, gated on the gateway's app id, with an info-level announcement that names the path to the new plaintext and never the value. - Six new tests covering rotate-on-denylisted (including 0600 modes and that the
.pwsibling verifies against the new hash), no-op-on-unique, no-op-on-unrecognised, no-op-on-absent, idempotence, and four bystander secrets left byte-identical.
Findings the plan asked for
Boot reconciler needs no separate call
boot_reconciler calls reconcile_all() → reconcile_all_with_mode() → per-manifest
ensure_running_with_mode() (prod_orchestrator.rs:1714) → resolve_dynamic_env()
(prod_orchestrator.rs:1914) → the rotation call. install_fresh reaches it by the same route.
So boot and reconcile funnel through one chokepoint and no boot-specific wiring was added;
boot_reconciler.rs is not in files_modified.
✅ RESOLVED 2026-08-02: the carve-out closes it, proven on the same node
The failure documented below was fixed and re-verified on archi-dev-box.
The fix. Rotation now records the app id in a credential_rotated set on the orchestrator, and
the drift check consumes that flag to recreate the container even when the app is
restart-sensitive, with a WARN naming the reason. It deliberately mirrors the published-port
carve-out sitting a few lines above it in the same function, which already makes exactly this trade
for exactly this reason: a container that is already non-functional (there) or already compromised
(here) is not protected by leaving it running. No teardown was hand-rolled — the existing recreate
path does the work.
The design call. Restart-sensitivity protects working services. A gateway answering to a
credential published in this repository is not working, it is compromised, and an attacker with
gateway admin can drain Lightning liquidity. Indefinite exposure loses to a few seconds of restart.
The alternative — rotate but only raise an operator alert — was rejected because the monitoring
system fires alerts from metric thresholds only (check_alerts(&MetricSnapshot)), so it would have
needed new event-alert plumbing to deliver something strictly weaker than just fixing it.
Before / after on the same box, same scenario:
| Before (15:51, first checkpoint) | After (06:39, re-run) | |
|---|---|---|
| Credential file | rotated ✅ | rotated ✅ |
| Container recreated | ❌ never — same PID 25 min later | ✅ 4s after rotation, PID 3923125 → 148426 |
| Running credential | the compromised default | matches the file (636d6031…) |
| Log | "leaving running restart-sensitive app untouched" ×4 and counting | "recreating restart-sensitive app: its admin credential was rotated off a publicly known default…" |
Post-recreate state: container healthy, same name, same ports (8176/9737), gatewayd.db
intact at 18 files with IDENTITY present, credential 0600 archipelago:archipelago, 32 containers
untouched, and zero repeat rotations on subsequent ticks (the rotated value is not denylisted, so
it is self-terminating — T-01-73 holds).
Covered by three new tests: rotating_a_compromised_credential_flags_the_app_for_recreate,
a_unique_credential_does_not_flag_the_app, a_second_pass_does_not_re_flag_the_app.
⛔ ORIGINAL CHECKPOINT RESULT (2026-08-01) — retained: this is what the fix had to defeat
Run on archi-dev-box, 2026-08-01. The code-reading conclusion below was wrong in practice, which is exactly why the plan made this a blocking checkpoint. The rotation works; the propagation to the running container does not.
Observed: rotation fired ~15s after restart, wrote a fresh unique credential, and logged correctly.
But 25 minutes later the running gatewayd process was still using the pre-rotation credential.
/proc/<pid>/environ for PID 1 held the old value while the file and the podman secret held the new
one. The orchestrator says why, in its own log lines:
Quadlet unit drift-synced — file rewritten, .service NOT restarted
(operator restart picks up new config) app_id=fedimint-gateway
container drift detected during boot reconcile;
leaving running restart-sensitive app untouched app_id=fedimint-gateway
Two independent guards, both deliberate:
- The Quadlet path rewrites the
.containerunit but never restarts the.service. The unit file was rewritten at 15:51:36 (same second as the rotation) carrying the newsecret-env-hash=de9870c642a515f7label — so the definition updated correctly. Systemd does not apply a changed unit to a running container without a restart. fedimint-gatewayis classified restart-sensitive, so the drift check detects the change on every reconcile tick and then deliberately leaves the container alone. That line repeated at 15:51, 15:53, 15:54, 15:56 — it will repeat forever.
Consequence: on a real affected node, rotation makes the credential file unique but the gateway
keeps answering to the compromised one indefinitely — until an unrelated reboot or a manual
restart. Worse, the operator reading fedimint-gateway-hash.pw gets a password the running gateway
does not accept, which is the lockout risk T-01-77 inverted.
Confirmed the fix works when applied: systemctl --user restart fedimint-gateway.service produced a
new PID whose FEDI_HASH is the rotated value, container healthy, same name, same ports, marker file
and gatewayd.db intact.
Deliberately NOT hand-rolled. The plan's action says: "If that path does not fire for this app for some reason you discover, do not hand-roll a remove-and-run; stop and record what you found." So this is recorded, not patched. The fix belongs in a follow-up and has to answer a real design question: a compromised credential is arguably the one case that should override restart-sensitivity — or, failing that, the rotation must raise an operator-facing "restart required" alert rather than logging into the void.
What the checkpoint DID prove
| Step | Result |
|---|---|
| 1. State recorded | Node was CLEAN; affected state seeded deliberately (plan's Planner Assumption) |
| 2. Deploy | install -m0755 to /usr/local/bin/archipelago + systemctl restart archipelago; rollback kept at archipelago.bak-pre-fed07 |
| 3. Rotation announced | ✅ One info line, fired once, names the .pw path, no credential value in the log |
| 4. Credential unique | ✅ Third distinct value (not the default, not the pre-test original), 0600 archipelago:archipelago |
| 5. Data preserved | ✅ Marker file and all 13 gatewayd.db files incl. IDENTITY intact; same container name; same ports 8176/9737 |
| 5. Container recreated | ❌ FAILED — see above |
| 6. Auth proof | ⚠️ Not obtainable: gateway-cli in this image returns the same "Invalid request" for a correct and an incorrect password, so it cannot distinguish them. Substituted PID-1 environ comparison, which is stronger evidence of which credential is in force. |
| 7. Fresh install differs | ⏸ Not run |
8. run-gate.sh |
⏸ Not run |
Also proved incidentally: restarting archipelago does not kill containers on this box — 29/29
and later 31/31 survived, and the orchestrator logged "Adopted 31 existing container(s)". The
CLAUDE.md "restart SIGKILLs containers" rule does not apply under ARCHIPELAGO_USE_QUADLET_BACKENDS=true
with podman in the user slice (the service is system.slice/KillMode=control-group; the containers
live in user-1000.slice/…/libpod-*, a different cgroup entirely).
Original code-reading conclusion (retained — it is what the checkpoint disproved)
resolve_dynamic_env computes secret_env_content_hash(&secret_bearing) over the resolved
secret-bearing env and stores it as manifest.app.container.secret_env_hash
(prod_orchestrator.rs:3309). The drift check (prod_orchestrator.rs:3374) inspects the running
container's SECRET_ENV_HASH_LABEL and returns "drifted" when it differs from the expected hash,
which drives the existing recreate. The gateway's FEDI_HASH comes from the rotated file, so a
rotation necessarily changes that hash and therefore the label comparison.
This is a code-reading conclusion. It has not been observed firing on a node — that is Task 2 step 5, and it is the single most important thing the checkpoint proves.
Operator recovery: the surface exists but does NOT cover this app — a real gap
- The UI path is live:
Apps.vuecallspackage.credentialswith anapp_idbefore launching an app and renders a credentials modal from the response. - The backend,
handle_package_credentialsincore/archipelago/src/api/rpc/package/install.rs:2093, is a hardcoded per-app if-chain covering onlyfilebrowserandphotoprism. Every other app, includingfedimint-gateway, falls through toOk(json!({ "credentials": [] })). - Consequence: after rotation the operator has no in-UI way to obtain the new gateway password.
The recovery path is the file the log line names:
/var/lib/archipelago/secrets/fedimint-gateway-hash.pw(0600, service user), readable over SSH. - Gap owner:
handle_package_credentialsincore/archipelago/src/api/rpc/package/install.rs. Adding afedimint-gatewayarm that reads the.pwsibling would close it; the UI needs no change. Deliberately not done here — this plan'sfiles_modifiedis scoped to two files, and that handler belongs to the app-credentials surface, not to FED-07's rotation.
Adjacent finding — NOT part of this plan, raised deliberately
apps/photoprism/manifest.yml:35 sets PHOTOPRISM_ADMIN_PASSWORD=archipelago, and
handle_package_credentials hands that same literal back to the UI. That is a shipped default
credential in a manifest — the same class of defect as FED-07, on a different app. Every node running
PhotoPrism answers to admin / archipelago.
It is out of scope here (this plan is the gateway migration) and was not touched. It wants its own
requirement and plan, and probably the same treatment: a generated_secrets entry plus a denylist
entry for the shipped value.
Deviations from Plan
Bcrypt generation had to be factored out first
Found during: Task 1
Issue: The plan says rotation should "generate a replacement pair through the same helper
ensure_gateway_credential uses". 01-11 never actually created such a helper — it left the bcrypt
arm inline in ensure_one and had ensure_gateway_credential call ensure_one. Rotation cannot
call ensure_gateway_credential, because ensure_one's idempotent fast path returns early when the
target files are present and non-empty, which is precisely the state rotation acts on.
Resolution: Extracted write_bcrypt_pair(dir, name) from the Bcrypt arm; ensure_one and
rotation both call it. Still exactly one generation implementation, which is what the instruction was
protecting.
Files modified: core/archipelago/src/container/secrets.rs
Known Stubs
None.
Threat Flags
- T-01-72 (critical, EoP) — mitigated in code, not yet proven on a node. Task 2 step 6 (old credential rejected, new one accepted) is the proof and has not been run.
- T-01-73 (DoS, rotation loop) — mitigated and unit-tested: the rotated value is not on the
denylist, so the next tick is a no-op (
rotation_is_idempotent). - T-01-74 (info disclosure) — mitigated structurally:
rotate_compromised_gateway_credentialreturnsbool, so the credential is not even in scope at the logging call site. - T-01-75 (tampering / data loss) — mitigated: no teardown primitives added (grep-verified), the
recreate goes through
secret_env_hash. On-node data-survival check is Task 2 step 5, not run. - T-01-76 (repudiation — signing off without exercising rotation) — live risk, unresolved. Whether archi-dev-box is affected or already clean is still unknown; the plan requires declaring which case it is and deliberately seeding the old value if the node is clean.
- T-01-77 (operator lockout) — partially mitigated: the plaintext exists at a named 0600 path and the log line points at it, but there is no UI retrieval path (see the gap above).
- T-01-SC — no crates added.
Self-Check
- CONFIRMED:
cargo test -p archipelago secrets→ 16 passed, 0 failed (thecontainer::secretsmodule holds 14#[test]fns, all six new rotation cases among them:rotates_a_denylisted_gateway_credential,leaves_a_unique_gateway_credential_alone,leaves_an_unrecognised_credential_alone,no_op_when_no_gateway_credential_exists,rotation_is_idempotent,rotation_touches_no_other_secret) - FOUND:
rotate_compromised_gateway_credentialinsecrets.rs(definition + 5 test uses) - FOUND: exactly 1 non-comment reference in
prod_orchestrator.rs - CONFIRMED: 0 added teardown primitives (
rm -f/remove_dir_all/podman rm/chown) in theprod_orchestrator.rsdiff - CONFIRMED:
cargo fmt --check -p archipelagoclean. It was not clean before this plan —install.rscarried drift introduced by 01-11's commit (42652547), fixed here. That check has blocked the release gate before (37d293be), so it is worth keeping green rather than discovering at ship time. - CONFIRMED:
cargo test -p archipelago(aftercargo clean -p archipelago) — 1008 passed, 1 failed. The failure iscontainer::boot_reconciler::tests::second_pass_fires_after_interval, the same wall-clock-timed test (50ms tick) that was flaky during 01-11; re-run in isolation it is 4 passed / 0 failed in 0.46s.boot_reconciler.rsis untouched by this plan. - NOT RUN: Task 2's eight-step on-node checkpoint, and
tests/lifecycle/run-gate.sh
A false alarm worth recording, because it cost an hour
An intermediate full-suite run reported credentials::operations::tests::test_list_credentials_filter_by_did
failing with "invalid utf-8 sequence of 1 bytes from index 2" — an identity-credentials test in a
module this plan does not touch, which had passed in the 01-11 run two hours earlier.
Cause: corrupted build artifacts, not a regression. Two duplicate cargo test runs had been
started against the same workspace lock and one was SIGTERMed to free it. The next compile surfaced
rust-lld: error: undefined hidden symbol — precisely the incremental-cache corruption CLAUDE.md
documents. After cargo clean -p archipelago the credentials test passes and the only failure is the
known timing flake above.
Lesson for the next executor on this box: do not kill an in-flight cargo to free the build lock —
let it finish. A corrupted target dir produces failures in modules you never touched, which reads
exactly like a real regression and is not one.