4.6 KiB
Legacy IndeeHub maintenance controller
Status: isolated source implementation. Twelve pure Python fake-runtime regressions pass; no live invocation or production qualification. The controller is not part of the already signed private app candidate and needs no new app image/API.
The supervised updater owns the lifecycle flock, seven durable holds, original Quadlets and private writable-layer recovery images. It records its destructive obligation before invoking the fixed controller with bounded JSON over stdin:
python3 /opt/archipelago/scripts/indeehub-maintenance-controller.py acquire
{ "operation_id": "<uuid>", "original_members": [
{ "name": "indeedhub", "container_id": "<64hex>", "image_id": "<64hex>",
"unit_sha256": "<64hex>", "config_sha256": "<64hex>", "running": true }
// All seven exact members; JSON does not include this illustrative comment.
], "recovery": false }
Other actions are verify with operation_id, and release with operation_id and
outcome committed/restored/aborted. Replies are <=4KiB and report drained, held,
released, or recovering for explicit recovery acquire. Inherited
ARCHY_UPDATE_LOCK_FD stays open and is passed to child commands; the script never
unlocks it. Journal: data/update-transactions/indeehub-maintenance//journal.json.
Forward sequence
- Validate exact original IDs/source-unit hashes and known port exposure. Only frontend127.0.0.1:7778 is supported; direct backend/S3 ports refuse before stop.
- Require deployed native AppGate and legacy nginx maintenance guards. Inspect every known legacy sublocation and any direct7778 proxy; unknown routes refuse. Save an operation-owned readable sentinel, then verify local ingress returns503.
- Record prior BullMQ transcode pause state, globally pause future job admission, retain queued/delayed/failed jobs. Gracefully stop frontend ingress; bounded polling waits for active transcodes to finish before stopping worker and API.
- Require successful systemd shutdown plus an exact original Podman died event with exit0. Forced exits and missing event evidence retain the hold and are never labelled completed writes.
- While PostgreSQL remains running, capture a fresh custom dump. Cleanly stop MinIO/Redis/relay/Postgres, then archive all four complete quiescent volumes (including SQLite WAL and Redis persistence) with metadata. No volume deletion or migration rollback. Archive hashes/size and per-step obligations are durable.
- Keep admission closed while the updater renders, starts and verifies targets.
Interrupted recovery
The node first records phase Restoring with boolean target_startup_began, then calls acquire with recovery:true. That path preserves the original failure and fence; it does not retry a killed original into a fictitious successful drain or claim missing backups exist. The node restores exact saved old runtime under the same hold. Release before any target startup can state only that original runtime was restored. If target startup/migration began, recorded data-compatibility verification is required before restored release; an old image alone does not prove compatibility with newly changed data. No automatic DB/media restore exists.
Qualification and remaining integration
python3 tests/regression/test_indeehub_maintenance_controller.py passes twelve
fake-runtime cases in temporary directories, without services/network/containers.
Source nginx template guard coverage also passes its parser check. Production
adapter compilation, actual Podman event format/systemd clean-exit behavior,
application writer shutdown, interrupted backup and supervised restart still need
isolated lifecycle fixtures and then coordinated node acceptance. A long-lived
WebSocket or active upload can exceed graceful-stop deadlines; the current code
refuses completion and preserves recovery obligations rather than silently
calling interrupted work finished.
The deployment must install the exact qualified controller script and record its hash alongside the backend artifact. Binary-only deployment does not install it. The backend must refuse missing/mismatched prerequisites before snapshots/stops. Native AppGate + nginx guards are separate node source changes owned by the supervised updater agent. The signed app catalog/private image receipts remain unchanged. Existing live stop/uninstall intent must not be rewritten as maintenance.
A pre-acquire snapshot/preflight failure may leave no controller journal. An Aborted node journal with target_startup_began=false then permits idempotent no-op acknowledgement, without touching any other operation’s admission fence. A matching fence without its controller journal requires recovery investigation.