Add operation-owned legacy IndeeHub maintenance controller draft
This commit is contained in:
@@ -0,0 +1,73 @@
|
||||
# Legacy IndeeHub maintenance controller
|
||||
|
||||
Status: isolated source implementation. Ten pure Python fake-runtime regressions
|
||||
pass; no live invocation or production qualification. The controller is not part
|
||||
of the already signed private app candidate and needs no new app image/API.
|
||||
|
||||
The supervised updater owns the lifecycle flock, seven durable holds, original
|
||||
Quadlets and private writable-layer recovery images. It records its destructive
|
||||
obligation before invoking the fixed controller with bounded JSON over stdin:
|
||||
|
||||
```
|
||||
python3 /opt/archipelago/scripts/indeehub-maintenance-controller.py acquire
|
||||
{ "operation_id": "<uuid>", "original_members": [
|
||||
{ "name": "indeedhub", "container_id": "<64hex>", "image_id": "<64hex>",
|
||||
"unit_sha256": "<64hex>", "config_sha256": "<64hex>", "running": true }
|
||||
// All seven exact members; JSON does not include this illustrative comment.
|
||||
], "recovery": false }
|
||||
```
|
||||
|
||||
Other actions are `verify` with operation_id, and `release` with operation_id and
|
||||
outcome committed/restored/aborted. Replies are <=4KiB and report drained, held,
|
||||
released, or recovering for explicit recovery acquire. Inherited
|
||||
ARCHY_UPDATE_LOCK_FD stays open and is passed to child commands; the script never
|
||||
unlocks it. Journal: data/update-transactions/indeehub-maintenance/<uuid>/journal.json.
|
||||
|
||||
## Forward sequence
|
||||
|
||||
- Validate exact original IDs/source-unit hashes and known port exposure. Only
|
||||
frontend127.0.0.1:7778 is supported; direct backend/S3 ports refuse before stop.
|
||||
- Require deployed native AppGate and legacy nginx maintenance guards. Inspect
|
||||
every known legacy sublocation and any direct7778 proxy; unknown routes refuse.
|
||||
Save an operation-owned readable sentinel, then verify local ingress returns503.
|
||||
- Record prior BullMQ transcode pause state, globally pause future job admission,
|
||||
retain queued/delayed/failed jobs. Gracefully stop frontend ingress; bounded
|
||||
polling waits for active transcodes to finish before stopping worker and API.
|
||||
- Require successful systemd shutdown plus an exact original Podman died event
|
||||
with exit0. Forced exits and missing event evidence retain the hold and are
|
||||
never labelled completed writes.
|
||||
- While PostgreSQL remains running, capture a fresh custom dump. Cleanly stop
|
||||
MinIO/Redis/relay/Postgres, then archive all four complete quiescent volumes
|
||||
(including SQLite WAL and Redis persistence) with metadata. No volume deletion
|
||||
or migration rollback. Archive hashes/size and per-step obligations are durable.
|
||||
- Keep admission closed while the updater renders, starts and verifies targets.
|
||||
|
||||
## Interrupted recovery
|
||||
|
||||
The node first records phase Restoring with boolean target_startup_began, then
|
||||
calls acquire with recovery:true. That path preserves the original failure and
|
||||
fence; it does not retry a killed original into a fictitious successful drain or
|
||||
claim missing backups exist. The node restores exact saved old runtime under the
|
||||
same hold. Release before any target startup can state only that original runtime
|
||||
was restored. If target startup/migration began, recorded data-compatibility
|
||||
verification is required before restored release; an old image alone does not
|
||||
prove compatibility with newly changed data. No automatic DB/media restore exists.
|
||||
|
||||
## Qualification and remaining integration
|
||||
|
||||
`python3 tests/regression/test_indeehub_maintenance_controller.py` passes ten
|
||||
fake-runtime cases in temporary directories, without services/network/containers.
|
||||
Source nginx template guard coverage also passes its parser check. Production
|
||||
adapter compilation, actual Podman event format/systemd clean-exit behavior,
|
||||
application writer shutdown, interrupted backup and supervised restart still need
|
||||
isolated lifecycle fixtures and then coordinated node acceptance. A long-lived
|
||||
WebSocket or active upload can exceed graceful-stop deadlines; the current code
|
||||
refuses completion and preserves recovery obligations rather than silently
|
||||
calling interrupted work finished.
|
||||
|
||||
The deployment must install the exact qualified controller script and record its
|
||||
hash alongside the backend artifact. Binary-only deployment does not install it.
|
||||
The backend must refuse missing/mismatched prerequisites before snapshots/stops.
|
||||
Native AppGate + nginx guards are separate node source changes owned by the
|
||||
supervised updater agent. The signed app catalog/private image receipts remain
|
||||
unchanged. Existing live stop/uninstall intent must not be rewritten as maintenance.
|
||||
Reference in New Issue
Block a user