Integrate retained updater candidate for release qualification

This commit is contained in:
archipelago
2026-10-07 13:50:26 -04:00
22 changed files with 5389 additions and 168 deletions
@@ -0,0 +1,114 @@
# Legacy IndeeHub maintenance controller
Status: isolated source implementation. Fourteen pure Python fake-runtime regressions
pass; no live invocation or production qualification. The controller is not part
of the already signed private app candidate and needs no new app image/API.
The supervised updater owns the lifecycle flock, seven durable holds, original
Quadlets and private writable-layer recovery images. It records its destructive
obligation before invoking the fixed controller with bounded JSON over stdin:
```
python3 /opt/archipelago/scripts/indeehub-maintenance-controller.py acquire
{ "operation_id": "<uuid>", "original_members": [
{ "name": "indeedhub", "container_id": "<64hex>", "image_id": "<64hex>",
"unit_sha256": "<64hex>", "config_sha256": "<64hex>", "running": true }
// All seven exact members; JSON does not include this illustrative comment.
], "recovery": false }
```
Other actions are `verify` with operation_id, and `release` with operation_id and
outcome committed/restored/aborted. Replies are <=4KiB and report drained, held,
released, or recovering for explicit recovery acquire. Inherited
ARCHY_UPDATE_LOCK_FD stays open and is passed to child commands; the script never
unlocks it. Journal: data/update-transactions/indeehub-maintenance/<uuid>/journal.json.
## Forward sequence
- Validate exact original IDs/source-unit hashes and known port exposure. Only
frontend127.0.0.1:7778 is supported; direct backend/S3 ports refuse before stop.
- Require deployed native AppGate and legacy nginx maintenance guards. Inspect
every known legacy sublocation and any direct7778 proxy; unknown routes refuse.
Save an operation-owned readable sentinel, then verify local ingress returns503.
- Record prior BullMQ transcode pause state, globally pause future job admission,
retain queued/delayed/failed jobs. Gracefully stop frontend ingress; bounded
polling waits for active transcodes to finish before stopping worker and API.
- Require successful systemd shutdown plus an exact original Podman died event
with exit0. Forced exits and missing event evidence retain the hold and are
never labelled completed writes.
- While PostgreSQL remains running, capture a fresh custom dump. Cleanly stop
MinIO/Redis/relay/Postgres, then archive all four complete quiescent volumes
(including SQLite WAL and Redis persistence) with metadata. No volume deletion
or migration rollback. Archive hashes/size and per-step obligations are durable.
- Keep admission closed while the updater renders, starts and verifies targets.
## Interrupted recovery
The node first records phase Restoring with boolean target_startup_began, then
calls acquire with recovery:true. That path preserves the original failure and
fence; it does not retry a killed original into a fictitious successful drain or
claim missing backups exist. The node restores exact saved old runtime under the
same hold. Release before any target startup can state only that original runtime
was restored. If target startup/migration began, recorded data-compatibility
verification is required before restored release; an old image alone does not
prove compatibility with newly changed data. No automatic DB/media restore exists.
## Qualification and remaining integration
`python3 tests/regression/test_indeehub_maintenance_controller.py` passes fourteen
fake-runtime cases in temporary directories, without services/network/containers.
Source nginx template guard coverage also passes its parser check. Production
adapter compilation, actual Podman event format/systemd clean-exit behavior,
application writer shutdown, interrupted backup and supervised restart still need
isolated lifecycle fixtures and then coordinated node acceptance. A long-lived
WebSocket or active upload can exceed graceful-stop deadlines; the current code
refuses completion and preserves recovery obligations rather than silently
calling interrupted work finished.
The deployment must install the exact qualified controller script and record its
hash alongside the backend artifact. Binary-only deployment does not install it.
The backend must refuse missing/mismatched prerequisites before snapshots/stops.
Native AppGate + nginx guards are separate node source changes owned by the
supervised updater agent. The signed app catalog/private image receipts remain
unchanged. Existing live stop/uninstall intent must not be rewritten as maintenance.
A pre-acquire snapshot/preflight failure may leave no controller journal. An
Aborted node journal with target_startup_began=false then permits idempotent
no-op acknowledgement, without touching any other operation’s admission fence.
A matching fence without its controller journal requires recovery investigation.
Read-only source evidence from actual old API/ffmpeg shows neither has SIGTERM
shutdown hooks. The controller permits worker143 only after a paused queue has
zero active jobs. Legacy API143 additionally requires closed/stopped frontend,
stopped worker, and a fresh empty projects/contents/payments/shareholders/
subscriptions/library_items store with no other active DB transaction. This is
a narrow first-upgrade compatibility path, not evidence populated work completed.
Populated or ambiguous legacy state remains a refused forward cutover.
### Operation-bound rollback data verification
Restored release after any target startup now performs its own PostgreSQL
compatibility check; it does not accept a manually asserted verification boolean.
Before the coherent backup, the controller captures a read-only, repeatable-read
transaction containing every original public table's columns, constraints,
indexes, triggers, row-security policies, row count and canonical row SHA-256,
plus the exact migration history. The private operation journal binds this
baseline to the original operation UUID.
After original runtime recovery, while ingress and worker admission remain
closed, the controller captures the same observations again. It requires original
tables and definitions unchanged, original non-migration rows unchanged and the
original migration-history prefix intact. Additional migration records must be
the exact ordered three migrations qualified for API commit `3b09b81`, and only
their five named new tables may appear, all empty. Any unexpected data or schema
change keeps ingress closed. Successful proof records before/after commitment
hashes and the operation UUID before release. No down migration or automatic
volume restoration is performed.
Sixteen pure Python regressions pass, including altered rows/schema/history,
foreign operation, nonempty added tables and durable proof before fence release.
The SQL transaction and Podman lifecycle still require isolated integration
qualification. These are table-level compatibility checks, not a claim that
arbitrary database extensions/functions, other writers, or changed application
code are safe. The candidate images, migration scope and admission barrier must
also match the reviewed operation.
@@ -0,0 +1,77 @@
# Managed update runtime recovery
Status: isolated source implementation; Rust and real Podman/systemd qualification
pending. No live update, snapshot, stop, backup or rollback has been performed by
this work. Active deployed source is unchanged.
The managed path captures original source Quadlet bytes, mode, immutable image,
container identity, launch configuration and running intent. It requires an
original-hash-bound reviewed forward plan and verifies every planned image against
the already prepared catalog image. New manifest configuration/hooks are applied
on the forward path; rollback uses the captured original recipe and a private,
local-only writable-layer image. AutoRemove rollback recreates containers and does
not claim to restore their original IDs. Stopped supervised stacks currently fail
before mutation; the retained-container and separate stopped-stage paths cover
only their respective supported cases.
The legacy IndeeHub controller runs under the same inherited lifecycle lock and
operation-owned reconciliation holds. Original writable layers are captured
before any destructive stop. The controller fences ingress, drains supported
legacy work, takes coherent quiescent volume/database backups and retains the
fence through cutover or recovery. Its exact source hash must match the separately
installed script; a backend binary alone does not install the controller.
Completed updates and verified runtime restorations publish exact unit recipes
before releasing holds. Routine drift reconciliation validates those recipes;
it does not regenerate them from a newer catalog. Catalog-driven pre-start file
and mount mutations are skipped for these pinned installations. Missing saved
units/images are explicit recovery failures, never permission to reconstruct a
different runtime. Explicit uninstall removes the installed recipe, and completed
old journals cannot recreate it. A new reviewed transaction replaces the recipe.
Opted-in API registration environments must match an already provisioned pin
and the existing node identity in both manifest and exact Quadlet. Administrative
plan preparation must use the existing installer provisioning code. Execution
never invents a node identity or accepts a browser-supplied unit or hook.
Runtime restoration does not establish database compatibility. The controller's
restored-release verifier compares original table schemas and row commitments,
permitting only the exact reviewed additive migration prefix and empty new
application tables. Any other data/schema change keeps ingress closed. This is
not automatic database rollback or a promise that arbitrary migrations are
reversible.
Qualification required before integration/activation:
- Isolated backend compile and injectable lifecycle/fault tests, including lost
replies, daemon interruption, foreign units/holds and preflight failures.
- Disposable real Podman/systemd and PostgreSQL execution of the controller and
adapter. Sixteen pure controller tests currently pass; SQL/runtime behavior is
not yet qualified.
- Final seven-member private plan with installer-resolved identity environment,
verified local images and source-unit provenance; review required changes and
retained operator configuration without exposing secret values.
- Disk-capacity and recovery-image retention checks, backup integrity and a
documented recovery path for missing runtime artifacts.
- Actual-node controlled deployment and acceptance, preserving persistent data.
## Resumed qualification — 2026-10-07
Backup verification now requires the full database/four-volume artifact set and
rechecks SHA256, including same-size corruption. Nineteen pure controller tests
pass. The owned, network-none PostgreSQL fixture passes unchanged/additive
commitments, rejects four data/schema/history mutations, restores a real custom
dump with matching original commitments, and rejects a truncated dump. It mounts
no live volume and removes only its own container. This does not qualify actual
application writer drain or the complete supervised systemd cutover.
The updater compiled and its full isolated suite ran: 1,958 passed, one failed,
five ignored. The failure is the old snake_case rental receipt JSON fixture;
`c2c4d915` already corrects that exact test on the release candidate branch.
Do not duplicate or suppress it here. Integrate and rerun the complete candidate
suite before claiming a green backend gate. The earlier interrupted compile
and PostgreSQL timeout remain failed/incomplete attempts, not acceptance.
Evidence: `/tmp/archy-resumed-20261007-updater-full-backend.log`,
`/tmp/archy-resumed-20261007-indeehub-controller-tests-final.log`, and
`/tmp/archy-resumed-20261007-indeehub-postgres-restore.log`.