Files
SluijsensandClaude Opus 5 5f3eda2680 Makes a redeploy safe for the key ring and the schema
Nothing here is visible in normal operation. Its whole purpose is that
swapping the release directory on deploy cannot silently destroy state.

Data Protection secures the API keys that authenticate master/slave
communication. Two separate defaults would each have destroyed them:
keys are held on the filesystem, which a release swap discards, and the
application discriminator is derived from the content root path, which
changes with every release directory — so even keys stored in a database
would have stopped being derivable. Keys now live in
ApplicationDbContext and the discriminator is a fixed constant.

Losing them produces no error. It produces stored keys that no longer
decrypt, which presents as an apparent network fault between a Master
and its slaves and is easily misdiagnosed. That is also why the tests
assert the resulting configuration rather than the registration: the
XmlRepository must be the EF one and the discriminator must be the
constant, plus a round-trip proving a value encrypted before a deploy is
readable after one. A test that only checked "Data Protection is
registered" would have passed in the broken case too.

Both modules previously called AddDataProtection() themselves. Module
registration runs after the host's, so those calls re-registered the
configuration chain and would have overridden the persistent store while
IDataProtector still resolved. They are removed, with a comment at each
site — the deletion otherwise looks like a regression. Each module's own
test project now guards against it being reintroduced.

ApplicationDbContext also migrates itself at startup. Deploy targets
offer no CLI, so migrations cannot be a manual step on the server.
Failures are classified rather than treated alike: a connection failure
means the database is not up yet, normal when the app and the database
start together after a reboot, and is retried with backoff; a migration
failure means something is broken and fails at once. Either way the
process does not start, which is what makes the liveness health check
trustworthy — an application that cannot reach its schema never answers
/health, so monitoring goes red instead of reporting a healthy instance
that cannot serve a request.

The cost of migrating without a human gate is that migrations must stay
forward-compatible and non-destructive, since rollback is "redeploy the
previous release". The new migration is purely additive.

Also wires this and the preceding hosting commit into both hosts, as
they touch the same lines of Program.cs.

Two constraints are enforced by documentation rather than code, and
belong in the deployment instructions: the key table must never be
pruned, and only one instance may migrate a given database at a time.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HHoJpxYXzHACSQguHrC5fw
2026-07-28 00:00:45 +02:00

9.1 KiB

Business Rules — U2 Data Durability


Migration Decision Logic

graph TD
    start["Startup: migrate ApplicationDbContext"]
    attempt["Attempt migration"]
    ok{"Succeeded ?"}
    done["Continue startup"]
    kind{"Failure kind ?"}
    attempts{"Retry attempts remaining ?"}
    wait["Wait with increasing delay"]
    logfail["Log the failure with context"]
    stop["Propagate: process does not start"]

    start --> attempt
    attempt --> ok
    ok -->|yes| done
    ok -->|no| kind
    kind -->|"connection failure"| attempts
    kind -->|"migration failure"| logfail
    attempts -->|yes| wait
    attempts -->|no| logfail
    wait --> attempt
    logfail --> stop

    classDef entry fill:#f6e05e,stroke:#c05621,stroke-width:1px,color:#000;
    classDef decision fill:#90cdf4,stroke:#2b6cb0,stroke-width:1px,color:#000;
    classDef good fill:#9ae6b4,stroke:#2f855a,stroke-width:1px,color:#000;
    classDef bad fill:#fbb6ce,stroke:#b83280,stroke-width:1px,color:#000;
    class start entry;
    class ok,kind,attempts decision;
    class attempt,wait,done good;
    class logfail,stop bad;

Text alternative: migration is retried with increasing delay only when the failure is a connection problem; a genuine migration failure is logged and stops the process immediately, as does exhausting the retry attempts.


Key Ring Durability Rules

ID Rule
BR-U2-01 Data Protection keys are persisted in the database, in the context that owns Identity — never on the filesystem.
BR-U2-02 The application discriminator is a fixed constant in code. It is not derived from any path, and it is not configurable.
BR-U2-03 Data Protection is configured exactly once, by the host, before module service registration.
BR-U2-04 No module may call AddDataProtection(). Modules consume IDataProtector only.
BR-U2-05 Key lifetime uses the framework default of 90 days, with automatic rotation.
BR-U2-06 Old keys are never deleted. Removing a key makes every value encrypted with it permanently unreadable, including stored slave API keys.
BR-U2-07 Keys are stored unencrypted at rest in the database. This is an accepted, documented deviation — see DEV-05.
BR-U2-08 The database connection must enforce TLS, and the database must not be publicly reachable. These are the compensating controls for BR-U2-07.

Rationale for BR-U2-02: the default discriminator derives from the content root path, which changes on every atomic release switch. Persisting keys in the database while letting the discriminator move produces keys that are stored but underivable — the same failure, quieter. A constant cannot be forgotten during a host migration or accidentally differ between two instances sharing a database.

Rationale for BR-U2-03 and BR-U2-04: this is the § 5.1 conflict. Because module registration runs after the host's, a module's bare AddDataProtection() would override the persistent store. IDataProtector still resolves, so registration tests pass — the defect appears only after the first release switch, as slave API keys that no longer decrypt.

Rationale for BR-U2-06: this is the single most destructive maintenance action available against this system. Pruning the keys table looks like harmless housekeeping and permanently breaks every Master↔slave relationship. It must be stated in the deployment documentation, not only in code comments (Q5 = C).


Migration Rules

ID Rule
BR-U2-09 ApplicationDbContext migrations are applied automatically at startup, before the application accepts traffic.
BR-U2-10 Core migrations run before module middleware installation, so the keys table exists before any module resolves an IDataProtector.
BR-U2-11 A connection failure is retried with increasing delay, up to a bounded number of attempts.
BR-U2-12 A migration failure — invalid, conflicting, or inapplicable — fails immediately with no retry.
BR-U2-13 When retries are exhausted, or on a migration failure, the exception propagates and the process does not start.
BR-U2-14 Before failing, the reason is logged with enough context to diagnose it — but never including the connection string, credentials, or any secret.
BR-U2-15 AvailabilityDbContext and MasterDbContext keep their existing self-migration in UseModule. U2 does not change them.
BR-U2-16 Migrations must be forward-compatible and non-destructive, so redeploying an earlier release remains a valid rollback.

Rationale for BR-U2-11 and BR-U2-12: the two failures mean different things. On the Pi the application and SQL Server may start together after a reboot, so a brief unavailability window is normal operation, not a fault. A broken migration is a fault, and retrying it only delays the inevitable while filling the log.

Rationale for BR-U2-13: this is what gives U1's liveness check meaning. A process that starts regardless would report healthy while being unusable; strict startup makes the absence of a /health response a trustworthy signal.

Rationale for BR-U2-16: rollback is "redeploy the previous release" (D-26). That is only safe if the older code can run against the newer schema. A destructive migration — dropping a column, narrowing a type — makes rollback impossible precisely when it is most needed.


Operational Constraint Rules

ID Rule
BR-U2-17 Exactly one application instance may migrate a given database at a time. Not enforced in code; documented as an operational constraint (Q4 = C).
BR-U2-18 Each instance has its own database. The Master and each slave never share one.

Rationale: Database.Migrate() is not safe under concurrency. Today the constraint holds by design — one instance per database — so a distributed lock would add failure modes without removing any. Should the deployment model ever change to multiple instances sharing a database, automatic startup migration must be revisited before that change is made.


Error and Edge-Case Scenarios

Scenario Expected behaviour
Fresh database, no tables All Core migrations applied, including the keys table; startup proceeds
Database up to date Migration is a no-op; startup proceeds
Database unreachable at startup, comes up within the retry window Retries succeed; startup proceeds; the delay is logged
Database unreachable for the whole retry window Logged, process does not start, /health silent, UptimeRobot red
Migration conflicts with existing schema Failed immediately, no retry, process does not start
Credentials wrong Treated as a connection failure — retried, then fails. Distinguishing bad credentials from an unreachable server is not worth the complexity; the outcome is identical
Keys table empty on first run Data Protection generates a key and persists it; normal first-run behaviour
Keys table populated from a previous release Existing keys are read; previously encrypted values remain readable — the purpose of the unit
Release directory changed since last start Irrelevant — the discriminator is a constant (BR-U2-02)
Someone deletes rows from the keys table Every value encrypted with those keys becomes permanently unreadable. Prevented by documentation only (BR-U2-06)
Two instances start simultaneously against one database Undefined. Prevented by the operational constraint (BR-U2-17), not by code
A module still calls AddDataProtection() after this unit The persistent store is silently overridden. Prevented by BR-U2-04 and asserted by a test

Security Compliance for U2

Rule Status Notes
SECURITY-01 Partially compliant — DEV-05 Encryption in transit enforced by BR-U2-08 (TLS on the connection). Encryption at rest for the keys themselves is deferred: keys are stored unencrypted, relying on the database's own at-rest encryption and network isolation. Accepted with a follow-up (Q2 = C)
SECURITY-03 Compliant BR-U2-14 requires diagnostic context without secrets
SECURITY-09 Compliant No default credentials; failure messages carry no connection details
SECURITY-13 Improved The key ring surviving redeploys is precisely a software-integrity property: without it, the encrypted API keys that authenticate Master↔slave communication silently become invalid
SECURITY-15 Compliant Fails closed — the process does not start rather than serving in an unknown schema state

New Documented Deviation

ID Deviation Rationale Decided
DEV-05 Data Protection keys are stored unencrypted at rest. SECURITY-01 requires encryption at rest for persisted data. DPAPI is unavailable on Linux, and X.509 certificate encryption relocates the loss problem to the certificate — reintroducing the failure mode this unit exists to eliminate. Compensating controls: TLS on the database connection, and the database not publicly reachable (BR-U2-08). Certificate-based encryption is recorded as a separate follow-up item. Q2 = C