Completes local-dev-master-slave-setup: dual-instance frontend tooling, module-capability gating, and master/slave protocol self-healing fixes

Frontend (Unit 2 completion): dual dev-server tooling (pnpm dev:slave,
pnpm dev:all), per-instance browser tab titles, and a backend
capability check (SystemController + useSystemCapabilities +
ModuleGuard) so a Master-only page is hidden on a slave instance
instead of assuming every backend has every module.

Master/slave protocol fixes surfaced by actually running master and
slave side by side locally:
- Deactivating a CMS instance (Inactive) now releases the slave's
  master gate instead of leaving it stuck on its last pushed status.
- The periodic integrity check now also re-pushes status to every
  reachable slave (previously URL-verification only) and runs once
  immediately on startup.
- Added the originally-specified (but never implemented) slave-pull
  path: a slave now periodically polls its own status from the master
  (GET /api/v1/SlaveStatus) and fails open to Available if the master
  is unreachable for too long, complementing the existing push.
- The slave's own Settings page can no longer "successfully" change
  local availability while the master controls it; it's now locked
  with an explanatory banner and the backend rejects the write with
  409 instead of silently no-op'ing it.
- CMS instance status badges now match the dashboard's color/icon
  styling instead of a plain grey badge.

Also corrected the master-cms-module design docs to match this
as-built behavior, and flagged (without a full rewrite) a larger,
pre-existing divergence between its inception-stage application
design and what construction actually built.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
2026-07-04 19:53:52 +02:00
co-authored by Claude Sonnet 5
parent 274946dbff
commit 0447993181
81 changed files with 2191 additions and 86 deletions
@@ -150,3 +150,108 @@ sequenceDiagram
```
Text alternative: Middleware first checks bypass paths, then admin JWT. If neither applies: checks static master cache; if master unavailable return 503. If master available: checks local availability service; if locally unavailable return 503. Otherwise pass through.
---
## Flow 5: Slave Poll (Slave → Master) — Added 2026-07-04
**Trigger**: `MasterStatusPollingBackgroundService` — one tick immediately on slave startup, then every `MasterPolling:PollIntervalSeconds` (default 30s).
```mermaid
sequenceDiagram
box rgba(76,175,80,0.15) Slave CMS
participant Timer as PeriodicTimer
participant BgSvc as MasterStatusPollingBackgroundService
participant Svc as MasterAvailabilityService
participant Client as MasterStatusPollClient
participant DB as AvailabilityDbContext
participant Cache as StaticCache
end
box rgba(33,150,243,0.15) Master CMS
participant MC as SlaveStatusController
end
Timer->>BgSvc: Tick
BgSvc->>Svc: GetPollTargetAsync()
Svc->>DB: GetRegistrationAsync()
alt No registration
DB-->>Svc: null
Svc-->>BgSvc: null
BgSvc-->>Timer: no-op, wait for next tick
else Registration exists
DB-->>Svc: MasterUrl, encrypted ApiKey
Svc-->>BgSvc: MasterUrl, plain ApiKey
BgSvc->>Client: GetStatusAsync(masterUrl, plainApiKey)
Client->>MC: GET /api/v1/SlaveStatus with X-Master-Api-Key header
alt Poll succeeds
MC-->>Client: 200 OK { IsAvailable, DisableMessage }
Client-->>BgSvc: PolledMasterStatus
BgSvc->>Svc: ApplyPolledStatusAsync(isAvailable, disableMessage)
Svc->>Cache: set _masterIsAvailable / _masterDisableMessage
Svc->>DB: Update LastPolledAt=now, LastContactedAt=now
else Poll fails (network error, timeout, 401, etc.)
Client-->>BgSvc: null
BgSvc->>Svc: RecordPollFailureAsync(failOpenAfter)
Svc->>DB: read LastPolledAt (or RegisteredAt if never polled)
alt Unreachable longer than failOpenAfter
Svc->>Cache: force _masterIsAvailable=true, _masterDisableMessage=null
Note over Svc: Fail-open — a dead/unreachable master must never permanently block this slave
else Still within grace period
Note over Svc: No change — leave the existing cached gate as-is
end
end
end
```
Text alternative: Background timer triggers the poller; if no master is registered, it's a no-op. Otherwise it calls `GET /api/v1/SlaveStatus` on the registered master. On success, the response overwrites the in-memory gate and updates `LastPolledAt`/`LastContactedAt`. On failure, the gate is left alone unless the master has been unreachable (via poll) for longer than `MasterPolling:FailOpenAfterMinutes`, in which case the gate is forced open (`Available`, no message).
---
## Flow 6: Local Availability Status Read/Write with Master-Gate Override — Added 2026-07-04
Applies to the slave's own `/settings` admin UI/API — `GET /api/v1/Availability/status` and `POST /api/v1/Availability/admin/status` — layered on top of `PersistentAvailabilityService`, which previously only ever reflected the locally-persisted `GlobalAvailabilityState` regardless of the master gate.
```mermaid
sequenceDiagram
box rgba(33,150,243,0.15) Frontend (SettingsPage)
participant FE as Owner Browser
end
box rgba(76,175,80,0.15) Slave CMS
participant Ctrl as AvailabilityController
participant Svc as PersistentAvailabilityService
participant MasterSvc as MasterAvailabilityService
participant DB as ApplicationDbContext
end
FE->>Ctrl: GET /api/v1/Availability/status
Ctrl->>Svc: GetStatusDetailsAsync()
Svc->>MasterSvc: GetMasterStatus()
alt Master gate closed (IsAvailable = false)
MasterSvc-->>Svc: NotAvailable, masterMessage
Svc-->>Ctrl: Status=NotAvailable, Message=masterMessage, IsMasterControlled=true
else Master gate open
MasterSvc-->>Svc: Available
Svc->>DB: read GlobalAvailabilityState
DB-->>Svc: local Status/Message
Svc-->>Ctrl: Status, Message, IsMasterControlled=false
end
Ctrl-->>FE: 200 OK
FE->>Ctrl: POST /api/v1/Availability/admin/status (attempted local change)
Ctrl->>Svc: UpdateStatusAsync(newStatus, reason, updatedBy)
Svc->>MasterSvc: GetMasterStatus()
alt Master gate closed
MasterSvc-->>Svc: NotAvailable
Svc-->>Ctrl: throw MasterControlledAvailabilityException
Ctrl-->>FE: 409 Conflict (ProblemDetails)
else Master gate open
MasterSvc-->>Svc: Available
Svc->>DB: persist newStatus/reason
Svc-->>Ctrl: success
Ctrl-->>FE: 200 OK
end
```
Text alternative: Reading status now checks the master gate first — if closed, the response reflects the master's forced status and message, with `IsMasterControlled=true`, regardless of what's persisted locally. Writing a status change now performs the same master-gate check first: if closed, the write is rejected outright with `409 Conflict` instead of silently succeeding with no visible effect, since the master gate would have overridden it on the very next read anyway.
> Previously (before 2026-07-04): `GetStatusDetailsAsync` ignored the master gate entirely and always returned the locally-persisted status; `UpdateStatusAsync` always wrote the requested change regardless of the master gate, giving a misleading "success" for a change that was immediately invisible.
@@ -69,6 +69,7 @@ Text alternative: Bypass path check first. Admin JWT next (bypasses both gates).
| `/api/v1/Auth/` | Already bypassed — login must always work |
| `/api/v1/Setup/status` | Already bypassed — frontend init check |
| `/api/v1/master/` | **NEW** — master management endpoints must bypass gate so master can always push status or re-register |
| `/api/v1/SlaveStatus` | **NEW, added 2026-07-04** — this is actually the *master's* incoming endpoint for slave pulls, but it's added to this same bypass list on any instance that also loads `Modules.Availability` (i.e. the master itself), so the master's own local-gate status never blocks a slave from reading it |
**Cache behavior rules**:
@@ -76,7 +77,7 @@ Text alternative: Bypass path check first. Admin JWT next (bypasses both gates).
|---|------|
| BR-SLAVE-08 | `_masterIsAvailable` defaults to `true` (fail-open) on process startup |
| BR-SLAVE-09 | `_masterDisableMessage` defaults to `null` on process startup |
| BR-SLAVE-10 | Cache has no expiry (Q4=A); only updated on `POST /status` with valid API key |
| BR-SLAVE-10 | **(Updated 2026-07-04)** Cache is updated on `POST /status` (push, valid API key) **and** periodically overwritten by `MasterStatusPollingBackgroundService` pulling `GET /api/v1/SlaveStatus` from the master (see Rule Set 5). It is no longer purely push-driven or expiry-free: a poll failure that persists past `MasterPolling:FailOpenAfterMinutes` (default 5 min, measured from `MasterRegistration.LastPolledAt`) forcibly resets the cache to `Available`/`null` regardless of the last pushed value. |
| BR-SLAVE-11 | 503 response from master gate includes `_masterDisableMessage` in `ProblemDetails.Detail` |
---
@@ -89,6 +90,7 @@ Text alternative: Bypass path check first. Admin JWT next (bypasses both gates).
| BR-SLAVE-13 | On `RegisterAsync`: if row with that Id exists → update. If not → insert. Never delete. |
| BR-SLAVE-14 | `RegisteredAt` is set once at creation and never updated |
| BR-SLAVE-15 | `LastContactedAt` is updated on every successful master call (register, status push, get-url) |
| BR-SLAVE-16 | **(Added 2026-07-04)** `LastPolledAt` is updated whenever this slave successfully polls the master via `MasterStatusPollingBackgroundService` (distinct from `LastContactedAt`, which tracks master-initiated contact) |
---
@@ -101,3 +103,31 @@ Text alternative: Bypass path check first. Admin JWT next (bypasses both gates).
| Get registered URL | `GET` | `/api/v1/master/registered-url` | None (API key in header) |
All three endpoints are unauthenticated from ASP.NET Core's perspective — they use the custom `X-Master-Api-Key` header validation implemented in `MasterAvailabilityService`. They are also in the middleware bypass list so the gate cannot block master management calls.
---
## Rule Set 5: Slave Pull + Fail-Open — Added 2026-07-04
Closes a gap versus the original inception requirements (`inception/requirements/requirements.md`, FR-MASTER-06 "Slave Pull Model", FR-MASTER-07 "Slave Fallback Behavior", NFR-MASTER-01 "Fail-Open Safety"): construction had implemented push-only, with fail-open surviving only as an in-memory startup default (BR-SLAVE-08/09) rather than an actively-reconciling pull. Added after a slave was observed remaining on a stale status through a restart and again after being deactivated on the master.
| # | Rule |
|---|------|
| BR-PULL-01 | `MasterStatusPollingBackgroundService` runs one tick immediately on slave startup, then every `MasterPolling:PollIntervalSeconds` (default `30`) |
| BR-PULL-02 | If no `MasterRegistration` exists yet, the poll tick is a no-op (nothing to poll) |
| BR-PULL-03 | On a successful poll (`GET /api/v1/SlaveStatus` on the registered master, header `X-Master-Api-Key`), the response overwrites the in-memory gate (`_masterIsAvailable`/`_masterDisableMessage`) and updates `MasterRegistration.LastPolledAt` + `LastContactedAt` |
| BR-PULL-04 | On a failed poll (network error, timeout, or non-success HTTP status), the gate is **not** changed immediately — instead `RecordPollFailureAsync` checks how long it's been since `LastPolledAt` (or `RegisteredAt` if never polled) |
| BR-PULL-05 | **Fail-open**: if that elapsed time exceeds `MasterPolling:FailOpenAfterMinutes` (default `5`), the gate is forced to `Available`/`null` — a master that is dead or unreachable must never permanently block a slave |
| BR-PULL-06 | Push (BR-SLAVE-01 through 11) and pull (this rule set) are independent and complementary: push gives instant reactivity to an explicit admin status change; pull is the self-healing safety net for everything push can miss (slave restarts, dropped pushes, local tampering with the in-memory gate) |
---
## Rule Set 6: Master-Controlled Availability Lock — Added 2026-07-04
Applies to the slave's own local availability admin UI/API (`PersistentAvailabilityService` / `AvailabilityController` — the endpoints an Owner uses on `/settings` to set *this instance's own* Maintenance/NotAvailable status), not the master↔slave protocol endpoints above. Added because an Owner on a master-disabled slave could previously "successfully" set local status to `Available` with no visible effect, since the master gate silently overrode the display.
| # | Rule |
|---|------|
| BR-LOCK-01 | `GET /api/v1/Availability/status` returns the master-gate status (not the locally-persisted one) whenever the master gate is closed (`IsAvailable = false`), and includes a new `IsMasterControlled: true` flag in that case |
| BR-LOCK-02 | `POST /api/v1/Availability/admin/status` (`PersistentAvailabilityService.UpdateStatusAsync`) throws `MasterControlledAvailabilityException` and makes **no** DB write when the master gate is closed, instead of silently persisting a change that would have no visible effect |
| BR-LOCK-03 | The controller translates that exception into `409 Conflict` (`ProblemDetails`) |
| BR-LOCK-04 | The frontend Settings page reads `isMasterControlled` and disables the mode selector, the reason field, and the save button, showing a banner explaining that the Master CMS controls this status |
@@ -40,9 +40,10 @@ Singleton row — at most one record exists per slave instance. Upserted on each
|-------|------|-------------|-------|
| `Id` | `Guid` | PK | Fixed value (e.g. `Guid.Empty`) enforces singleton |
| `MasterUrl` | `string` | Required, max 500 | URL of the master CMS that registered this slave |
| `ApiKey` | `string` | Required, max 1000 | Plain-text API key sent in first registration; used for subsequent validation |
| `ApiKey` | `string` | Required, max 1000 | API key sent in first registration, encrypted via `IMasterApiKeyProtector` (ASP.NET Core Data Protection) before storage, decrypted for each subsequent validation **corrected 2026-07-04**, this was previously (incorrectly) documented as stored plain-text |
| `RegisteredAt` | `DateTimeOffset` | Required | Timestamp of first registration |
| `LastContactedAt` | `DateTimeOffset?` | Optional | Updated on every successful master call (register, status push, get-url) |
| `LastContactedAt` | `DateTimeOffset?` | Optional | Updated on every successful master-initiated call (register, status push, get-url) |
| `LastPolledAt` | `DateTimeOffset?` | Optional | **Added 2026-07-04**. Updated on every successful *slave-initiated* poll (`MasterStatusPollingBackgroundService`) — the counterpart to `LastContactedAt`, tracking the opposite direction of contact. Also the basis for the fail-open timeout (see below). |
> **Singleton enforcement**: The `Id` is a fixed known value (`Guid.Parse("00000000-0000-0000-0000-000000000001")`). On first `POST /api/v1/master/register` the row is created; on re-registration the same row is updated in-place. This avoids a composite unique constraint and makes EF upsert trivial.
@@ -54,10 +55,10 @@ Not a DB entity — lives in memory on the slave process.
| Field | Type | Default | Notes |
|-------|------|---------|-------|
| `_masterIsAvailable` | `bool` | `true` | Set by status push; `true` = pass gate |
| `_masterDisableMessage` | `string?` | `null` | Message forwarded from master to 503 response |
| `_masterIsAvailable` | `bool` | `true` | Set by status push **or** successful poll; forced back to `true` on prolonged poll failure (fail-open) |
| `_masterDisableMessage` | `string?` | `null` | Message forwarded from master to 503 response; cleared to `null` on fail-open |
**No expiry** (Q4=A): cache is valid indefinitely until the master pushes again. If the master goes offline permanently, the last known status is used. Default = `true` (Available) = fail-open.
> **Updated 2026-07-04**: originally documented as having "no expiry" (Q4=A), valid indefinitely until the next push. This is no longer accurate now that `MasterStatusPollingBackgroundService` actively polls the master (see Rule Set 5 / Flow 5 in the sibling docs) and forces the cache back to `Available`/`null` if the master has been unreachable via poll for longer than `MasterPolling:FailOpenAfterMinutes` (default 5 min, measured from `MasterRegistration.LastPolledAt`). Push-driven updates (this section's original description) are unchanged and still apply; the poll is an additional, independent path that can also write to this same cache.
---