Completes local-dev-master-slave-setup: dual-instance frontend tooling, module-capability gating, and master/slave protocol self-healing fixes
Frontend (Unit 2 completion): dual dev-server tooling (pnpm dev:slave, pnpm dev:all), per-instance browser tab titles, and a backend capability check (SystemController + useSystemCapabilities + ModuleGuard) so a Master-only page is hidden on a slave instance instead of assuming every backend has every module. Master/slave protocol fixes surfaced by actually running master and slave side by side locally: - Deactivating a CMS instance (Inactive) now releases the slave's master gate instead of leaving it stuck on its last pushed status. - The periodic integrity check now also re-pushes status to every reachable slave (previously URL-verification only) and runs once immediately on startup. - Added the originally-specified (but never implemented) slave-pull path: a slave now periodically polls its own status from the master (GET /api/v1/SlaveStatus) and fails open to Available if the master is unreachable for too long, complementing the existing push. - The slave's own Settings page can no longer "successfully" change local availability while the master controls it; it's now locked with an explanatory banner and the backend rejects the write with 409 instead of silently no-op'ing it. - CMS instance status badges now match the dashboard's color/icon styling instead of a plain grey badge. Also corrected the master-cms-module design docs to match this as-built behavior, and flagged (without a full rewrite) a larger, pre-existing divergence between its inception-stage application design and what construction actually built. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
+11
@@ -61,3 +61,14 @@ This generates the `Migrations/` folder contents. The migration is applied autom
|
||||
- `Microsoft.Extensions.Http.Resilience` version `9.6.0` — verify/update during `dotnet restore` if a newer version is available for .NET 10
|
||||
- `IntegrityCheckIntervalMinutes = 0` in tests forces immediate PeriodicTimer ticks (valid for test scenarios only)
|
||||
- Slave-side endpoints (`/api/v1/master/register`, `/api/v1/master/status`, `/api/v1/master/registered-url`) are implemented in Unit 2 (slave-availability-extension)
|
||||
|
||||
## Addendum — 2026-07-04 (added outside this unit's original scope, in `local-dev-master-slave-setup` follow-up fixes)
|
||||
|
||||
This unit predates the following; see `slave-availability-extension/functional-design/business-rules.md` Rule Set 5 and `application-design/application-design.md` for the full picture:
|
||||
|
||||
| File | Description |
|
||||
|------|-------------|
|
||||
| `Controllers/SlaveStatusController.cs` | New. `[AllowAnonymous]` `GET /api/v1/SlaveStatus`; authenticates via `X-Master-Api-Key` header matched against each active `CmsInstance`'s decrypted key; lets a slave pull its own status instead of relying solely on the master's push |
|
||||
| `Services/ICmsInstanceService.cs` / `CmsInstanceService.cs` | `GetStatusForApiKeyAsync(plainApiKey)` added; `UpdateStatusAsync`'s `Inactive` branch now pushes `Available`/null to release the gate (previously a no-op); `VerifyIntegrityAsync` now also re-pushes persisted status to every reachable active slave each cycle |
|
||||
| `BackgroundServices/IntegrityCheckBackgroundService.cs` | Now runs one tick immediately on startup, in addition to the periodic timer |
|
||||
| `Models/SlaveStatusPollResponse` (in `ICmsInstanceService.cs`) | New record: `IsAvailable`, `DisableMessage` |
|
||||
|
||||
+31
-4
@@ -97,7 +97,19 @@ sequenceDiagram
|
||||
Svc->>Repo: UpdateAsync (Status=Inactive, DisableMessage=null)
|
||||
Svc->>Repo: SaveChangesAsync()
|
||||
Repo->>DB: UPDATE CmsInstances
|
||||
Svc-->>Ctrl: UpdateStatusResult(Success=true, SlaveContactSuccess=true)
|
||||
Svc->>DP: Unprotect(entity.ApiKey)
|
||||
DP-->>Svc: plainApiKey
|
||||
Svc->>Client: PushStatusAsync(slaveUrl, plainApiKey, isAvailable=true, disableMessage=null)
|
||||
Note over Svc: Releases the master gate — the master no longer manages this slave, so it must not stay stuck on its last pushed status
|
||||
alt Release success
|
||||
Client-->>Svc: true
|
||||
Svc->>Repo: UpdateAsync (LastStatusPushedAt = UtcNow)
|
||||
Svc->>Repo: SaveChangesAsync()
|
||||
Svc-->>Ctrl: UpdateStatusResult(Success=true, SlaveContactSuccess=true)
|
||||
else Release failed
|
||||
Client-->>Svc: false
|
||||
Svc-->>Ctrl: UpdateStatusResult(Success=true, SlaveContactSuccess=false)
|
||||
end
|
||||
else newStatus = Available or NotAvailable
|
||||
Svc->>Repo: UpdateAsync (Status, DisableMessage)
|
||||
Svc->>Repo: SaveChangesAsync()
|
||||
@@ -119,13 +131,15 @@ sequenceDiagram
|
||||
Ctrl-->>Ctrl: return 200 OK with UpdateStatusResult
|
||||
```
|
||||
|
||||
Text alternative: Controller calls service with id and new status; service loads entity, validates, updates DB, then for non-Inactive transitions decrypts ApiKey and pushes status to slave; returns SlaveContactSuccess=false if push fails but DB is always the authority.
|
||||
Text alternative: Controller calls service with id and new status; service loads entity, validates, updates DB, then decrypts ApiKey and pushes status to slave — for `Inactive` this push is always `isAvailable=true, disableMessage=null` (releasing the gate); for `Available`/`NotAvailable` it pushes the new status as-is. Returns SlaveContactSuccess=false if the push fails, but the DB write is always the authority.
|
||||
|
||||
> **Updated 2026-07-04**: the `Inactive` branch previously did not push anything to the slave at all (see history below) — this left the slave stuck on whatever status it had last received, indefinitely. Fixed by always releasing the gate on deactivation.
|
||||
|
||||
---
|
||||
|
||||
## Flow 3 — VerifyIntegrityAsync (Background Integrity Check)
|
||||
|
||||
**Trigger**: `IntegrityCheckBackgroundService` periodic timer (every `IntegrityCheckIntervalMinutes`)
|
||||
**Trigger**: `IntegrityCheckBackgroundService` — one tick immediately on host startup, then every `IntegrityCheckIntervalMinutes` **(startup tick added 2026-07-04)**
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
@@ -163,6 +177,7 @@ sequenceDiagram
|
||||
Svc->>Repo: UpdateAsync (LastIntegrityCheckFailedAt = UtcNow)
|
||||
Svc->>Repo: SaveChangesAsync()
|
||||
Repo->>DB: UPDATE CmsInstances
|
||||
Note over Svc: Unreachable for URL check — skip the status re-push for this instance this cycle
|
||||
else Slave reachable
|
||||
Client-->>Svc: registeredMasterUrl
|
||||
alt URLs match
|
||||
@@ -183,9 +198,21 @@ sequenceDiagram
|
||||
Repo->>DB: UPDATE CmsInstances
|
||||
end
|
||||
end
|
||||
Note over Svc: Status re-push (added 2026-07-04) — runs whenever the slave was reachable, independent of the URL-match outcome
|
||||
Svc->>Client: PushStatusAsync(slaveUrl, plainApiKey, isAvailable=(Status==Available), disableMessage)
|
||||
alt Push success
|
||||
Client-->>Svc: true
|
||||
Svc->>Repo: UpdateAsync (LastStatusPushedAt = UtcNow)
|
||||
Svc->>Repo: SaveChangesAsync()
|
||||
else Push failed
|
||||
Client-->>Svc: false
|
||||
Note over Svc: Logged; no flag change — next cycle (or the immediate startup tick) will retry
|
||||
end
|
||||
end
|
||||
end
|
||||
Svc-->>BgSvc: done
|
||||
```
|
||||
|
||||
Text alternative: Background timer triggers integrity service; for each non-Inactive slave: decrypts key, retrieves registered master URL, clears failure flag on match, re-registers on mismatch, sets LastIntegrityCheckFailedAt when slave is unreachable or re-registration fails.
|
||||
Text alternative: Background timer triggers integrity service; for each non-Inactive slave: decrypts key, retrieves registered master URL, clears failure flag on match, re-registers on mismatch, sets LastIntegrityCheckFailedAt when slave is unreachable or re-registration fails. **(Added 2026-07-04)** For every slave that was reachable, the service additionally re-pushes the master's currently persisted `Status`/`DisableMessage` to that slave — this is what lets a slave that reset its in-memory gate (e.g. after a restart) catch up without waiting for the next explicit admin status change. Combined with the new immediate startup tick on `IntegrityCheckBackgroundService`, this reconciliation now also runs right after the master process (re)starts.
|
||||
|
||||
> **Updated 2026-07-04**: previously this flow only verified/re-registered the master URL and never re-pushed status (see history below) — a restarted slave (whose in-memory master-gate defaults to `Available`) would show the wrong status until the master's next explicit UI-driven change.
|
||||
|
||||
+45
-14
@@ -10,7 +10,11 @@ graph TD
|
||||
CheckMsg{"newStatus = NotAvailable\nAND disableMessage\nis null or empty?"}
|
||||
ValidationErr["Throw ValidationException\nDisableMessage required"]
|
||||
CheckInactive{"newStatus\n= Inactive?"}
|
||||
SetInactive["Status = Inactive\nDisableMessage = null\nNo HTTP push\nSlaveContactSuccess = true"]
|
||||
SetInactive["Status = Inactive\nDisableMessage = null\nPersist to MasterDbContext"]
|
||||
ReleaseGate["Decrypt ApiKey\nPushStatusAsync(isAvailable=true,\ndisableMessage=null)\n— release the master gate"]
|
||||
ReleaseOk{"Release push\nsucceeded?"}
|
||||
ReleaseDone["LastStatusPushedAt = UtcNow\nSave"]
|
||||
ReturnRelease["Return UpdateStatusResult\nSuccess=true\nSlaveContactSuccess=(release result)"]
|
||||
PersistStatus["Persist Status + DisableMessage\nto MasterDbContext"]
|
||||
DecryptKey["Decrypt ApiKey\nvia IDataProtector"]
|
||||
PushSlave["PushStatusAsync\nto slave endpoint"]
|
||||
@@ -25,7 +29,8 @@ graph TD
|
||||
CheckExists -->|"yes"| CheckMsg
|
||||
CheckMsg -->|"yes — invalid"| ValidationErr
|
||||
CheckMsg -->|"no — valid"| CheckInactive
|
||||
CheckInactive -->|"yes"| SetInactive --> Done
|
||||
CheckInactive -->|"yes"| SetInactive --> ReleaseGate --> ReleaseOk
|
||||
ReleaseOk -->|"yes/no"| ReleaseDone --> ReturnRelease --> Done
|
||||
CheckInactive -->|"no"| PersistStatus --> DecryptKey --> PushSlave --> PushOk
|
||||
PushOk -->|"yes"| UpdatePushed --> ReturnOk --> Done
|
||||
PushOk -->|"no"| ReturnWarn --> Done
|
||||
@@ -34,13 +39,15 @@ graph TD
|
||||
classDef action fill:#63b3ed,stroke:#2b6cb0,stroke-width:1px,color:#000
|
||||
classDef terminal fill:#9ae6b4,stroke:#2f855a,stroke-width:2px,color:#000
|
||||
classDef error fill:#FC8181,stroke:#C53030,stroke-width:2px,color:#000
|
||||
class CheckExists,CheckMsg,CheckInactive,PushOk decision
|
||||
class PersistStatus,DecryptKey,PushSlave,UpdatePushed,SetInactive action
|
||||
class CheckExists,CheckMsg,CheckInactive,PushOk,ReleaseOk decision
|
||||
class PersistStatus,DecryptKey,PushSlave,UpdatePushed,SetInactive,ReleaseGate,ReleaseDone action
|
||||
class Start,Done terminal
|
||||
class NotFound,ValidationErr error
|
||||
```
|
||||
|
||||
Text alternative: Load entity (404 if missing) → validate DisableMessage required for NotAvailable → for Inactive skip push → for others persist, decrypt key, push to slave, set SlaveContactSuccess based on push result.
|
||||
Text alternative: Load entity (404 if missing) → validate DisableMessage required for NotAvailable → for Inactive, persist the Inactive status **and then explicitly push `Available` (disableMessage=null) to the slave** to release the master gate (the master no longer manages this slave, so it must not leave the slave stuck on a stale status) → for other statuses persist, decrypt key, push to slave, set SlaveContactSuccess based on push result.
|
||||
|
||||
> **Updated 2026-07-04**: originally `Inactive` skipped the HTTP push entirely (see history below); this was found to leave slaves permanently stuck on their last pushed status (e.g. `NotAvailable`) after being detached from the master, and was fixed to always release the gate on deactivation.
|
||||
|
||||
---
|
||||
|
||||
@@ -58,26 +65,34 @@ graph TD
|
||||
RegOk{"Re-registration\nsucceeded?"}
|
||||
ClearAfterReg["LastIntegrityCheckFailedAt = null\nLastContactedAt = UtcNow\nSave"]
|
||||
SetFailedReg["LastIntegrityCheckFailedAt = UtcNow\nSave"]
|
||||
RePushStatus["PushStatusAsync\n(re-push persisted Status/DisableMessage\nto this active slave)"]
|
||||
RePushOk{"Push\nsucceeded?"}
|
||||
UpdatePushedAt["LastStatusPushedAt = UtcNow\nSave"]
|
||||
Next(["Next instance"])
|
||||
|
||||
Start --> GetUrl --> Reachable
|
||||
Reachable -->|"no"| SetFailed --> Next
|
||||
Reachable -->|"yes"| UrlMatch
|
||||
UrlMatch -->|"match"| ClearOk --> Next
|
||||
UrlMatch -->|"match"| ClearOk --> RePushStatus
|
||||
UrlMatch -->|"mismatch"| ReRegister --> RegOk
|
||||
RegOk -->|"yes"| ClearAfterReg --> Next
|
||||
RegOk -->|"yes"| ClearAfterReg --> RePushStatus
|
||||
RegOk -->|"no"| SetFailedReg --> Next
|
||||
RePushStatus --> RePushOk
|
||||
RePushOk -->|"yes"| UpdatePushedAt --> Next
|
||||
RePushOk -->|"no"| Next
|
||||
|
||||
classDef decision fill:#FFC107,stroke:#F57F17,stroke-width:2px,color:#000
|
||||
classDef action fill:#63b3ed,stroke:#2b6cb0,stroke-width:1px,color:#000
|
||||
classDef terminal fill:#9ae6b4,stroke:#2f855a,stroke-width:2px,color:#000
|
||||
classDef error fill:#FC8181,stroke:#C53030,stroke-width:2px,color:#000
|
||||
class Reachable,UrlMatch,RegOk decision
|
||||
class GetUrl,SetFailed,ClearOk,ReRegister,ClearAfterReg,SetFailedReg action
|
||||
class Reachable,UrlMatch,RegOk,RePushOk decision
|
||||
class GetUrl,SetFailed,ClearOk,ReRegister,ClearAfterReg,SetFailedReg,RePushStatus,UpdatePushedAt action
|
||||
class Start,Next terminal
|
||||
```
|
||||
|
||||
Text alternative: For each active slave — attempt to get its registered master URL; if unreachable set failure flag; if reachable and URL matches clear flag; if mismatch re-register; clear flag on success, set flag on failure.
|
||||
Text alternative: For each active slave — attempt to get its registered master URL; if unreachable set failure flag and move on; if reachable and URL matches clear the flag, if mismatch re-register (clearing the flag on success, setting it on failure); then — regardless of the URL check's outcome, as long as the slave was reachable — **re-push the master's currently persisted Status/DisableMessage to that slave** (this is the reconciliation path for a slave that reset its in-memory gate, e.g. after a restart, or missed an earlier push). `IntegrityCheckBackgroundService` also now runs one tick immediately on host startup, in addition to its periodic interval, so this resync happens right after the master (re)starts rather than waiting a full cycle.
|
||||
|
||||
> **Updated 2026-07-04**: originally this check only verified/re-registered the master URL and never re-pushed status (see history below); this left restarted slaves stuck on a stale in-memory status (reset to `Available` by default) until the next explicit status change — fixed by adding the re-push step and the immediate startup run.
|
||||
|
||||
---
|
||||
|
||||
@@ -101,10 +116,12 @@ Text alternative: For each active slave — attempt to get its registered master
|
||||
|------|----|---------------|-----------|-------|
|
||||
| Any | `Available` | Clear to null | Yes | Slave re-enabled |
|
||||
| Any | `NotAvailable` | Required, non-empty | Yes | Slave disabled with message |
|
||||
| Any | `Inactive` | Clear to null | **No** | Master stops all contact |
|
||||
| Any | `Inactive` | Clear to null | **Yes** — pushes `Available`/null | Master stops managing the slave, but must first release the gate so the slave doesn't stay stuck on its last pushed status |
|
||||
| `Inactive` | `Available` | Clear to null | Yes | Reactivation |
|
||||
| `Inactive` | `NotAvailable` | Required, non-empty | Yes | Reactivation with disable |
|
||||
|
||||
> **Updated 2026-07-04**: the `Inactive` row previously said "No" HTTP push (see history below) — corrected after the no-push behavior was found to leave slaves permanently stuck on their last status.
|
||||
|
||||
---
|
||||
|
||||
## ApiKey Encryption Rules
|
||||
@@ -132,9 +149,9 @@ Text alternative: For each active slave — attempt to get its registered master
|
||||
|
||||
| Rule | Description |
|
||||
|------|-------------|
|
||||
| BR-CONTACT-01 | Instances with `Status = Inactive` are excluded from `GetActiveAsync` and never contacted via HTTP |
|
||||
| BR-CONTACT-02 | Status push is skipped when transitioning any status → `Inactive` |
|
||||
| BR-CONTACT-03 | Integrity check runs only against instances where `Status != Inactive` |
|
||||
| BR-CONTACT-01 | Instances with `Status = Inactive` are excluded from `GetActiveAsync` and are not contacted by the periodic integrity check / re-push cycle |
|
||||
| BR-CONTACT-02 | **(Updated 2026-07-04)** The transition to `Inactive` itself always performs exactly one status push — `Available`, `disableMessage=null` — to release the master gate on the slave before the instance drops out of `GetActiveAsync` for good. Previously this push was skipped entirely; that left the slave stuck on its last pushed status indefinitely. |
|
||||
| BR-CONTACT-03 | Integrity check (and its new status re-push, see BR-02) runs only against instances where `Status != Inactive` |
|
||||
|
||||
---
|
||||
|
||||
@@ -146,3 +163,17 @@ Text alternative: For each active slave — attempt to get its registered master
|
||||
| BR-BG-02 | Each tick creates and disposes its own `IServiceScope` |
|
||||
| BR-BG-03 | Exceptions within a single slave's integrity check are caught, logged, and do not abort processing for remaining slaves |
|
||||
| BR-BG-04 | If `MasterModuleOptions.MasterUrl` is null or empty, the background service logs a warning and skips the entire integrity check for that cycle |
|
||||
| BR-BG-05 | **(Added 2026-07-04)** `IntegrityCheckBackgroundService` runs one tick immediately on host startup (in addition to its periodic `PeriodicTimer` cycle), so a freshly (re)started master resyncs slave statuses right away instead of waiting up to `IntegrityCheckIntervalMinutes` |
|
||||
|
||||
---
|
||||
|
||||
## Slave Pull (Status Poll) Rules — Added 2026-07-04
|
||||
|
||||
This closes a design gap: the original inception requirements (`inception/requirements/requirements.md`, FR-MASTER-06/07, NFR-MASTER-01) specified a slave-initiated periodic pull with fail-open, but construction implemented push-only. The pull side was added alongside the existing push mechanism (not instead of it) after a slave was observed staying on a stale status through a restart and a deactivation.
|
||||
|
||||
| Rule | Description |
|
||||
|------|-------------|
|
||||
| BR-PULL-01 | `SlaveStatusController` exposes `GET /api/v1/SlaveStatus`, authenticated via the `X-Master-Api-Key` header — no `[Authorize]`/JWT, since the caller is a slave process, not a logged-in user |
|
||||
| BR-PULL-02 | `CmsInstanceService.GetStatusForApiKeyAsync` identifies the calling slave by decrypting each active `CmsInstance.ApiKey` and comparing it to the caller's plain key (no separate slave-identity field exists; the shared key is the only credential) — returns `null` (→ 401) when no match is found |
|
||||
| BR-PULL-03 | A successful poll updates `CmsInstance.LastContactedAt`, mirroring the existing convention used by push-based master↔slave calls |
|
||||
| BR-PULL-04 | `/api/v1/SlaveStatus` is added to `AvailabilityMiddleware`'s bypass list on the master's own instance, so it stays reachable regardless of the master's own local availability status |
|
||||
|
||||
+19
-2
@@ -71,7 +71,9 @@ public enum CmsInstanceStatus
|
||||
|-------|---------|----------------------|
|
||||
| `Available` | Slave is enabled; normal operation | Yes (status push + integrity checks) |
|
||||
| `NotAvailable` | Slave is disabled; `DisableMessage` served to end-users | Yes (status push + integrity checks) |
|
||||
| `Inactive` | Soft-removed; greyed out in UI | **No** — all HTTP contact is halted |
|
||||
| `Inactive` | Soft-removed; greyed out in UI | On the transition **into** `Inactive`: one final push (`Available`, no message) to release the gate. Afterwards: **No** further contact — excluded from `GetActiveAsync`, so no more pushes/integrity checks/re-pushes |
|
||||
|
||||
> **Updated 2026-07-04**: previously `Inactive` meant no HTTP contact at all, including on the transition itself — this left slaves stuck on their last pushed status after being detached. See BR-CONTACT-02 in `business-rules.md`.
|
||||
|
||||
---
|
||||
|
||||
@@ -129,4 +131,19 @@ public enum CmsInstanceStatus
|
||||
| Property | Type | Notes |
|
||||
|----------|------|-------|
|
||||
| `Success` | `bool` | Always `true` when status persisted to DB (DB is the authority) |
|
||||
| `SlaveContactSuccess` | `bool` | `true` if HTTP push to slave succeeded; `false` if push failed (slave unreachable); not applicable for `Inactive` transitions (returns `true`) |
|
||||
| `SlaveContactSuccess` | `bool` | `true` if the HTTP push to the slave succeeded; `false` if it failed (slave unreachable). Applies to `Inactive` transitions too — reflects whether the gate-release push succeeded, not a hardcoded `true` |
|
||||
|
||||
> **Updated 2026-07-04**: `SlaveContactSuccess` for `Inactive` used to always be hardcoded `true` (no push happened, so nothing could fail) — now reflects the real result of the release push.
|
||||
|
||||
---
|
||||
|
||||
## SlaveStatusPollResponse (API Response) — Added 2026-07-04
|
||||
|
||||
Response shape for the new slave-pull endpoint (`GET /api/v1/SlaveStatus`, `SlaveStatusController`), used by `MasterStatusPollingBackgroundService` on the slave side (see `slave-availability-extension/functional-design/domain-entities.md`). This is the counterpart of the push-based flow above — added to close a gap versus the original inception requirements (FR-MASTER-06/07), which specified a slave-initiated pull in addition to the push that construction actually implemented.
|
||||
|
||||
| Property | Type | Notes |
|
||||
|----------|------|-------|
|
||||
| `IsAvailable` | `bool` | `true` when `CmsInstance.Status == Available` |
|
||||
| `DisableMessage` | `string?` | `CmsInstance.DisableMessage` |
|
||||
|
||||
Identified by matching the caller's plain API key (header `X-Master-Api-Key`) against each active `CmsInstance`'s decrypted key — there is no separate slave-identity field, the shared key doubles as the credential. No `[Authorize]`/JWT on this endpoint.
|
||||
|
||||
Reference in New Issue
Block a user