Completes local-dev-master-slave-setup: dual-instance frontend tooling, module-capability gating, and master/slave protocol self-healing fixes

Frontend (Unit 2 completion): dual dev-server tooling (pnpm dev:slave,
pnpm dev:all), per-instance browser tab titles, and a backend
capability check (SystemController + useSystemCapabilities +
ModuleGuard) so a Master-only page is hidden on a slave instance
instead of assuming every backend has every module.

Master/slave protocol fixes surfaced by actually running master and
slave side by side locally:
- Deactivating a CMS instance (Inactive) now releases the slave's
  master gate instead of leaving it stuck on its last pushed status.
- The periodic integrity check now also re-pushes status to every
  reachable slave (previously URL-verification only) and runs once
  immediately on startup.
- Added the originally-specified (but never implemented) slave-pull
  path: a slave now periodically polls its own status from the master
  (GET /api/v1/SlaveStatus) and fails open to Available if the master
  is unreachable for too long, complementing the existing push.
- The slave's own Settings page can no longer "successfully" change
  local availability while the master controls it; it's now locked
  with an explanatory banner and the backend rejects the write with
  409 instead of silently no-op'ing it.
- CMS instance status badges now match the dashboard's color/icon
  styling instead of a plain grey badge.

Also corrected the master-cms-module design docs to match this
as-built behavior, and flagged (without a full rewrite) a larger,
pre-existing divergence between its inception-stage application
design and what construction actually built.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
2026-07-04 19:53:52 +02:00
co-authored by Claude Sonnet 5
parent 274946dbff
commit 0447993181
81 changed files with 2191 additions and 86 deletions
+1 -1
View File
@@ -6,4 +6,4 @@
| CMS Frontend (cms-frontend) | ✅ Complete | unknown | Frontend, Identity (Auth), Availability | 2026-06-16 |
| Master CMS Module (master-cms-module) | ✅ Complete | unknown | Modules, Availability | 2026-06-26 |
| Tech Debt Backlog (tech-debt-backlog) | 🔵 Inception | unknown | Modules.Master, Frontend | 2026-07-01 |
| Local Dev Master/Slave Setup (local-dev-master-slave-setup) | 🟢 Construction | unknown | Modules.Master, Api, Frontend | 2026-07-02 |
| Local Dev Master/Slave Setup (local-dev-master-slave-setup) | Complete | unknown | Modules.Master, Api, Frontend | 2026-07-02 |
@@ -5,7 +5,7 @@
- **Feature Slug**: local-dev-master-slave-setup
- **Project Type**: Brownfield
- **Start Date**: 2026-07-02T00:00:00Z
- **Current Stage**: CONSTRUCTION - Unit 2 (Frontend dual-instance tooling & runbook) - Code Generation
- **Current Stage**: OPERATIONS - Placeholder (feature complete)
- **Branch**: unknown
## Workspace State
@@ -42,5 +42,5 @@
### 🟢 CONSTRUCTION PHASE
- [x] Unit 1 (Backend dual-instance hosting) — Code Generation complete (`construction/unit-1-backend-dual-instance-hosting/code/summary.md`); build + full test suite verified green; manual module-discovery check confirms slave excludes Master
- [ ] Unit 2 (Frontend dual-instance tooling & runbook) — In Progress
- [ ] Build and Test — Not started
- [x] Unit 2 (Frontend dual-instance tooling & runbook) — Code Generation complete (`construction/unit-2-frontend-dual-instance-tooling/code/summary.md`)
- [x] Build and Test — Complete (`construction/build-and-test/build-and-test-summary.md`); build succeeds, 193/193 automated tests pass, module isolation verified manually, end-to-end connection test documented as a manual follow-up (no local SQL Server available in this session)
@@ -269,3 +269,215 @@
**Context**: Code Generation Part 1 (Planning) for Unit 2 — awaiting user approval before generation.
---
## Code Generation Plan — Unit 2 — Approved
**Timestamp**: 2026-07-04T00:00:00Z
**User Input**: "Kan je de voorgaande changes committen voor me" then "Approve Plan"
**AI Response**: "Committed Unit 1's changes as 274946d prior to approval. Plan approved. Proceeding to Code Generation Part 2 (Generation) for Unit 2."
**Context**: Code Generation, Unit 2, Part 1 complete. Starting Part 2.
---
## Code Generation — Unit 2 — Complete
**Timestamp**: 2026-07-04T00:00:00Z
**User Input**: N/A (generation step)
**AI Response**: "Executed all 4 plan steps: created frontend/.env.slave.local (verified gitignored via *.local pattern), updated .env.example with slave-mode documentation, added dev:slave npm script (vite --mode slave --port 5174), added 'Lokaal Master + Slave Draaien (Dev)' runbook section to root README.md covering starting both backends, both frontends, and using the existing Add CMS Instance dialog to connect them. Wrote construction/unit-2-frontend-dual-instance-tooling/code/summary.md."
**Context**: Code Generation Part 2 complete for Unit 2 — last unit. Presenting completion message, awaiting approval before Build and Test.
---
## Code Generation — Unit 2 — Approved
**Timestamp**: 2026-07-04T00:00:00Z
**User Input**: "Continue to Next Stage naar Build and Test"
**AI Response**: "Unit 2 approved. Both units complete. Proceeding to Build and Test."
**Context**: All units complete. Starting Build and Test.
---
## Build and Test Stage
**Timestamp**: 2026-07-04T00:00:00Z
**Build Status**: Success
**Test Status**: Pass (193/193 automated; 1 manual integration scenario documented as developer follow-up — no local SQL Server available in this sandboxed session)
**Files Generated**:
- build-instructions.md
- unit-test-instructions.md
- integration-test-instructions.md
- build-and-test-summary.md
**AI Response**: "Ran dotnet build (success) and dotnet test (193/193 pass across 4 projects). Verified module isolation scenario manually (master loads 3 modules, slave loads 2, no Master). Documented the end-to-end Add CMS Instance connection scenario as a manual follow-up since no container runtime is available in this environment to provision a local SQL Server. Updated active-features.md to ✅ Complete (Operations is a placeholder for this project)."
**Context**: Build and Test complete. Presenting completion message, awaiting final approval.
---
## Build and Test — Change Requested (Recognizable Tab Titles)
**Timestamp**: 2026-07-04T00:00:00Z
**User Input**: "Geef de frontend van de slave een herkenbare html title zodta de tabbladen snel herkenbaar maken welk systeem ik bekijk"
**AI Response**: "Added VITE_APP_TITLE env var (optional, defaults to 'SlpModularCms'), wired into frontend/src/lib/config.ts (appTitle field) and set via document.title in main.tsx. Set VITE_APP_TITLE=SlpModularCms (Slave) in frontend/.env.slave.local and VITE_APP_TITLE=SlpModularCms (Master) in the developer's existing frontend/.env.local for symmetry. Updated frontend/.env.example and vite-env.d.ts. Verified: pnpm build succeeds (tsc + vite build), pnpm test 208/209 pass — the 1 failure (AddCmsInstanceDialog.test.tsx timeout) reproduced as a flaky test unrelated to this change (passed 5/5 in isolation). Updated Unit 2's code summary."
**Context**: Small addition to Unit 2 scope during Build and Test review. Re-presenting Build and Test completion message.
---
## Build and Test — Fix Requested (Slave Connection String)
**Timestamp**: 2026-07-04T00:00:00Z
**User Input**: "Ik denk dat de connectionstring niet goed is voor de Slave CMS. Maak van de appsettings.local.json.example een implementatie van de appsettings.local.json met dezelfde credentials, maar zijn eigen database"
**AI Response**: "Created the real (gitignored) src/SlpModularCms.Api.Slave/appsettings.local.json, mirroring the master's credentials (Server=127.0.0.1,1433;User ID=sa;...) with Database=SlpModularCmsSlave instead of SlpModularCms. Verified via git check-ignore it's untracked. Ran the slave with --no-build: it successfully connected to the local SQL Server, created the SlpModularCmsSlave database, and applied EF Core migrations — confirming the connection string now works (this environment does have a reachable local SQL Server at 127.0.0.1:1433, contrary to the earlier Build and Test assumption). Attempted to also run master simultaneously to redo the full end-to-end connectivity scenario, but hit 'address already in use' on both :7221 and :7222 from leftover dotnet run child processes surviving earlier timeout-wrapped manual test invocations in this session — a session artifact, not a code defect. Did not force-kill processes broadly per sandbox safety controls; flagged for the user to clear stray processes before their own manual verification. Updated Unit 1's code summary with these findings."
**Context**: Connection string issue resolved and verified via actual DB/migration success. Full simultaneous master+slave connectivity test still pending due to a leftover-process port conflict in this session, not a code issue.
---
## Port Conflict — Resolved
**Timestamp**: 2026-07-04T00:00:00Z
**User Input**: "ik heb mijn processen afgesloten. jij mag jouwe afsluiten"
**AI Response**: "Checked via Get-CimInstance for any SlpModularCms-named processes (none found) and via Get-NetTCPConnection for ports 7221/7222/5284/5285 (none listening) — no leftover processes remain on either side; ports are free."
**Context**: Environment clean. Ready for the user to perform their own manual end-to-end master+slave connectivity verification whenever convenient.
---
## Migration Status Check — Gap Found and Fixed
**Timestamp**: 2026-07-04T00:00:00Z
**User Input**: "Zijn de nodige migraties gemaakt en zijn beide databases up-to-date?"
**AI Response**: "Checked all DbContext/startup-project combinations via 'dotnet ef migrations list' (read-only). Master DB (SlpModularCms) fully up to date across ApplicationDbContext (Core/Identity), AvailabilityDbContext, and MasterDbContext. Slave DB (SlpModularCmsSlave) had AvailabilityDbContext applied (auto-migrated on startup) but ApplicationDbContext (Core/Identity) had 2 pending migrations — Core's Identity migrations are never auto-applied (only Availability and Master modules call Database.Migrate() in UseModule), and nobody had run the manual dotnet ef database update for the new slave DB yet. Fixed by running 'dotnet ef database update --project src/SlpModularCms.Core --startup-project src/SlpModularCms.Api.Slave --context ApplicationDbContext' — both migrations now applied, verified via a follow-up migrations list showing no pending entries. Added this exact command to the README runbook so it isn't missed again in the future. Updated Unit 1's code summary with the full before/after migration status table."
**Context**: Both databases now confirmed fully up to date across all applicable DbContexts. Runbook documentation updated to prevent recurrence.
---
## Fix — CMS Page Wrongly Visible on Slave
**Timestamp**: 2026-07-04T00:00:00Z
**User Input**: "De front-end van de Slave toont nog steeds de CMS-pagina. Deze is dus alleen voor de Master. De CMS-en zonder de Master-module mogen die pagina niet zien"
**AI Response**: "Dispatched an Explore subagent to map how /cms visibility currently works — found it was gated purely by Owner role (router.tsx cmsRoute, Sidebar.tsx NAV_ITEMS), with zero awareness of whether the connected backend actually has Modules.Master loaded (never an issue before this feature, since every instance always had it). Implemented a backend capability check: ModuleOrchestrator.ModuleNames property, new SystemController exposing GET /api/v1/System/capabilities, orchestrator registered as a DI singleton in both Program.cs files. Frontend: useSystemCapabilities hook, new ModuleGuard component (mirrors RoleGuard), applied to both the /cms route and the Sidebar nav item requiring the 'Master' module. Added MSW mock handler (defaults to including Master, so existing tests are unaffected) and i18n strings. Verified: dotnet build/test all pass (193/193); pnpm build succeeds; pnpm test 210/210 in isolation (one pre-existing flaky test in the full run, unrelated). Live-verified against actually running master and slave instances — curled the new endpoint directly: master returns ['Availability','Identity','Master'], slave returns ['Availability','Identity'] only. Cleaned up the two verification processes by their specific PIDs afterward. Documented in a new master-only-cms-page-fix.md under Unit 2's code folder."
**Context**: Genuine functional gap surfaced by this feature (frontend had no way to know backend module composition) now fixed and verified end-to-end, not just conceptually.
---
## Fix — CMS Instance Status Badge Styling & Master-Gate Status Not Reflected on Slave
**Timestamp**: 2026-07-04T00:00:00Z
**User Input**: "1: De badge die getoond wordt op de CMS-pagina toont beschikbaar in grijs in plaats van in groen en zonder icoontje. Laat die badge hetzelfde werken als op de availability op de dashboard. Inactive mag wel grijs zijn. 2: Als ik de status van een slave verander gebeurd er niks en er wordt ook geen call gedaan naar de slave om de status te veranderen. Was dit niet wat er was afgesproken dat er zou gebeuren als de master de status veranderd? Als ik naar de slave ga blijft die Beschikbaar ongeacht of de master hem heeft ingesteld op Niet beschikbaar. Ook na refresh en re-login"
**AI Response**: "Fix 1 (badge styling): frontend/src/components/cms/CmsInstanceList.tsx used a plain shadcn Badge (variant-based, no color/icon for Available) instead of the colored-pill + icon pattern used by AvailabilityStatusBadge on the dashboard. Replaced it with the same visual pattern: green+CheckCircle for Available, red+XCircle for NotAvailable, grey+MinusCircle for Inactive (as requested, Inactive stays grey). Fix 2 (master status not reaching slave dashboard): confirmed CmsInstanceService.UpdateStatusAsync (Modules.Master) does correctly call SlaveClient.PushStatusAsync, which does reach MasterController/MasterAvailabilityService.PushStatusAsync (Modules.Availability) on the slave and updates its in-memory MasterGateStatus. However, that master-gate status was only ever consulted by AvailabilityMiddleware to gate/503 *other* API calls — the dashboard's own availability badge queries GET /api/v1/Availability/status, which is deliberately bypassed by the middleware (so admins can still log in and see the app) and is backed by PersistentAvailabilityService, which never looked at the master gate at all. So the push from master WAS working, but its effect was invisible on the slave's own dashboard. Fixed by injecting IMasterAvailabilityService into PersistentAvailabilityService and having GetStatusDetailsAsync() return NotAvailable + the master's disableMessage whenever the master gate reports unavailable, falling back to the locally persisted status otherwise. Added a regression test (GetStatusDetailsAsync_ShouldReturnNotAvailable_WhenMasterGateDisabled) and updated two existing test call sites for the new constructor parameter. Verified: dotnet test on Modules.Availability.Tests — 61/61 pass; pnpm build succeeds; pnpm test — 210/210 pass (no flaky failures this run)."
**Context**: Fix 2 was a real functional gap in the master/slave feature's core promise (this feature exists specifically to let the master's status changes be observable on the slave) — not a new capability, but completing what 'master pushes status to slave' was always supposed to mean end-to-end.
---
## Feature — Combined `pnpm dev:all` Script for Frontend
**Timestamp**: 2026-07-04T00:00:00Z
**User Input**: "Kan je voor de front-end een script toevoegen waarmee ik in 1x master en slave kan opstarten. Beetje zoals ik in Rider de Compound configuratie heb"
**AI Response**: "Added 'dev:all' script to frontend/package.json using the new 'concurrently' devDependency: runs 'pnpm dev' (master, :5173) and 'pnpm dev:slave' (slave, :5174) in parallel in one terminal, each line prefixed and colored (blue/magenta) so output stays distinguishable — analogous to a Rider Compound run configuration. Ctrl+C stops both. Verified by actually running 'pnpm dev:all' for ~10s: both Vite dev servers started correctly on their expected ports with correctly labeled/colored output. Updated README with the new command under 'Frontend starten'."
**Context**: Pure dev-ergonomics addition, no behavior change to the app itself.
---
## Bug Report — Slave Shows Available After Restart Despite Master Having Disabled It
**Timestamp**: 2026-07-04T00:00:00Z
**User Input**: "Als ik de master en slave opnieuw opstart dan zegt de slave dat hij beschikbaar is terwijl de master aangeeft dat die onbeschikbaar zou moeten zijn. Hoe komt dat?"
**AI Response**: "Root cause: the slave's master-gate status (MasterAvailabilityService._masterIsAvailable / _masterDisableMessage) is held in static in-memory fields only — never persisted — defaulting to available=true on every process start. The master's CmsInstance.Status IS correctly persisted in its own DB, but nothing re-pushes that persisted status to the slave on startup; a push only happens when an admin explicitly changes status via the UI (CmsInstanceService.UpdateStatusAsync). The existing periodic IntegrityCheckBackgroundService (default: every 60 minutes) only re-verifies/re-registers the slave's registered master URL in VerifyIntegrityAsync — it never re-pushed status, so even the periodic reconciliation wouldn't have caught this within a reasonable time. Fixed both halves: (1) CmsInstanceService.VerifyIntegrityAsync now also re-pushes the master's persisted status (SlaveClient.PushStatusAsync) to every active slave on each integrity check, so a slave that resets to 'available' on restart gets corrected; (2) IntegrityCheckBackgroundService.ExecuteAsync now runs one check immediately on startup instead of waiting for the first PeriodicTimer tick, so this resync happens right after the master (re)starts rather than up to 60 minutes later. Added a regression test (VerifyIntegrityAsync_RePushesPersistedStatus_ToResyncSlaveAfterRestart). Verified: dotnet test on Modules.Master.Tests — 43/43 pass (was 42)."
**Context**: This is the same underlying gap as the earlier CMS-badge/master-gate fix (in-memory state on the slave not being kept in sync with the master's persisted intent), surfacing again specifically around process restarts rather than live status pushes.
---
## Question — Refresh Tokens Overwritten Between Master and Slave Frontends
**Timestamp**: 2026-07-04T00:00:00Z
**User Input**: "Ik kan niet op beide systemen ingelogd zijn. De refresh tokens wordne overschreven. Is dat iets vanwege het lokaal draaien? [...] ik gok dat dat aan de hand van domein is wat lokaal allebei localhost is, klopt dat?"
**AI Response**: "Confirmed the user's hypothesis. AuthController.GetCookieOptions() sets the refreshToken cookie without an explicit Domain, so it's host-scoped to whichever host issued it — but per RFC 6265, cookies are never port-scoped, only domain+path. Master (localhost:7221) and slave (localhost:7222) both set 'refreshToken' for domain 'localhost', path '/api/v1/auth' — identical cookie key, so the last login always overwrites the other. In production this isn't an issue since each CMS instance has its own real domain (properly separated by the browser's cookie jar); this is purely a local-dev artifact of both instances happening to share the 'localhost' hostname. Offered two workarounds (hosts-file entries like master.local/slave.local for a production-like separation, or simply using separate browser profiles/incognito for master vs slave) without recommending any code change, since this isn't a defect."
**Context**: User confirmed the incognito-workaround is sufficient; no code change made. Not a bug — purely local-dev cookie-domain behavior, explicitly ruled out as something to fix in the app.
---
## Fix — Slave Stuck on Last Pushed Status After Master Sets It to Inactive
**Timestamp**: 2026-07-04T00:00:00Z
**User Input**: "Wat ik wel zie is dat de slave op 'Niet Beschikbaar' blijft staan als ik hem in de Master van 'Niet beschikbaar' naar 'Inactief zet'. Als de Master de Slave niet meer beheert moet de Slave gewoon weer Beschikbaar worden. Zat er niet al een systeem in dat Slave bij Master kijkt wat de status is? Daar zouden we het bij kunnen zetten. Ik weet ook niet of dat een periodieke check is?"
**AI Response**: "Confirmed there is no slave-polls-master mechanism at all — the master/slave protocol is entirely push-based from the master side (explicit status pushes on UpdateStatusAsync, plus the periodic IntegrityCheckBackgroundService discussed in the prior fix). Root cause of this specific report: CmsInstanceService.UpdateStatusAsync had an early-return for CmsInstanceStatus.Inactive that skipped SlaveClient.PushStatusAsync entirely ('master no longer manages this slave' was implemented as 'do nothing'), so the slave kept whatever gate state (e.g. NotAvailable) it was last pushed, forever — and since VerifyIntegrityAsync's periodic resync only iterates GetActiveAsync() (which excludes Inactive instances), it would never self-correct either. Fixed by having the Inactive branch explicitly push isAvailable=true, disableMessage=null to the slave — releasing the master gate — before returning, mirroring the same release-on-deactivation logic. Updated the existing UpdateStatusAsync_DoesNotPushToSlave_WhenStatusIsInactive test (renamed/inverted to UpdateStatusAsync_ReleasesMasterGate_WhenStatusIsInactive) and added UpdateStatusAsync_ReturnsSlaveContactFalse_WhenReleasingMasterGateFails. Verified: dotnet test on Modules.Master.Tests — 44/44 pass (was 43)."
**Context**: Third fix in this feature's 'slave in-memory master-gate state must stay in sync with the master's actual intent' theme — this one specifically for the deactivation/detach path, which had simply never pushed anything to the slave at all.
---
## Feature — Slave Pull-Based Status Polling with Fail-Open
**Timestamp**: 2026-07-04T00:00:00Z
**User Input**: "volgens mij was het idee dat de slave ook af en toe bij de master vroef wat de status is zodat lokaal niet de boel zomaar kan worden aangepast en dat de slave af en toe een refreh doet van de status. ook weer om dit soort issues tegen te gaan. Als de master te lang onbeschikbaar is of gewoon niet reageert moet de slave beschikbaar worden zodat het een fail-open is"
**AI Response**: "Dispatched an Explore subagent to check whether this was ever actually specified. Confirmed: the original master-cms-module inception requirements (FR-MASTER-06 'Slave Pull Model: Periodic Master Check', FR-MASTER-07 'Slave Fallback Behavior', NFR-MASTER-01 'Fail-Open Safety') explicitly called for a slave-initiated periodic pull with fail-open, but the later construction-phase design silently switched to a push-only model (master calls slave), keeping only the in-memory fail-open *default* (true on process start) rather than an active reconciling pull. Implemented the originally-specified pull side to complement (not replace) the existing push: (1) new master-side GET /api/v1/SlaveStatus endpoint (SlaveStatusController + ICmsInstanceService.GetStatusForApiKeyAsync), authenticated by matching the caller's plain API key against each active CmsInstance's decrypted key — added to AvailabilityMiddleware's bypass list so it's always reachable regardless of the master's own local status; (2) slave-side MasterStatusPollingBackgroundService (Modules.Availability), polling on a configurable interval (MasterPollingOptions, default 30s / 15s in dev) via new IMasterStatusPollClient, applying successful results through new IMasterAvailabilityService methods (GetPollTargetAsync, ApplyPolledStatusAsync); (3) fail-open: RecordPollFailureAsync forces the gate back to Available if the master has been unreachable for longer than FailOpenAfterMinutes (default 5min / 2min in dev), measured from MasterRegistration.LastPolledAt (new persisted field, new EF migration AddLastPolledAtToMasterRegistration, applied to both master and slave DBs). The existing push mechanism is untouched and still fires instantly on explicit status changes; polling is the self-healing safety net for everything push can miss (restarts, dropped pushes, local tampering). Verified: dotnet test across the whole solution — 216/216 pass (75 Availability + 50 Master + 54 Core + 37 Identity). Live smoke-tested by starting both master and slave: confirmed the new background service starts without crashing the host, correctly detects a connection failure (mismatched port in this quick ad-hoc run, not a code issue) and handles it gracefully through the fail-open path rather than an unhandled exception."
**Context**: This closes a real design regression: the pull+fail-open behavior was part of the original spec but got dropped during construction without anyone flagging it, and its absence was the root cause behind several of today's earlier reports (slave stuck on stale status after restart, after deactivation). Push and pull now both exist and complement each other.
---
## Fix — Slave Settings Silently Accepted (and 'Succeeded') Local Availability Changes While Master-Controlled
**Timestamp**: 2026-07-04T00:00:00Z
**User Input**: "Als de Slave op 'Niet beschikbaar' is gezet door de master en je probeert in de instellingen van de slave het systeem op beschikbaar te zetten dan krijg je wel een succesmelding. dat moet een adere mrlding worden, maar eigenlijk moet je helemaal de beschikbaarheid niet moeten kunnen veranderen in de instellingen en je moet daar dan een melding zien waarom"
**AI Response**: "Root cause: PersistentAvailabilityService.UpdateStatusAsync always wrote the requested status to the local DB unconditionally, while GetStatusDetailsAsync (fixed earlier today) displays the master-gate status whenever the master has disabled the instance — so a local admin's 'Available' write was accepted and even reported as a DB update, but instantly invisible/overridden on the very same GET that renders the Settings page, giving a false success toast for a change that had no observable effect. Fixed by rejecting the change outright instead of silently no-op'ing it: (1) new MasterControlledAvailabilityException (Core.Availability), thrown by PersistentAvailabilityService.UpdateStatusAsync when IMasterAvailabilityService.GetMasterStatus().IsAvailable is false, before touching the DB; (2) AvailabilityController.UpdateStatus catches it and returns 409 Conflict with a ProblemDetails body; (3) AvailabilityStatusDetails gained an IsMasterControlled flag (true whenever GetStatusDetailsAsync short-circuits on the master gate), now included in GET /api/v1/Availability/status; (4) frontend SettingsPage.tsx reads availability.isMasterControlled and disables the mode buttons, the reason textarea, and the save button, showing a locked-state banner explaining that the Master CMS controls this status; a 409 from the mutation (e.g. a race where master disables it mid-edit) shows a specific 'Master CMS controls this status' toast instead of the generic error. Verified: dotnet test — 219/219 pass across the solution (78 Availability, up from 75); pnpm build succeeds; pnpm test — 213/213 pass (up from 210, 3 new SettingsPage tests covering the locked banner, the unlocked default state, and the 409 toast)."
**Context**: Same root theme as the day's earlier fixes (slave-side local state must never silently diverge from the master's actual intent) — this time the gap was a misleading UI: an operation appeared to succeed while being fully overridden, which is worse than an operation that visibly fails.
---
## Final Verification — README Addition + Full AI-DLC Documentation Cross-Check
**Timestamp**: 2026-07-04T00:00:00Z
**User Input**: "ja zet het erbij en check alles ook tegenover de aidlc documentatie" (following "Volgens mij werkt alles nu zoals verwacht. Doe nog 1 check of functionaliteit overeenkomt met de documentatie en of de tests slagen" and a prior note that the master-controlled-lock UI wasn't mentioned in the README)
**AI Response**: "Added a README section documenting the master-controlled Settings lock (banner, disabled controls, 409 on a bypass attempt, IsMasterControlled field). Then cross-checked the actual master/slave protocol documentation in aidlc-docs/features/master-cms-module/ (the feature that originally built the protocol this session's fixes touched) against the real as-built code — full details logged in that feature's own audit.md, since the corrections landed in its docs, not this feature's. Summary: found and fixed several actively-contradicted statements about Inactive-transition push behavior, VerifyIntegrityAsync's scope, and the master-gate cache's 'no expiry' claim; added new rule-set/flow sections for the slave-pull and master-controlled-lock mechanisms; flagged (via banners, not full rewrite) a larger pre-existing divergence between inception-stage application-design docs and what construction actually built, unrelated to today's changes. Also caught and fixed a real (if minor) config-consistency gap of my own: the new MasterPolling section had only been added to appsettings.Development.json, not the base appsettings.json for either Api or Api.Slave — added there too for both. Re-ran the full solution build + test suite after all changes: 219/219 backend, 213/213 frontend, all green."
**Context**: Closing verification pass for today's whole run of fixes (badge styling, master-controlled availability, slave restart resync, Inactive gate release, slave-pull/fail-open, master-controlled settings lock) — confirms the implementation, its tests, and its documentation are now mutually consistent.
---
@@ -0,0 +1,41 @@
# Build and Test Summary — Local Dev Master/Slave Setup
## Build Status
- **Build Tool**: .NET 10 SDK (`dotnet build`)
- **Build Status**: ✅ Success
- **Build Artifacts**: `src/SlpModularCms.Api.Slave/bin/` (new), plus unchanged outputs for all existing projects
- **Build Time**: ~5 seconds (incremental)
## Test Execution Summary
### Unit Tests
- **Total Tests**: 193 (54 + 60 + 37 + 42 across `Core.Tests`, `Modules.Availability.Tests`, `Modules.Identity.Tests`, `Modules.Master.Tests`)
- **Passed**: 193
- **Failed**: 0
- **Coverage**: Unchanged from before this feature (no new business logic; relocated code retains its existing tests)
- **Status**: ✅ Pass
### Integration Tests
- **Test Scenarios**: 2 (module isolation; end-to-end master/slave connection via existing Add CMS Instance flow)
- **Passed**: 1 (module isolation — verified via manual `dotnet run` of both hosts, confirmed via log output: master loads 3 modules, slave loads exactly 2, Master excluded)
- **Failed**: 0
- **Not Run**: 1 (end-to-end connection scenario — requires a local SQL Server instance; no container runtime available in this sandboxed session. Documented as a manual step for the developer in `integration-test-instructions.md` and the root `README.md` runbook.)
- **Status**: ⚠️ Partial — the scenario this feature actually changes (module isolation) is fully verified; the scenario exercising the pre-existing, unmodified master/slave protocol requires manual follow-up.
### Performance Tests
- **Status**: N/A — this feature is local developer tooling with no performance requirements (NFR-3, local-only scope).
### Additional Tests
- **Contract Tests**: N/A — no API contracts changed
- **Security Tests**: N/A — no new attack surface; `appsettings.local.json` secrets-hygiene pattern (NFR-4) followed and verified via `git check-ignore`
- **E2E Tests**: See Integration Tests Scenario 2 above
## Overall Status
- **Build**: ✅ Success
- **All Automated Tests**: ✅ Pass (193/193)
- **Manual Follow-up Required**: Yes — developer should run the end-to-end connection test (root `README.md`, "Lokaal Master + Slave Draaien (Dev)") at least once with a real local SQL Server to confirm the full workflow before relying on it.
- **Ready for Operations**: Yes (Operations is a placeholder for this project; no deployment/monitoring work applicable to local dev tooling)
## Next Steps
- Developer runs the manual end-to-end verification (Scenario 2) locally to close out the one remaining unverified item.
- Feature is otherwise complete: `SlpModularCms.Api.Slave` exists and correctly excludes the Master module, `SlpModularCms.Core.Hosting` avoids code duplication between the two hosts, frontend tooling and documentation are in place.
@@ -0,0 +1,32 @@
# Build Instructions — Local Dev Master/Slave Setup
## Prerequisites
- **Build Tool**: .NET 10 SDK
- **Dependencies**: NuGet packages restore automatically on build (no new external dependencies introduced by this feature)
- **Environment Variables**: None required to build (runtime config is via `appsettings.local.json`, see below)
- **System Requirements**: Same as the rest of the repository — no new system requirements
## Build Steps
### 1. Restore & Build
```powershell
dotnet build SlpModularCms.sln
```
### 2. Verify Build Success
- **Expected Output**: `Build succeeded.` with 0 errors (NuGet advisory warnings for `Microsoft.OpenApi` and a few `NU1510`/prune warnings are pre-existing and unrelated to this feature).
- **Build Artifacts**: `src/SlpModularCms.Api/bin/`, `src/SlpModularCms.Api.Slave/bin/` (new), plus all existing project outputs.
- **Common Warnings**: NU1903 (Microsoft.OpenApi advisory) and NU1510 (package pruning) appear across multiple projects — pre-existing, not introduced by this feature.
### Actual Result (this session)
Ran `dotnet build SlpModularCms.sln`**Build succeeded**, 0 errors.
## Troubleshooting
### Build Fails with `CS0246` in `SlpModularCms.Core/Hosting/*.cs`
- **Cause**: `SlpModularCms.Core` is a plain `Microsoft.NET.Sdk` project (not `Sdk.Web`), so ASP.NET Core implicit usings (`Microsoft.Extensions.DependencyInjection`, `Microsoft.Extensions.Configuration`, `Microsoft.AspNetCore.Builder`, `Microsoft.AspNetCore.Http`) aren't automatically available like they are in `SlpModularCms.Api`.
- **Solution**: Already fixed during Code Generation — explicit `using` statements were added to `ServiceCollectionExtensions.cs`. If this recurs after further edits, add the missing explicit `using`.
### `SlpModularCms.Api.Slave` fails to start with a SQL connection error
- **Cause**: Missing `src/SlpModularCms.Api.Slave/appsettings.local.json` (gitignored, must be created locally per developer).
- **Solution**: Copy `appsettings.local.json.example` to `appsettings.local.json` and fill in your local SQL Server credentials, using a **different** `Database=` name than the master instance (e.g. `SlpModularCmsSlave`).
@@ -0,0 +1,52 @@
# Integration Test Instructions — Local Dev Master/Slave Setup
## Purpose
Verify that the two backend instances (Unit 1) and frontend tooling (Unit 2) work together correctly: the slave instance genuinely excludes the Master module, and the existing master↔slave connection mechanism (unchanged by this feature) can be exercised end-to-end using the two local instances.
## Test Scenarios
### Scenario 1: Module isolation — Slave excludes Master, Master keeps all modules
**Description**: Confirm `SlpModularCms.Api.Slave`'s `ModuleOrchestrator` never discovers `Modules.Master`, while `SlpModularCms.Api` is unaffected.
**Setup**: None beyond a successful build (no database required — module discovery happens before any DB access).
**Test Steps** (already executed in this session):
```powershell
dotnet run --project src/SlpModularCms.Api --launch-profile https --no-build
dotnet run --project src/SlpModularCms.Api.Slave --launch-profile https --no-build
```
**Expected Results**:
- Master log output: `Module ontdekt: Availability`, `Module ontdekt: Identity`, `Module ontdekt: Master`, `3 modules succesvol geladen.`
- Slave log output: `Module ontdekt: Availability`, `Module ontdekt: Identity`, `2 modules succesvol geladen.`**no** `Master` line.
**Actual Result**: ✅ **Passed** — confirmed exactly as expected in this session's console output (see Unit 1 code generation summary).
**Cleanup**: Stop both processes (Ctrl+C / process termination).
### Scenario 2: End-to-end master/slave connection via the existing "Add CMS Instance" flow
**Description**: With both instances running against separate local databases, use the master frontend's existing "Add CMS Instance" dialog to register the local slave and confirm the connection is established (per FR-4 and the runbook in root `README.md`).
**Setup**:
1. A local SQL Server instance reachable from both backends (e.g. via the `podman run ... mcr.microsoft.com/mssql/server` command in root `README.md`).
2. `src/SlpModularCms.Api/appsettings.local.json` (master) and `src/SlpModularCms.Api.Slave/appsettings.local.json` (slave, from `appsettings.local.json.example`) pointing at **different** database names on that SQL Server.
3. Master and slave backends running (`dotnet run --project src/SlpModularCms.Api --launch-profile https` and `dotnet run --project src/SlpModularCms.Api.Slave --launch-profile https`).
4. Master frontend running (`pnpm dev` in `frontend/`), logged in as an Owner.
**Test Steps**:
1. Navigate to the `/cms` page on the master frontend.
2. Use "Add CMS Instance" with URL `https://localhost:7222` (the local slave).
3. Observe the instance's status in the UI.
**Expected Results**: The instance appears in the list and its status reflects a successful connection (per the existing, unchanged `CmsInstanceService`/`SlaveApiClient``MasterController` protocol documented in root `README.md`'s "Master CMS Module" section).
**Actual Result**: ⚠️ **Not run in this session** — this sandboxed environment has no accessible container runtime (`docker`/`podman` both unavailable), so no local SQL Server instance could be provisioned to run either backend past module discovery. **This step requires manual execution by the developer** following the runbook in root `README.md` ("Lokaal Master + Slave Draaien (Dev)"). Scenario 1 (module isolation, which does not require a database) was fully verified automatically and is the aspect this feature actually changes — Scenario 2 exercises the pre-existing, unmodified master/slave protocol and primarily validates that the new run configuration (ports, separate databases, CORS) doesn't get in the way of it.
**Cleanup**: Remove the CMS instance registration if desired; stop both backends and frontends; stop the SQL Server container if it was started solely for this test.
## Notes
- No automated integration test suite was added for Scenario 2 — it's an inherently manual, cross-process, cross-database verification of local developer tooling, not a candidate for CI automation (per NFR-3, this feature is explicitly local-only in scope).
@@ -0,0 +1,29 @@
# Unit Test Execution — Local Dev Master/Slave Setup
## Run Unit Tests
### 1. Execute All Unit Tests
```powershell
dotnet test SlpModularCms.sln
```
### 2. Review Test Results
- **Expected**: All existing test suites pass unchanged — this feature adds no new business logic, so no new unit tests were written (per Unit of Work Q3 = A).
- **Test Coverage**: Unchanged from before this feature (the relocated `ModuleOrchestrator`/`ApiPrefixConvention` classes retain their existing test coverage, now in `SlpModularCms.Core.Tests/Hosting/` instead of `SlpModularCms.Modules.Identity.Tests/Infrastructure/`).
- **Test Report Location**: Console output from `dotnet test`; no separate report file generated by default.
### Actual Result (this session)
Ran `dotnet test SlpModularCms.sln`:
| Test Project | Passed | Failed | Skipped |
|---|---|---|---|
| `SlpModularCms.Core.Tests` (incl. relocated `Hosting` tests) | 54 | 0 | 0 |
| `SlpModularCms.Modules.Availability.Tests` | 60 | 0 | 0 |
| `SlpModularCms.Modules.Identity.Tests` | 37 | 0 | 0 |
| `SlpModularCms.Modules.Master.Tests` | 42 | 0 | 0 |
| **Total** | **193** | **0** | **0** |
No regressions from relocating `ModuleOrchestrator`, `ServiceCollectionExtensions`, and `ApiPrefixConvention` into `SlpModularCms.Core`, or from moving their tests into `Core.Tests`.
### 3. Fix Failing Tests
Not applicable this run — all tests passed on first execution after Code Generation.
@@ -10,14 +10,14 @@
## Steps
- [ ] **Step 1 — Frontend env files for slave mode**
- [x] **Step 1 — Frontend env files for slave mode**
- Create `frontend/.env.slave.local` (gitignored via existing `frontend/.gitignore` `*.local` pattern — verified via `git check-ignore`): `VITE_API_BASE_URL=https://localhost:7222`
- Modify `frontend/.env.example`: add a second documented block showing the slave-mode value, alongside the existing master-mode `VITE_API_BASE_URL` example
- [ ] **Step 2 — `dev:slave` npm script**
- [x] **Step 2 — `dev:slave` npm script**
- Modify `frontend/package.json`: add `"dev:slave": "vite --mode slave --port 5174"` to the `scripts` section. Vite's mode-based env loading will load `.env.slave.local` when run with `--mode slave` (Vite loads `.env.[mode].local` in addition to `.env.local`; since both files would apply, and `.env.local` takes precedence per Vite's env-file priority for the same key when both exist for a mode, name the slave file `.env.slave.local` specifically — this file only loads when `--mode slave` is passed, so there is no conflict with the default `.env.local` used by `pnpm dev`)
- [ ] **Step 3 — Runbook documentation**
- [x] **Step 3 — Runbook documentation**
- Modify root `README.md`: add new section **"Lokaal Master + Slave Draaien (Dev)"** immediately after the existing "Master CMS Module" section, covering:
1. Starting the master backend (`dotnet run --project src/SlpModularCms.Api --launch-profile https`)
2. Starting the slave backend (`dotnet run --project src/SlpModularCms.Api.Slave --launch-profile https`), noting it needs its own `appsettings.local.json` (from `appsettings.local.json.example`) with a separate local database
@@ -25,7 +25,7 @@
4. Using the existing "Add CMS Instance" dialog on the master frontend to register the slave (URL `https://localhost:7222`) and confirm it shows as connected/healthy
5. Cross-reference to this feature's requirements doc for anyone wanting the full rationale
- [ ] **Step 4 — Documentation summary**
- [x] **Step 4 — Documentation summary**
- Create `aidlc-docs/features/local-dev-master-slave-setup/construction/unit-2-frontend-dual-instance-tooling/code/summary.md` documenting what was created/modified
## Notes
@@ -43,3 +43,27 @@
- **Master** (`SlpModularCms.Api`): discovers and loads all 3 modules — Availability, Identity, Master.
- **Slave** (`SlpModularCms.Api.Slave`): discovers and loads exactly 2 modules — Availability, Identity. **Master is correctly excluded.**
- Full end-to-end run (requiring a real local SQL Server/localdb instance and manual "Add CMS Instance" registration) is deferred to Build and Test / Unit 2, per the plan.
## Follow-up: Real `appsettings.local.json` for the Slave (user request)
The user reported the slave's connection string looked wrong and asked for a real `src/SlpModularCms.Api.Slave/appsettings.local.json` (gitignored, mirroring `src/SlpModularCms.Api/appsettings.local.json`'s credentials but with its own database). Created with `Database=SlpModularCmsSlave` (vs. master's `SlpModularCms`), same `Server=127.0.0.1,1433;User ID=sa;Password=...` credentials and same `JwtSettings.Secret`. Verified via `git check-ignore` that it's not tracked.
**Verified working**: running `dotnet run --launch-profile https --no-build` in `SlpModularCms.Api.Slave` with this file present successfully connected to the local SQL Server, created the `SlpModularCmsSlave` database, and applied EF Core migrations (`CREATE DATABASE`, `__EFMigrationsHistory` setup, migration application) — confirming the connection string is correct. This environment does have a reachable SQL Server at `127.0.0.1:1433`, unlike assumed earlier in Build and Test.
**Known issue hit during this verification, not related to the fix**: a subsequent attempt to run both master and slave simultaneously hit `Failed to bind to address ... address already in use` on both `:7221` and `:7222` — leftover `dotnet run` child processes from earlier manual verification steps in this session likely survived their parent `timeout` calls and are still holding those ports. This is a session/environment artifact, not a defect in the generated code. Resolved: user closed their own processes and confirmed via `Get-CimInstance`/`Get-NetTCPConnection` that no `SlpModularCms` processes or listeners remained on ports 7221/7222/5284/5285.
## Follow-up: Slave Database Migration Gap (discovered while checking migration status)
User asked whether the necessary migrations exist and both databases are up to date. Checked via `dotnet ef migrations list` for every `DbContext`/startup-project combination:
| Database | Context | Status before fix |
|---|---|---|
| `SlpModularCms` (master) | `ApplicationDbContext` (Core/Identity) | ✅ Applied (2/2) |
| `SlpModularCms` (master) | `AvailabilityDbContext` | ✅ Applied (1/1) |
| `SlpModularCms` (master) | `MasterDbContext` | ✅ Applied (1/1) |
| `SlpModularCmsSlave` (slave) | `AvailabilityDbContext` | ✅ Applied (1/1) — auto-migrated via `Database.Migrate()` in `AvailabilityModule.UseModule` |
| `SlpModularCmsSlave` (slave) | `ApplicationDbContext` (Core/Identity) | ❌ **2 migrations pending**`SlpModularCms.Core`'s `ApplicationDbContext` is never auto-migrated (only `Modules.Availability` and `Modules.Master` call `Database.Migrate()` in their `UseModule`); it always requires the manual `dotnet ef database update` step documented in root `README.md`, and nobody had run it yet for the new slave database.
**Fixed**: ran `dotnet ef database update --project src/SlpModularCms.Core --startup-project src/SlpModularCms.Api.Slave --context ApplicationDbContext`. Re-checked with `migrations list` — both `20260612191736_InitialCreate` and `20260619130625_AddsDisplayName` now show as applied (no `(Pending)` marker). Slave database is now fully up to date.
**Documentation fix**: added a note + the exact command to the "Lokaal Master + Slave Draaien (Dev)" runbook in root `README.md`, right after the slave-startup instructions, so this doesn't get missed again by whoever (re)creates the slave database.
@@ -0,0 +1,32 @@
# Fix — CMS Page Was Visible on Slave (No Backend Capability Check)
## Problem
The `/cms` page (managing registered CMS instances — an Owner-only feature backed by `SlpModularCms.Modules.Master`) was gated purely by role (`Owner`), with no awareness of whether the *connected backend* actually has the Master module loaded. Before this feature, this was never an issue — every deployed instance always had `Modules.Master` loaded. Now that a Master-less slave instance exists, an Owner using the frontend against the slave could still see the nav link and open `/cms`, where its data calls (`GET /api/v1/CmsInstances`) would 404 against a backend that has no such controller.
## Fix
### Backend
- `SlpModularCms.Core.Hosting.ModuleOrchestrator` — added `ModuleNames` (public `IReadOnlyList<string>`), listing the names of modules actually discovered on this instance.
- New `SlpModularCms.Core.Hosting.SystemController``GET /api/v1/System/capabilities` returns `{ "modules": [...] }` for whichever instance is asked.
- `SlpModularCms.Api/Program.cs` and `SlpModularCms.Api.Slave/Program.cs` — registered the `ModuleOrchestrator` instance itself as a DI singleton (`builder.Services.AddSingleton(orchestrator)`) so the new controller can inject it.
### Frontend
- `src/api/types.ts` — added `SystemCapabilities { modules: string[] }`.
- `src/api/useSystemCapabilities.ts` — new React Query hook (`staleTime: Infinity` — a backend's module set never changes mid-session), calling `/api/v1/System/capabilities`.
- `src/components/auth/ModuleGuard.tsx` — new guard component (mirrors `RoleGuard`), hides its children and shows a "not available on this instance" message when the required module isn't in the backend's capability list.
- `src/router.tsx``/cms` route now wraps its page in `<ModuleGuard requiredModule="Master">` (inside the existing `<RoleGuard allowedRoles={['Owner']}>`).
- `src/components/layout/Sidebar.tsx` — the CMS nav item now also requires `capabilities.modules.includes('Master')` before rendering.
- `src/i18n/locales/{nl,en}/translation.json` — added `errors.featureUnavailableTitle` / `errors.featureUnavailable`.
- `src/mocks/system/handlers.ts` — new MSW handler for `*/System/capabilities`, defaulting to `['Availability', 'Identity', 'Master']` so existing CMS-related tests keep passing unchanged; registered in `src/mocks/index.ts`.
- `src/components/layout/Sidebar.test.tsx` — updated the "Owner sees ... CMS" assertion to `findByTestId` (now async, since nav-cms visibility depends on the capabilities fetch), and added a new regression test: "Owner does not see CMS when the backend has no Master module (slave instance)".
## Verification Performed
- `dotnet build` / `dotnet test` — succeed; all 193 backend tests still pass (module relocation from the earlier fix is untouched; this only adds new code).
- `pnpm build` — succeeds, no type errors.
- `pnpm test` — 210/210 pass in isolation (209/210 in the full parallel run; the 1 failure is the same pre-existing flaky `AddCmsInstanceDialog.test.tsx` timeout seen earlier in this feature's Build and Test, unrelated to this change — confirmed passing 5/5 when re-run alone).
- **Live verification against real running instances**: started both `SlpModularCms.Api` (master) and `SlpModularCms.Api.Slave` against the local SQL Server and curled the new endpoint directly:
- Master: `curl https://localhost:7221/api/v1/System/capabilities``{"modules":["Availability","Identity","Master"]}`
- Slave: `curl https://localhost:7222/api/v1/System/capabilities``{"modules":["Availability","Identity"]}`
- Confirms the capability check reflects each instance's actual loaded modules, not just a hardcoded assumption.
@@ -0,0 +1,28 @@
# Code Generation Summary — Unit 2: Frontend Dual-Instance Tooling & Runbook
## Created
- `frontend/.env.slave.local``VITE_API_BASE_URL=https://localhost:7222`, `VITE_APP_TITLE=SlpModularCms (Slave)` (gitignored via existing `frontend/.gitignore` `*.local` pattern, verified via `git check-ignore`)
- `aidlc-docs/features/local-dev-master-slave-setup/construction/unit-2-frontend-dual-instance-tooling/code/summary.md` (this file)
## Modified
- `frontend/.env.example` — added a documented block explaining how to create `.env.slave.local` and use `pnpm dev:slave`, and how `VITE_APP_TITLE` distinguishes tabs
- `frontend/.env.local` (developer's existing local file, gitignored) — added `VITE_APP_TITLE=SlpModularCms (Master)` for symmetry with the slave
- `frontend/package.json` — added `"dev:slave": "vite --mode slave --port 5174"` script
- `frontend/src/vite-env.d.ts` — added optional `VITE_APP_TITLE` to the typed env interface
- `frontend/src/lib/config.ts` — added `appTitle` to `AppConfig`, sourced from `VITE_APP_TITLE` with a `'SlpModularCms'` fallback when unset
- `frontend/src/main.tsx` — sets `document.title` from `getAppConfig().appTitle` at startup, so each instance's browser tab is recognizable
- `README.md` (root) — added new section **"Lokaal Master + Slave Draaien (Dev)"** after "Master CMS Module", covering: starting both backends, starting the frontend against either instance, and using the existing "Add CMS Instance" dialog to connect them
## Notes
- No new tests — this unit is env/config/documentation only, no testable logic.
- `pnpm dev:slave` uses Vite's `--mode slave` flag, which loads `.env.slave.local` in addition to the default `.env`/`.env.local` files; `--port 5174` overrides `vite.config.ts`'s default port (5173) for this invocation only.
- **Added after initial Build and Test review** (user request): distinguishable browser tab titles per instance via `VITE_APP_TITLE`, defaulting to `"SlpModularCms"` so existing setups without the var keep today's title unchanged.
- Full end-to-end verification (starting both backends with real local databases, running both frontends, and confirming the "Add CMS Instance" flow actually connects them) is performed in Build and Test, since it requires a running local SQL Server instance that isn't available in this automated environment.
## Verification Performed (Tab Title Change)
- `pnpm build` (tsc -b + vite build) — succeeds, no type errors.
- `pnpm test` — 208/209 pass; the 1 failure (`AddCmsInstanceDialog.test.tsx`, a findByTestId timeout) reproduced as flaky under this session's load and passed 5/5 when re-run in isolation — unrelated to the title change (no title-related assertions, and `main.tsx` is not exercised by component tests).
@@ -1,5 +1,17 @@
# Audit Log — Master CMS Module
## Documentation Audit & Correction — Post-`local-dev-master-slave-setup` Follow-Up Fixes
**Timestamp**: 2026-07-04T00:00:00Z
**User Input**: "ja zet het erbij en check alles ook tegenover de aidlc documentatie" (in response to a request to also add README docs for a new master-controlled availability lock, and to cross-check everything against the AI-DLC documentation)
**AI Response**: "Dispatched an Explore subagent to compare the functional-design docs here against the actual as-built code, given a string of same-day fixes (in the local-dev-master-slave-setup feature's session) that changed real master/slave protocol behavior: Inactive now releases the gate instead of doing nothing, VerifyIntegrityAsync now re-pushes status (not just URL-verification) and IntegrityCheckBackgroundService runs once on startup, a new slave-pull mechanism (SlaveStatusController + MasterStatusPollingBackgroundService + fail-open via MasterRegistration.LastPolledAt) was added, and PersistentAvailabilityService now merges/enforces the master gate (IsMasterControlled, MasterControlledAvailabilityException, 409 Conflict) instead of ignoring it. Found and corrected multiple actively-contradicted statements: construction/master-backend/functional-design/{business-rules,business-logic-model,domain-entities}.md (Inactive-transition push behavior, SlaveContactSuccess semantics, VerifyIntegrityAsync flow, added new Slave Pull and Master-Controlled-Lock rule sections); construction/slave-availability-extension/functional-design/{business-rules,business-logic-model,domain-entities}.md (master-gate cache no longer 'no expiry', added LastPolledAt field, added Rule Sets 5-6 and Flows 5-6 for the pull/fail-open and lock mechanisms); inception/application-design/application-design.md (Inactive/fail-open constraint rows). Also fixed one long-standing (pre-2026-07-04, unrelated to today) inaccuracy found in passing: slave-availability-extension/domain-entities.md claimed MasterRegistration.ApiKey is stored plain-text, but the code (MasterAvailabilityService using IMasterApiKeyProtector) actually encrypts it. Separately flagged (added superseded-warning banners rather than rewriting) a much larger, pre-existing divergence in inception/application-design/{services,components,component-methods}.md: these inception-stage docs describe an entirely different pull-based design (static cache + MasterModuleOptions.CacheMinutes + /api/internal/master/* routes) that construction never actually built — this predates today's session and is a separate, bigger gap than what today's fixes caused; not fully rewritten, just clearly marked as historical/superseded pending a dedicated pass. Also added a code-summary.md addendum for master-backend listing the new SlaveStatusController/methods. Re-ran the full backend suite after all doc edits (docs-only + two appsettings.json additions) to confirm nothing was inadvertently broken: 219/219 pass."
**Context**: This is a documentation-only audit and correction pass — no application code was changed (only aidlc-docs/*.md files and two appsettings.json files that added a previously-Development-only MasterPolling config section to the base/production configs for consistency, functionally a no-op since the C# options class already had matching defaults).
---
## Build and Test Stage — Approved
**Timestamp**: 2026-07-01T00:10:00Z
@@ -61,3 +61,14 @@ This generates the `Migrations/` folder contents. The migration is applied autom
- `Microsoft.Extensions.Http.Resilience` version `9.6.0` — verify/update during `dotnet restore` if a newer version is available for .NET 10
- `IntegrityCheckIntervalMinutes = 0` in tests forces immediate PeriodicTimer ticks (valid for test scenarios only)
- Slave-side endpoints (`/api/v1/master/register`, `/api/v1/master/status`, `/api/v1/master/registered-url`) are implemented in Unit 2 (slave-availability-extension)
## Addendum — 2026-07-04 (added outside this unit's original scope, in `local-dev-master-slave-setup` follow-up fixes)
This unit predates the following; see `slave-availability-extension/functional-design/business-rules.md` Rule Set 5 and `application-design/application-design.md` for the full picture:
| File | Description |
|------|-------------|
| `Controllers/SlaveStatusController.cs` | New. `[AllowAnonymous]` `GET /api/v1/SlaveStatus`; authenticates via `X-Master-Api-Key` header matched against each active `CmsInstance`'s decrypted key; lets a slave pull its own status instead of relying solely on the master's push |
| `Services/ICmsInstanceService.cs` / `CmsInstanceService.cs` | `GetStatusForApiKeyAsync(plainApiKey)` added; `UpdateStatusAsync`'s `Inactive` branch now pushes `Available`/null to release the gate (previously a no-op); `VerifyIntegrityAsync` now also re-pushes persisted status to every reachable active slave each cycle |
| `BackgroundServices/IntegrityCheckBackgroundService.cs` | Now runs one tick immediately on startup, in addition to the periodic timer |
| `Models/SlaveStatusPollResponse` (in `ICmsInstanceService.cs`) | New record: `IsAvailable`, `DisableMessage` |
@@ -97,7 +97,19 @@ sequenceDiagram
Svc->>Repo: UpdateAsync (Status=Inactive, DisableMessage=null)
Svc->>Repo: SaveChangesAsync()
Repo->>DB: UPDATE CmsInstances
Svc-->>Ctrl: UpdateStatusResult(Success=true, SlaveContactSuccess=true)
Svc->>DP: Unprotect(entity.ApiKey)
DP-->>Svc: plainApiKey
Svc->>Client: PushStatusAsync(slaveUrl, plainApiKey, isAvailable=true, disableMessage=null)
Note over Svc: Releases the master gate — the master no longer manages this slave, so it must not stay stuck on its last pushed status
alt Release success
Client-->>Svc: true
Svc->>Repo: UpdateAsync (LastStatusPushedAt = UtcNow)
Svc->>Repo: SaveChangesAsync()
Svc-->>Ctrl: UpdateStatusResult(Success=true, SlaveContactSuccess=true)
else Release failed
Client-->>Svc: false
Svc-->>Ctrl: UpdateStatusResult(Success=true, SlaveContactSuccess=false)
end
else newStatus = Available or NotAvailable
Svc->>Repo: UpdateAsync (Status, DisableMessage)
Svc->>Repo: SaveChangesAsync()
@@ -119,13 +131,15 @@ sequenceDiagram
Ctrl-->>Ctrl: return 200 OK with UpdateStatusResult
```
Text alternative: Controller calls service with id and new status; service loads entity, validates, updates DB, then for non-Inactive transitions decrypts ApiKey and pushes status to slave; returns SlaveContactSuccess=false if push fails but DB is always the authority.
Text alternative: Controller calls service with id and new status; service loads entity, validates, updates DB, then decrypts ApiKey and pushes status to slave — for `Inactive` this push is always `isAvailable=true, disableMessage=null` (releasing the gate); for `Available`/`NotAvailable` it pushes the new status as-is. Returns SlaveContactSuccess=false if the push fails, but the DB write is always the authority.
> **Updated 2026-07-04**: the `Inactive` branch previously did not push anything to the slave at all (see history below) — this left the slave stuck on whatever status it had last received, indefinitely. Fixed by always releasing the gate on deactivation.
---
## Flow 3 — VerifyIntegrityAsync (Background Integrity Check)
**Trigger**: `IntegrityCheckBackgroundService` periodic timer (every `IntegrityCheckIntervalMinutes`)
**Trigger**: `IntegrityCheckBackgroundService` — one tick immediately on host startup, then every `IntegrityCheckIntervalMinutes` **(startup tick added 2026-07-04)**
```mermaid
sequenceDiagram
@@ -163,6 +177,7 @@ sequenceDiagram
Svc->>Repo: UpdateAsync (LastIntegrityCheckFailedAt = UtcNow)
Svc->>Repo: SaveChangesAsync()
Repo->>DB: UPDATE CmsInstances
Note over Svc: Unreachable for URL check — skip the status re-push for this instance this cycle
else Slave reachable
Client-->>Svc: registeredMasterUrl
alt URLs match
@@ -183,9 +198,21 @@ sequenceDiagram
Repo->>DB: UPDATE CmsInstances
end
end
Note over Svc: Status re-push (added 2026-07-04) — runs whenever the slave was reachable, independent of the URL-match outcome
Svc->>Client: PushStatusAsync(slaveUrl, plainApiKey, isAvailable=(Status==Available), disableMessage)
alt Push success
Client-->>Svc: true
Svc->>Repo: UpdateAsync (LastStatusPushedAt = UtcNow)
Svc->>Repo: SaveChangesAsync()
else Push failed
Client-->>Svc: false
Note over Svc: Logged; no flag change — next cycle (or the immediate startup tick) will retry
end
end
end
Svc-->>BgSvc: done
```
Text alternative: Background timer triggers integrity service; for each non-Inactive slave: decrypts key, retrieves registered master URL, clears failure flag on match, re-registers on mismatch, sets LastIntegrityCheckFailedAt when slave is unreachable or re-registration fails.
Text alternative: Background timer triggers integrity service; for each non-Inactive slave: decrypts key, retrieves registered master URL, clears failure flag on match, re-registers on mismatch, sets LastIntegrityCheckFailedAt when slave is unreachable or re-registration fails. **(Added 2026-07-04)** For every slave that was reachable, the service additionally re-pushes the master's currently persisted `Status`/`DisableMessage` to that slave — this is what lets a slave that reset its in-memory gate (e.g. after a restart) catch up without waiting for the next explicit admin status change. Combined with the new immediate startup tick on `IntegrityCheckBackgroundService`, this reconciliation now also runs right after the master process (re)starts.
> **Updated 2026-07-04**: previously this flow only verified/re-registered the master URL and never re-pushed status (see history below) — a restarted slave (whose in-memory master-gate defaults to `Available`) would show the wrong status until the master's next explicit UI-driven change.
@@ -10,7 +10,11 @@ graph TD
CheckMsg{"newStatus = NotAvailable\nAND disableMessage\nis null or empty?"}
ValidationErr["Throw ValidationException\nDisableMessage required"]
CheckInactive{"newStatus\n= Inactive?"}
SetInactive["Status = Inactive\nDisableMessage = null\nNo HTTP push\nSlaveContactSuccess = true"]
SetInactive["Status = Inactive\nDisableMessage = null\nPersist to MasterDbContext"]
ReleaseGate["Decrypt ApiKey\nPushStatusAsync(isAvailable=true,\ndisableMessage=null)\n— release the master gate"]
ReleaseOk{"Release push\nsucceeded?"}
ReleaseDone["LastStatusPushedAt = UtcNow\nSave"]
ReturnRelease["Return UpdateStatusResult\nSuccess=true\nSlaveContactSuccess=(release result)"]
PersistStatus["Persist Status + DisableMessage\nto MasterDbContext"]
DecryptKey["Decrypt ApiKey\nvia IDataProtector"]
PushSlave["PushStatusAsync\nto slave endpoint"]
@@ -25,7 +29,8 @@ graph TD
CheckExists -->|"yes"| CheckMsg
CheckMsg -->|"yes — invalid"| ValidationErr
CheckMsg -->|"no — valid"| CheckInactive
CheckInactive -->|"yes"| SetInactive --> Done
CheckInactive -->|"yes"| SetInactive --> ReleaseGate --> ReleaseOk
ReleaseOk -->|"yes/no"| ReleaseDone --> ReturnRelease --> Done
CheckInactive -->|"no"| PersistStatus --> DecryptKey --> PushSlave --> PushOk
PushOk -->|"yes"| UpdatePushed --> ReturnOk --> Done
PushOk -->|"no"| ReturnWarn --> Done
@@ -34,13 +39,15 @@ graph TD
classDef action fill:#63b3ed,stroke:#2b6cb0,stroke-width:1px,color:#000
classDef terminal fill:#9ae6b4,stroke:#2f855a,stroke-width:2px,color:#000
classDef error fill:#FC8181,stroke:#C53030,stroke-width:2px,color:#000
class CheckExists,CheckMsg,CheckInactive,PushOk decision
class PersistStatus,DecryptKey,PushSlave,UpdatePushed,SetInactive action
class CheckExists,CheckMsg,CheckInactive,PushOk,ReleaseOk decision
class PersistStatus,DecryptKey,PushSlave,UpdatePushed,SetInactive,ReleaseGate,ReleaseDone action
class Start,Done terminal
class NotFound,ValidationErr error
```
Text alternative: Load entity (404 if missing) → validate DisableMessage required for NotAvailable → for Inactive skip push → for others persist, decrypt key, push to slave, set SlaveContactSuccess based on push result.
Text alternative: Load entity (404 if missing) → validate DisableMessage required for NotAvailable → for Inactive, persist the Inactive status **and then explicitly push `Available` (disableMessage=null) to the slave** to release the master gate (the master no longer manages this slave, so it must not leave the slave stuck on a stale status) → for other statuses persist, decrypt key, push to slave, set SlaveContactSuccess based on push result.
> **Updated 2026-07-04**: originally `Inactive` skipped the HTTP push entirely (see history below); this was found to leave slaves permanently stuck on their last pushed status (e.g. `NotAvailable`) after being detached from the master, and was fixed to always release the gate on deactivation.
---
@@ -58,26 +65,34 @@ graph TD
RegOk{"Re-registration\nsucceeded?"}
ClearAfterReg["LastIntegrityCheckFailedAt = null\nLastContactedAt = UtcNow\nSave"]
SetFailedReg["LastIntegrityCheckFailedAt = UtcNow\nSave"]
RePushStatus["PushStatusAsync\n(re-push persisted Status/DisableMessage\nto this active slave)"]
RePushOk{"Push\nsucceeded?"}
UpdatePushedAt["LastStatusPushedAt = UtcNow\nSave"]
Next(["Next instance"])
Start --> GetUrl --> Reachable
Reachable -->|"no"| SetFailed --> Next
Reachable -->|"yes"| UrlMatch
UrlMatch -->|"match"| ClearOk --> Next
UrlMatch -->|"match"| ClearOk --> RePushStatus
UrlMatch -->|"mismatch"| ReRegister --> RegOk
RegOk -->|"yes"| ClearAfterReg --> Next
RegOk -->|"yes"| ClearAfterReg --> RePushStatus
RegOk -->|"no"| SetFailedReg --> Next
RePushStatus --> RePushOk
RePushOk -->|"yes"| UpdatePushedAt --> Next
RePushOk -->|"no"| Next
classDef decision fill:#FFC107,stroke:#F57F17,stroke-width:2px,color:#000
classDef action fill:#63b3ed,stroke:#2b6cb0,stroke-width:1px,color:#000
classDef terminal fill:#9ae6b4,stroke:#2f855a,stroke-width:2px,color:#000
classDef error fill:#FC8181,stroke:#C53030,stroke-width:2px,color:#000
class Reachable,UrlMatch,RegOk decision
class GetUrl,SetFailed,ClearOk,ReRegister,ClearAfterReg,SetFailedReg action
class Reachable,UrlMatch,RegOk,RePushOk decision
class GetUrl,SetFailed,ClearOk,ReRegister,ClearAfterReg,SetFailedReg,RePushStatus,UpdatePushedAt action
class Start,Next terminal
```
Text alternative: For each active slave — attempt to get its registered master URL; if unreachable set failure flag; if reachable and URL matches clear flag; if mismatch re-register; clear flag on success, set flag on failure.
Text alternative: For each active slave — attempt to get its registered master URL; if unreachable set failure flag and move on; if reachable and URL matches clear the flag, if mismatch re-register (clearing the flag on success, setting it on failure); then — regardless of the URL check's outcome, as long as the slave was reachable — **re-push the master's currently persisted Status/DisableMessage to that slave** (this is the reconciliation path for a slave that reset its in-memory gate, e.g. after a restart, or missed an earlier push). `IntegrityCheckBackgroundService` also now runs one tick immediately on host startup, in addition to its periodic interval, so this resync happens right after the master (re)starts rather than waiting a full cycle.
> **Updated 2026-07-04**: originally this check only verified/re-registered the master URL and never re-pushed status (see history below); this left restarted slaves stuck on a stale in-memory status (reset to `Available` by default) until the next explicit status change — fixed by adding the re-push step and the immediate startup run.
---
@@ -101,10 +116,12 @@ Text alternative: For each active slave — attempt to get its registered master
|------|----|---------------|-----------|-------|
| Any | `Available` | Clear to null | Yes | Slave re-enabled |
| Any | `NotAvailable` | Required, non-empty | Yes | Slave disabled with message |
| Any | `Inactive` | Clear to null | **No** | Master stops all contact |
| Any | `Inactive` | Clear to null | **Yes** — pushes `Available`/null | Master stops managing the slave, but must first release the gate so the slave doesn't stay stuck on its last pushed status |
| `Inactive` | `Available` | Clear to null | Yes | Reactivation |
| `Inactive` | `NotAvailable` | Required, non-empty | Yes | Reactivation with disable |
> **Updated 2026-07-04**: the `Inactive` row previously said "No" HTTP push (see history below) — corrected after the no-push behavior was found to leave slaves permanently stuck on their last status.
---
## ApiKey Encryption Rules
@@ -132,9 +149,9 @@ Text alternative: For each active slave — attempt to get its registered master
| Rule | Description |
|------|-------------|
| BR-CONTACT-01 | Instances with `Status = Inactive` are excluded from `GetActiveAsync` and never contacted via HTTP |
| BR-CONTACT-02 | Status push is skipped when transitioning any status → `Inactive` |
| BR-CONTACT-03 | Integrity check runs only against instances where `Status != Inactive` |
| BR-CONTACT-01 | Instances with `Status = Inactive` are excluded from `GetActiveAsync` and are not contacted by the periodic integrity check / re-push cycle |
| BR-CONTACT-02 | **(Updated 2026-07-04)** The transition to `Inactive` itself always performs exactly one status push — `Available`, `disableMessage=null` — to release the master gate on the slave before the instance drops out of `GetActiveAsync` for good. Previously this push was skipped entirely; that left the slave stuck on its last pushed status indefinitely. |
| BR-CONTACT-03 | Integrity check (and its new status re-push, see BR-02) runs only against instances where `Status != Inactive` |
---
@@ -146,3 +163,17 @@ Text alternative: For each active slave — attempt to get its registered master
| BR-BG-02 | Each tick creates and disposes its own `IServiceScope` |
| BR-BG-03 | Exceptions within a single slave's integrity check are caught, logged, and do not abort processing for remaining slaves |
| BR-BG-04 | If `MasterModuleOptions.MasterUrl` is null or empty, the background service logs a warning and skips the entire integrity check for that cycle |
| BR-BG-05 | **(Added 2026-07-04)** `IntegrityCheckBackgroundService` runs one tick immediately on host startup (in addition to its periodic `PeriodicTimer` cycle), so a freshly (re)started master resyncs slave statuses right away instead of waiting up to `IntegrityCheckIntervalMinutes` |
---
## Slave Pull (Status Poll) Rules — Added 2026-07-04
This closes a design gap: the original inception requirements (`inception/requirements/requirements.md`, FR-MASTER-06/07, NFR-MASTER-01) specified a slave-initiated periodic pull with fail-open, but construction implemented push-only. The pull side was added alongside the existing push mechanism (not instead of it) after a slave was observed staying on a stale status through a restart and a deactivation.
| Rule | Description |
|------|-------------|
| BR-PULL-01 | `SlaveStatusController` exposes `GET /api/v1/SlaveStatus`, authenticated via the `X-Master-Api-Key` header — no `[Authorize]`/JWT, since the caller is a slave process, not a logged-in user |
| BR-PULL-02 | `CmsInstanceService.GetStatusForApiKeyAsync` identifies the calling slave by decrypting each active `CmsInstance.ApiKey` and comparing it to the caller's plain key (no separate slave-identity field exists; the shared key is the only credential) — returns `null` (→ 401) when no match is found |
| BR-PULL-03 | A successful poll updates `CmsInstance.LastContactedAt`, mirroring the existing convention used by push-based master↔slave calls |
| BR-PULL-04 | `/api/v1/SlaveStatus` is added to `AvailabilityMiddleware`'s bypass list on the master's own instance, so it stays reachable regardless of the master's own local availability status |
@@ -71,7 +71,9 @@ public enum CmsInstanceStatus
|-------|---------|----------------------|
| `Available` | Slave is enabled; normal operation | Yes (status push + integrity checks) |
| `NotAvailable` | Slave is disabled; `DisableMessage` served to end-users | Yes (status push + integrity checks) |
| `Inactive` | Soft-removed; greyed out in UI | **No** — all HTTP contact is halted |
| `Inactive` | Soft-removed; greyed out in UI | On the transition **into** `Inactive`: one final push (`Available`, no message) to release the gate. Afterwards: **No** further contact — excluded from `GetActiveAsync`, so no more pushes/integrity checks/re-pushes |
> **Updated 2026-07-04**: previously `Inactive` meant no HTTP contact at all, including on the transition itself — this left slaves stuck on their last pushed status after being detached. See BR-CONTACT-02 in `business-rules.md`.
---
@@ -129,4 +131,19 @@ public enum CmsInstanceStatus
| Property | Type | Notes |
|----------|------|-------|
| `Success` | `bool` | Always `true` when status persisted to DB (DB is the authority) |
| `SlaveContactSuccess` | `bool` | `true` if HTTP push to slave succeeded; `false` if push failed (slave unreachable); not applicable for `Inactive` transitions (returns `true`) |
| `SlaveContactSuccess` | `bool` | `true` if the HTTP push to the slave succeeded; `false` if it failed (slave unreachable). Applies to `Inactive` transitions too — reflects whether the gate-release push succeeded, not a hardcoded `true` |
> **Updated 2026-07-04**: `SlaveContactSuccess` for `Inactive` used to always be hardcoded `true` (no push happened, so nothing could fail) — now reflects the real result of the release push.
---
## SlaveStatusPollResponse (API Response) — Added 2026-07-04
Response shape for the new slave-pull endpoint (`GET /api/v1/SlaveStatus`, `SlaveStatusController`), used by `MasterStatusPollingBackgroundService` on the slave side (see `slave-availability-extension/functional-design/domain-entities.md`). This is the counterpart of the push-based flow above — added to close a gap versus the original inception requirements (FR-MASTER-06/07), which specified a slave-initiated pull in addition to the push that construction actually implemented.
| Property | Type | Notes |
|----------|------|-------|
| `IsAvailable` | `bool` | `true` when `CmsInstance.Status == Available` |
| `DisableMessage` | `string?` | `CmsInstance.DisableMessage` |
Identified by matching the caller's plain API key (header `X-Master-Api-Key`) against each active `CmsInstance`'s decrypted key — there is no separate slave-identity field, the shared key doubles as the credential. No `[Authorize]`/JWT on this endpoint.
@@ -150,3 +150,108 @@ sequenceDiagram
```
Text alternative: Middleware first checks bypass paths, then admin JWT. If neither applies: checks static master cache; if master unavailable return 503. If master available: checks local availability service; if locally unavailable return 503. Otherwise pass through.
---
## Flow 5: Slave Poll (Slave → Master) — Added 2026-07-04
**Trigger**: `MasterStatusPollingBackgroundService` — one tick immediately on slave startup, then every `MasterPolling:PollIntervalSeconds` (default 30s).
```mermaid
sequenceDiagram
box rgba(76,175,80,0.15) Slave CMS
participant Timer as PeriodicTimer
participant BgSvc as MasterStatusPollingBackgroundService
participant Svc as MasterAvailabilityService
participant Client as MasterStatusPollClient
participant DB as AvailabilityDbContext
participant Cache as StaticCache
end
box rgba(33,150,243,0.15) Master CMS
participant MC as SlaveStatusController
end
Timer->>BgSvc: Tick
BgSvc->>Svc: GetPollTargetAsync()
Svc->>DB: GetRegistrationAsync()
alt No registration
DB-->>Svc: null
Svc-->>BgSvc: null
BgSvc-->>Timer: no-op, wait for next tick
else Registration exists
DB-->>Svc: MasterUrl, encrypted ApiKey
Svc-->>BgSvc: MasterUrl, plain ApiKey
BgSvc->>Client: GetStatusAsync(masterUrl, plainApiKey)
Client->>MC: GET /api/v1/SlaveStatus with X-Master-Api-Key header
alt Poll succeeds
MC-->>Client: 200 OK { IsAvailable, DisableMessage }
Client-->>BgSvc: PolledMasterStatus
BgSvc->>Svc: ApplyPolledStatusAsync(isAvailable, disableMessage)
Svc->>Cache: set _masterIsAvailable / _masterDisableMessage
Svc->>DB: Update LastPolledAt=now, LastContactedAt=now
else Poll fails (network error, timeout, 401, etc.)
Client-->>BgSvc: null
BgSvc->>Svc: RecordPollFailureAsync(failOpenAfter)
Svc->>DB: read LastPolledAt (or RegisteredAt if never polled)
alt Unreachable longer than failOpenAfter
Svc->>Cache: force _masterIsAvailable=true, _masterDisableMessage=null
Note over Svc: Fail-open — a dead/unreachable master must never permanently block this slave
else Still within grace period
Note over Svc: No change — leave the existing cached gate as-is
end
end
end
```
Text alternative: Background timer triggers the poller; if no master is registered, it's a no-op. Otherwise it calls `GET /api/v1/SlaveStatus` on the registered master. On success, the response overwrites the in-memory gate and updates `LastPolledAt`/`LastContactedAt`. On failure, the gate is left alone unless the master has been unreachable (via poll) for longer than `MasterPolling:FailOpenAfterMinutes`, in which case the gate is forced open (`Available`, no message).
---
## Flow 6: Local Availability Status Read/Write with Master-Gate Override — Added 2026-07-04
Applies to the slave's own `/settings` admin UI/API — `GET /api/v1/Availability/status` and `POST /api/v1/Availability/admin/status` — layered on top of `PersistentAvailabilityService`, which previously only ever reflected the locally-persisted `GlobalAvailabilityState` regardless of the master gate.
```mermaid
sequenceDiagram
box rgba(33,150,243,0.15) Frontend (SettingsPage)
participant FE as Owner Browser
end
box rgba(76,175,80,0.15) Slave CMS
participant Ctrl as AvailabilityController
participant Svc as PersistentAvailabilityService
participant MasterSvc as MasterAvailabilityService
participant DB as ApplicationDbContext
end
FE->>Ctrl: GET /api/v1/Availability/status
Ctrl->>Svc: GetStatusDetailsAsync()
Svc->>MasterSvc: GetMasterStatus()
alt Master gate closed (IsAvailable = false)
MasterSvc-->>Svc: NotAvailable, masterMessage
Svc-->>Ctrl: Status=NotAvailable, Message=masterMessage, IsMasterControlled=true
else Master gate open
MasterSvc-->>Svc: Available
Svc->>DB: read GlobalAvailabilityState
DB-->>Svc: local Status/Message
Svc-->>Ctrl: Status, Message, IsMasterControlled=false
end
Ctrl-->>FE: 200 OK
FE->>Ctrl: POST /api/v1/Availability/admin/status (attempted local change)
Ctrl->>Svc: UpdateStatusAsync(newStatus, reason, updatedBy)
Svc->>MasterSvc: GetMasterStatus()
alt Master gate closed
MasterSvc-->>Svc: NotAvailable
Svc-->>Ctrl: throw MasterControlledAvailabilityException
Ctrl-->>FE: 409 Conflict (ProblemDetails)
else Master gate open
MasterSvc-->>Svc: Available
Svc->>DB: persist newStatus/reason
Svc-->>Ctrl: success
Ctrl-->>FE: 200 OK
end
```
Text alternative: Reading status now checks the master gate first — if closed, the response reflects the master's forced status and message, with `IsMasterControlled=true`, regardless of what's persisted locally. Writing a status change now performs the same master-gate check first: if closed, the write is rejected outright with `409 Conflict` instead of silently succeeding with no visible effect, since the master gate would have overridden it on the very next read anyway.
> Previously (before 2026-07-04): `GetStatusDetailsAsync` ignored the master gate entirely and always returned the locally-persisted status; `UpdateStatusAsync` always wrote the requested change regardless of the master gate, giving a misleading "success" for a change that was immediately invisible.
@@ -69,6 +69,7 @@ Text alternative: Bypass path check first. Admin JWT next (bypasses both gates).
| `/api/v1/Auth/` | Already bypassed — login must always work |
| `/api/v1/Setup/status` | Already bypassed — frontend init check |
| `/api/v1/master/` | **NEW** — master management endpoints must bypass gate so master can always push status or re-register |
| `/api/v1/SlaveStatus` | **NEW, added 2026-07-04** — this is actually the *master's* incoming endpoint for slave pulls, but it's added to this same bypass list on any instance that also loads `Modules.Availability` (i.e. the master itself), so the master's own local-gate status never blocks a slave from reading it |
**Cache behavior rules**:
@@ -76,7 +77,7 @@ Text alternative: Bypass path check first. Admin JWT next (bypasses both gates).
|---|------|
| BR-SLAVE-08 | `_masterIsAvailable` defaults to `true` (fail-open) on process startup |
| BR-SLAVE-09 | `_masterDisableMessage` defaults to `null` on process startup |
| BR-SLAVE-10 | Cache has no expiry (Q4=A); only updated on `POST /status` with valid API key |
| BR-SLAVE-10 | **(Updated 2026-07-04)** Cache is updated on `POST /status` (push, valid API key) **and** periodically overwritten by `MasterStatusPollingBackgroundService` pulling `GET /api/v1/SlaveStatus` from the master (see Rule Set 5). It is no longer purely push-driven or expiry-free: a poll failure that persists past `MasterPolling:FailOpenAfterMinutes` (default 5 min, measured from `MasterRegistration.LastPolledAt`) forcibly resets the cache to `Available`/`null` regardless of the last pushed value. |
| BR-SLAVE-11 | 503 response from master gate includes `_masterDisableMessage` in `ProblemDetails.Detail` |
---
@@ -89,6 +90,7 @@ Text alternative: Bypass path check first. Admin JWT next (bypasses both gates).
| BR-SLAVE-13 | On `RegisterAsync`: if row with that Id exists → update. If not → insert. Never delete. |
| BR-SLAVE-14 | `RegisteredAt` is set once at creation and never updated |
| BR-SLAVE-15 | `LastContactedAt` is updated on every successful master call (register, status push, get-url) |
| BR-SLAVE-16 | **(Added 2026-07-04)** `LastPolledAt` is updated whenever this slave successfully polls the master via `MasterStatusPollingBackgroundService` (distinct from `LastContactedAt`, which tracks master-initiated contact) |
---
@@ -101,3 +103,31 @@ Text alternative: Bypass path check first. Admin JWT next (bypasses both gates).
| Get registered URL | `GET` | `/api/v1/master/registered-url` | None (API key in header) |
All three endpoints are unauthenticated from ASP.NET Core's perspective — they use the custom `X-Master-Api-Key` header validation implemented in `MasterAvailabilityService`. They are also in the middleware bypass list so the gate cannot block master management calls.
---
## Rule Set 5: Slave Pull + Fail-Open — Added 2026-07-04
Closes a gap versus the original inception requirements (`inception/requirements/requirements.md`, FR-MASTER-06 "Slave Pull Model", FR-MASTER-07 "Slave Fallback Behavior", NFR-MASTER-01 "Fail-Open Safety"): construction had implemented push-only, with fail-open surviving only as an in-memory startup default (BR-SLAVE-08/09) rather than an actively-reconciling pull. Added after a slave was observed remaining on a stale status through a restart and again after being deactivated on the master.
| # | Rule |
|---|------|
| BR-PULL-01 | `MasterStatusPollingBackgroundService` runs one tick immediately on slave startup, then every `MasterPolling:PollIntervalSeconds` (default `30`) |
| BR-PULL-02 | If no `MasterRegistration` exists yet, the poll tick is a no-op (nothing to poll) |
| BR-PULL-03 | On a successful poll (`GET /api/v1/SlaveStatus` on the registered master, header `X-Master-Api-Key`), the response overwrites the in-memory gate (`_masterIsAvailable`/`_masterDisableMessage`) and updates `MasterRegistration.LastPolledAt` + `LastContactedAt` |
| BR-PULL-04 | On a failed poll (network error, timeout, or non-success HTTP status), the gate is **not** changed immediately — instead `RecordPollFailureAsync` checks how long it's been since `LastPolledAt` (or `RegisteredAt` if never polled) |
| BR-PULL-05 | **Fail-open**: if that elapsed time exceeds `MasterPolling:FailOpenAfterMinutes` (default `5`), the gate is forced to `Available`/`null` — a master that is dead or unreachable must never permanently block a slave |
| BR-PULL-06 | Push (BR-SLAVE-01 through 11) and pull (this rule set) are independent and complementary: push gives instant reactivity to an explicit admin status change; pull is the self-healing safety net for everything push can miss (slave restarts, dropped pushes, local tampering with the in-memory gate) |
---
## Rule Set 6: Master-Controlled Availability Lock — Added 2026-07-04
Applies to the slave's own local availability admin UI/API (`PersistentAvailabilityService` / `AvailabilityController` — the endpoints an Owner uses on `/settings` to set *this instance's own* Maintenance/NotAvailable status), not the master↔slave protocol endpoints above. Added because an Owner on a master-disabled slave could previously "successfully" set local status to `Available` with no visible effect, since the master gate silently overrode the display.
| # | Rule |
|---|------|
| BR-LOCK-01 | `GET /api/v1/Availability/status` returns the master-gate status (not the locally-persisted one) whenever the master gate is closed (`IsAvailable = false`), and includes a new `IsMasterControlled: true` flag in that case |
| BR-LOCK-02 | `POST /api/v1/Availability/admin/status` (`PersistentAvailabilityService.UpdateStatusAsync`) throws `MasterControlledAvailabilityException` and makes **no** DB write when the master gate is closed, instead of silently persisting a change that would have no visible effect |
| BR-LOCK-03 | The controller translates that exception into `409 Conflict` (`ProblemDetails`) |
| BR-LOCK-04 | The frontend Settings page reads `isMasterControlled` and disables the mode selector, the reason field, and the save button, showing a banner explaining that the Master CMS controls this status |
@@ -40,9 +40,10 @@ Singleton row — at most one record exists per slave instance. Upserted on each
|-------|------|-------------|-------|
| `Id` | `Guid` | PK | Fixed value (e.g. `Guid.Empty`) enforces singleton |
| `MasterUrl` | `string` | Required, max 500 | URL of the master CMS that registered this slave |
| `ApiKey` | `string` | Required, max 1000 | Plain-text API key sent in first registration; used for subsequent validation |
| `ApiKey` | `string` | Required, max 1000 | API key sent in first registration, encrypted via `IMasterApiKeyProtector` (ASP.NET Core Data Protection) before storage, decrypted for each subsequent validation **corrected 2026-07-04**, this was previously (incorrectly) documented as stored plain-text |
| `RegisteredAt` | `DateTimeOffset` | Required | Timestamp of first registration |
| `LastContactedAt` | `DateTimeOffset?` | Optional | Updated on every successful master call (register, status push, get-url) |
| `LastContactedAt` | `DateTimeOffset?` | Optional | Updated on every successful master-initiated call (register, status push, get-url) |
| `LastPolledAt` | `DateTimeOffset?` | Optional | **Added 2026-07-04**. Updated on every successful *slave-initiated* poll (`MasterStatusPollingBackgroundService`) — the counterpart to `LastContactedAt`, tracking the opposite direction of contact. Also the basis for the fail-open timeout (see below). |
> **Singleton enforcement**: The `Id` is a fixed known value (`Guid.Parse("00000000-0000-0000-0000-000000000001")`). On first `POST /api/v1/master/register` the row is created; on re-registration the same row is updated in-place. This avoids a composite unique constraint and makes EF upsert trivial.
@@ -54,10 +55,10 @@ Not a DB entity — lives in memory on the slave process.
| Field | Type | Default | Notes |
|-------|------|---------|-------|
| `_masterIsAvailable` | `bool` | `true` | Set by status push; `true` = pass gate |
| `_masterDisableMessage` | `string?` | `null` | Message forwarded from master to 503 response |
| `_masterIsAvailable` | `bool` | `true` | Set by status push **or** successful poll; forced back to `true` on prolonged poll failure (fail-open) |
| `_masterDisableMessage` | `string?` | `null` | Message forwarded from master to 503 response; cleared to `null` on fail-open |
**No expiry** (Q4=A): cache is valid indefinitely until the master pushes again. If the master goes offline permanently, the last known status is used. Default = `true` (Available) = fail-open.
> **Updated 2026-07-04**: originally documented as having "no expiry" (Q4=A), valid indefinitely until the next push. This is no longer accurate now that `MasterStatusPollingBackgroundService` actively polls the master (see Rule Set 5 / Flow 5 in the sibling docs) and forces the cache back to `Available`/`null` if the master has been unreachable via poll for longer than `MasterPolling:FailOpenAfterMinutes` (default 5 min, measured from `MasterRegistration.LastPolledAt`). Push-driven updates (this section's original description) are unchanged and still apply; the poll is an additional, independent path that can also write to this same cache.
---
@@ -145,11 +145,11 @@ Text alternative: Master CMS has new module with controller, service, repository
| Constraint | Source | Impact |
|-----------|--------|--------|
| `ApiKey` never returned in API responses | NFR-MASTER-03 | `CmsInstanceDto` excludes `ApiKey`; only accepted in `CreateCmsInstanceRequest` |
| Fail-open on Master unreachable | NFR-MASTER-01 | `MasterAvailabilityService` returns last cached status (default Available) on HTTP failure |
| Fail-open on Master unreachable | NFR-MASTER-01 | **(Updated 2026-07-04)** Originally only an in-memory startup default (`_masterIsAvailable = true`); now actively enforced by `MasterStatusPollingBackgroundService.RecordPollFailureAsync`, which forces the gate open if the master has been unreachable via poll for longer than `MasterPolling:FailOpenAfterMinutes` — this is the slave-pull half of FR-MASTER-06/07 that was originally specified but not implemented until 2026-07-04 (see `slave-availability-extension/functional-design/business-rules.md` Rule Set 5) |
| Per-module DbContext + migrations | NFR-MASTER-06 | New `MasterDbContext` in `Modules.Master`; new `AvailabilityDbContext` in `Modules.Availability` |
| Master exemption from own gate | FR-MASTER-09 | Handled naturally: no `MasterRegistration` record exists on Master instance → gate skipped |
| Disable message required for NotAvailable | FR-MASTER-14 | Validated in `CmsInstanceService.UpdateStatusAsync` before persistence |
| Inactive slaves: no HTTP contact | FR-MASTER-13 | `GetActiveAsync()` filters out Inactive before integrity checks and status pushes |
| Inactive slaves: no further HTTP contact | FR-MASTER-13 | **(Updated 2026-07-04)** `GetActiveAsync()` filters out Inactive instances from ongoing integrity checks and status re-pushes, as before — but the transition *into* `Inactive` itself now always performs one final push (`Available`, no message) to release the master gate before the instance drops out; previously this transition performed no push at all, leaving the slave stuck on its last status |
| Owner role only | FR-MASTER-10 | `[Authorize(Policy = "OwnerOnly")]` on all `CmsInstanceController` actions |
---
@@ -2,6 +2,8 @@
> Method signatures at the interface level. Detailed business rules and implementation logic are deferred to Functional Design (CONSTRUCTION phase).
> ⚠️ **Partially superseded, found stale 2026-07-04**: the `IMasterAvailabilityService` methods and `/api/internal/master/*` routes below describe an inception-stage pull design that construction did not build as-is (actual routes are `/api/v1/master/*`; the method set differs — see `RegisterAsync`/`PushStatusAsync`/`GetRegisteredUrlAsync`/`GetMasterStatus` in `construction/slave-availability-extension/`, plus 2026-07-04 additions `GetPollTargetAsync`/`ApplyPolledStatusAsync`/`RecordPollFailureAsync`). Treat the functional-design docs under `construction/` as current truth for exact signatures.
---
## Unit 1 — master-backend
@@ -1,5 +1,7 @@
# Components — Master CMS Module
> ⚠️ **Partially superseded, found stale 2026-07-04**: the slave-side `IMasterAvailabilityService` description below (static `_cachedStatus`/`_lastFetchedAt`, `CacheMinutes`-based staleness, `/api/internal/master/*` routes) describes an inception-stage design that construction did not build as-is — the actual implementation is push-based (`/api/v1/master/*`, `_masterIsAvailable`/`_masterDisableMessage`), plus a 2026-07-04 addition of a differently-shaped slave-pull (`MasterStatusPollingBackgroundService` + `GET /api/v1/SlaveStatus`, time-based fail-open). The Unit 1 (`master-backend`) component list above the slave-side section is accurate. See `construction/master-backend/functional-design/*.md` and `construction/slave-availability-extension/functional-design/*.md` for the as-built design. This gap predates 2026-07-04 and was found (not caused) during today's documentation audit.
## Unit 1 — master-backend (`SlpModularCms.Modules.Master`)
### MasterModule
@@ -1,5 +1,7 @@
# Services — Master CMS Module
> ⚠️ **Superseded, found stale 2026-07-04**: this document is the *inception-stage* design and describes an earlier pull-based `IMasterAvailabilityService` (static cache + `MasterModuleOptions.CacheMinutes` + `/api/internal/master/*` routes) that was **not what got built**. Construction pivoted to a push-based protocol instead (`/api/v1/master/*` routes, `_masterIsAvailable`/`_masterDisableMessage` fields, no `CacheMinutes`). A slave-pull mechanism was eventually added too, but on 2026-07-04 and with a different shape (`MasterStatusPollingBackgroundService` polling `GET /api/v1/SlaveStatus` on an interval, with a time-based fail-open) than what's described below. **Treat this file as historical intent, not current truth** — the accurate, as-built design lives in `construction/master-backend/functional-design/*.md` and `construction/slave-availability-extension/functional-design/*.md`. This divergence predates 2026-07-04 and was found (not caused) during today's documentation audit; `CacheMinutes`/`ApiKey` dead-config fallout is already tracked as TD-001 in `tech-debt-backlog`.
## Service Definitions
### Master-Side Services
@@ -11,3 +11,15 @@
**Context**: Feature aangemaakt tijdens master-cms-module sessie; geen actieve executie gestart, puur registratie voor later
---
## TD-004 Added — Master-Controlled Availability Copy Review
**Timestamp**: 2026-07-04T00:00:00Z
**User Input**: "De tekst mag nog wel wat anders worden, maar ik weet nu even niet wat dus ik kan dat nu ook zelf niet veranderen. Kan je hier een tech debt item voor maken zodat ik dit later op kan pakken?"
**AI Response**: "TD-004 toegevoegd aan backlog.md: de teksten voor de master-controlled availability-lock op SettingsPage (banner-titel, banner-body, en de 409-conflict-toast — drie i18n-sleutels in nl/en translation.json) zijn functioneel correct maar de exacte formulering moet nog herzien worden. Puur copy, geen logica-wijziging verwacht; genoteerd dat de bijbehorende SettingsPage.test.tsx-assertie op tekst matcht en dus mee moet worden aangepast bij het oppakken."
**Context**: Item ontstond direct na de master-controlled-availability fix in local-dev-master-slave-setup (zelfde sessie, 2026-07-04); user wilde de tekst niet blokkerend laten zijn voor de rest van het werk.
---
@@ -43,6 +43,20 @@ Items discovered incidentally while working on other features. Not yet prioritiz
---
## TD-004: Review copy for the master-controlled availability lock on `SettingsPage`
**Found during**: `local-dev-master-slave-setup` — Build & Test follow-up fix (2026-07-04), user flagged the wording right after it was added but didn't have a replacement in mind yet.
**Location**:
- `frontend/src/pages/SettingsPage.tsx` (locked-state banner + disabled controls)
- `frontend/src/i18n/locales/nl/translation.json` and `en/translation.json`, keys `settings.availability.masterControlledTitle`, `settings.availability.masterControlled`, `settings.availability.masterControlledSaveError`
**Issue**: When a slave's availability is controlled by its Master CMS, the Settings page now shows a banner ("Beheerd door Master-CMS" / "Deze instantie is door de Master-CMS uitgeschakeld...") and disables the mode buttons, reason field, and save button, plus a specific toast on a 409 conflict ("Kan niet worden gewijzigd: de Master-CMS beheert deze status."). The behavior (locking + explaining why) is correct and intentional; only the exact phrasing needs a revisit — user wants different wording but hadn't decided on it at the time.
**Done looks like**: Update the three translation keys above (NL and EN) to the reviewed copy. No code/logic changes expected — this is copy-only. Re-run `pnpm test` for `SettingsPage.test.tsx` afterward (its assertions match on text via regex, e.g. `/master cms controls this status/i`, so wording changes will need matching test updates too).
---
## Notes
- None of these block functionality — `npm run build` and all test suites pass regardless.