Files
SluijsensandClaude Sonnet 5 56f4f6fe0f
Continuous Integration / config (pull_request) Successful in 9s
Continuous Integration / backend-build (pull_request) Successful in 4m27s
Continuous Integration / vulnerability-scan (pull_request) Successful in 4m10s
Continuous Integration / backend-test (pull_request) Canceled after 0s
Continuous Integration / frontend-build (pull_request) Canceled after 0s
Continuous Integration / frontend-test (pull_request) Canceled after 0s
Continuous Integration / frontend-lint (pull_request) Canceled after 0s
Continuous Integration / publish-test (pull_request) Canceled after 0s
Continuous Integration / publish-production (pull_request) Canceled after 0s
Continuous Integration / deploy-test (pull_request) Canceled after 0s
Continuous Integration / deploy-production (pull_request) Canceled after 0s
Continuous Integration / frontend-prepare (pull_request) Canceled after 50s
Adds Monitoring Setup docs and deploy-scp troubleshooting/debug fixes
Monitoring Setup: operations/plans/monitoring-setup-plan.md and
operations/monitoring/monitoring-instructions.md, covering Sentry alert
rules on the security_event tag, UptimeRobot's 6 liveness monitors, and
the two new Umami website entries for the admin SPA.

deployment-instructions.md gains a missing Observability__Environment
host var (without it, both environments would tag Sentry events as
"Production"), the nginx client_max_body_size fix for the 413 seen on
publish-test/production artifact uploads, and two troubleshooting notes
on Gitea Actions re-run behaviour: re-running deploy-test/production
alone loses the run's uploaded artifact, and re-running all jobs on an
existing (rather than a brand new) run can replay stale secrets.

deploy-scp.yaml: step names no longer show literal unresolved
${{ inputs.* }} text (Gitea doesn't interpolate that context in step
names), and a temporary debug step logs PI_MAIN_USERNAME/PASSWORD
length plus a username equality check to diagnose a persistent
Permission denied during the SSH steps, without ever logging the
secret values themselves.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015FffvxxJp5wG34Ru48GBig
2026-07-29 19:39:44 +02:00

22 KiB
Raw Permalink Blame History

AI-DLC State Tracking

Project Information

  • Feature Name: Gitea Deployment Workflow
  • Feature Slug: gitea-deployment-workflow
  • Project Type: Brownfield
  • Start Date: 2026-07-27T00:00:00Z
  • Current Stage: CONSTRUCTION - Code Generation complete for Round 2 (U3 + U4)
  • Branch: feature/gitea-deployment-workflow

Workspace State

  • Existing Code: Yes
  • Reverse Engineering Needed: Completed — full rerun on 2026-07-27 (user chose Q3 = B)
  • Workspace Root: K:\Development\Projects\SlpModularCms

Reverse Engineering Status

  • Reverse Engineering — Completed on 2026-07-27
  • Artifacts Location: aidlc-docs/_shared/reverse-engineering/ (all 8 artifacts regenerated + timestamp)
  • Verified by execution: Release build 0 errors / 50 warnings; 219 backend tests pass; 213 frontend tests pass; pnpm run lint fails (5 errors, 1 warning); 2 high-severity transitive package advisories

Code Location Rules

  • Application Code: Workspace root (NEVER in aidlc-docs/)
  • Feature Documentation: aidlc-docs/features/gitea-deployment-workflow/ only
  • Shared Artifacts: aidlc-docs/_shared/
  • Structure patterns: See code-generation.md Critical Rules

Language Configuration

  • Documentation Language: English
  • Conversation Language: User Language (Dutch)

Extension Configuration

Extension Enabled Decided At
Security Baseline Yes (blocking) Requirements Analysis
Property-Based Testing No Requirements Analysis

Operations Configuration

  • Include Operations Phase: Yes
  • Decided At: Requirements Analysis

Deployment Setup

  • Included: Yes
  • Method: CI/CD (Gitea Actions — confirms what U5/U6 already built)
  • Completed: 2026-07-28. Real domains: test.slpsoftware.nl / slpsoftware.nl. Ports 5100/5101 chosen for the Pi's local Kestrel bindings. Database backup script drafted. Full FTPS future-switch procedure drafted (Q5 = B), with an explicit caveat that shared hosting is very likely IIS-based, so systemd-restart and atomic-symlink-switch do not carry over unchanged — treated as a starting brief for a future Infrastructure Design pass, not a ready-to-execute procedure
  • Revised after user feedback (real host facts): webadmin (the Pi's FileZilla/SFTP account, root /mnt/storage1/www/) cannot SSH in and is not the deploy account — it's the WEBSITE_WORKSPACE.md website-author role. A separate, dedicated SSH-capable account now runs the deploy pipeline (PI_MAIN_USERNAME), with deploy paths under that account's own home directory rather than under /mnt/storage1/www/html/, so the CMS's release structure never interferes with the other hosted websites. shared/wwwroot-web is now a cross-account symlink to wherever webadmin uploads this customer's site, requiring a one-time shared-group permission setup — documented, not automated
  • Artifacts: operations/deployment/deployment-plan.md, deployment-instructions.md, rollback-plan.md

Database Provider Migration — SQL Server → MariaDB (2026-07-29)

Discovered while finalizing Deployment Setup: the target Pi only runs MariaDB, and Microsoft SQL Server has no ARM64 build at all (the mcr.microsoft.com/mssql/server image is linux/amd64 only; Azure SQL Edge, the former ARM answer, is retired). This invalidates ASM-04. Confirmed no other machine is available, and the eventual production host (mijnhostingpartner.nl) will also run MariaDB, plus the workload is small enough that MariaDB's performance is not a concern — user decided: switch the database provider, not the deployment target.

What changed (application code, not just Operations docs):

  • SlpModularCms.Core.csproj: Microsoft.EntityFrameworkCore.SqlServerMySql.EntityFrameworkCore 10.0.7 (Oracle's official provider)
  • UseSqlServer(...)UseMySQL(...) in ServiceCollectionExtensions.cs, AvailabilityModule.cs, MasterModule.cs
  • DatabaseMigrationExtensions.cs + its test: Microsoft.Data.SqlClient.SqlExceptionMySql.Data.MySqlClient.MySqlException for the transient-failure classifier
  • All three DbContexts' migrations deleted and regenerated (ApplicationDbContext, AvailabilityDbContext, MasterDbContext) — no MySQL/MariaDB model-compatibility issues surfaced (no index-length problems, no raw SQL anywhere in the codebase to translate)
  • Connection strings updated across appsettings.json (Api + Api.Slave, all three tiers) from SQL Server format to MySQL format
  • README.md dev setup: MariaDB container command replacing the SQL Server one, with a migration note for anyone returning to old instructions
  • deployment-instructions.md / rollback-plan.md: connection string format, backup script rewritten around mariadb-dump (was sqlcmd/BACKUP DATABASE), restore procedure rewritten

Provider selection — two failed attempts before the working one, both reproduced against a real MariaDB instance, not just reasoned about:

  1. Pomelo.EntityFrameworkCore.MySql (the usual first choice, explicit first-class MariaDB support) — caps at EF Core 9.x, no EF Core 10 release exists. A Pomelo maintainer states (PR #2017) that Pomelo 9 / EF Core 9 packages work fine on a net10.0 TFM — confirmed true only when nothing else forces EF Core 10 packages. This project's Microsoft.AspNetCore.Identity.EntityFrameworkCore and Microsoft.AspNetCore.DataProtection.EntityFrameworkCore are versioned in lockstep with the .NET 10 runtime and hard-require EF Core >= 10.0.9, so the resolved Microsoft.EntityFrameworkCore.Abstractions ends up at 10.0.9 regardless. Restore only warns (NU1608), and it builds — but fails at runtime with MissingMethodException: AbstractionsStrings.ArgumentIsEmpty the moment EF tooling touches a DbContext: Pomelo's compiled assembly calls an internal EF Core 9 helper that no longer exists in the 10.0.9 assembly actually loaded.
  2. MySql.EntityFrameworkCore 10.0.7 (Oracle's official provider) — its net10.0 dependency group targets EF Core 10.0.7, compatible with 10.0.9. Builds and migrates cleanly, but dotnet ef database update against the real MariaDB throws InvalidCastException: Unable to cast object of type 'System.DBNull' to type 'System.Int64' in MySQLHistoryRepository.AcquireDatabaseLock() — a confirmed MariaDB-incompatibility bug (MariaDB's GET_LOCK() apparently returns NULL in a case Oracle's code doesn't handle, and Oracle's provider is tested against real MySQL Server, not MariaDB). This isn't limited to CLI tooling — the same code path runs on every application startup via MigrateCoreDatabase().

Working solution: kept Oracle's MySql.EntityFrameworkCore 10.0.7 (otherwise fully compatible) and added NonLockingMySQLHistoryRepository (SlpModularCms.Core/Hosting/), wired in via options.ReplaceService<IHistoryRepository, NonLockingMySQLHistoryRepository>() on all three AddDbContext registrations. Oracle's internal MySQLHistoryRepository class can't be subclassed (it's internal, only its constructor is public), so the workaround constructs a real instance of it via reflection and forwards every IHistoryRepository member to it except AcquireDatabaseLock/AcquireDatabaseLockAsync, which return a no-op lock instead of ever reaching the broken GET_LOCK call. Accepted as safe because this deployment model never runs migrations from more than one place at a time (MigrateCoreDatabase() at startup, one deploy at a time via the atomic-release sequence) — a genuinely concurrent multi-instance migration race is not a scenario this architecture produces.

Verified end to end: all three migrations applied successfully to the user's real local MariaDB (dotnet ef database update, all expected tables present including DataProtectionKeys and the renamed Identity tables); both SlpModularCms.Api and SlpModularCms.Api.Slave start cleanly against it (/health → 200, MigrateCoreDatabase() logs "already up to date" on the second run); full backend suite re-confirmed at 372/372 passed, 0 build errors.

Reverse Proxy Topology Correction (2026-07-29)

Infrastructure Design's Q10 (infrastructure-design.md § 1) established that an nginx reverse proxy exists — right about that, wrong about where: it runs on a separate, dedicated Pi ("the proxy Pi"), not on pi-main (where this deployment's release directories and systemd units live). The proxy Pi terminates TLS and forwards plain HTTP to pi-main over the LAN.

Corrected:

  • ASPNETCORE_URLS binds 0.0.0.0, not localhost (deployment-instructions.md § 1.5) — the proxy Pi must reach Kestrel over the network, not loopback
  • pi-main runs no nginx and holds no certificates for these domains at all — the entire nginx/certbot procedure in deployment-instructions.md § 1.7 happens on the proxy Pi instead
  • Added a firewall requirement on pi-main (§ 1.7.1 step 4): since Kestrel now listens on all interfaces, the backend ports must be restricted to just the proxy Pi's address, or the TLS termination the whole design depends on is trivially bypassable by anything else on the LAN hitting pi-main directly

Also corrected in infrastructure-design.md § 1 (erratum note, history preserved rather than rewritten). No code changes — this is host topology and Operations documentation only, the application itself has no opinion on where TLS terminates.

Scope Decisions (from feature-selection.md)

  • Public website: documentation/instructions only — where the website build lands in wwwroot/, how it coexists with wwwroot/admin/, and what a per-website workspace must deliver. The website's own build/deploy workflow stays out of scope (Q4 = A).
  • Environments: local, test, production only.
  • Observability stack: UptimeRobot (uptime), Umami (analytics), console logging + Sentry (logging/errors).
  • Deployment constraint: upload as a published .NET application; no server configuration may be required.
  • Reference: existing working Gitea Actions setup at K:\Development\SlpSoftware\Projects\SlpSoftware (React/Vite) is the starting point.
  • Health check endpoint: IN SCOPE (decided 2026-07-27). Liveness onlyAddHealthChecks() + MapHealthChecks("/health"), no package needed and no database check (Q17 = A / D-21, superseding the earlier note that AddDbContextCheck might be included). /health must be added to AvailabilityMiddleware._bypassPrefixes so the availability gate cannot return 503 for it. Health = infrastructure liveness; Availability/capabilities = CMS domain state — these stay strictly separate.

Note

: this section records the earliest scope decisions. The authoritative and complete decision set is inception/requirements/requirements.md § 3 (D-01…D-32) — in particular, Q4 = C changed the public website from living directly in wwwroot/ to wwwroot/web/.

Stage Progress

INCEPTION

  • Workspace Detection — Complete
  • Reverse Engineering — Complete, approved 2026-07-27 (full rerun of all 8 _shared/ artifacts)
  • Requirements Analysis — Complete, approved 2026-07-27. 24 FRs (FR-24 added at Application Design), 10 NFRs, 32 decisions, 7 assumptions, 4 open items, 4 documented security deviations. Two question rounds: requirement-verification-questions.md (25 Q) and requirement-clarification-questions.md (5 Q).
  • User Stories — SKIP (infrastructure/operations work; no new end-user functionality or persona. Offered at Requirements Analysis approval, not requested.)
  • Workflow Planning — Complete, approved 2026-07-27. Artifact: inception/plans/execution-plan.md
  • Application Design — Complete, approved 2026-07-27. 14 code components (9 new, 5 modified) + 2 workflow components. Artifacts in inception/application-design/. Two composition conflicts found and carried to Unit 2. Added FR-24, closed OPEN-02.
  • Units Generation — Complete (awaiting approval). 7 units in 4 execution rounds. Artifacts: unit-of-work.md, unit-of-work-dependency.md, unit-of-work-story-map.md

CONSTRUCTION

Units finalised at Units Generation (see inception/application-design/unit-of-work.md): U1 Hosting & Serving · U2 Data Durability · U3 Security Headers & CSP · U4 Observability · U5 CI Workflow & Gates · U6 Deploy Workflow · U7 Documentation

Execution rounds (Q4 = B): R1 = U1 + U2 · R2 = U3 + U4 · R3 = U5 + U6 · R4 = U7. One commit per unit; single PR at the end (Q6 = A).

  • Functional Design — EXECUTE for U1, U2, U3, U4; SKIP for U5, U6, U7. U1 U2 approved 2026-07-27 · U3 U4 2026-07-28
  • NFR Requirements — SKIP (all units) — already comprehensively captured in requirements.md § 5 and § 6
  • NFR Design — EXECUTE for U3, U4; SKIP for the rest. Deliberate deviation from the default NFR-Requirements/NFR-Design coupling — rationale in the execution plan. U3 U4 2026-07-28 — 11 patterns for U3, 10 for U4. Closed OPEN-01; raised REF-U3-01
  • Infrastructure Design — EXECUTE for U6, U7 per the original plan; SKIP for the rest. U6 2026-07-28 — single Pi, directory-only environment split, systemd --user (no sudo), releases/current/shared layout, 2-release retention, health check via public URL. Raised INFRA-U6-01 (linger requirement, carried to Operations). U7 did not need it in practice: by Units Generation, U7's scope had narrowed to documentation-only (Q8 = A — operational/infrastructure decisions moved to Operations), so it went straight to Code Generation with no infrastructure to design
  • Code Generation — EXECUTE (all 7 units, each built and tested before its completion message). U1 U2 2026-07-27 (253 backend tests). U3 2026-07-28 (315). U4 2026-07-28 (366 backend + 237 frontend). U6 U5 2026-07-28 — Round 3: deploy-scp.yaml + continuous_integration.yaml, 372 backend + 237 frontend tests (unchanged from Round 2, confirming no regressions from FR-21/FR-22 fixes), 0 vulnerable packages, lint clean. Raised REF-U5-01 (Umami-origin gate needs a parallel Gitea variable, since the backend side is a host env var per D-16). U7 2026-07-28 — Round 4: WEBSITE_WORKSPACE.md, README.md and .env.example updated, all claims re-verified against actual U1U6 source. All 7 units complete.
  • Build and Test — EXECUTE — Complete 2026-07-28. 609 unit tests (372 backend + 237 frontend), 0 vulnerable packages, 0 build errors. Plus live-host integration verification: static content/SPA-fallback/header-scoping, real master↔slave communication proving the persistent key ring survives a process restart, and W3C trace-id correlation in real log output. Artifacts in construction/build-and-test/

OPERATIONS

  • Deployment Setup — EXECUTE — Complete 2026-07-28, extended 2026-07-29 (MariaDB migration, reverse proxy topology correction, artifact-upload 413 fix in deployment-instructions.md § 1.7.3, missing Observability__Environment host var added § 1.5). Artifacts: operations/deployment/
  • Monitoring Setup — EXECUTE — Complete 2026-07-29. Plan answered (Sentry project reused, UptimeRobot account reused, Umami website entries created by user, 20-events/5-min default alert threshold). User has applied the real values in Gitea variables, UptimeRobot and Sentry directly (DSN, Umami website IDs, alert rules, 6 monitors) — not verified by execution from this session (no access to those external accounts), taken on the user's report. Artifacts: operations/plans/monitoring-setup-plan.md, operations/monitoring/monitoring-instructions.md
  • Production Readiness Validation — EXECUTE (includes the dotnet-appsettings compliance gate)

Execution Plan Summary

  • Risk Level: High — three destructive-and-silent failure modes (customer website loss, Data Protection key-ring loss, automatic migration against production)
  • Stages to Execute: Functional Design (×4: U1U4), NFR Design (×2: U3, U4), Infrastructure Design (×2: U6, U7), Code Generation (×7), Build and Test, Deployment Setup, Monitoring Setup, Production Readiness Validation
  • Stages to Skip: User Stories (no end-user functionality), NFR Requirements (already captured), plus per-unit skips as listed above

Current Status

  • Lifecycle Phase: OPERATIONS
  • Current Stage: Deployment Setup complete (2026-07-28, extended 2026-07-29). Monitoring Setup complete 2026-07-29 — user applied real values in Gitea, UptimeRobot and Sentry
  • Next Stage: Production Readiness Validation (includes the dotnet-appsettings compliance gate)
  • Status: CI pipeline built and merging via an open PR; workflow_dispatch manual testing of deploy-test/deploy-production (and therefore a real end-to-end run touching the monitors/alerts just configured) blocked until that PR merges to master (Gitea Actions only shows manual-dispatch workflows that exist on the default branch)

Round 2 Design Record (2026-07-28)

  • Functional Design U3 + U4 complete and committed (357d395)
  • NFR Design U3 — construction/u3-security-headers/nfr-design/ — 11 patterns. All new types in Core/Hosting/Security/, so Core.Tests can reach them (avoids repeating U1's Step 11 deviation)
  • NFR Design U4 — construction/u4-observability/nfr-design/ — 10 patterns. Sentry.AspNetCore 6.8.0 into Core; @sentry/react ^10.68.0 into frontend
  • OPEN-01 CLOSED: correlation ID = W3C trace ID from the ambient Activity, TraceIdentifier as fallback. Rationale: propagates master→slave via traceparent, and equals the traceId ASP.NET Core's ProblemDetails already returns
  • REF-U3-01 raised: BR-U3-22's Umami-origin startup warning is not implementable — the backend cannot read VITE_UMAMI_WEBSITE_ID. Withdrawn from U3 and replaced by a blocking U5 CI gate comparing the frontend build variable against that environment's SecurityHeaders:AllowedScriptOrigins
  • Additions beyond the functional design, each with rationale in the pattern docs: Set-Cookie added to the scrub list; a sentry-tunnel rate limiter; OnRejected on the existing rate limiter (today a 429 leaves no trace anywhere); SentrySdk.FlushAsync before the migration-failure rethrow (otherwise the one Critical event dies with the process)
  • Three defaults chosen rather than escalated (each one line to change, listed at the end of U4's pattern doc): JSON console outside Development, TracesSampleRate 0.1, tunnel cap 200 KB
  • U4 modifies two files from already-committed units — DatabaseMigrationExtensions (U2) and AdminTokenValidator (U1). Both additive; to be named in the U4 commit message

Round 1 Verification Record (2026-07-27)

  • dotnet build SlpModularCms.sln -c Release — 0 errors
  • Backend tests — 253 passed, 0 failed (Core 83, Availability 82, Identity 37, Master 51); baseline was 219
  • New EF migration 20260727203036_AddDataProtectionKeys — verified purely additive
  • Embedded placeholder resource name verified against the compiled assembly manifest
  • Carried to phase-level Build and Test: composed-startup behaviour that needs a running host and a real database — /admin trailing-slash redirect, 404-vs-HTML for missing assets, SPA fallback and placeholder resolution, /health while availability-disabled, MigrateCoreDatabase against SQL Server, and both hosts starting
  • Deviation: U1 plan Step 11 (StaticContentTests) not implemented — the code lives in SlpModularCms.Api, which has no test project by convention; behaviour carried to Build and Test instead. Recorded in the unit's generation-summary.md

Round 2 Verification Record (2026-07-28)

  • dotnet build SlpModularCms.sln -c Release — 0 errors (70 warnings, all pre-existing package advisories)
  • Backend tests — 372 passed, 0 failed at stage close (Core 196, Availability 82, Master 57, Identity 37); was 253 after Round 1, 366 after U4, +6 from the integrity-check fix
  • Frontend tests — 237 passed, 0 failed; baseline 213
  • npx tsc -b clean; eslint on every changed frontend file reports 0 problems; full pnpm run lint unchanged at the pre-existing 5 errors / 1 warning (FR-21, U5)
  • Sentry.AspNetCore 6.8.0 ships a native net10.0 asset — the carried-forward compatibility question is closed
  • @sentry/react 10.68.0; pnpm-lock.yaml diff is additions only
  • Two findings: z.string().url() accepts htp:// in Zod 4 (URL constructor accepts any scheme), so the pre-existing frontend validation never caught the typo BR-U4-24 names — now z.url({ protocol: /^https?$/ }). And appsettings.json comments are verified by DeployedConfigurationTests against the real provider rather than assumed, because the failure mode is both hosts refusing to start
  • One deviation: IAdminTokenValidator collapsed to a single ValidateAdminTokenResult instead of adding an overload. Two methods with an invisible difference at the call site let a substitute silently invert the access decision while both compiled — see U4's generation-summary.md
  • Two defects found in local testing after Round 2, both fixed in this branch rather than filed, per the standing rule that tech debt is for large or high-impact changes only:
    1. Ciphertext written before U2 cannot be decrypted by the database key ring (980dc80). Recorded as ASM-08; local rows cleared and re-registered
    2. VerifyIntegrityAsync could not tell an unreachable slave from one that does not recognise the master, so the only recoverable state was never repaired (6957ec7). Now four distinct outcomes; automatic registration on a rejected key is safe because the slave refuses any key that does not match an existing registration
  • Note: fix 2 is master/slave domain behaviour, not deployment work. It sits in this branch by explicit decision, not because it belongs to the feature
  • Carried to phase-level Build and Test: trace-ID propagation master → slave, TraceId present in rendered console output, tunnel status codes, the security_event tag on a real Sentry event, threshold behaviour end to end, and CSP/HSTS header presence on real static assets and error responses