Chapter 25 · Production Architecture, DDD/CQRS Integration, Reliability, and Capstone
Production Deployment: Migrations, Health Checks, Retries, Pool Sizing, Observability, and Incident Runbooks
Assemble ServiceHub deployment and operations around reviewed migrations, bounded readiness probes, retry-safe transactions, measured pool/timeout settings, observability and concrete incident runbooks rather than startup magic.
Learning outcomes
Separate migration authoring from production deployment ownership and use reviewed scripts/bundles according to change-control needs.
Design liveness/readiness probes that prove required dependencies without generating self-inflicted database load.
Apply execution strategies only to transient faults and make retried operations idempotent.
Treat DbContext pool size, connection-pool size and command timeout as measured independent controls.
Build a minimum production dashboard for EF/database failures and migration state.
Write incident runbooks for pool exhaustion, timeouts, bad migrations and query regressions with rollback/escalation evidence.
1. Production reliability is a system, not one EF option
A production ServiceHub request depends on application capacity,
EF context lifecycle, driver connections, database locks/plans,
schema compatibility, tenant authorization and external
services. A single setting such as
EnableRetryOnFailure() cannot compensate for a bad
migration, unbounded query, exhausted pool or permanent
authentication failure.
2. Migration ownership: build once, review, apply once
# CI authoring/check dotnet ef migrations has-pending-model-changes dotnet ef migrations script --output artifacts/servicehub.sql dotnet ef migrations bundle --output artifacts/efbundle# deployment uses a separate schema-capable credential ./artifacts/efbundle --connection "$DEPLOYMENT_CONNECTION_STRING"
Use scripts where a DBA/security review must inspect SQL; use bundles for coordinated automated deployment when appropriate. Keep the normal runtime identity DML-only where practical. Do not let every app replica race to migrate on startup.
3. Expand/contract rollout keeps old and new binaries compatible
For a breaking rename, first expand: add the
new column/table and dual-read/write or backfill while old code
still works. Deploy new application versions. Then
contract in a later reviewed release after old
versions are gone. Down() is not a backup strategy;
restore procedures and data migration reversibility are separate
concerns.
4. Health checks: liveness is not a database query loop
public static async Task<bool> DatabaseReadyAsync( ServiceHubContext db, CancellationToken ct){ if (!await db.Database.CanConnectAsync(ct)) return false; var pending = await db.Database.GetPendingMigrationsAsync(ct); return !pending.Any();}
A liveness check answers whether the process should be restarted; it should not normally fail just because the database is momentarily unavailable. A readiness check can gate traffic when required dependencies are unavailable. Keep probes cheap, cached/rate-limited where needed, and separate deep diagnostics from high-frequency health endpoints.
5. Retries need transaction delegates and idempotency
var strategy = db.Database.CreateExecutionStrategy();await strategy.ExecuteAsync(async () =>{ await using var tx = await db.Database.BeginTransactionAsync(ct); // operation must be safe to replay or verify after ambiguous commit await handler.ExecuteOnceAsync(db, ct); await tx.CommitAsync(ct);});
Retry transient connection/service errors; do not retry syntax errors, authorization failures, invalid schema or business concurrency conflicts as though they were transient. Ambiguous commit still requires idempotency/client-generated identifiers or explicit verification.
6. Three knobs that teams wrongly collapse into “the pool”
| Control | What it governs | Failure signal |
|---|---|---|
| DbContext pool | Reusable EF context instances | allocation/setup overhead or stale custom state if mismanaged |
| ADO.NET connection pool | Reusable physical/provider connections | waits/timeouts/pool exhaustion, server-session pressure |
| Command timeout | How long one command may execute/wait | timeout exceptions; can hide bad plans/locks if simply increased |
Do not copy pool-size or timeout values from tutorials. Measure concurrency, database limits, connection lifetime, query latency, lock waits and failover behavior in the target topology.
7. Observable production acceptance signals
EF/application: command latency + error rate by operation/provider (low-cardinality) retry attempts/failures; optimistic concurrency conflicts active DbContexts; query-cache hit/miss trend where exposed outbox backlog age/count; dispatcher failuresDatabase/driver: connection usage/waits; lock/deadlock/busy evidence slow-query exemplars + execution plans CPU/IO/storage and transaction-log/WAL pressureDeployment: current migration ID; pending-model-change CI result release version + schema compatibility window
Do not attach raw parameters, connection strings or high-cardinality tenant IDs to every metric. Chapter 24’s data-minimization rules apply to operations dashboards too.
8. Incident runbooks turn telemetry into action
| Incident | First evidence | Unsafe reflex | Runbook direction |
|---|---|---|---|
| Connection exhaustion | pool waits, open sessions, request concurrency | only increase pool max | find leaks/long transactions/load; compare DB session capacity |
| Command timeouts | query tag, plan, locks, duration | raise timeout globally | classify lock vs plan vs network; bound/rewrite/index as evidence shows |
| Bad migration | migration ID, deployment logs, schema/data state | run Down blindly | stop rollout, use tested forward-fix/restore plan, preserve evidence |
| Query regression | tagged SQL, plan change, rows/bytes | disable EF globally | reproduce provider/data distribution, fix query/index/model |
| Outbox backlog | oldest age, failure class, broker health | delete backlog | restore relay, dedupe/idempotency, scale only after cause known |
9. Mandatory lab: rehearse deployment and one incident
- Generate a reviewed SQLite migration artifact and record the expected migration history before/after.
- Implement a cheap readiness check and prove a missing database makes readiness fail without crashing liveness logic.
- Create a deterministic lock/timeout or query-regression scenario and tag the query.
- Capture EF logs/metrics and database evidence without sensitive values.
- Follow a written runbook to diagnose and repair the issue; record rollback/cleanup steps.
- Document why your pool/timeout settings are lab values rather than universal production recommendations.
10. Production judgment and bridge
Reliable EF operations require coordinated schema ownership, compatibility-aware rollout, bounded health checks, retry-safe transactions, measured resource settings, secure observability and rehearsed incident actions. The final lesson asks the learner to defend all of those choices together in a capstone instead of presenting a happy-path CRUD demo.
Check your understanding
- Why avoid runtime migrations from every app instance?
- What is the difference between liveness and readiness?
- Can execution strategies retry business concurrency conflicts?
- Is DbContext pooling the same as connection pooling?
- Why not solve timeouts by increasing the timeout first?
- What makes an incident runbook useful?
Review the answers
1. Schema changes need coordinated ownership, review and permissions; multiple replicas add race/blast-radius risk even with migration locking.
2. Liveness asks whether the process should be restarted; readiness asks whether it can currently serve traffic with required dependencies.
3. They target transient provider failures; concurrency conflicts require business resolution policy.
4. No. EF reuses context objects; the ADO.NET provider separately pools physical connections.
5. Timeouts may be symptoms of locks, bad plans, unbounded work or resource saturation; increasing the value can worsen capacity.
6. Specific evidence, decision points, safe actions, escalation/rollback and verification rather than generic advice.
Authoritative references
- Applying migrations — production scripts, bundles, migration locking and deployment identity guidance
- ASP.NET Core health checks — liveness/readiness and EF DbContext probes
- Connection resiliency — execution strategies, user transactions and retry semantics
- Advanced performance topics — DbContext versus connection pooling
- EF Core metrics — current EF observability signals
- Simple logging — structured diagnostics and sensitive-data constraints