Chapter 25 · Production Architecture, DDD/CQRS Integration, Reliability, and Capstone

Production Deployment: Migrations, Health Checks, Retries, Pool Sizing, Observability, and Incident Runbooks

Assemble ServiceHub deployment and operations around reviewed migrations, bounded readiness probes, retry-safe transactions, measured pool/timeout settings, observability and concrete incident runbooks rather than startup magic.

Advanced210–300 minutesdeployment + incident labEF Core 10.0.11 · Microsoft.EntityFrameworkCore.Sqlite 10.0.11 · dotnet-ef 10.0.11 · .NET 10.0.11 · SDK 10.0.400Mandatory free local SQLite path · Dapper 2.1.79 optional · production-like server provider optionalArchitecture/package/platform status reviewed: August 27, 2026

Learning outcomes

01

Separate migration authoring from production deployment ownership and use reviewed scripts/bundles according to change-control needs.

02

Design liveness/readiness probes that prove required dependencies without generating self-inflicted database load.

03

Apply execution strategies only to transient faults and make retried operations idempotent.

04

Treat DbContext pool size, connection-pool size and command timeout as measured independent controls.

05

Build a minimum production dashboard for EF/database failures and migration state.

06

Write incident runbooks for pool exhaustion, timeouts, bad migrations and query regressions with rollback/escalation evidence.

1. Production reliability is a system, not one EF option

A production ServiceHub request depends on application capacity, EF context lifecycle, driver connections, database locks/plans, schema compatibility, tenant authorization and external services. A single setting such as EnableRetryOnFailure() cannot compensate for a bad migration, unbounded query, exhausted pool or permanent authentication failure.

2. Migration ownership: build once, review, apply once

Deployment artifact examples
# CI authoring/check dotnet ef migrations has-pending-model-changes dotnet ef migrations script --output artifacts/servicehub.sql dotnet ef migrations bundle --output artifacts/efbundle# deployment uses a separate schema-capable credential ./artifacts/efbundle --connection "$DEPLOYMENT_CONNECTION_STRING"

Use scripts where a DBA/security review must inspect SQL; use bundles for coordinated automated deployment when appropriate. Keep the normal runtime identity DML-only where practical. Do not let every app replica race to migrate on startup.

3. Expand/contract rollout keeps old and new binaries compatible

For a breaking rename, first expand: add the new column/table and dual-read/write or backfill while old code still works. Deploy new application versions. Then contract in a later reviewed release after old versions are gone. Down() is not a backup strategy; restore procedures and data migration reversibility are separate concerns.

4. Health checks: liveness is not a database query loop

C# · bounded readiness concept
public static async Task<bool> DatabaseReadyAsync(    ServiceHubContext db, CancellationToken ct){    if (!await db.Database.CanConnectAsync(ct)) return false;    var pending = await db.Database.GetPendingMigrationsAsync(ct);    return !pending.Any();}

A liveness check answers whether the process should be restarted; it should not normally fail just because the database is momentarily unavailable. A readiness check can gate traffic when required dependencies are unavailable. Keep probes cheap, cached/rate-limited where needed, and separate deep diagnostics from high-frequency health endpoints.

5. Retries need transaction delegates and idempotency

C# · execution strategy owns replayable transaction
var strategy = db.Database.CreateExecutionStrategy();await strategy.ExecuteAsync(async () =>{    await using var tx = await db.Database.BeginTransactionAsync(ct);    // operation must be safe to replay or verify after ambiguous commit    await handler.ExecuteOnceAsync(db, ct);    await tx.CommitAsync(ct);});

Retry transient connection/service errors; do not retry syntax errors, authorization failures, invalid schema or business concurrency conflicts as though they were transient. Ambiguous commit still requires idempotency/client-generated identifiers or explicit verification.

6. Three knobs that teams wrongly collapse into “the pool”

Control What it governs Failure signal
DbContext pool Reusable EF context instances allocation/setup overhead or stale custom state if mismanaged
ADO.NET connection pool Reusable physical/provider connections waits/timeouts/pool exhaustion, server-session pressure
Command timeout How long one command may execute/wait timeout exceptions; can hide bad plans/locks if simply increased

Do not copy pool-size or timeout values from tutorials. Measure concurrency, database limits, connection lifetime, query latency, lock waits and failover behavior in the target topology.

7. Observable production acceptance signals

Dashboard / alert dimensions
EF/application:  command latency + error rate by operation/provider (low-cardinality)  retry attempts/failures; optimistic concurrency conflicts  active DbContexts; query-cache hit/miss trend where exposed  outbox backlog age/count; dispatcher failuresDatabase/driver:  connection usage/waits; lock/deadlock/busy evidence  slow-query exemplars + execution plans  CPU/IO/storage and transaction-log/WAL pressureDeployment:  current migration ID; pending-model-change CI result  release version + schema compatibility window

Do not attach raw parameters, connection strings or high-cardinality tenant IDs to every metric. Chapter 24’s data-minimization rules apply to operations dashboards too.

8. Incident runbooks turn telemetry into action

Incident First evidence Unsafe reflex Runbook direction
Connection exhaustion pool waits, open sessions, request concurrency only increase pool max find leaks/long transactions/load; compare DB session capacity
Command timeouts query tag, plan, locks, duration raise timeout globally classify lock vs plan vs network; bound/rewrite/index as evidence shows
Bad migration migration ID, deployment logs, schema/data state run Down blindly stop rollout, use tested forward-fix/restore plan, preserve evidence
Query regression tagged SQL, plan change, rows/bytes disable EF globally reproduce provider/data distribution, fix query/index/model
Outbox backlog oldest age, failure class, broker health delete backlog restore relay, dedupe/idempotency, scale only after cause known

9. Mandatory lab: rehearse deployment and one incident

  1. Generate a reviewed SQLite migration artifact and record the expected migration history before/after.
  2. Implement a cheap readiness check and prove a missing database makes readiness fail without crashing liveness logic.
  3. Create a deterministic lock/timeout or query-regression scenario and tag the query.
  4. Capture EF logs/metrics and database evidence without sensitive values.
  5. Follow a written runbook to diagnose and repair the issue; record rollback/cleanup steps.
  6. Document why your pool/timeout settings are lab values rather than universal production recommendations.

10. Production judgment and bridge

Reliable EF operations require coordinated schema ownership, compatibility-aware rollout, bounded health checks, retry-safe transactions, measured resource settings, secure observability and rehearsed incident actions. The final lesson asks the learner to defend all of those choices together in a capstone instead of presenting a happy-path CRUD demo.

Check your understanding

  1. Why avoid runtime migrations from every app instance?
  2. What is the difference between liveness and readiness?
  3. Can execution strategies retry business concurrency conflicts?
  4. Is DbContext pooling the same as connection pooling?
  5. Why not solve timeouts by increasing the timeout first?
  6. What makes an incident runbook useful?
Review the answers

1. Schema changes need coordinated ownership, review and permissions; multiple replicas add race/blast-radius risk even with migration locking.

2. Liveness asks whether the process should be restarted; readiness asks whether it can currently serve traffic with required dependencies.

3. They target transient provider failures; concurrency conflicts require business resolution policy.

4. No. EF reuses context objects; the ADO.NET provider separately pools physical connections.

5. Timeouts may be symptoms of locks, bad plans, unbounded work or resource saturation; increasing the value can worsen capacity.

6. Specific evidence, decision points, safe actions, escalation/rollback and verification rather than generic advice.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.

\n