Chapter 18 · Diagnostics, Logging, Interceptors, Metrics, and Observability

Build an EF Core Production Dashboard with Error Rates, Latency, Pool Pressure, and Migration State

Turn the chapter into an operational dashboard/runbook: define low-cardinality EF and provider signals, correlate errors and slow-query exemplars, track pool and migration state, and create alerts that point engineers toward evidence rather than storing sensitive SQL or parameter payloads.

Advanced150–190 minutesdashboard + runbook labEF Core 10.0.11 · .NET 10.0.11SQLite provider 10.0.11 mandatory baselinedotnet-ef 10.0.11 · SDK 10.0.400Observability reviewed: August 2026

Learning outcomes

01

Design an EF Core production dashboard from user-impact and failure questions instead of collecting every available counter.

02

Combine EF metrics with provider/database signals for latency, errors, retries/concurrency failures, context health, connection pressure, and migration state.

03

Use low-cardinality dimensions and sampled trace exemplars rather than SQL text, parameter values, trace IDs, or tenant/work-order identifiers as metric labels.

04

Differentiate DbContext pool indicators from ADO.NET connection-pool/server-session pressure and document provider-specific gaps.

05

Track deployed migration state and detect unexpected schema/version divergence without granting the runtime unnecessary DDL permissions.

06

Write a runbook that moves from dashboard symptom to trace/log/plan/database evidence and records when a signal is insufficient.

1. A dashboard is an operational decision surface, not a telemetry inventory

The ServiceHub team can now emit safe logs, EF metrics, trace correlation, interceptor observations, query tags, and database-plan evidence. The temptation is to put all of it on one screen. A useful dashboard instead begins with operational questions: Are users seeing database-related errors? Are requests getting slower? Are retries/concurrency conflicts rising? Are contexts or connections leaking/saturating? Did a deployment change migration state? Which operation should an engineer inspect next?

Frozen lab baseline

Mandatory baseline: .NET runtime 10.0.11, SDK 10.0.400, Microsoft.EntityFrameworkCore / Design / Sqlite and dotnet-ef 10.0.11. SQLite is the free local database. EF Core 11 preview APIs are out of scope. OpenTelemetry core/hosting 1.18.0 may be used for an optional console telemetry path; OpenTelemetry.Instrumentation.EntityFrameworkCore 1.18.0-beta.1 remains prerelease/beta and is never required for the mandatory labs.

2. Start with four symptom groups

Group Primary signals Next evidence
User-impact latency request/job p50/p95/p99; EF command duration distribution; slow-operation count trace exemplar → command/log → database plan/lock metrics
Errors/reliability EF command failures; timeout class; execution-strategy failures; optimistic-concurrency failures exception type/provider code; trace; retry/concurrency policy
Resource pressure active DbContexts; driver pool wait/active/idle where provider exposes it; DB sessions/locks pool/provider counters, server metrics, context lifetime code
Deployment/schema deployed app version; expected migration; __EFMigrationsHistory state; drift check outcome deployment record, migration artifact, DB schema inspection

EF does not expose all driver/database pool metrics itself. The dashboard must visibly label the source of each signal so engineers do not interpret active_dbcontexts as “database connections in use.”

3. Build a low-cardinality ServiceHub telemetry contract

text · recommended bounded dimensions
service.name=servicehub-apiservice.version=2026.08.27.1environment=productionoperation=dispatch.queue | workorder.update | outbox.dispatchprovider=sqlite | sqlserver | postgresqloutcome=success | error | timeout | concurrency_conflicterror.type=<bounded exception/failure class>db.role=primary | read-replica   # only if topology actually has these rolesNever metric labels:trace_id, span_id, work_order_number, customer_name,tenant_id, raw SQL, parameter values, exception message

A metric backend creates time-series from dimension combinations. Unbounded user/business identifiers create explosive cardinality and can leak sensitive data. Unique trace IDs belong in the trace/log system and may be attached as exemplars when the backend supports that pattern.

4. EF Core health metrics: what to graph and how to interpret them

EF metric Useful dashboard view Interpretation trap
active_dbcontexts gauge/trend per process instance with context pooling, idle pooled contexts are included; not DB connections
queries rate and total by service instance does not indicate rows, cost, or N+1 by itself
savechanges rate one SaveChanges may persist many entity changes
compiled query cache hits/misses derived hit ratio after warm-up low ratio may indicate dynamic query shapes; still measure user impact
execution strategy failures rate + trace exemplars individual retry failures may occur before eventual success
optimistic concurrency failures rate + operation business conflicts are not transient infrastructure failures

Do not invent one universal alert threshold. Establish a baseline from real ServiceHub traffic, then alert on sustained deviations tied to user impact or resource risk.

5. Connection pressure must come from the driver/database layer

EF opens connections late and closes them after operations; ADO.NET providers generally pool physical connections separately. A rising EF command latency can therefore reflect waiting for a driver pool slot, database lock, server CPU, or network path. Instrument the exact provider in production. SQLite is in-process/file-based; SQL Server and PostgreSQL have very different connection/session metrics and network topology.

Do not fabricate “EF connection pool usage”

EF Core does not own a universal cross-provider connection pool metric. Use provider/ADO.NET/database metrics and label the source/version.

6. Migration state belongs on the operational surface

Chapter 15 separated migration authoring from deployment. A dashboard can expose an application build's expected migration identifier and a read-only probe of __EFMigrationsHistory. The runtime account does not need DDL permission just to report state.

csharp · read-only migration-state probe
var applied = await db.Database    .GetAppliedMigrationsAsync(cancellationToken);var latestApplied = applied.LastOrDefault();logger.LogInformation(    "Migration state observed; LatestApplied={LatestApplied}; AppVersion={AppVersion}",    latestApplied,    typeof(ServiceHubContext).Assembly.GetName().Version?.ToString() ?? "unknown");

Do not run Database.Migrate() from the dashboard health check. Report divergence; let the coordinated migration owner/pipeline resolve it. For stronger drift detection, use a separately permissioned deployment/verification job rather than granting production application instances broad schema rights.

7. Slow-query panel: aggregate first, exemplars second

Graph command latency by stable operation and provider. Keep a bounded set of sampled slow traces as exemplars. A click should lead to a trace that includes the operation, command event, and safe query tag. From there, the runbook tells the engineer how to obtain an execution plan against an authorized environment. Do not store every raw SQL statement and parameter list in the metric system.

text · example dashboard rows
Panel: EF command latency by operation  dispatch.queue    p95=...  error_rate=...  slow_exemplars=[trace A, trace B]  workorder.update  p95=...  concurrency_conflicts=...  outbox.dispatch   p95=...  retry_failures=...Panel: process / pool health  active_dbcontexts=...  driver_pool_wait=...      # provider metric if available  database_sessions=...     # DB metric if availablePanel: deployment  app_version=...  expected_migration=...  latest_applied_migration=...  drift_check=pass|fail|unknown

8. Error taxonomy determines the runbook branch

Failure class Example Runbook response
business concurrency DbUpdateConcurrencyException / optimistic conflict follow Chapter 13 merge/retry/business policy; do not infrastructure-retry blindly
transient provider/infrastructure provider-classified network/transient fault inspect execution strategy/retry count and infrastructure health
timeout/lock provider timeout/busy/lock inspect lock/pool/plan and operation deadline; do not just increase timeout
translation/model InvalidOperationException before SQL inspect query/model change, deployment version, logs; database plan is irrelevant
constraint/update DbUpdateException / provider constraint inspect domain validation, constraint metadata, SQL state/code; avoid exposing row data in alert

9. Failure case: alert on every EF exception at page severity

A concurrency conflict, one expected unique-constraint rejection, a brief transient retry, and a widespread database outage all page the same on-call engineer. Alert fatigue follows, and real outages are ignored.

Repair: derive alerting from user impact, rate, duration, failure class, and saturation. Concurrency conflicts may be a business KPI; transient retry failures may warn only when eventual operation failures or latency rise; constraint errors may be application-quality signals. Page on sustained user-impacting failure or critical resource exhaustion, not on every exception object.

10. Runbook: dashboard symptom → evidence chain

text · ServiceHub EF incident runbook skeleton
1. Identify affected operation, provider, deployment version, time window.2. Check request/job error and latency impact.3. Check EF metric trends: queries, SaveChanges, cache misses,   execution-strategy failures, concurrency failures, active contexts.4. Open one sampled trace exemplar.5. Inspect safe EF command/log event: duration, failure class, tag.6. Check provider pool/network and database sessions/locks/CPU/I/O.7. Capture/compare the database execution plan for the tagged query.8. Verify recent migration/schema/index changes.9. Choose repair based on evidence; document rollback/mitigation.10. Verify metrics return to baseline and close the loop with a regression test.

11. Mandatory lab: create a local dashboard model without a commercial backend

You do not need Grafana, Azure Monitor, Datadog, New Relic, or another hosted product to learn the model. Use console-exported metrics/logs plus a small local JSON/Markdown snapshot. The task is to define signals and runbook logic, not to require a vendor.

  1. Execute deterministic read/write/concurrency scenarios against the disposable SQLite database.
  2. Capture EF meter values before/after each scenario.
  3. Produce safe command events with operation tags and trace IDs.
  4. Record SQLite busy/locked behavior in a dedicated timeout/lock scenario.
  5. Read the latest applied migration identifier.
  6. Create a dashboard snapshot with only bounded dimensions.
  7. Run the runbook from one “slow dispatch.queue” symptom to its SQLite plan evidence.
  8. Search the snapshot for fake sensitive values and fail the lab if any are present.

12. Production judgment and bridge to provider deep dives

A production dashboard should help an engineer decide what evidence to collect next. EF metrics explain application data-access behavior; provider/driver metrics explain connection and transport behavior; database metrics/plans explain execution and contention. Keep those layers separate but correlated. Use traces for unique request exemplars, metrics for aggregates, logs for discrete structured events, and plans for database mechanics. Chapter 19 now applies this observability discipline to SQL Server and Azure SQL provider-specific type mappings, temporal/JSON features, resiliency, indexes, DDL, and places where EF abstractions intentionally leak.

Check your understanding

  1. Why is active_dbcontexts not a connection-pool utilization metric?
  2. Which dimensions should never become ordinary metric labels?
  3. How should migration divergence be handled?
  4. Why should optimistic concurrency failures have their own classification?
  5. What is the purpose of a slow-query exemplar?
  6. What makes an observability alert actionable?
Review the answers

1. It counts active DbContext instances and, with context pooling, includes pooled contexts not currently in use; database connections are a separate driver/server resource.

2. Unbounded/unique or sensitive values such as trace IDs, work-order/customer/tenant identifiers, raw SQL and parameter values.

3. Report it with read-only state and route it to the coordinated deployment/migration owner; do not silently run migrations from every application instance.

4. They represent business-state conflicts and require Chapter 13 conflict policy, not blind transient infrastructure retries.

5. To link an aggregated latency signal to a small number of representative traces/logs for deeper evidence collection without high-cardinality metric labels.

6. It ties a sustained symptom to user impact/resource risk and provides a runbook path to traces, provider/database metrics, plans, deployment state and verification.

Authoritative references

Observability and diagnostics are version-, provider-, and topology-sensitive. Re-check these primary sources before standardizing a production telemetry contract.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.

\n