Chapter 18 · Diagnostics, Logging, Interceptors, Metrics, and Observability
Build an EF Core Production Dashboard with Error Rates, Latency, Pool Pressure, and Migration State
Turn the chapter into an operational dashboard/runbook: define low-cardinality EF and provider signals, correlate errors and slow-query exemplars, track pool and migration state, and create alerts that point engineers toward evidence rather than storing sensitive SQL or parameter payloads.
Learning outcomes
Design an EF Core production dashboard from user-impact and failure questions instead of collecting every available counter.
Combine EF metrics with provider/database signals for latency, errors, retries/concurrency failures, context health, connection pressure, and migration state.
Use low-cardinality dimensions and sampled trace exemplars rather than SQL text, parameter values, trace IDs, or tenant/work-order identifiers as metric labels.
Differentiate DbContext pool indicators from ADO.NET connection-pool/server-session pressure and document provider-specific gaps.
Track deployed migration state and detect unexpected schema/version divergence without granting the runtime unnecessary DDL permissions.
Write a runbook that moves from dashboard symptom to trace/log/plan/database evidence and records when a signal is insufficient.
1. A dashboard is an operational decision surface, not a telemetry inventory
The ServiceHub team can now emit safe logs, EF metrics, trace correlation, interceptor observations, query tags, and database-plan evidence. The temptation is to put all of it on one screen. A useful dashboard instead begins with operational questions: Are users seeing database-related errors? Are requests getting slower? Are retries/concurrency conflicts rising? Are contexts or connections leaking/saturating? Did a deployment change migration state? Which operation should an engineer inspect next?
Mandatory baseline: .NET runtime 10.0.11, SDK 10.0.400, Microsoft.EntityFrameworkCore / Design / Sqlite and dotnet-ef 10.0.11. SQLite is the free local database. EF Core 11 preview APIs are out of scope. OpenTelemetry core/hosting 1.18.0 may be used for an optional console telemetry path; OpenTelemetry.Instrumentation.EntityFrameworkCore 1.18.0-beta.1 remains prerelease/beta and is never required for the mandatory labs.
2. Start with four symptom groups
| Group | Primary signals | Next evidence |
|---|---|---|
| User-impact latency | request/job p50/p95/p99; EF command duration distribution; slow-operation count | trace exemplar → command/log → database plan/lock metrics |
| Errors/reliability | EF command failures; timeout class; execution-strategy failures; optimistic-concurrency failures | exception type/provider code; trace; retry/concurrency policy |
| Resource pressure | active DbContexts; driver pool wait/active/idle where provider exposes it; DB sessions/locks | pool/provider counters, server metrics, context lifetime code |
| Deployment/schema |
deployed app version; expected migration;
__EFMigrationsHistory state; drift check
outcome
|
deployment record, migration artifact, DB schema inspection |
EF does not expose all driver/database pool metrics itself. The
dashboard must visibly label the source of each signal so
engineers do not interpret active_dbcontexts as
“database connections in use.”
3. Build a low-cardinality ServiceHub telemetry contract
service.name=servicehub-apiservice.version=2026.08.27.1environment=productionoperation=dispatch.queue | workorder.update | outbox.dispatchprovider=sqlite | sqlserver | postgresqloutcome=success | error | timeout | concurrency_conflicterror.type=<bounded exception/failure class>db.role=primary | read-replica # only if topology actually has these rolesNever metric labels:trace_id, span_id, work_order_number, customer_name,tenant_id, raw SQL, parameter values, exception message
A metric backend creates time-series from dimension combinations. Unbounded user/business identifiers create explosive cardinality and can leak sensitive data. Unique trace IDs belong in the trace/log system and may be attached as exemplars when the backend supports that pattern.
4. EF Core health metrics: what to graph and how to interpret them
| EF metric | Useful dashboard view | Interpretation trap |
|---|---|---|
| active_dbcontexts | gauge/trend per process instance | with context pooling, idle pooled contexts are included; not DB connections |
| queries | rate and total by service instance | does not indicate rows, cost, or N+1 by itself |
| savechanges | rate | one SaveChanges may persist many entity changes |
| compiled query cache hits/misses | derived hit ratio after warm-up | low ratio may indicate dynamic query shapes; still measure user impact |
| execution strategy failures | rate + trace exemplars | individual retry failures may occur before eventual success |
| optimistic concurrency failures | rate + operation | business conflicts are not transient infrastructure failures |
Do not invent one universal alert threshold. Establish a baseline from real ServiceHub traffic, then alert on sustained deviations tied to user impact or resource risk.
5. Connection pressure must come from the driver/database layer
EF opens connections late and closes them after operations; ADO.NET providers generally pool physical connections separately. A rising EF command latency can therefore reflect waiting for a driver pool slot, database lock, server CPU, or network path. Instrument the exact provider in production. SQLite is in-process/file-based; SQL Server and PostgreSQL have very different connection/session metrics and network topology.
EF Core does not own a universal cross-provider connection pool metric. Use provider/ADO.NET/database metrics and label the source/version.
6. Migration state belongs on the operational surface
Chapter 15 separated migration authoring from deployment. A
dashboard can expose an application build's expected migration
identifier and a read-only probe of
__EFMigrationsHistory. The runtime account does not
need DDL permission just to report state.
var applied = await db.Database .GetAppliedMigrationsAsync(cancellationToken);var latestApplied = applied.LastOrDefault();logger.LogInformation( "Migration state observed; LatestApplied={LatestApplied}; AppVersion={AppVersion}", latestApplied, typeof(ServiceHubContext).Assembly.GetName().Version?.ToString() ?? "unknown");
Do not run Database.Migrate() from the dashboard
health check. Report divergence; let the coordinated migration
owner/pipeline resolve it. For stronger drift detection, use a
separately permissioned deployment/verification job rather than
granting production application instances broad schema rights.
7. Slow-query panel: aggregate first, exemplars second
Graph command latency by stable operation and
provider. Keep a bounded set of sampled slow traces
as exemplars. A click should lead to a trace that includes the
operation, command event, and safe query tag. From there, the
runbook tells the engineer how to obtain an execution plan
against an authorized environment. Do not store every raw SQL
statement and parameter list in the metric system.
Panel: EF command latency by operation dispatch.queue p95=... error_rate=... slow_exemplars=[trace A, trace B] workorder.update p95=... concurrency_conflicts=... outbox.dispatch p95=... retry_failures=...Panel: process / pool health active_dbcontexts=... driver_pool_wait=... # provider metric if available database_sessions=... # DB metric if availablePanel: deployment app_version=... expected_migration=... latest_applied_migration=... drift_check=pass|fail|unknown
8. Error taxonomy determines the runbook branch
| Failure class | Example | Runbook response |
|---|---|---|
| business concurrency |
DbUpdateConcurrencyException / optimistic
conflict
|
follow Chapter 13 merge/retry/business policy; do not infrastructure-retry blindly |
| transient provider/infrastructure | provider-classified network/transient fault | inspect execution strategy/retry count and infrastructure health |
| timeout/lock | provider timeout/busy/lock | inspect lock/pool/plan and operation deadline; do not just increase timeout |
| translation/model | InvalidOperationException before SQL | inspect query/model change, deployment version, logs; database plan is irrelevant |
| constraint/update | DbUpdateException / provider constraint | inspect domain validation, constraint metadata, SQL state/code; avoid exposing row data in alert |
9. Failure case: alert on every EF exception at page severity
A concurrency conflict, one expected unique-constraint rejection, a brief transient retry, and a widespread database outage all page the same on-call engineer. Alert fatigue follows, and real outages are ignored.
Repair: derive alerting from user impact, rate, duration, failure class, and saturation. Concurrency conflicts may be a business KPI; transient retry failures may warn only when eventual operation failures or latency rise; constraint errors may be application-quality signals. Page on sustained user-impacting failure or critical resource exhaustion, not on every exception object.
10. Runbook: dashboard symptom → evidence chain
1. Identify affected operation, provider, deployment version, time window.2. Check request/job error and latency impact.3. Check EF metric trends: queries, SaveChanges, cache misses, execution-strategy failures, concurrency failures, active contexts.4. Open one sampled trace exemplar.5. Inspect safe EF command/log event: duration, failure class, tag.6. Check provider pool/network and database sessions/locks/CPU/I/O.7. Capture/compare the database execution plan for the tagged query.8. Verify recent migration/schema/index changes.9. Choose repair based on evidence; document rollback/mitigation.10. Verify metrics return to baseline and close the loop with a regression test.
11. Mandatory lab: create a local dashboard model without a commercial backend
You do not need Grafana, Azure Monitor, Datadog, New Relic, or another hosted product to learn the model. Use console-exported metrics/logs plus a small local JSON/Markdown snapshot. The task is to define signals and runbook logic, not to require a vendor.
- Execute deterministic read/write/concurrency scenarios against the disposable SQLite database.
- Capture EF meter values before/after each scenario.
- Produce safe command events with operation tags and trace IDs.
- Record SQLite busy/locked behavior in a dedicated timeout/lock scenario.
- Read the latest applied migration identifier.
- Create a dashboard snapshot with only bounded dimensions.
- Run the runbook from one “slow dispatch.queue” symptom to its SQLite plan evidence.
- Search the snapshot for fake sensitive values and fail the lab if any are present.
12. Production judgment and bridge to provider deep dives
A production dashboard should help an engineer decide what evidence to collect next. EF metrics explain application data-access behavior; provider/driver metrics explain connection and transport behavior; database metrics/plans explain execution and contention. Keep those layers separate but correlated. Use traces for unique request exemplars, metrics for aggregates, logs for discrete structured events, and plans for database mechanics. Chapter 19 now applies this observability discipline to SQL Server and Azure SQL provider-specific type mappings, temporal/JSON features, resiliency, indexes, DDL, and places where EF abstractions intentionally leak.
Check your understanding
- Why is active_dbcontexts not a connection-pool utilization metric?
- Which dimensions should never become ordinary metric labels?
- How should migration divergence be handled?
- Why should optimistic concurrency failures have their own classification?
- What is the purpose of a slow-query exemplar?
- What makes an observability alert actionable?
Review the answers
1. It counts active DbContext instances and, with context pooling, includes pooled contexts not currently in use; database connections are a separate driver/server resource.
2. Unbounded/unique or sensitive values such as trace IDs, work-order/customer/tenant identifiers, raw SQL and parameter values.
3. Report it with read-only state and route it to the coordinated deployment/migration owner; do not silently run migrations from every application instance.
4. They represent business-state conflicts and require Chapter 13 conflict policy, not blind transient infrastructure retries.
5. To link an aggregated latency signal to a small number of representative traces/logs for deeper evidence collection without high-cardinality metric labels.
6. It ties a sustained symptom to user impact/resource risk and provides a runbook path to traces, provider/database metrics, plans, deployment state and verification.
Authoritative references
Observability and diagnostics are version-, provider-, and topology-sensitive. Re-check these primary sources before standardizing a production telemetry contract.
- Metrics in EF Core — current EF Core meter instruments and EventCounters
- Performance Diagnosis - EF Core — logs, query tags and database-plan correlation
- Advanced Performance Topics - EF Core — context pooling, query cache and performance evidence
- Migrations overview - EF Core — migration history and deployment state context
- .NET observability with OpenTelemetry — logs, metrics and traces as complementary signals
- Microsoft.Data.Sqlite database errors — SQLite lock/busy behavior and command timeout evidence