Chapter 16 · Reverse Engineering and Database-First Workflows
Re-Scaffolding, Schema Drift, Handwritten Changes, and Sustainable Database-First Ownership
Treat re-scaffolding as a repeatable code-generation pipeline: stage output, diff it, classify schema drift, preserve handwritten extensions, and decide explicitly whether the database or the EF model owns future schema change.
Learning outcomes
Define an explicit source-of-truth policy for database-first schema ownership.
Re-scaffold into a scratch directory and review generated-code diffs before replacing committed generated output.
Distinguish intentional database change, unreviewed database drift, generated-code drift, and application-only custom behavior.
Use repeatable schema fingerprints and runtime model probes as evidence rather than relying on timestamps or “it compiles.”
Demonstrate how direct edits are lost and how partial/separate code survives schema evolution.
Plan a deliberate transition if the team decides to move from database-owned schema to EF migrations instead of mixing ownership models accidentally.
1. The practical problem: six months later, nobody knows what changed
The ServiceHub legacy database gains
external_reference directly through an operations
change. One developer manually adds a C# property instead of
re-scaffolding. Another re-scaffolds from production with a
newer provider. A third begins creating EF migrations. Now
source code, database schema, generated model, and migration
history tell different stories. This is
schema drift plus ownership drift.
Course baseline: .NET 10 runtime 10.0.11, SDK 10.0.400, EF Core/dotnet-ef/Microsoft.EntityFrameworkCore.Design/Microsoft.EntityFrameworkCore.Sqlite 10.0.11. SQLite is the mandatory free/local provider. EF Core 11 preview APIs are out of scope unless clearly labeled.
2. Choose the authority before choosing the command
| Ownership model | Authoritative schema change | Generated EF role | Main risk |
|---|---|---|---|
| Database-first | DBA/schema pipeline/database project | Derivative adapter; re-scaffold after approved DB change | Hand edits drift from database |
| Migrations-first | Reviewed EF migration/model evolution | Primary schema-change artifact | Out-of-band DB hotfixes drift from snapshot/history |
| Shared/legacy transition | Explicitly coordinated process | Temporary bridge | Both sides believe they own DDL |
Chapter 16 assumes database-first ownership for this legacy
integration. Chapter 15 remains valid for the course's
model-first ServiceHubContext. Do not use both
ownership models against the same production schema without a
written transition plan.
3. Re-scaffold into scratch first
A safe refresh does not start with --force against
committed output. Generate into a clean scratch directory using
the same tool/provider/options recorded in the project, then
diff.
rm -rf .scaffold-nextmkdir -p .scaffold-nextdotnet ef dbcontext scaffold "Data Source=servicehub-legacy.db" Microsoft.EntityFrameworkCore.Sqlite --table work_orders --table technicians --table open_work_order_summary --context ServiceHubLegacyContext --context-dir .scaffold-next/Context --output-dir .scaffold-next/Entities --namespace ServiceHub.LegacyScaffold.Persistence.Generated.Entities --context-namespace ServiceHub.LegacyScaffold.Persistence.Generated --no-onconfiguring --forcegit diff --no-index Persistence/Generated .scaffold-next || true
On Windows, use a temporary directory and your normal diff tool. What matters is that the candidate generated output is reviewable before it replaces the committed generated contract.
4. Classify every diff before accepting it
| Observed diff | Likely cause | Action |
|---|---|---|
| New property + column mapping | Approved DB column/table change | Review semantics, tests, then adopt generated update |
| Hundreds of whitespace/naming changes | Tool/provider/template/options changed | Stop; reproduce with pinned versions/options before attributing to schema |
| Navigation disappears | Filtered table omitted or FK changed | Verify selected model boundary and FK metadata |
| Concurrency flag changes | Provider metadata/model customization changed | Verify provider behavior and handwritten partial configuration |
| Handwritten method disappears | Direct edit in generated file | Move to partial/separate code; do not “restore” it inside generated source |
| DB changed but generated diff is empty | Unsupported/unmodeled metadata or wrong source DB | Inspect schema/provider connection and model metadata directly |
A large diff is not automatically “schema drift.” Code generation version drift can be equally large. Record command line, tool version, provider version, and database target in the review evidence.
5. Capture a lightweight database schema fingerprint
A database-first CI process benefits from a deterministic
representation of the source schema. SQLite's
sqlite_master can provide a simple lab fingerprint.
Production systems should use engine-appropriate catalog queries
or a reviewed schema artifact/database project.
SELECT type, name, tbl_name, sqlFROM sqlite_masterWHERE name NOT LIKE 'sqlite_%'ORDER BY type, name;
import hashlib, sqlite3con = sqlite3.connect("servicehub-legacy.db")rows = con.execute("""SELECT type, name, tbl_name, coalesce(sql, '')FROM sqlite_masterWHERE name NOT LIKE 'sqlite_%'ORDER BY type, name""").fetchall()canonical = "\n".join("|".join(map(str, row)) for row in rows)print(hashlib.sha256(canonical.encode()).hexdigest())
The hash is evidence that the observed schema representation changed. It does not explain whether the change is safe, authorized, or compatible.
6. Pair schema evidence with EF model evidence
After re-scaffolding, compile the candidate model and emit a normalized list of entities, properties, keys, foreign keys, indexes, and store mappings. The database fingerprint and model fingerprint answer different questions; comparing both catches connection mistakes and generation drift earlier.
await using var db = new ServiceHubLegacyContext(options);var lines = db.Model.GetEntityTypes() .OrderBy(e => e.Name) .SelectMany(e => new[] { $"ENTITY|{e.DisplayName()}|{e.GetTableName()}|{e.GetViewName()}" } .Concat(e.GetProperties() .OrderBy(p => p.Name) .Select(p => $"PROP|{e.DisplayName()}|{p.Name}|{p.ClrType.FullName}|nullable={p.IsNullable}|concurrency={p.IsConcurrencyToken}")));foreach (var line in lines) Console.WriteLine(line);
Do not hard-code a hash into production logic. This is build/review evidence for an integration contract.
7. Demonstrate drift safely: add a legacy column
ALTER TABLE work_orders ADD COLUMN external_reference TEXT NULL;CREATE INDEX ix_work_orders_external_reference ON work_orders(external_reference);
Re-run the scratch scaffold. Expected diff: one property, column mapping, and index metadata. If the generated diff also rewrites unrelated types or names, investigate tool/options before accepting it.
Adding ExternalReference manually may make
compilation pass, but it bypasses the repeatable generation
contract and can miss provider facets/indexes/defaults. In
database-first ownership, fix the schema/generation pipeline,
then regenerate.
8. Re-scaffold acceptance workflow
- Identify the exact database/schema artifact and approved change ticket/version.
- Restore/create a disposable database at that schema version.
-
Verify
dotnet-ef, provider, Design package, templates, and scaffold options. - Generate into scratch—not committed output.
- Diff and classify every change.
- Run model metadata probes and representative query/write tests.
- Run security checks: no connection strings/secrets, no accidental new tables/DbSets outside ownership boundary.
- Replace committed generated output only after review.
- Confirm partial/handwritten extensions still compile and tests pass.
- Record database schema version/fingerprint in CI artifacts.
9. Failure case: mix database-first and migrations-first ownership casually
A developer re-scaffolds from production, then runs
dotnet ef migrations add SyncLegacy and assumes the
migration will “capture what production changed.” That
assumption is wrong. Migrations compare the EF model with a
migration snapshot, while reverse engineering builds a model
from database metadata. If this scaffolded context does not
already have a coherent migrations history/snapshot representing
the same ownership model, the generated migration can be
misleading or destructive.
If ServiceHub decides EF migrations will own this legacy
schema going forward, baseline the schema deliberately, review
history/snapshot strategy, coordinate with DBAs/other
applications, test deployment/rollback, and establish the
cutover point. Do not let migrations add silently
become a second schema owner.
10. Lab: drift, diff, repair, and preserve custom code
- Start from a clean Lesson 3 scaffold and committed/generated baseline.
- Capture SQLite schema and EF model fingerprints.
-
Add
external_referenceplus its index to the disposable database. -
Scaffold to
.scaffold-nextwith the exact same tool/options. - Diff and explain each changed line category.
-
Verify the partial
RequiresSupervisorReview()method and concurrency-token extension remain outside generated output. - Replace generated output only after the diff is understood.
- Run query/write tests, regenerate fingerprints, and clean up the disposable DB/scratch folder.
11. Production judgment and bridge
A sustainable database-first workflow is deterministic: known schema input, pinned tools, controlled selection/naming, reproducible generated output, reviewed diff, preserved custom code, and provider-real integration tests. Version generated code if it improves deployment traceability and code review; the important point is that it is still generated and replaceable.
The final lesson applies the same honesty to imperfect legacy objects. Some database objects are excellent EF entities; others are read-only keyless views; some stored procedures belong behind raw SQL; and a table with no stable key is not made update-safe merely by inventing a C# identifier.
Check your understanding
-
Why scaffold to scratch instead of immediately using
--force? - What can cause a large generated diff besides schema changes?
- Does a schema hash prove a change is authorized or safe?
- Why is adding a missing generated property by hand a bad database-first repair?
-
Why can
migrations addbe dangerous on a casually scaffolded legacy context? - What survives a correct re-scaffold?
Review the answers
1. It makes candidate generated changes reviewable before replacing the committed/generated contract.
2. Tool, provider, template, naming/filter option, target-database, or nullable-reference configuration changes.
3. No. It only proves the canonicalized observed schema representation changed.
4. It breaks reproducibility and may omit provider/model metadata; change the schema/generation workflow and regenerate.
5. Migrations need a coherent model snapshot/history/ownership model; they are not a generic live-database diff.
6. Handwritten partials, extension configuration, services, and templates stored outside overwrite-prone generated output.
Authoritative references
Reverse engineering is provider- and version-sensitive. Re-check these primary sources when regenerating the chapter or adapting it to another database engine.
- Reverse Engineering - EF Core — scaffold workflow and inference model
- Migrations Overview — model snapshot and migration ownership semantics
- Custom Reverse Engineering Templates — generated-code customization lifecycle
- Testing against production databases — provider-real testing rationale
- SQLite schema provider guidance — provider-specific database behavior