Chapter 16 · Reverse Engineering and Database-First Workflows

Re-Scaffolding, Schema Drift, Handwritten Changes, and Sustainable Database-First Ownership

Treat re-scaffolding as a repeatable code-generation pipeline: stage output, diff it, classify schema drift, preserve handwritten extensions, and decide explicitly whether the database or the EF model owns future schema change.

Advanced150–190 minutesre-scaffold/diff/drift labEF Core 10.0.11 · .NET 10.0.11SQLite provider 10.0.11 mandatory baselinedotnet-ef 10.0.11 · SDK 10.0.400Last reviewed: August 2026

Learning outcomes

01

Define an explicit source-of-truth policy for database-first schema ownership.

02

Re-scaffold into a scratch directory and review generated-code diffs before replacing committed generated output.

03

Distinguish intentional database change, unreviewed database drift, generated-code drift, and application-only custom behavior.

04

Use repeatable schema fingerprints and runtime model probes as evidence rather than relying on timestamps or “it compiles.”

05

Demonstrate how direct edits are lost and how partial/separate code survives schema evolution.

06

Plan a deliberate transition if the team decides to move from database-owned schema to EF migrations instead of mixing ownership models accidentally.

1. The practical problem: six months later, nobody knows what changed

The ServiceHub legacy database gains external_reference directly through an operations change. One developer manually adds a C# property instead of re-scaffolding. Another re-scaffolds from production with a newer provider. A third begins creating EF migrations. Now source code, database schema, generated model, and migration history tell different stories. This is schema drift plus ownership drift.

Frozen lab baseline

Course baseline: .NET 10 runtime 10.0.11, SDK 10.0.400, EF Core/dotnet-ef/Microsoft.EntityFrameworkCore.Design/Microsoft.EntityFrameworkCore.Sqlite 10.0.11. SQLite is the mandatory free/local provider. EF Core 11 preview APIs are out of scope unless clearly labeled.

2. Choose the authority before choosing the command

Ownership model Authoritative schema change Generated EF role Main risk
Database-first DBA/schema pipeline/database project Derivative adapter; re-scaffold after approved DB change Hand edits drift from database
Migrations-first Reviewed EF migration/model evolution Primary schema-change artifact Out-of-band DB hotfixes drift from snapshot/history
Shared/legacy transition Explicitly coordinated process Temporary bridge Both sides believe they own DDL

Chapter 16 assumes database-first ownership for this legacy integration. Chapter 15 remains valid for the course's model-first ServiceHubContext. Do not use both ownership models against the same production schema without a written transition plan.

3. Re-scaffold into scratch first

A safe refresh does not start with --force against committed output. Generate into a clean scratch directory using the same tool/provider/options recorded in the project, then diff.

shell · repeatable scratch scaffold
rm -rf .scaffold-nextmkdir -p .scaffold-nextdotnet ef dbcontext scaffold "Data Source=servicehub-legacy.db" Microsoft.EntityFrameworkCore.Sqlite   --table work_orders --table technicians --table open_work_order_summary   --context ServiceHubLegacyContext   --context-dir .scaffold-next/Context   --output-dir .scaffold-next/Entities   --namespace ServiceHub.LegacyScaffold.Persistence.Generated.Entities   --context-namespace ServiceHub.LegacyScaffold.Persistence.Generated   --no-onconfiguring   --forcegit diff --no-index Persistence/Generated .scaffold-next || true

On Windows, use a temporary directory and your normal diff tool. What matters is that the candidate generated output is reviewable before it replaces the committed generated contract.

4. Classify every diff before accepting it

Observed diff Likely cause Action
New property + column mapping Approved DB column/table change Review semantics, tests, then adopt generated update
Hundreds of whitespace/naming changes Tool/provider/template/options changed Stop; reproduce with pinned versions/options before attributing to schema
Navigation disappears Filtered table omitted or FK changed Verify selected model boundary and FK metadata
Concurrency flag changes Provider metadata/model customization changed Verify provider behavior and handwritten partial configuration
Handwritten method disappears Direct edit in generated file Move to partial/separate code; do not “restore” it inside generated source
DB changed but generated diff is empty Unsupported/unmodeled metadata or wrong source DB Inspect schema/provider connection and model metadata directly

A large diff is not automatically “schema drift.” Code generation version drift can be equally large. Record command line, tool version, provider version, and database target in the review evidence.

5. Capture a lightweight database schema fingerprint

A database-first CI process benefits from a deterministic representation of the source schema. SQLite's sqlite_master can provide a simple lab fingerprint. Production systems should use engine-appropriate catalog queries or a reviewed schema artifact/database project.

sql · SQLite schema fingerprint
SELECT type, name, tbl_name, sqlFROM sqlite_masterWHERE name NOT LIKE 'sqlite_%'ORDER BY type, name;
python · hash the canonicalized schema in a local script
import hashlib, sqlite3con = sqlite3.connect("servicehub-legacy.db")rows = con.execute("""SELECT type, name, tbl_name, coalesce(sql, '')FROM sqlite_masterWHERE name NOT LIKE 'sqlite_%'ORDER BY type, name""").fetchall()canonical = "\n".join("|".join(map(str, row)) for row in rows)print(hashlib.sha256(canonical.encode()).hexdigest())

The hash is evidence that the observed schema representation changed. It does not explain whether the change is safe, authorized, or compatible.

6. Pair schema evidence with EF model evidence

After re-scaffolding, compile the candidate model and emit a normalized list of entities, properties, keys, foreign keys, indexes, and store mappings. The database fingerprint and model fingerprint answer different questions; comparing both catches connection mistakes and generation drift earlier.

csharp · candidate model fingerprint sketch
await using var db = new ServiceHubLegacyContext(options);var lines = db.Model.GetEntityTypes()    .OrderBy(e => e.Name)    .SelectMany(e =>        new[] { $"ENTITY|{e.DisplayName()}|{e.GetTableName()}|{e.GetViewName()}" }        .Concat(e.GetProperties()            .OrderBy(p => p.Name)            .Select(p => $"PROP|{e.DisplayName()}|{p.Name}|{p.ClrType.FullName}|nullable={p.IsNullable}|concurrency={p.IsConcurrencyToken}")));foreach (var line in lines) Console.WriteLine(line);

Do not hard-code a hash into production logic. This is build/review evidence for an integration contract.

7. Demonstrate drift safely: add a legacy column

sql · approved database-first change
ALTER TABLE work_orders ADD COLUMN external_reference TEXT NULL;CREATE INDEX ix_work_orders_external_reference  ON work_orders(external_reference);

Re-run the scratch scaffold. Expected diff: one property, column mapping, and index metadata. If the generated diff also rewrites unrelated types or names, investigate tool/options before accepting it.

Do not hide schema drift by hand-editing generated C#

Adding ExternalReference manually may make compilation pass, but it bypasses the repeatable generation contract and can miss provider facets/indexes/defaults. In database-first ownership, fix the schema/generation pipeline, then regenerate.

8. Re-scaffold acceptance workflow

  1. Identify the exact database/schema artifact and approved change ticket/version.
  2. Restore/create a disposable database at that schema version.
  3. Verify dotnet-ef, provider, Design package, templates, and scaffold options.
  4. Generate into scratch—not committed output.
  5. Diff and classify every change.
  6. Run model metadata probes and representative query/write tests.
  7. Run security checks: no connection strings/secrets, no accidental new tables/DbSets outside ownership boundary.
  8. Replace committed generated output only after review.
  9. Confirm partial/handwritten extensions still compile and tests pass.
  10. Record database schema version/fingerprint in CI artifacts.

9. Failure case: mix database-first and migrations-first ownership casually

A developer re-scaffolds from production, then runs dotnet ef migrations add SyncLegacy and assumes the migration will “capture what production changed.” That assumption is wrong. Migrations compare the EF model with a migration snapshot, while reverse engineering builds a model from database metadata. If this scaffolded context does not already have a coherent migrations history/snapshot representing the same ownership model, the generated migration can be misleading or destructive.

Transition is a project, not a command

If ServiceHub decides EF migrations will own this legacy schema going forward, baseline the schema deliberately, review history/snapshot strategy, coordinate with DBAs/other applications, test deployment/rollback, and establish the cutover point. Do not let migrations add silently become a second schema owner.

10. Lab: drift, diff, repair, and preserve custom code

  1. Start from a clean Lesson 3 scaffold and committed/generated baseline.
  2. Capture SQLite schema and EF model fingerprints.
  3. Add external_reference plus its index to the disposable database.
  4. Scaffold to .scaffold-next with the exact same tool/options.
  5. Diff and explain each changed line category.
  6. Verify the partial RequiresSupervisorReview() method and concurrency-token extension remain outside generated output.
  7. Replace generated output only after the diff is understood.
  8. Run query/write tests, regenerate fingerprints, and clean up the disposable DB/scratch folder.

11. Production judgment and bridge

A sustainable database-first workflow is deterministic: known schema input, pinned tools, controlled selection/naming, reproducible generated output, reviewed diff, preserved custom code, and provider-real integration tests. Version generated code if it improves deployment traceability and code review; the important point is that it is still generated and replaceable.

The final lesson applies the same honesty to imperfect legacy objects. Some database objects are excellent EF entities; others are read-only keyless views; some stored procedures belong behind raw SQL; and a table with no stable key is not made update-safe merely by inventing a C# identifier.

Check your understanding

  1. Why scaffold to scratch instead of immediately using --force?
  2. What can cause a large generated diff besides schema changes?
  3. Does a schema hash prove a change is authorized or safe?
  4. Why is adding a missing generated property by hand a bad database-first repair?
  5. Why can migrations add be dangerous on a casually scaffolded legacy context?
  6. What survives a correct re-scaffold?
Review the answers

1. It makes candidate generated changes reviewable before replacing the committed/generated contract.

2. Tool, provider, template, naming/filter option, target-database, or nullable-reference configuration changes.

3. No. It only proves the canonicalized observed schema representation changed.

4. It breaks reproducibility and may omit provider/model metadata; change the schema/generation workflow and regenerate.

5. Migrations need a coherent model snapshot/history/ownership model; they are not a generic live-database diff.

6. Handwritten partials, extension configuration, services, and templates stored outside overwrite-prone generated output.

Authoritative references

Reverse engineering is provider- and version-sensitive. Re-check these primary sources when regenerating the chapter or adapting it to another database engine.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.

\n