Chapter 28Lesson 04~195 minutes

Monorepos, Path Filters, Changed-File Detection, and Selective Pipelines: Diagnostics, Failure Modes, and Production Practices

Diagnose under-selection, shallow-history errors, missing required checks, rename/deletion gaps and unsafe path handling while preserving first-failure evidence.

DiagnosticsRenamesShallow cloneUnsafe pathsRecovery

Learning objectives

  • Use an evidence-first sequence to diagnose wrong or incomplete component selection.
  • Recognize shallow-history, transitive-dependency, required-check, rename/deletion and path-injection failures.
  • Preserve run/attempt, revisions, selector evidence and matrix conclusions before rerunning.
  • Repair the narrow causal layer instead of broadening permissions or disabling checks.
  • Explain when the correct recovery is a full-suite run rather than another selective retry.

1. Evidence-first diagnostic sequence

When a monorepo check is wrong, preserve the run ID/attempt and first-failure evidence before rerunning or editing glob patterns. Confirm event, base/head SHAs, workflow revision, permissions, checkout depth, merge base, changed paths, ownership/dependency mapping, generated matrix, component conclusions and aggregate check.

  1. Preserve run ID, attempt, source/workflow SHA and first-failure logs/artifacts.
  2. Confirm event type and exact base/head revisions.
  3. Confirm runner/checkout/action versions and fetched Git objects.
  4. Recompute merge base and raw changed-file set without mutation.
  5. Inspect ownership and transitive dependency reasons.
  6. Inspect selected/skipped jobs and matrix conclusions.
  7. Inspect branch-protection required check and whether a workflow existed.
  8. Apply the least-destructive selector/history repair.
  9. Rerun the smallest equivalent change set, then run the independent full suite if correctness was uncertain.

2. Failure: missing transitive dependents

Evidence: changed-paths.txt contains libs/shared/contract.txt, but selection.json contains only ["api"]. Git history is correct; the dependency graph is not. Repair shared fan-out, add a regression fixture, and preserve the original run as proof that Web was omitted.

3. Failure: shallow checkout cannot prove the merge base

A common broken example uses depth 1 and immediately diffs event SHAs. It may fail with an unknown revision or cannot prove intended branch history. The first question is whether both commits and their merge base are present.

# BROKEN — do not use in production
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1
  with:
    fetch-depth: 1
- run: |
    BASE="${{ github.event.pull_request.base.sha }}"
    HEAD="${{ github.event.pull_request.head.sha }}"
    git diff --name-only "$BASE...$HEAD"

Repair by fetching sufficient history for the exact base/head pair. On the small lab, fetch-depth: 0 is clear; large repositories can deepen or fetch exact refs until git merge-base succeeds.

4. Failure: path filter leaves a required check Pending

Evidence: the pull request expects Monorepo CI / required, but there is no run because workflow-level paths excluded the change. This is not a runner outage—the workflow never reached run creation. Start the lightweight required workflow on every applicable PR and perform selection inside it.

5. Failure: rename or deletion crosses ownership boundaries

If apps/api/schema.json moves to libs/shared/schema.json, a parser that sees only one side can omit required tests. Preserve raw Git status, then evaluate both old and new paths. The lab’s --no-renames strategy converts the move into deletion + addition.

6. Failure: changed paths are interpolated into shell syntax

A contributor can create filenames containing spaces, glob characters, leading dashes or shell metacharacters. Word-splitting is incorrect; eval is dangerous because path bytes can become executable syntax.

# BROKEN — do not use in production
FILES="$(git diff --name-only "$BASE...$HEAD")"
for f in $FILES; do
  eval "python check.py $f"
done

Repair by keeping filenames NUL-delimited and parsing them as data. When invoking tools, use explicit argument arrays or -- where supported.

7. Failure: optimization removed the only broad integration signal

A selective pipeline can miss cross-component behavior the graph does not model. If the only full integration suite was removed because it was slow, there is no independent detector for selector drift. Restore a scheduled/manual/release full suite with an authoritative component inventory separate from the selector.

8. Failure: aggregate check is green despite a failed selected job

This usually means the aggregate job ran with always() but never inspected needs.*.result, or a step normalized failure with continue-on-error. The aggregate must explicitly translate selector + selected conclusions into policy.

9. Separate causal layers before fixing

Symptom Likely layer Do not “fix” by
No workflow/check exists event/path filter / governance adding write token.
merge-base unknown checkout/history changing component globs.
changed paths correct, matrix wrong ownership/dependency selector increasing runner size.
matrix right, tests fail component build/test forcing selector success.
cache artifact stale cache/data layer treating cache as change source.
provider/deploy issue after CI deployment/external target rewriting diff logic.

10. Recovery rule: preserve uncertainty, then broaden evidence

If evidence shows the selector may have skipped necessary work, do not rerun only the same selected matrix. Preserve the defective evidence, repair the causal rule, prove the equivalent change set, then run the full suite to re-establish confidence.

11. Production practices

  • Version-control ownership/dependency rules and test them with synthetic change sets.
  • Unknown executable/global paths select all rather than none.
  • Keep the aggregate job low privilege and free of deployment side effects.
  • Treat changed filenames/event metadata as untrusted data.
  • Retain selector evidence long enough to investigate merge regressions.
  • Use periodic full coverage independent of the incremental selector.

12. Diagnostic summary

Monorepo failures become tractable when you prove the revision pair, Git history, changed paths and dependency closure before touching YAML conditions. Preserve first-failure evidence, fix the narrow layer, and broaden to a full suite whenever selection correctness was in doubt.

Next lesson

Checkpoint Lab — Monorepos, Path Filters, Changed-File Detection, and Selective Pipelines

Continue with the next lesson to build on the current concepts, evidence, security boundaries, and operational practices.

Knowledge check

The raw changed-file list is correct but Web was skipped after a shared change. What failed?

A required check is Pending and there is no workflow run. What should you inspect?

Why is eval unsafe for changed filenames?

When should recovery include a full suite?

Does always() make an aggregate check correct?

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.