Chapter 28Lesson 01~190 minutes

Monorepos, Path Filters, Changed-File Detection, and Selective Pipelines: Core Concepts and Mental Model

Model selective monorepo CI as an evidence-backed optimization over exact base/head revisions, complete changed-file sets and an explicit component dependency graph.

Monorepo modelBase/head SHAPath filtersDependency graphRequired checks

Learning objectives

  • Explain why selective CI is an optimization layer, not the source of correctness.
  • Separate event base/head identity, changed-file derivation, component ownership, dependency fan-out, job selection and required-check aggregation.
  • Inspect a monorepo change set read-only before deciding what to skip.
  • Explain current native path-filter semantics and their required-check hazard.
  • Define a fallback policy that makes selector mistakes observable rather than silent.

1. The practical problem: monorepos make “run everything” expensive

A monorepo can contain many applications, libraries, infrastructure directories and documentation trees under one Git history. Running every build and test for every pull request is simple and conservative, but as the repository grows it increases queue time, feedback latency and billed runner minutes. Selective CI tries to spend work only on components that a change can affect.

The danger is asymmetric: running one unnecessary job costs time, but skipping one necessary job can merge a defect. Therefore the selector is not merely a performance script. It is part of the repository’s correctness boundary and must have its own inputs, dependency model, evidence and tests.

2. Mental model: exact revisions → complete change set → dependency closure → stable check

Begin with the event’s base and head revisions. A changed-file algorithm derives paths from those revisions. Path ownership maps files to components. The component dependency graph expands direct changes to transitive dependents. Only then should the workflow create selected jobs or a matrix. A final aggregate check reports whether the selection process and every selected test reached an acceptable conclusion.

Selective CI causality
flowchart TD
  A[Event base SHA] --> C[Changed-file algorithm]
  B[Event head SHA] --> C
  C --> D[Changed paths]
  D --> E[Ownership rules]
  E --> F[Dependency closure]
  F --> G[Selected jobs / matrix]
  G --> H[Test evidence]
  C --> I[Selector evidence]
  H --> J[Stable required check]
  I --> J
  K[Scheduled/manual full suite] --> L[Selector contract + all components]
  L --> J

The key causal arrow is ownership → dependency closure. A change under libs/shared/ may not belong to either application directly, yet both applications may consume it. A selector that only asks “which top-level directory changed?” is incomplete even if its Git diff is perfect.

3. Define every state layer before skipping work

Layer Evidence to record Failure if confused
Event/revision event name, base SHA, head SHA, workflow revision, run ID/attempt Diff the wrong revisions or test a merge commit while reporting the PR head.
Git history fetch depth, merge base, comparison mode A shallow checkout cannot prove the intended merge base.
Changed-file set algorithm, raw path list, rename/deletion treatment Selector silently omits relevant files.
Ownership path patterns → component Direct changes map to wrong component.
Dependency graph library/component edges and closure version Shared changes fail to fan out.
Job graph selected/skipped components and matrix JSON “Skipped” is mistaken for “tested.”
Checks/governance stable required check name and fallback/full-run policy Branch protection waits for a workflow that never started.
Artifacts/cache selector/test evidence IDs; cache only as optimization A cache hit is mistaken for component correctness.
External/deployment normally none in PR selection Selective CI must not implicitly authorize deployment.

4. Native path filters are workflow-run selectors

GitHub evaluates paths and paths-ignore before creating the filtered workflow run for push and pull_request. Patterns are repository-root relative. Positive and negative patterns are order-sensitive when ! is used. If branch and path filters are both present, both conditions must match. Tag pushes do not use path filters.

# Appropriate for an optional, non-required workflow.
on:
  pull_request:
    paths:
      - 'apps/**'
      - 'libs/**'
      - '!apps/**/docs/**'

As verified on September 10, 2026, GitHub.com derives path-filter changes with three-dot diffs for pull requests and two-dot diffs for pushes. The current GitHub.com documentation describes a first-3,000-files evaluation boundary; a push containing more than 1,000 commits or a diff-generation timeout runs the workflow rather than trusting the filter. These are platform behaviors, not a substitute for your dependency graph.

5. Required workflows and path filters can deadlock a pull request

If an entire workflow is skipped because of path, branch or commit-message filtering, checks associated with that workflow can remain Pending. If branch protection requires that check, the pull request can be blocked indefinitely even though no runner job exists to turn it green. This is why a selective monorepo design should be deliberate about which workflow/check is required.

A safer pattern is to let one lightweight workflow start for every relevant pull request, compute the selector inside the workflow, skip component jobs with job-level conditions, and always produce one stable aggregate job such as Monorepo CI / required. The aggregate job must fail if selection failed or any selected component failed; it may succeed when the selector explicitly proves no component work is required.

6. Pull-request change detection is a three-dot question

For a pull request, the useful question is “what did this topic branch introduce since its common ancestor with the base?” GitHub native path filtering uses a three-dot comparison. A custom selector should reproduce that intent explicitly: identify the base/head commits, compute the merge base with sufficient history, then diff merge-base → head.

set -euo pipefail
BASE="$PR_BASE_SHA"
HEAD="$PR_HEAD_SHA"
git cat-file -e "$BASE^{commit}"
git cat-file -e "$HEAD^{commit}"
MERGE_BASE="$(git merge-base "$BASE" "$HEAD")"
git diff --name-only --no-renames -z "$MERGE_BASE" "$HEAD" > changed-paths.z

The example uses NUL-delimited paths and disables rename detection so a rename is represented conservatively as old-path removal plus new-path addition. That lets ownership logic see both sides of a move and also handles spaces/newlines without evaluating filenames as shell source.

7. Direct ownership is not enough: calculate dependency closure

Assume apps/api and apps/web both import libs/shared. A change under apps/api/** selects API. A change under apps/web/** selects Web. A change under libs/shared/** must select both because both depend on the library. Repository-wide files such as dependency locks, build configuration, workflow code or selector rules should normally fan out to all components.

This dependency map belongs in version-controlled source with tests. If the architectural dependency graph changes but the selector graph does not, the optimization becomes stale. Treat that drift as a CI bug, not a harmless performance issue.

8. Correctness needs a fallback/full-run policy

No selector should be trusted only because it is fast. Use at least one independent mechanism that regularly runs broader coverage: a scheduled full suite, a manual “run all” dispatch, merge-queue/full-integration lane, or a selector contract test that feeds known change sets into the mapping logic. Unknown paths and selector errors should fail closed toward more testing, not silently select nothing.

A useful invariant is: uncertainty increases work. If merge-base computation fails, the component graph cannot be loaded, the changed-file set crosses a configured safety threshold, or an unowned executable path appears, select the full suite and preserve the selector warning.

9. Changed paths are data, not shell syntax

File names in a pull request can be attacker-controlled. Do not concatenate changed paths into eval, generate shell commands from them, or interpolate event-derived path strings directly into a run: program. Write paths to a NUL-delimited file and parse them as data in Python or another language with explicit path rules. Keep PR jobs on GitHub-hosted runners with read-only/minimal permissions as established in Chapter 22.

10. Evidence that makes a skip independently reviewable

Question Evidence
Which revision pair was compared? base SHA, head SHA, merge-base SHA, event name.
Was history sufficient? checkout action SHA, fetch-depth policy, successful cat-file/merge-base.
What changed? NUL-safe raw path list plus human-readable escaped summary.
Why was a component selected? ownership rule and dependency-edge reason.
Why was one skipped? selector output showing no direct/transitive dependency.
Did selected work pass? matrix cell conclusions and test logs/artifacts.
What protects selector correctness? full-run cadence, selector-contract tests, unknown-path fail-safe.
What does branch protection require? one stable aggregate check name, not a workflow that can disappear.

11. Lesson summary

Selective monorepo CI is safe when skipped work is justified by exact base/head history, a complete changed-file set and an explicit dependency graph, and when one stable required check plus a broader fallback suite makes selector failure visible. Performance is the outcome; evidence-backed correctness is the design constraint.

Next lesson

Monorepos, Path Filters, Changed-File Detection, and Selective Pipelines: Guided Hands-On Workflow

Continue with the next lesson to build on the current concepts, evidence, security boundaries, and operational practices.

Knowledge check

Why should a shared-library change often select more than one component?

What is the main branch-protection risk of putting paths on a required workflow?

Why use a three-dot comparison for pull-request selection?

What should happen when the selector cannot determine a safe change set?

Why are changed filenames security-sensitive input?

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.