Chapter 31Lesson 03~190 minutes

Large Repositories, Monorepos, Search, Actions Cost Controls, and Platform Performance: Configuration, Design Choices, and Tradeoffs

Choose repository boundaries, large-file storage, validation scope, checkout depth, runner class, concurrency, and retention based on governance and measured workload.

ArchitectureGit LFSRunnersRetention

Learning objectives

  • Choose monorepo or multi-repo based on governance, release, ownership, and operational boundaries.
  • Choose Git blobs, Git LFS, releases/packages, or external artifact storage by lifecycle and distribution needs.
  • Balance selective CI against always-run validation and understand runner/concurrency/storage tradeoffs.
  • Apply current plan and deployment boundaries without requiring paid features.

1. Configuration begins with the boundary you want to own

There is no universal “monorepo is faster” or “multi-repo is cleaner” answer. Repository boundaries determine permission scope, ruleset scope, release independence, search scope, issue/PR workflow, dependency change coordination, and the failure radius of automation. The correct design is the one whose governance boundaries match the product.

2. Monorepo versus multi-repo

Dimension Monorepo Multi-repo
Atomic cross-component change Strong: one commit/PR can move interfaces and consumers together. Requires coordinated PRs/releases or compatibility windows.
Permissions Repository read visibility is shared; write/review can be path-governed but not true per-directory repository ACL. Repository boundary gives stronger isolation.
CI Requires trustworthy affected-project routing and global dependency modeling. Smaller default scope, but cross-repo integration requires orchestration.
Issues/releases Shared planning/release metadata can simplify platform work. Independent cadence and ownership are clearer.
Search/history One Git history and repository-scoped search. Cross-repo search/API inventory becomes more important.
Failure blast radius One broken workflow/rule can affect many components. Failures may isolate better but duplicated policy can drift.

A useful heuristic: if components need hard confidentiality/administrative isolation or radically independent lifecycle, separate repositories are often more honest. If teams make frequent atomic cross-component changes and share one policy/release surface, a monorepo can be efficient—provided CI and ownership are engineered rather than improvised.

3. Normal Git blobs versus LFS, packages, releases, and external storage

Need Best-fit starting point Tradeoff
Diffable source/history Normal Git Large blobs permanently increase history cost.
Large versioned source asset Git LFS Separate payload storage/bandwidth, plan-dependent limits, client support.
Versioned build output consumed by tooling Package/container registry Package identity/access/lifecycle must be governed (Chapter 21).
Human-downloadable release bundle Release asset Release lifecycle rather than source-history semantics.
Huge generated dataset/cache External object/artifact storage External identity, retention, access, provenance must be integrated.

The decision should answer “who consumes this, how is it versioned, how long is it retained, and what provenance must be verified?” before it answers “where can I upload it?”

4. Separate local, cross-cutting, and policy validation

Path-aware CI is safest when jobs are classified:

  • Component-local: unit tests/lint for one directory.
  • Cross-cutting: shared libraries, schemas, lockfiles, base images, global build config, workflow changes.
  • Always-run policy gate: cheap router and aggregate signal that proves the routing decision itself executed.

Do not require a component workflow that sometimes never starts because a top-level path filter skipped it. Require the stable gate that starts on every relevant PR and understands intentional skips.

5. Choose checkout semantics per job

Job History Tree Reason
Affected router Enough base/head ancestry; often full in simple labs Router/config files Must calculate exact diff reliably.
Component unit test Usually shallow Sparse component + declared shared dependencies Minimize materialization after scope known.
Release/version calculation Tags/history may be required Depends on build Do not shallow away required version evidence.
History audit Full Potentially no full checkout needed Information requirement is history, not working tree.

actions/checkout supports sparse checkout, configurable fetch depth, and optional LFS fetching. Use those controls because the job's data requirements justify them—not as a blanket “performance recipe.”

6. Runner size and concurrency: throughput is not the same as cost efficiency

A larger runner can shorten a CPU-bound job, but it can also cost more per minute and hide inefficient parallelization. Current GitHub larger runners are an organization/enterprise feature for Team/Enterprise Cloud and are billable; included standard-runner minutes do not apply to them. For the mandatory path, stay on standard public runners.

Choice Benefit Risk Measure
More matrix shards Lower wall time More startup/duplicate setup/storage. Total runner-minutes + wall time.
concurrency + cancel stale PR runs Avoids work on superseded commits. Wrong key can cancel unrelated work. Cancelled vs completed runs by PR/ref.
Larger runner More CPU/RAM/disk/network options. Paid and can overprovision. Cost per successful change, not just job seconds.
Self-hosted Custom environment/local cache potential. Security/operations burden; not a free optimization. Utilization, queue, maintenance and incident cost.

7. Cache and artifact retention are different policies

A cache exists to accelerate a future run and should tolerate eviction. An artifact exists because somebody needs an output or evidence. Confusing them creates expensive, unreliable designs. Cache keys should have controlled cardinality and useful hit rates. Artifacts should have an owner, consumer, and retention period.

Current GitHub defaults include a 10 GB cache allowance per repository and eviction for entries unused for more than seven days. Artifact/log retention defaults to 90 days, with public repositories configurable down to one day. Optional higher cache size/retention settings require eligible billing/plan configuration, so they are not part of the mandatory lab.

8. Feature and plan boundaries

Capability Mandatory/free path Constraint
Standard hosted Actions in public repo Yes Standard public-runner execution is free.
Private repository hosted Actions Optional Plan-specific included minutes/artifact storage; overage may bill.
Larger runners No Organization/enterprise Team or Enterprise Cloud; billed including public repos.
CODEOWNERS in public repository Yes Owners must have appropriate write access for review requests.
Git LFS Optional Per-file maximum and storage/bandwidth allowances are plan-dependent.
Cache over default 10 GB No Opt-in higher limits can become billable.
GHES Conceptual Host/version/admin policy and available features/limits differ; inspect instance docs.

9. Worked design: a 70-team product platform

Suppose one product platform contains 40 services, six client applications, shared schemas, and common build tooling. Teams need atomic schema migrations, but payroll and security components require separate confidentiality. A sensible design may be hybrid:

  1. Keep ordinary product services sharing one release/governance cadence in a monorepo.
  2. Keep confidential or independently administered components in separate repositories.
  3. Model shared/, schema registries, lockfiles, and workflow changes as cross-cutting dependencies.
  4. Use team-oriented CODEOWNERS for human review, not thousands of user lines.
  5. Use one always-run router/gate and selective component jobs.
  6. Store release binaries/packages outside Git history, with immutable identity/provenance.
  7. Track cost per successful PR and storage growth, not raw workflow count alone.

This design accepts a little orchestration complexity to preserve real authorization and lifecycle boundaries.

10. Summary

Platform configuration is about aligning repository boundaries, artifact lifecycles, validation classes, checkout depth, runner choice, concurrency, and retention with actual requirements. Lesson 4 now deliberately breaks these assumptions and diagnoses the evidence.

Knowledge check

When is multi-repo a stronger choice than monorepo?

Why is Git LFS not a substitute for artifact retention policy?

Which check should typically be required: every conditionally selected component workflow or one stable aggregate gate?

Does a larger runner automatically reduce cost?

What is the essential difference between a cache and an artifact?

Next lesson

Large Repositories, Monorepos, Search, Actions Cost Controls, and Platform Performance: Diagnostics, Failure Modes, Security, and Performance

Further reading — current primary sources

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.