Large Repositories, Monorepos, Search, Actions Cost Controls, and Platform Performance: Configuration, Design Choices, and Tradeoffs
Choose repository boundaries, large-file storage, validation scope, checkout depth, runner class, concurrency, and retention based on governance and measured workload.
Learning objectives
- Choose monorepo or multi-repo based on governance, release, ownership, and operational boundaries.
- Choose Git blobs, Git LFS, releases/packages, or external artifact storage by lifecycle and distribution needs.
- Balance selective CI against always-run validation and understand runner/concurrency/storage tradeoffs.
- Apply current plan and deployment boundaries without requiring paid features.
1. Configuration begins with the boundary you want to own
There is no universal “monorepo is faster” or “multi-repo is cleaner” answer. Repository boundaries determine permission scope, ruleset scope, release independence, search scope, issue/PR workflow, dependency change coordination, and the failure radius of automation. The correct design is the one whose governance boundaries match the product.
2. Monorepo versus multi-repo
| Dimension | Monorepo | Multi-repo |
|---|---|---|
| Atomic cross-component change | Strong: one commit/PR can move interfaces and consumers together. | Requires coordinated PRs/releases or compatibility windows. |
| Permissions | Repository read visibility is shared; write/review can be path-governed but not true per-directory repository ACL. | Repository boundary gives stronger isolation. |
| CI | Requires trustworthy affected-project routing and global dependency modeling. | Smaller default scope, but cross-repo integration requires orchestration. |
| Issues/releases | Shared planning/release metadata can simplify platform work. | Independent cadence and ownership are clearer. |
| Search/history | One Git history and repository-scoped search. | Cross-repo search/API inventory becomes more important. |
| Failure blast radius | One broken workflow/rule can affect many components. | Failures may isolate better but duplicated policy can drift. |
A useful heuristic: if components need hard confidentiality/administrative isolation or radically independent lifecycle, separate repositories are often more honest. If teams make frequent atomic cross-component changes and share one policy/release surface, a monorepo can be efficient—provided CI and ownership are engineered rather than improvised.
3. Normal Git blobs versus LFS, packages, releases, and external storage
| Need | Best-fit starting point | Tradeoff |
|---|---|---|
| Diffable source/history | Normal Git | Large blobs permanently increase history cost. |
| Large versioned source asset | Git LFS | Separate payload storage/bandwidth, plan-dependent limits, client support. |
| Versioned build output consumed by tooling | Package/container registry | Package identity/access/lifecycle must be governed (Chapter 21). |
| Human-downloadable release bundle | Release asset | Release lifecycle rather than source-history semantics. |
| Huge generated dataset/cache | External object/artifact storage | External identity, retention, access, provenance must be integrated. |
The decision should answer “who consumes this, how is it versioned, how long is it retained, and what provenance must be verified?” before it answers “where can I upload it?”
4. Separate local, cross-cutting, and policy validation
Path-aware CI is safest when jobs are classified:
- Component-local: unit tests/lint for one directory.
- Cross-cutting: shared libraries, schemas, lockfiles, base images, global build config, workflow changes.
- Always-run policy gate: cheap router and aggregate signal that proves the routing decision itself executed.
Do not require a component workflow that sometimes never starts because a top-level path filter skipped it. Require the stable gate that starts on every relevant PR and understands intentional skips.
5. Choose checkout semantics per job
| Job | History | Tree | Reason |
|---|---|---|---|
| Affected router | Enough base/head ancestry; often full in simple labs | Router/config files | Must calculate exact diff reliably. |
| Component unit test | Usually shallow | Sparse component + declared shared dependencies | Minimize materialization after scope known. |
| Release/version calculation | Tags/history may be required | Depends on build | Do not shallow away required version evidence. |
| History audit | Full | Potentially no full checkout needed | Information requirement is history, not working tree. |
actions/checkout supports sparse checkout, configurable
fetch depth, and optional LFS fetching. Use those controls because
the job's data requirements justify them—not as a blanket
“performance recipe.”
6. Runner size and concurrency: throughput is not the same as cost efficiency
A larger runner can shorten a CPU-bound job, but it can also cost more per minute and hide inefficient parallelization. Current GitHub larger runners are an organization/enterprise feature for Team/Enterprise Cloud and are billable; included standard-runner minutes do not apply to them. For the mandatory path, stay on standard public runners.
| Choice | Benefit | Risk | Measure |
|---|---|---|---|
| More matrix shards | Lower wall time | More startup/duplicate setup/storage. | Total runner-minutes + wall time. |
concurrency + cancel stale PR runs |
Avoids work on superseded commits. | Wrong key can cancel unrelated work. | Cancelled vs completed runs by PR/ref. |
| Larger runner | More CPU/RAM/disk/network options. | Paid and can overprovision. | Cost per successful change, not just job seconds. |
| Self-hosted | Custom environment/local cache potential. | Security/operations burden; not a free optimization. | Utilization, queue, maintenance and incident cost. |
7. Cache and artifact retention are different policies
A cache exists to accelerate a future run and should tolerate eviction. An artifact exists because somebody needs an output or evidence. Confusing them creates expensive, unreliable designs. Cache keys should have controlled cardinality and useful hit rates. Artifacts should have an owner, consumer, and retention period.
Current GitHub defaults include a 10 GB cache allowance per repository and eviction for entries unused for more than seven days. Artifact/log retention defaults to 90 days, with public repositories configurable down to one day. Optional higher cache size/retention settings require eligible billing/plan configuration, so they are not part of the mandatory lab.
8. Feature and plan boundaries
| Capability | Mandatory/free path | Constraint |
|---|---|---|
| Standard hosted Actions in public repo | Yes | Standard public-runner execution is free. |
| Private repository hosted Actions | Optional | Plan-specific included minutes/artifact storage; overage may bill. |
| Larger runners | No | Organization/enterprise Team or Enterprise Cloud; billed including public repos. |
| CODEOWNERS in public repository | Yes | Owners must have appropriate write access for review requests. |
| Git LFS | Optional | Per-file maximum and storage/bandwidth allowances are plan-dependent. |
| Cache over default 10 GB | No | Opt-in higher limits can become billable. |
| GHES | Conceptual | Host/version/admin policy and available features/limits differ; inspect instance docs. |
9. Worked design: a 70-team product platform
Suppose one product platform contains 40 services, six client applications, shared schemas, and common build tooling. Teams need atomic schema migrations, but payroll and security components require separate confidentiality. A sensible design may be hybrid:
- Keep ordinary product services sharing one release/governance cadence in a monorepo.
- Keep confidential or independently administered components in separate repositories.
-
Model
shared/, schema registries, lockfiles, and workflow changes as cross-cutting dependencies. - Use team-oriented CODEOWNERS for human review, not thousands of user lines.
- Use one always-run router/gate and selective component jobs.
- Store release binaries/packages outside Git history, with immutable identity/provenance.
- Track cost per successful PR and storage growth, not raw workflow count alone.
This design accepts a little orchestration complexity to preserve real authorization and lifecycle boundaries.
10. Summary
Platform configuration is about aligning repository boundaries, artifact lifecycles, validation classes, checkout depth, runner choice, concurrency, and retention with actual requirements. Lesson 4 now deliberately breaks these assumptions and diagnoses the evidence.
Knowledge check
When is multi-repo a stronger choice than monorepo?
When components require hard repository-level confidentiality/administrative isolation or genuinely independent lifecycle/governance that path-level conventions cannot provide.
Why is Git LFS not a substitute for artifact retention policy?
LFS is a source-versioning mechanism with separate payload storage/bandwidth. Build outputs and temporary/generated data usually need release/package/artifact lifecycles instead.
Which check should typically be required: every conditionally selected component workflow or one stable aggregate gate?
A stable aggregate gate is safer when component jobs can be intentionally skipped, because conditionally absent workflows can create confusing pending required checks.
Does a larger runner automatically reduce cost?
No. It may reduce wall time but is billed at a different rate and can overprovision. Compare cost per successful change and total resource time, not only job duration.
What is the essential difference between a cache and an artifact?
A cache is disposable acceleration state that should tolerate eviction; an artifact is retained output/evidence with a known consumer and retention requirement.
Further reading — current primary sources
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.