Chapter 36Lesson 03~285 minutes

Capstone: Build a Secure Reusable Enterprise CI/CD Platform: Configuration, Design Patterns, and Trade-Offs

Choose platform boundaries deliberately across centralization, runner ownership, OIDC, reusable contracts, policy, exceptions, performance and cost.

ArchitectureOIDCRunner trustGolden pathsCost

Learning objectives

  • Choose centralization and autonomy boundaries intentionally.
  • Compare hosted and self-hosted/ARC runner trust models.
  • Prefer OIDC where federation can replace long-lived static credentials.
  • Design composable golden-path contracts, strict policy and time-bounded exceptions.
  • Optimize speed/cost without weakening coverage or isolation.

1. Architecture decisions are trust and ownership decisions

The capstone is intentionally small, but production platforms must decide where to centralize behavior and where teams retain control. A “golden path” is useful when it makes safe behavior easy and observable. It becomes harmful when it hides assumptions, forces every workload into one giant workflow, or removes a documented escape hatch for legitimate constraints.

2. Decision table for the production platform

Decision Prefer when Main risk Evidence required
Central reusable workflow many repos share stable CI/deploy contract blast radius of breaking platform change caller inventory, immutable version, compatibility tests, rollback version
Repository-owned workflow workload is unique or policy only sets guardrails drift and duplicated fixes policy result, dependency pins, required checks
GitHub-hosted runner public internet access and standard tooling suffice image/tool drift and external egress explicit image label, tool versions, run metadata
Self-hosted / ARC private network, special hardware or fleet economics justify it persistent trust boundary and capacity ownership runner group, image provenance, patching, logs, queue metrics
OIDC provider supports federation and target trust can be constrained over-broad audience/subject policy actual claims, provider trust rule, short-lived session identity
Static secret legacy integration cannot federate long-lived theft/rotation risk scope, expiry/rotation, consumer inventory, no log exposure
Strict policy violation creates systemic security/compliance risk blocking legitimate exceptions policy owner, scope, denial evidence, documented exception path
Time-bounded exception business need is real and control cannot be met immediately exception becomes permanent owner, reason, scope, expiry, compensating control, review date

3. Centralization versus autonomy

Centralize stable cross-repository contracts: minimal permissions, artifact identity conventions, provenance requirements, environment naming, required evidence and safe runner classes. Keep application-specific build commands, test selection and rollout semantics close to the repository unless the platform can model them explicitly. The platform should expose a small typed interface rather than dozens of escape-hatch booleans.

Version every central component. A caller pinned to a full commit SHA is reproducible; an upgrade channel can still exist through automated pull requests that map reviewed releases to new SHAs. Test the platform against representative consumers before moving the channel. The rollback is the previous known platform SHA, not main from yesterday.

4. Hosted versus self-hosted/ARC

Hosted runners minimize platform operations and are the default for untrusted pull-request CI in this capstone. Self-hosted/ARC becomes attractive when the workload requires private networking, specialized hardware, predictable warm dependencies or organization-controlled images. That convenience is purchased with a larger trust boundary: runner registration, image lifecycle, network segmentation, log retention, capacity, autoscaling and emergency isolation all become your responsibility.

Use runner groups to encode access boundaries where the plan and topology support them. Do not expose a privileged runner group to public repositories or untrusted fork code. If a deployment requires a private-network runner, separate the untrusted build from the trusted deploy and promote a verified immutable artifact across the boundary.

5. Static secrets versus OIDC

OIDC is the preferred cloud identity model when the provider supports it because GitHub can mint a short-lived token containing repository/workflow/event claims, and the provider can exchange it for a narrowly scoped session. The job needs id-token: write only where federation occurs. That permission does not itself grant cloud access; the external provider trust policy decides whether the token is accepted.

Static secrets remain necessary for some legacy systems. Treat them as expiring inventory: name, owner, target, scope, rotation cadence, consumers and revocation method. Never broaden a repository secret to organization scope simply to make a reusable workflow easier to call.

6. One golden path versus composable contracts

A single mega-workflow can standardize everything but creates a large blast radius, slow upgrades and an interface full of unrelated knobs. Composable reusable workflows and actions create smaller contracts but can become deep dependency graphs. Current GitHub.com permits ten connected reusable-workflow levels and fifty unique reusable workflows per top-level call tree; use those as hard ceilings, not design targets.

A practical platform usually has a thin caller template, a reusable CI workflow, one or more specialized deployment workflows, and small custom actions only for logic that truly benefits from code reuse. Templates are copied scaffolding; reusable workflows/actions remain referenced dependencies. That distinction determines how updates propagate.

7. Strict policy versus exception process

Organization/enterprise policy should enforce non-negotiable boundaries such as allowed action sources, full-SHA action pinning, workflow permission defaults and runner access where appropriate. Repository YAML cannot legitimately “override” a higher-level deny. A paved road must accompany enforcement: versioned examples, migration tooling, observable failure messages and an exception process with owner and expiry.

Exceptions are governance state, not comments. Record the exact repository/workflow/control waived, business reason, compensating control, approver, expiry date and remediation owner. A permanent exception with no review date is policy drift.

8. Speed and cost versus coverage and isolation

Optimize from run evidence, not intuition. Measure queue time, setup time, critical path, cache hit rate, artifact transfer and rerun waste. Public repositories using standard hosted runners are currently free; private repositories consume plan allowances and then billable usage. Larger runners add cost and do not fix poor dependency graphs or unnecessary work by themselves.

Do not save minutes by skipping the only cross-component or security test. Prefer safe parallelism, precise caching, cancellation of superseded non-release work and component selection with a full-suite fallback. A faster pipeline that changes correctness is not an optimization.

9. Plan and product boundaries

Capability Free/disposable path Production note
Standard hosted CI public repository free/unlimited standard hosted usage for public repos; private billing differs
Environment gate public disposable environment required reviewers on Free/Pro/Team are public-repository capability
Artifact attestation public repository private/internal require Enterprise Cloud
OIDC claim inspection public or permitted repository job provider federation requires provider-side trust configuration
Runner groups simulate locally if unavailable availability/access differs by organization/enterprise and runner type
Enterprise Actions policy local policy fixture real enforcement depends on org/enterprise controls
ARC architecture exercise only Kubernetes + controller operations belong to platform team

10. Architecture decisions must produce observable contracts

Lesson 4 assumes one of these controls fails. The diagnostic method will not start with “rerun it” or “add permissions.” It starts from the preserved run/attempt, source revision, evaluated permissions, runner and artifact identities, then traces forward to deployment and governance evidence.

Next lesson

Capstone: Build a Secure Reusable Enterprise CI/CD Platform: Diagnostics, Failure Modes, and Production Practices

Continue with the next lesson to build on the current concepts, evidence, security boundaries, and operational practices.

Knowledge check

What should be centralized first in a platform?

When does a self-hosted runner become justified?

Does id-token: write grant cloud permissions by itself?

Why should platform callers pin immutable SHAs even when an upgrade channel exists?

What makes an exception governable?

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.