CI/CD Build Patterns, Caching, Test Sharding, Artifact Promotion, and Ephemeral Build Agents: Diagnostics, Failure Modes, Security, and Performance
Diagnose false-green and unsafe pipelines: poisoned cache namespaces, stale build-output restoration, shard gaps/overlap, checksum drift during promotion, Wrapper/JDK mismatches between stages, and secrets leaked through logs or command lines.
Learning objectives
- Use an evidence-preserving diagnostic sequence for CI build failures and false greens.
- Recognize cache poisoning and stale build-output restoration as trust/modeling defects.
- Detect shard gaps/duplicates independently of individual shard success.
- Prove promotion checksum drift and Wrapper/JDK drift between stages.
- Contain secret exposure without deleting the evidence needed for incident response.
Do not fix evidence by deleting it. If a cache,
artifact, or credential may be compromised, preserve the run ID,
checksum, relevant logs, cache key, source revision, and tool
identity before containment. Never begin by deleting normal
~/.gradle or ~/.m2 state.
1. Diagnostic sequence
- Preserve concise failing evidence and exact pipeline/run identity.
- Confirm source revision, Wrapper/build-tool, JDK/toolchain, and runner identity.
- Inspect build configuration, dependency/repository policy, and shard definitions.
- Inspect task/test graph and all shard report sets.
- Inspect cache keys, cache producer trust, workspace outputs, and artifact-store entries.
- Inspect compiler/test/plugin failure without suppressing it.
- Apply the least destructive correction.
- Rerun in a controlled clean agent and compare evidence.
2. Failure: cache key mixes trusted and untrusted branches
Symptom: protected builds sometimes restore unexpected outputs or dependencies after a pull-request build. Cause: both trust classes write the same cache namespace. Wrong fix: periodically wipe the shared cache and keep the key unchanged.
Repair: separate trust classes in keys/policies; prevent untrusted runs from publishing entries consumed by protected releases; retain verification/lock controls on restored dependencies. If the suspect entry may have affected a release, treat it as an incident and rebuild in fresh isolated state.
3. Failure: CI restores build outputs as if they were dependencies
Symptom: build/libs/app.jar exists
before package tasks run, or a job can “succeed” without recreating
expected outputs. Cause: a provider cache archived
the project build/ directory.
Repair: remove project outputs from
dependency-cache restore rules. If task-output reuse is desired, use
Gradle's modeled Build Cache instead.
# Harmless local demonstration of the wrong state class:
mkdir -p .ci/bad-cache/build/libs
printf 'stale bytes\n' > .ci/bad-cache/build/libs/app.jar
# A correct clean build should not treat that file as release truth.
rm -rf build
./gradlew ciTest jar
sha256sum build/libs/*.jar
5. Failure: promotion rebuilds and checksum changes
Symptom: staging and production JARs carry the same
version but different SHA-256 values. Cause: the
deployment stage ran jar/package again or
mutated the binary. Repair: return to the
build-stage exported artifact. Verify its checksum and promote that
object. If a production artifact already differs, record both hashes
and investigate before replacing it.
6. Failure: Wrapper or JDK differs between stages
Symptom: tests run under JDK 21 but packaging or
signing runs under another runtime; Gradle/Maven versions differ
despite a single repository commit. Cause: later
stages bypass the Wrapper or use a different runner image without
evidence. Repair: execute the Wrapper in every
build-tool stage and record --version/java -version. For pure promotion, do not run the build tool at all.
7. Failure: a secret appears in logs, command arguments, or telemetry
Symptom: a repository token appears in a shell trace, error message, process command line, or published telemetry. Immediate action: preserve the relevant evidence securely, revoke/rotate the credential, determine which jobs/artifacts/logs received it, and remove unnecessary exposure. Repair: inject secrets through the CI secret mechanism at the smallest required stage and avoid echoing/embedding them in build metadata.
Do not merely redact the current log and continue using the same token. Exposure is a credential incident.
8. Distinguish cold dependency resolution, build-cache reuse, and stale workspace state
Performance symptoms have different causes. A cold dependency cache increases network/resolution time. A Build Cache miss reruns cacheable tasks. A stale persistent workspace may skip work via up-to-date state or leave unrelated files behind. Capture task outcome labels and cache state before changing multiple knobs at once.
| Observed evidence | Likely layer |
|---|---|
| Many dependency downloads | Dependency cache / repository resolution. |
FROM-CACHE changes to executed tasks |
Build Cache key or availability. |
UP-TO-DATE on persistent workspace only |
Workspace incremental state. |
| Old JAR exists before build starts | Improper provider cache or persistent workspace residue. |
9. Controlled rebuild
When correctness is in doubt, create a fresh project copy plus fresh isolated Gradle/Maven state. Keep the suspect cache untouched for evidence. Rebuild from the same commit with the same verified toolchain, then compare dependency graph, tests, JAR checksum, and report set. The clean run answers whether cached/shared state caused the discrepancy without destroying the suspect state first.
10. Performance guardrail
Never “optimize” by setting tests to ignored, restoring
build/, disabling dependency verification, broadening
repository credentials, or rebuilding fewer artifacts without
changing the evidence contract. Measure CI throughput by critical
path and total agent minutes, but keep correctness controls
invariant while tuning workers, shards, caches, or agent size.
Knowledge check
Why is deleting the suspect cache first a poor incident response?
It destroys evidence about which namespace/key/producer may have affected the build and can make root-cause analysis harder.
What proves a shard plan is complete?
An independent comparison between the intended test inventory and the union of shard assignments/reports.
What should happen if staging and production hashes differ for the “same” release?
Stop and investigate. Reuse the original verified artifact; do not normalize the difference by rebuilding again.
Why should promotion usually not need a JDK?
A pure promotion step verifies and copies/publishes existing artifact bytes; it does not compile or test.
What is the correct response to a leaked CI token?
Preserve evidence securely, revoke/rotate, determine exposure scope, and reduce future injection/logging—not merely hide the log line.
Official references and version notes
- Gradle 9.7.1 release notes — pinned Gradle baseline.
- Gradle dependency caching — dependency-cache state and guidance for ephemeral builds.
- Gradle Build Cache — local/remote task-output cache and CI push/read trust model.
- Build Cache use cases — CI-produced cache entries and cross-machine reuse.
-
Gradle task outcomes
—
UP-TO-DATEversusFROM-CACHE. - Gradle on GitLab CI — current Wrapper-first CI guidance and cache considerations.
- Maven release history — Maven 3.9.16 GA baseline.
- Maven Wrapper 3.3.4 — current stable Wrapper baseline.
- Maven Surefire test goal — explicit test selection and failure behavior.
Version-sensitive statements were rechecked against primary documentation on 2026-08-24. Mandatory labs remain local/free; hosted CI, remote caches, artifact repositories, and secret stores are represented as optional production mappings only.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.