Chapter 29Lesson 04~310 minutes

CI/CD Build Patterns, Caching, Test Sharding, Artifact Promotion, and Ephemeral Build Agents: Diagnostics, Failure Modes, Security, and Performance

Diagnose false-green and unsafe pipelines: poisoned cache namespaces, stale build-output restoration, shard gaps/overlap, checksum drift during promotion, Wrapper/JDK mismatches between stages, and secrets leaked through logs or command lines.

Cache poisoningFalse greenChecksum driftToolchain driftSecrets

Learning objectives

  • Use an evidence-preserving diagnostic sequence for CI build failures and false greens.
  • Recognize cache poisoning and stale build-output restoration as trust/modeling defects.
  • Detect shard gaps/duplicates independently of individual shard success.
  • Prove promotion checksum drift and Wrapper/JDK drift between stages.
  • Contain secret exposure without deleting the evidence needed for incident response.

Do not fix evidence by deleting it. If a cache, artifact, or credential may be compromised, preserve the run ID, checksum, relevant logs, cache key, source revision, and tool identity before containment. Never begin by deleting normal ~/.gradle or ~/.m2 state.

1. Diagnostic sequence

  1. Preserve concise failing evidence and exact pipeline/run identity.
  2. Confirm source revision, Wrapper/build-tool, JDK/toolchain, and runner identity.
  3. Inspect build configuration, dependency/repository policy, and shard definitions.
  4. Inspect task/test graph and all shard report sets.
  5. Inspect cache keys, cache producer trust, workspace outputs, and artifact-store entries.
  6. Inspect compiler/test/plugin failure without suppressing it.
  7. Apply the least destructive correction.
  8. Rerun in a controlled clean agent and compare evidence.

2. Failure: cache key mixes trusted and untrusted branches

Symptom: protected builds sometimes restore unexpected outputs or dependencies after a pull-request build. Cause: both trust classes write the same cache namespace. Wrong fix: periodically wipe the shared cache and keep the key unchanged.

Repair: separate trust classes in keys/policies; prevent untrusted runs from publishing entries consumed by protected releases; retain verification/lock controls on restored dependencies. If the suspect entry may have affected a release, treat it as an incident and rebuild in fresh isolated state.

3. Failure: CI restores build outputs as if they were dependencies

Symptom: build/libs/app.jar exists before package tasks run, or a job can “succeed” without recreating expected outputs. Cause: a provider cache archived the project build/ directory. Repair: remove project outputs from dependency-cache restore rules. If task-output reuse is desired, use Gradle's modeled Build Cache instead.

# Harmless local demonstration of the wrong state class:
mkdir -p .ci/bad-cache/build/libs
printf 'stale bytes\n' > .ci/bad-cache/build/libs/app.jar

# A correct clean build should not treat that file as release truth.
rm -rf build
./gradlew ciTest jar
sha256sum build/libs/*.jar

4. Failure: every shard job is green but the suite is incomplete

Symptom: shard A and B both pass, yet a newly added test never appears in any XML report. Cause: the partition manifest/rules were not updated. Repair: validate the expected test inventory against the union of shard assignments before or after execution.

cat > expected-tests.txt <<'EOF'
ShardATest
ShardBTest
EOF
cat > shard-a.txt <<'EOF'
ShardATest
EOF
cat > shard-b.txt <<'EOF'
# Intentionally broken: ShardBTest missing
EOF

grep -v '^#' shard-a.txt shard-b.txt | sed '/^$/d;s/^[^:]*://' | sort -u > assigned-tests.txt
if ! diff -u expected-tests.txt assigned-tests.txt; then
  echo "Shard plan is incomplete" >&2
  exit 1
fi

Also check duplicates when the policy expects exactly-once assignment. A duplicate may waste time or hide ordering/shared-state bugs.

5. Failure: promotion rebuilds and checksum changes

Symptom: staging and production JARs carry the same version but different SHA-256 values. Cause: the deployment stage ran jar/package again or mutated the binary. Repair: return to the build-stage exported artifact. Verify its checksum and promote that object. If a production artifact already differs, record both hashes and investigate before replacing it.

6. Failure: Wrapper or JDK differs between stages

Symptom: tests run under JDK 21 but packaging or signing runs under another runtime; Gradle/Maven versions differ despite a single repository commit. Cause: later stages bypass the Wrapper or use a different runner image without evidence. Repair: execute the Wrapper in every build-tool stage and record --version/java -version. For pure promotion, do not run the build tool at all.

7. Failure: a secret appears in logs, command arguments, or telemetry

Symptom: a repository token appears in a shell trace, error message, process command line, or published telemetry. Immediate action: preserve the relevant evidence securely, revoke/rotate the credential, determine which jobs/artifacts/logs received it, and remove unnecessary exposure. Repair: inject secrets through the CI secret mechanism at the smallest required stage and avoid echoing/embedding them in build metadata.

Do not merely redact the current log and continue using the same token. Exposure is a credential incident.

8. Distinguish cold dependency resolution, build-cache reuse, and stale workspace state

Performance symptoms have different causes. A cold dependency cache increases network/resolution time. A Build Cache miss reruns cacheable tasks. A stale persistent workspace may skip work via up-to-date state or leave unrelated files behind. Capture task outcome labels and cache state before changing multiple knobs at once.

Observed evidence Likely layer
Many dependency downloads Dependency cache / repository resolution.
FROM-CACHE changes to executed tasks Build Cache key or availability.
UP-TO-DATE on persistent workspace only Workspace incremental state.
Old JAR exists before build starts Improper provider cache or persistent workspace residue.

9. Controlled rebuild

When correctness is in doubt, create a fresh project copy plus fresh isolated Gradle/Maven state. Keep the suspect cache untouched for evidence. Rebuild from the same commit with the same verified toolchain, then compare dependency graph, tests, JAR checksum, and report set. The clean run answers whether cached/shared state caused the discrepancy without destroying the suspect state first.

10. Performance guardrail

Never “optimize” by setting tests to ignored, restoring build/, disabling dependency verification, broadening repository credentials, or rebuilding fewer artifacts without changing the evidence contract. Measure CI throughput by critical path and total agent minutes, but keep correctness controls invariant while tuning workers, shards, caches, or agent size.

Knowledge check

Why is deleting the suspect cache first a poor incident response?

What proves a shard plan is complete?

What should happen if staging and production hashes differ for the “same” release?

Why should promotion usually not need a JDK?

What is the correct response to a leaked CI token?

Official references and version notes

Version-sensitive statements were rechecked against primary documentation on 2026-08-24. Mandatory labs remain local/free; hosted CI, remote caches, artifact repositories, and secret stores are represented as optional production mappings only.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.