Checkpoint Lab — Nodes, Agents, Labels, Executors, Workspaces, Offline Causes, and Agent Capacity Design
Design and test a two-pool topology under controlled queue pressure, repair a deliberate routing/capacity defect, and produce a value-neutral evidence packet that explains utilization and trust tradeoffs.
Learning objectives
- Design a two-pool topology with explicit capabilities, trust intent and executor counts.
- Predict queue/node/workspace state changes before generating pressure.
- Diagnose a deliberate wait using queue reason, labels, offline state and executor occupancy.
- Rebalance one setting safely and compare before/after utilization evidence.
- Produce an evidence packet that contains no agent/API credential values.
1. Scenario and success criteria
You operate a small Jenkins installation with ordinary CI and a simulated protected release workflow. Both run on Linux, but they must not share the same trust pool. CI demand occasionally creates a queue. Your task is to prove whether the wait is caused by routing or capacity, then make one reversible improvement.
| Pool | Initial node | Labels | Executors | Trust |
|---|---|---|---|---|
| CI | agent-ci-1 |
linux linux-ci shared |
1 | Ordinary synthetic CI |
| Release | agent-release-1 |
linux release-trusted |
1 | Reviewed synthetic release |
Success means the built-in node remains at zero executors, untrusted/ordinary CI cannot land on the release pool, queue behavior is explained by evidence, and the final tuning does not hide host contention.
2. Baseline assumptions and preflight
- Jenkins
2.568.3 LTS; Java 21. -
Pipeline: Nodes and Processes
1479.v56e587f413a_7. - Two disposable agents online, separate remote roots/process identities where feasible.
- Built-in node executors =
0. - No real deployment/signing credentials or proprietary source.
- If the authenticated API is used, token stays in an environment variable and is excluded from logs/artifacts.
Capture controller/core/Java/plugin versions, job full name, source revision, node labels, executor counts, online/offline state, host CPU/RAM/disk summary and assumption timestamp before workload starts.
3. Predict before acting
Write two predictions in predictions.md:
-
Triggering two 90-second
linux-cibuilds with one CI executor will leave one build running and one queue item waiting for matching capacity; the release executor remains free. -
Temporarily taking
agent-ci-1offline will make a newlinux-citask ineligible and preserve an explicit offline cause; bringing the node back should allow the queued task to proceed without changing the Jenkinsfile.
Also predict how NODE_NAME and
WORKSPACE will differ between CI and release pool runs.
4. Create the bounded test Pipelines
CI pressure job:
pipeline {
agent { label 'linux-ci' }
options { timeout(time: 3, unit: 'MINUTES') }
stages {
stage('Pressure') {
steps {
sh '''
set -eu
mkdir -p out
printf 'build=%s\nnode=%s\nlabels=%s\nworkspace=%s\n' \
"$BUILD_NUMBER" "$NODE_NAME" "$NODE_LABELS" "$WORKSPACE" \
| tee out/context.txt
'''
sleep time: 90, unit: 'SECONDS'
archiveArtifacts artifacts: 'out/context.txt', fingerprint: true
}
}
}
}
Release routing proof:
pipeline {
agent { label 'release-trusted' }
stages {
stage('Trust routing evidence') {
steps {
sh '''
set -eu
test "$NODE_NAME" = 'agent-release-1'
printf 'node=%s\nlabels=%s\nworkspace=%s\n' \
"$NODE_NAME" "$NODE_LABELS" "$WORKSPACE"
'''
}
}
}
}
5. Generate and capture queue pressure
- Trigger CI build A.
- Within ten seconds, trigger CI build B.
- Capture build A number/node/executor and build B queue item ID/reason.
-
Capture
agent-ci-1busy/total executors and host resource evidence while A runs. - Run the release routing proof and verify it uses the release pool independently.
Do not create more than the two bounded CI runs for this step.
6. Inject one deliberate failure
After A/B complete, temporarily mark agent-ci-1 offline
with reason Chapter17 checkpoint maintenance. Trigger
CI build C. Preserve the queue item and offline cause before repair.
Alternative if you cannot change node state: create a disposable
copy that requests linux-ci && arm64 when no
node carries arm64. Preserve the zero-eligibility queue
reason, then restore the original label expression.
7. Diagnose by layers
- Preserve job/build/queue/source identities.
- Confirm controller/Java/plugin baseline.
- Confirm requested label expression.
- List matching nodes and online/offline causes.
- Inspect free/busy executor counts.
- Inspect agent connection/workspace/tool evidence only after scheduling eligibility is understood.
- Inspect host resource pressure before increasing concurrency.
State the root cause in one sentence before changing anything.
8. Make one reversible capacity change
If your host showed clear spare CPU/memory/disk capacity during A/B,
temporarily change agent-ci-1 from one executor to two.
Trigger exactly two shorter 45-second CI builds and compare queue
delay, host utilization and build duration against the one-executor
baseline.
If the host was already pressured, do not add an executor. Instead document that the correct capacity action is another worker/ephemeral capacity or workload optimization. The checkpoint rewards correct evidence, not a mandatory increase.
9. Verification checklist
- Built-in node has zero executors.
-
CI builds use only
linux-ci; release proof uses onlyrelease-trusted. - Original queue/offline evidence is retained.
-
Each completed build records
NODE_NAME,NODE_LABELSandWORKSPACE. - No agent secret/API token is present in console output, artifacts or evidence files.
- Executor change, if made, has measurable before/after queue and host-resource evidence.
- No real release side effect or real privileged credential is used.
10. Required evidence packet
-
baseline.md: Jenkins/Java/plugin versions and built-in-node executor setting. -
topology.md: node names, labels, executor counts, usage modes and trust classification. -
predictions.md: predictions and later pass/fail comparison. -
queue-before.json: reviewed queue IDs/reasons during initial pressure. -
nodes-before.json: reviewed online/offline/executor/label state. -
ci-context-a.txt,ci-context-b.txtand release routing context. -
failure.md: deliberate offline/label defect, first evidence and root cause. -
capacity-change.md: executor change or explicit decision not to change, with resource evidence. -
queue-after.jsonand host summary for comparison. -
assumptions-limitations.md: local disposable agents, synthetic work, no production credentials, and metrics limitations.
11. Cleanup and rollback
-
Restore
agent-ci-1to its original executor count unless the measured change is intentionally retained in your lab. - Bring any temporarily-offline agent online after confirming the drill condition is cleared.
- Restore labels changed for deliberate mismatch testing.
- Delete disposable jobs/nodes only after reviewed evidence is saved.
- Stop disposable agent processes and remove their work directories.
- Revoke/delete lab-only API or inbound-agent bootstrap credentials.
12. What Chapter 17 adds to the production operating model
You can now explain execution capacity causally: queue requirements select eligible nodes; online state and usage policy narrow them; free executors determine immediate capacity; workspaces carry ephemeral execution state; host resources bound useful concurrency; and trust classifications determine which source and credentials may enter each pool.
Chapter 18 builds on this capacity model by focusing on how agents actually connect and live: SSH, inbound/WebSocket, Windows and ephemeral launch patterns, Remoting connectivity and agent lifecycle.
Knowledge check
Answer before revealing the explanation.
1. What are the two primary outputs of the checkpoint?
A working two-pool scheduling design and an evidence packet explaining queue demand, labels, executor counts, workspace/trust boundaries and the repair made under controlled pressure.
2. Why does the checkpoint avoid “just add executors until the queue disappears”?
Queue length alone does not reveal host capacity or trust/isolation requirements. Unlimited concurrency can move the bottleneck into CPU, memory, disk, licenses or downstream systems.
3. How do you prove an offline cause rather than infer it?
Capture the node/computer state or UI/API evidence that records the cause and timestamp, then correlate it with the queue item that became ineligible.
4. What proves a rerouted build used the intended pool?
Build number/cause plus NODE_NAME, NODE_LABELS, WORKSPACE, label expression and the node configuration captured at that time.
5. What is the safe production lesson from this lab?
Capacity and trust must be modeled together: schedule only onto eligible, authorized pools; measure queue and host pressure; keep controller executors at zero; and use reversible evidence-based tuning.
Official references and version notes
-
Jenkins LTS changelog
and
Java support policy
— lab baseline
Jenkins 2.568.3 LTS, Java 21; this LTS line supports Java 21 and 25. - Managing Nodes — controller, node, agent, executor and node-monitor concepts.
- Using Jenkins agents — labels, executor counts, usage modes and distributed builds.
-
Controller Isolation
— do not run builds on the built-in node; set its executor count
to
0once agents exist. - Securing Builds — isolate builds and separate trust domains instead of treating every shared agent as equivalent.
-
Pipeline: Nodes and Processes step reference
—
node,ws, workspace context and supported label-expression operators. -
Pipeline: Nodes and Processes plugin
— baseline
1479.v56e587f413a_7, requires Jenkins 2.479.3 or newer. -
SSH Build Agents plugin
— optional launcher baseline
3.1097.v868116049892; use verified host keys and least-privilege agent accounts. -
Built-In Node Name and Label Migration
— terminology and
NODE_NAME/NODE_LABELSbehavior.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.