Chapter 17Lesson 05~180 minutes

Checkpoint Lab — Nodes, Agents, Labels, Executors, Workspaces, Offline Causes, and Agent Capacity Design

Design and test a two-pool topology under controlled queue pressure, repair a deliberate routing/capacity defect, and produce a value-neutral evidence packet that explains utilization and trust tradeoffs.

Checkpoint labTwo poolsCapacity tuningUtilizationTrust tradeoffsCleanup

Learning objectives

  • Design a two-pool topology with explicit capabilities, trust intent and executor counts.
  • Predict queue/node/workspace state changes before generating pressure.
  • Diagnose a deliberate wait using queue reason, labels, offline state and executor occupancy.
  • Rebalance one setting safely and compare before/after utilization evidence.
  • Produce an evidence packet that contains no agent/API credential values.

1. Scenario and success criteria

You operate a small Jenkins installation with ordinary CI and a simulated protected release workflow. Both run on Linux, but they must not share the same trust pool. CI demand occasionally creates a queue. Your task is to prove whether the wait is caused by routing or capacity, then make one reversible improvement.

Pool Initial node Labels Executors Trust
CI agent-ci-1 linux linux-ci shared 1 Ordinary synthetic CI
Release agent-release-1 linux release-trusted 1 Reviewed synthetic release

Success means the built-in node remains at zero executors, untrusted/ordinary CI cannot land on the release pool, queue behavior is explained by evidence, and the final tuning does not hide host contention.

2. Baseline assumptions and preflight

  • Jenkins 2.568.3 LTS; Java 21.
  • Pipeline: Nodes and Processes 1479.v56e587f413a_7.
  • Two disposable agents online, separate remote roots/process identities where feasible.
  • Built-in node executors = 0.
  • No real deployment/signing credentials or proprietary source.
  • If the authenticated API is used, token stays in an environment variable and is excluded from logs/artifacts.

Capture controller/core/Java/plugin versions, job full name, source revision, node labels, executor counts, online/offline state, host CPU/RAM/disk summary and assumption timestamp before workload starts.

3. Predict before acting

Write two predictions in predictions.md:

  1. Triggering two 90-second linux-ci builds with one CI executor will leave one build running and one queue item waiting for matching capacity; the release executor remains free.
  2. Temporarily taking agent-ci-1 offline will make a new linux-ci task ineligible and preserve an explicit offline cause; bringing the node back should allow the queued task to proceed without changing the Jenkinsfile.

Also predict how NODE_NAME and WORKSPACE will differ between CI and release pool runs.

4. Create the bounded test Pipelines

CI pressure job:

pipeline {
  agent { label 'linux-ci' }
  options { timeout(time: 3, unit: 'MINUTES') }
  stages {
    stage('Pressure') {
      steps {
        sh '''
          set -eu
          mkdir -p out
          printf 'build=%s\nnode=%s\nlabels=%s\nworkspace=%s\n' \
            "$BUILD_NUMBER" "$NODE_NAME" "$NODE_LABELS" "$WORKSPACE" \
            | tee out/context.txt
        '''
        sleep time: 90, unit: 'SECONDS'
        archiveArtifacts artifacts: 'out/context.txt', fingerprint: true
      }
    }
  }
}

Release routing proof:

pipeline {
  agent { label 'release-trusted' }
  stages {
    stage('Trust routing evidence') {
      steps {
        sh '''
          set -eu
          test "$NODE_NAME" = 'agent-release-1'
          printf 'node=%s\nlabels=%s\nworkspace=%s\n' \
            "$NODE_NAME" "$NODE_LABELS" "$WORKSPACE"
        '''
      }
    }
  }
}

5. Generate and capture queue pressure

  1. Trigger CI build A.
  2. Within ten seconds, trigger CI build B.
  3. Capture build A number/node/executor and build B queue item ID/reason.
  4. Capture agent-ci-1 busy/total executors and host resource evidence while A runs.
  5. Run the release routing proof and verify it uses the release pool independently.

Do not create more than the two bounded CI runs for this step.

6. Inject one deliberate failure

After A/B complete, temporarily mark agent-ci-1 offline with reason Chapter17 checkpoint maintenance. Trigger CI build C. Preserve the queue item and offline cause before repair.

Alternative if you cannot change node state: create a disposable copy that requests linux-ci && arm64 when no node carries arm64. Preserve the zero-eligibility queue reason, then restore the original label expression.

7. Diagnose by layers

  1. Preserve job/build/queue/source identities.
  2. Confirm controller/Java/plugin baseline.
  3. Confirm requested label expression.
  4. List matching nodes and online/offline causes.
  5. Inspect free/busy executor counts.
  6. Inspect agent connection/workspace/tool evidence only after scheduling eligibility is understood.
  7. Inspect host resource pressure before increasing concurrency.

State the root cause in one sentence before changing anything.

8. Make one reversible capacity change

If your host showed clear spare CPU/memory/disk capacity during A/B, temporarily change agent-ci-1 from one executor to two. Trigger exactly two shorter 45-second CI builds and compare queue delay, host utilization and build duration against the one-executor baseline.

If the host was already pressured, do not add an executor. Instead document that the correct capacity action is another worker/ephemeral capacity or workload optimization. The checkpoint rewards correct evidence, not a mandatory increase.

9. Verification checklist

  • Built-in node has zero executors.
  • CI builds use only linux-ci; release proof uses only release-trusted.
  • Original queue/offline evidence is retained.
  • Each completed build records NODE_NAME, NODE_LABELS and WORKSPACE.
  • No agent secret/API token is present in console output, artifacts or evidence files.
  • Executor change, if made, has measurable before/after queue and host-resource evidence.
  • No real release side effect or real privileged credential is used.

10. Required evidence packet

  • baseline.md: Jenkins/Java/plugin versions and built-in-node executor setting.
  • topology.md: node names, labels, executor counts, usage modes and trust classification.
  • predictions.md: predictions and later pass/fail comparison.
  • queue-before.json: reviewed queue IDs/reasons during initial pressure.
  • nodes-before.json: reviewed online/offline/executor/label state.
  • ci-context-a.txt, ci-context-b.txt and release routing context.
  • failure.md: deliberate offline/label defect, first evidence and root cause.
  • capacity-change.md: executor change or explicit decision not to change, with resource evidence.
  • queue-after.json and host summary for comparison.
  • assumptions-limitations.md: local disposable agents, synthetic work, no production credentials, and metrics limitations.

11. Cleanup and rollback

  1. Restore agent-ci-1 to its original executor count unless the measured change is intentionally retained in your lab.
  2. Bring any temporarily-offline agent online after confirming the drill condition is cleared.
  3. Restore labels changed for deliberate mismatch testing.
  4. Delete disposable jobs/nodes only after reviewed evidence is saved.
  5. Stop disposable agent processes and remove their work directories.
  6. Revoke/delete lab-only API or inbound-agent bootstrap credentials.

12. What Chapter 17 adds to the production operating model

You can now explain execution capacity causally: queue requirements select eligible nodes; online state and usage policy narrow them; free executors determine immediate capacity; workspaces carry ephemeral execution state; host resources bound useful concurrency; and trust classifications determine which source and credentials may enter each pool.

Chapter 18 builds on this capacity model by focusing on how agents actually connect and live: SSH, inbound/WebSocket, Windows and ephemeral launch patterns, Remoting connectivity and agent lifecycle.

Next chapter

Chapter 18 — SSH, Inbound, WebSocket, Windows, and Ephemeral Agents: Connectivity, Remoting, and Agent Lifecycle

Keep the capacity and trust model from this checkpoint, then examine launcher choices, connection security, reconnection, platform differences and ephemeral worker lifecycle.

Knowledge check

Answer before revealing the explanation.

1. What are the two primary outputs of the checkpoint?

2. Why does the checkpoint avoid “just add executors until the queue disappears”?

3. How do you prove an offline cause rather than infer it?

4. What proves a rerouted build used the intended pool?

5. What is the safe production lesson from this lab?

Official references and version notes

Assumption timestamp: 2026-09-17. Recheck Jenkins LTS/Java support, agent/remoting compatibility, launcher-plugin advisories, and current node/monitor behavior before reproducing the lab later.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.