Chapter 11Lesson 01~155 minutes

Job Dependencies, needs, Concurrency, Cancellation, and Deployment Serialization: Core Concepts and Mental Model

Chapter 10 supplied elastic runner capacity. Chapter 11 decides which eligible work may use that capacity together, which work must wait, and when newer work may safely replace older work. The goal is to make ordering and cancellation a delivery invariant rather than an accident of YAML layout.

needs DAGJob resultsConcurrencyCancellationSerialization

Learning objectives

  • Explain why textual job order does not create an execution dependency and model workflows as a directed acyclic graph (DAG).
  • Use needs, direct dependency results and outputs to express fan-out and fan-in explicitly.
  • Distinguish workflow-level concurrency from job-level concurrency and derive a stable, collision-resistant group key.
  • Predict the difference between queueing, pending replacement and cancel-in-progress behavior.
  • Explain why cancellation cannot roll back external side effects and why cleanup/finalization conditions need a trust-aware design.

1. The practical problem: capacity is not ordering

Suppose a repository has four jobs: lint, unit tests, package, and deploy. If the YAML lists them in that order, GitHub does not infer a pipeline. Jobs without dependencies are independently eligible and may run in parallel on different runners. A deployment that must happen only after successful validation therefore needs an explicit graph.

Now add a second workflow run while the first deployment is still changing the same target. More runners make the conflict easier to create. The problem is no longer “can GitHub execute both?” but “should these operations overlap, wait, or cancel?”

2. Causal model: DAG first, arbitration second

Start with job eligibility. needs creates dependency edges. Completed dependencies expose conclusions and outputs. Only then does a job that reaches a concurrency boundary compete for its concurrency group. Arbitration may let it run, make it pending, replace an older pending item, queue it, or cancel in-progress work depending on configuration.

Job graph and concurrency arbitration
flowchart TD
  A[Workflow run created] --> B[Independent jobs become eligible]
  B --> C1[lint]
  B --> C2[test]
  C1 --> D[needs fan-in: aggregate]
  C2 --> D
  D --> E{aggregate result / outputs}
  E -->|ready| F[deployment job reaches concurrency group]
  E -->|not ready| G[deployment skipped]
  F --> H{group arbitration}
  H -->|free| I[run critical section]
  H -->|busy + queue| J[pending / queued]
  H -->|cancel old| K[cancel in-progress work]
  I --> L[external side effect + evidence]
  K --> M[side effects already committed remain]

The DAG answers what must finish before this job is eligible? Concurrency answers which eligible jobs or runs may occupy this critical section at the same time? Cancellation answers what happens to older work when a replacement arrives? None of those mechanisms automatically reverses an external change already committed.

3. Name the states that must be observable

State Example Owned by Evidence
Job identity lint, test, aggregate, deploy Workflow revision workflow YAML + job list
DAG edge aggregate needs [lint, test] Workflow revision needs declaration
Dependency result needs.test.result == 'success' GitHub run state job conclusion / needs context
Concurrency group deploy-staging Repository concurrency state resolved group expression
Pending/queue state run B waits for run A GitHub Actions run/job status and timestamps
Cancellation policy cancel-in-progress: true Workflow/job config workflow revision + evaluated group
External target fake staging revision External system/simulator before/after target record
Committed side effect target changed to run A External system target audit/log, not merely job status

4. needs is an explicit DAG edge

Without needs, jobs are peers. With needs, a job waits for the listed direct dependencies. If a needed job fails or is skipped, the dependent job is skipped by default unless its own if explicitly allows continuation.

jobs:
  lint:
    runs-on: ubuntu-24.04
    steps:
      - run: echo "lint ok"

  test:
    runs-on: ubuntu-24.04
    steps:
      - run: echo "tests ok"

  aggregate:
    needs: [lint, test]
    runs-on: ubuntu-24.04
    steps:
      - run: echo "both direct dependencies succeeded"

This is a fan-out/fan-in graph: lint and test can run in parallel; aggregate waits for both.

5. needs exposes only direct dependencies

The needs context contains each direct dependency’s outputs and result. Results are success, failure, cancelled, or skipped. It does not recursively contain dependencies of dependencies.

aggregate:
  if: ${{ always() }}
  needs: [lint, test]
  runs-on: ubuntu-24.04
  steps:
    - shell: bash
      run: |
        printf 'lint=%s\n' '${{ needs.lint.result }}'
        printf 'test=%s\n' '${{ needs.test.result }}'
Use always() narrowly

This read-only aggregation is a reasonable use because it preserves evidence after failures. Do not infer that always() is appropriate for deployments, credential changes, destructive cleanup, or other privileged work. Cancellation semantics can keep an always() job or step running.

6. Concurrency is repository-wide mutual exclusion

A concurrency group is a repository-level key shared by jobs or workflow runs that declare the same resolved value. Group names are case-insensitive. Workflow-level concurrency arbitrates entire runs; job-level concurrency arbitrates only the job that reaches that boundary.

concurrency:
  group: ${{ github.workflow }}-${{ github.ref }}
  cancel-in-progress: true

This is a good pattern for stale CI on a branch: a newer run can replace older in-progress work. It is usually a bad default for an irreversible deployment, because a canceled deployment may already have changed the external target.

7. Pending replacement, queueing and cancellation are different policies

Configuration Running Pending When useful
default group only 1 1; newer pending replaces older pending latest-only waiting behavior without canceling current work
queue: max 1 up to 100 serialized deployments where every accepted request should wait
cancel-in-progress: true new work can cancel current default replacement rules discard stale validation/build work

queue: max and cancel-in-progress: true are mutually exclusive. A queue is appropriate when each deployment must be observed and serialized. Cancellation is appropriate only when older work is genuinely disposable or has a safe compensation protocol.

8. Cancellation is a control signal, not rollback

When GitHub cancels a run, it re-evaluates conditions on running jobs and steps. Work whose condition remains true can continue; other work receives a cancellation signal and is eventually terminated. If an external API call already committed a release, database migration, DNS update, or fake target marker, process termination does not reverse it.

run A timeline
00:00 prepare
00:20 external target changed to revision A   <-- committed
00:25 run B arrives, cancels run A
00:26 runner process stops

external target is still revision A unless compensation changes it.

9. Read-only inspection before changing policy

Before adding a lock or cancellation rule, capture the current workflow revision and run graph. In GitHub UI or gh, record run ID/attempt/SHA, job IDs, conclusions, timestamps and whether overlapping runs currently exist. Do not cancel anything merely to “see what happens” in a production repository.

# Optional read-only inspection in an authorized repository.
gh run list --limit 10
gh run view RUN_ID --json databaseId,attempt,headSha,event,status,conclusion,jobs

Then state the delivery invariant in plain language, such as “only one deployment to staging may execute at a time, and accepted deployments must not be dropped.” That sentence determines the concurrency design.

Knowledge check

If two jobs appear one after another in YAML but neither has needs, what ordering guarantee exists?

What does needs.test.result report?

Why is queue: max different from cancel-in-progress: true?

A deploy job with always() is running when the workflow is canceled. Why is that dangerous?

Does canceling a runner process roll back an already committed external API change?

Next chapter concept

From model to overlapping-run evidence

Lesson 2 builds a disposable workflow with parallel jobs, fan-in evidence, serialized fake deployment and a separate cancel-old experiment.

Official references and version notes

Version and compatibility note

Version-sensitive behavior was rechecked against current GitHub-maintained documentation on 2026-09-09. Current documentation states that needs contains only direct dependencies and exposes result as success, failure, cancelled or skipped. A failed or skipped dependency normally skips downstream jobs unless an explicit job condition permits continuation. Concurrency groups are repository-wide and case-insensitive. By default at most one item may be running and one pending in a group; a newer pending item replaces an older pending item. Current GitHub Actions also supports queue: max to allow up to 100 pending items, processed FIFO by time waiting on the group; this mode cannot be combined with cancel-in-progress: true. Cancellation re-evaluates job/step conditions, so always() can keep work running during cancellation and must not be used casually for privileged or irreversible operations. Mandatory labs use ubuntu-24.04, permissions: {}, no Marketplace action and no real credential.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.