Chapter 11Lesson 04~190 minutes

Job Dependencies, needs, Concurrency, Cancellation, and Deployment Serialization: Diagnostics, Failure Modes, and Production Practices

Most concurrency incidents are misdiagnosed because teams look only at the job that is currently visible. This lesson preserves the original run/attempt and reconstructs the job graph, resolved concurrency group, cancellation timeline and external target before changing YAML.

DiagnosticsFirst-failure evidenceGroup collisionsCancellationCompensation

Learning objectives

  • Diagnose missing dependencies, skip propagation and direct needs results before blaming runners.
  • Identify concurrency-group collisions, over-broad locks and invalid queue/cancellation combinations.
  • Interpret cancellation from run/job status and process termination evidence rather than assuming rollback.
  • Separate read-only always() evidence work from unsafe privileged finalizers.
  • Apply the least destructive correction and rerun the smallest equivalent scope.

1. Evidence-first diagnostic sequence

  1. Preserve run ID, attempt, SHA, workflow revision and first-failure/cancellation logs.
  2. Draw the actual needs DAG from job IDs; do not infer it from file order.
  3. Record each direct dependency result and any outputs consumed downstream.
  4. Resolve the concurrency group exactly as GitHub evaluates it and record queue/cancel-in-progress.
  5. Inspect which jobs were queued, pending, running, skipped, canceled or completed.
  6. Confirm runner/image/toolchain only after selection/arbitration is understood.
  7. Inspect external target state separately from GitHub conclusions.
  8. Apply one smallest correction; do not simultaneously rewrite the DAG, group key and deployment logic.

2. Failure: assuming YAML order creates dependencies

jobs:
  build:
    runs-on: ubuntu-24.04
    steps:
      - run: sleep 20 && echo build-complete

  deploy:
    runs-on: ubuntu-24.04
    steps:
      - run: echo "deploy started"

The broken assumption is textual ordering. Both jobs are immediately eligible. The smallest repair is deploy: needs: build, followed by a target guard. Do not “fix” it by adding arbitrary sleeps.

3. Failure: dependency skip propagates farther than expected

If test fails and package needs it, package is skipped by default. If deploy needs package, deploy is also skipped. A diagnostic summary that must still run needs its own explicit condition.

diagnose:
  if: ${{ always() }}
  needs: [test, package]
  runs-on: ubuntu-24.04
  steps:
    - run: |
        echo "test=${{ needs.test.result }}"
        echo "package=${{ needs.package.result }}"

Keep this job read-only. Do not convert skip propagation into “continue deployment anyway.”

4. Failure: a group key collides across workflows

# Workflow A
concurrency:
  group: production
  cancel-in-progress: true

# Workflow B — unrelated but same repository and same key
concurrency:
  group: Production
  cancel-in-progress: true

Group names are case-insensitive, so these collide. A new run from B can cancel A. Preserve both run IDs and resolved group names, then namespace the groups by workflow/application/target.

5. Intentionally broken example: contradictory queue policy

The current concurrency schema rejects the following combination:

concurrency:
  group: deploy-production
  queue: max
  cancel-in-progress: true

queue: max means keep pending work; cancel-in-progress: true means replace active work. GitHub documents the combination as invalid. Preserve the workflow validation error, then decide which delivery invariant is actually intended instead of toggling syntax until validation passes.

6. Failure: cancellation happens after an irreversible side effect

12:00:00 run 810 deploy starts
12:00:20 provider accepted deployment revision A
12:00:25 run 811 enters same cancel-in-progress group
12:00:26 run 810 receives cancellation
12:00:30 job concludes cancelled
12:01:00 external target still reports revision A

The causal error is treating job cancellation as transaction rollback. The repair belongs at the external layer: idempotent release identifiers, optimistic version checks, a compensation/rollback API, or queueing so a deployment is not preempted after commitment.

7. Failure: always() keeps privileged work alive

cleanup-production:
  if: ${{ always() }}
  needs: deploy
  permissions:
    deployments: write
  # privileged external cleanup omitted

During cancellation, GitHub re-evaluates job conditions. Because always() remains true, this job may continue when the operator expected the run to stop. A safe design separates read-only evidence collection from explicit compensation, and compensation has target identity, cancellation/failure predicates and idempotency guards.

8. Failure: global serialization creates an operational traffic jam

A group such as deploy shared by every service and environment can make unrelated deployments wait. With queue: max, the effect may look like a deadlock even though GitHub is correctly serializing a huge queue. Inspect the resolved group key and waiting timestamps before blaming runner capacity.

Repair by expressing the smallest real exclusion domain: for example deploy-${{ inputs.service }}-${{ inputs.environment }}.

9. Failure: reusable workflow cancels its caller

If the caller and called workflow both use a group derived from github.workflow with cancel-in-progress: true, the called workflow can share the caller’s group. The correction is a distinct reusable-workflow namespace or a caller-provided group contract.

10. Layered failure matrix

Symptom First layer to inspect Evidence
deploy started before test DAG/workflow syntax needs edges + timestamps
new run canceled unrelated workflow concurrency key resolved group + both run IDs
job never reaches runner needs/if/concurrency queue job status, dependency results, pending state
job canceled but target changed external side effect provider/target audit + cancellation time
all deployments waiting over-broad group/queue group key + waiting jobs
cleanup keeps running after cancel job/step condition if evaluation + cancellation log

11. Rerun discipline

Do not blindly rerun a canceled or failed deployment. First verify external target state and whether the prior attempt committed anything. If the issue is only a validation job, rerun that smallest safe scope. If a deployment side effect is ambiguous, reconcile the target before any retry.

Knowledge check

A deploy job starts before build finishes. What is the first thing to inspect?

Why do group names production and Production collide?

What should you do with a validation error caused by queue: max plus cancel-in-progress: true?

A canceled job changed the external target before cancellation. What evidence matters most?

Why can a huge deployment queue look like a runner-capacity problem?

Next chapter concept

Checkpoint: predict the whole graph and target

Lesson 5 combines DAG evidence, queueing, cancel-old behavior and a deterministic shared-target simulator into one operating-model checkpoint.

Official references and version notes

Version and compatibility note

Version-sensitive behavior was rechecked against current GitHub-maintained documentation on 2026-09-09. Current documentation states that needs contains only direct dependencies and exposes result as success, failure, cancelled or skipped. A failed or skipped dependency normally skips downstream jobs unless an explicit job condition permits continuation. Concurrency groups are repository-wide and case-insensitive. By default at most one item may be running and one pending in a group; a newer pending item replaces an older pending item. Current GitHub Actions also supports queue: max to allow up to 100 pending items, processed FIFO by time waiting on the group; this mode cannot be combined with cancel-in-progress: true. Cancellation re-evaluates job/step conditions, so always() can keep work running during cancellation and must not be used casually for privileged or irreversible operations. Mandatory labs use ubuntu-24.04, permissions: {}, no Marketplace action and no real credential.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.