Job Dependencies, needs, Concurrency, Cancellation, and Deployment Serialization: Diagnostics, Failure Modes, and Production Practices
Most concurrency incidents are misdiagnosed because teams look only at the job that is currently visible. This lesson preserves the original run/attempt and reconstructs the job graph, resolved concurrency group, cancellation timeline and external target before changing YAML.
Learning objectives
-
Diagnose missing dependencies, skip propagation and direct
needsresults before blaming runners. - Identify concurrency-group collisions, over-broad locks and invalid queue/cancellation combinations.
- Interpret cancellation from run/job status and process termination evidence rather than assuming rollback.
-
Separate read-only
always()evidence work from unsafe privileged finalizers. - Apply the least destructive correction and rerun the smallest equivalent scope.
1. Evidence-first diagnostic sequence
- Preserve run ID, attempt, SHA, workflow revision and first-failure/cancellation logs.
-
Draw the actual
needsDAG from job IDs; do not infer it from file order. - Record each direct dependency result and any outputs consumed downstream.
-
Resolve the concurrency group exactly as GitHub evaluates it and
record
queue/cancel-in-progress. - Inspect which jobs were queued, pending, running, skipped, canceled or completed.
- Confirm runner/image/toolchain only after selection/arbitration is understood.
- Inspect external target state separately from GitHub conclusions.
- Apply one smallest correction; do not simultaneously rewrite the DAG, group key and deployment logic.
2. Failure: assuming YAML order creates dependencies
jobs:
build:
runs-on: ubuntu-24.04
steps:
- run: sleep 20 && echo build-complete
deploy:
runs-on: ubuntu-24.04
steps:
- run: echo "deploy started"
The broken assumption is textual ordering. Both jobs are immediately
eligible. The smallest repair is deploy: needs: build,
followed by a target guard. Do not “fix” it by adding arbitrary
sleeps.
3. Failure: dependency skip propagates farther than expected
If test fails and package needs it,
package is skipped by default. If deploy needs package,
deploy is also skipped. A diagnostic summary that must still run
needs its own explicit condition.
diagnose:
if: ${{ always() }}
needs: [test, package]
runs-on: ubuntu-24.04
steps:
- run: |
echo "test=${{ needs.test.result }}"
echo "package=${{ needs.package.result }}"
Keep this job read-only. Do not convert skip propagation into “continue deployment anyway.”
4. Failure: a group key collides across workflows
# Workflow A
concurrency:
group: production
cancel-in-progress: true
# Workflow B — unrelated but same repository and same key
concurrency:
group: Production
cancel-in-progress: true
Group names are case-insensitive, so these collide. A new run from B can cancel A. Preserve both run IDs and resolved group names, then namespace the groups by workflow/application/target.
5. Intentionally broken example: contradictory queue policy
The current concurrency schema rejects the following combination:
concurrency:
group: deploy-production
queue: max
cancel-in-progress: true
queue: max means keep pending work;
cancel-in-progress: true means replace active work.
GitHub documents the combination as invalid. Preserve the workflow
validation error, then decide which delivery invariant is actually
intended instead of toggling syntax until validation passes.
6. Failure: cancellation happens after an irreversible side effect
12:00:00 run 810 deploy starts
12:00:20 provider accepted deployment revision A
12:00:25 run 811 enters same cancel-in-progress group
12:00:26 run 810 receives cancellation
12:00:30 job concludes cancelled
12:01:00 external target still reports revision A
The causal error is treating job cancellation as transaction rollback. The repair belongs at the external layer: idempotent release identifiers, optimistic version checks, a compensation/rollback API, or queueing so a deployment is not preempted after commitment.
7. Failure: always() keeps privileged work alive
cleanup-production:
if: ${{ always() }}
needs: deploy
permissions:
deployments: write
# privileged external cleanup omitted
During cancellation, GitHub re-evaluates job conditions. Because
always() remains true, this job may continue when the
operator expected the run to stop. A safe design separates read-only
evidence collection from explicit compensation, and compensation has
target identity, cancellation/failure predicates and idempotency
guards.
8. Failure: global serialization creates an operational traffic jam
A group such as deploy shared by every service and
environment can make unrelated deployments wait. With
queue: max, the effect may look like a deadlock even
though GitHub is correctly serializing a huge queue. Inspect the
resolved group key and waiting timestamps before blaming runner
capacity.
Repair by expressing the smallest real exclusion domain: for example
deploy-${{ inputs.service }}-${{ inputs.environment }}.
9. Failure: reusable workflow cancels its caller
If the caller and called workflow both use a group derived from
github.workflow with
cancel-in-progress: true, the called workflow can share
the caller’s group. The correction is a distinct reusable-workflow
namespace or a caller-provided group contract.
10. Layered failure matrix
| Symptom | First layer to inspect | Evidence |
|---|---|---|
| deploy started before test | DAG/workflow syntax | needs edges + timestamps |
| new run canceled unrelated workflow | concurrency key | resolved group + both run IDs |
| job never reaches runner | needs/if/concurrency queue | job status, dependency results, pending state |
| job canceled but target changed | external side effect | provider/target audit + cancellation time |
| all deployments waiting | over-broad group/queue | group key + waiting jobs |
| cleanup keeps running after cancel | job/step condition | if evaluation + cancellation log |
11. Rerun discipline
Do not blindly rerun a canceled or failed deployment. First verify external target state and whether the prior attempt committed anything. If the issue is only a validation job, rerun that smallest safe scope. If a deployment side effect is ambiguous, reconcile the target before any retry.
Knowledge check
A deploy job starts before build finishes. What is the first thing to inspect?
The explicit needs graph, not runner speed or YAML order.
Why do group names production and Production collide?
Concurrency group names are case-insensitive within the repository.
What should you do with a validation error caused by queue: max plus cancel-in-progress: true?
Preserve the error, decide whether the invariant is queue-every-request or cancel-stale-work, then configure one coherent policy.
A canceled job changed the external target before cancellation. What evidence matters most?
The external target/provider audit correlated with run ID and cancellation timestamp, because the GitHub canceled conclusion does not prove rollback.
Why can a huge deployment queue look like a runner-capacity problem?
Jobs may be intentionally waiting on an over-broad concurrency group before runner assignment; inspect arbitration before runner capacity.
Official references and version notes
-
GitHub Docs — workflow syntax:
needs— explicit job dependencies, skip propagation and job-level conditions. -
GitHub Docs —
needscontext — direct-dependency results and outputs. -
GitHub Docs — control workflow/job concurrency
— concurrency groups, cancellation,
queuebehavior and expression contexts. - GitHub Docs — concurrency concepts — simultaneous execution, pending replacement and serialized queues.
- GitHub Docs — workflow cancellation reference — server condition re-evaluation, runner signals and forced termination.
- GitHub Docs — reusable workflow configuration — caller/called-workflow concurrency interaction.
Version-sensitive behavior was rechecked against current
GitHub-maintained documentation on 2026-09-09.
Current documentation states that needs contains only
direct dependencies and exposes result as
success, failure,
cancelled or skipped. A failed or
skipped dependency normally skips downstream jobs unless an
explicit job condition permits continuation. Concurrency groups
are repository-wide and case-insensitive. By default at most one
item may be running and one pending in a group; a newer pending
item replaces an older pending item. Current GitHub Actions also
supports queue: max to allow up to 100 pending items,
processed FIFO by time waiting on the group; this mode cannot be
combined with cancel-in-progress: true. Cancellation
re-evaluates job/step conditions, so always() can
keep work running during cancellation and must not be used
casually for privileged or irreversible operations. Mandatory labs
use ubuntu-24.04, permissions: {}, no
Marketplace action and no real credential.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.