Chapter 31Lesson 04~235 minutes

GitHub CLI, REST/GraphQL APIs, Workflow Dispatch, and Automation Control: Diagnostics, Failure Modes, and Production Practices

Diagnose ambiguous dispatches, wrong-run mutations, blind POST retries, pagination gaps, rate limits, token leakage, and false assumptions about asynchronous acceptance without destroying first-failure evidence.

DiagnosticsRate limitsWrong-run guardPOST retryRecovery

Learning objectives

  • Diagnose wrong-run selection, blind mutation retries and incomplete pagination from preserved API/run evidence.
  • Separate authentication/authorization failures from resource identity, API version, rate-limit and workflow-execution failures.
  • Prevent Authorization/header leakage while still retaining useful request/response metadata.
  • Interpret dispatch/cancel/rerun response status separately from the final workflow conclusion.
  • Apply the least-destructive correction and rerun only the smallest equivalent scope.

1. Evidence-first diagnostic sequence

When a controller misbehaves, resist the urge to “try the command again.” Preserve the controller ledger and first run/attempt first. Then confirm repository/workflow/run identity, API version, request method/body (without credentials), response status, run event/ref/SHA, jobs/runner, artifacts and any external side effects. Only then choose the smallest correction.

  1. Preserve controller intent/request ID, REST status and exact run/attempt evidence.
  2. Confirm repository, workflow path/ID and run ID still match the intended resource.
  3. Confirm event/ref/SHA, actor, evaluated permissions and workflow inputs.
  4. Inspect job graph/queue/runner before blaming API control.
  5. Inspect failing action/script/network and retained artifacts/log links.
  6. Inspect environment/deployment/external target separately if the run had side effects.
  7. Apply one least-destructive correction and rerun only the exact required scope.

2. BROKEN — “latest run” becomes the mutation target

# BROKEN — race-prone and has no ownership proof.
gh workflow run control-target.yml -f request_id=lab-a -f behavior=wait
RUN_ID=$(gh run list --workflow control-target.yml --limit 1 --json databaseId --jq '.[0].databaseId')
gh run cancel "$RUN_ID"

The race is between dispatch and list: another actor or automation can create a newer run. The fix is not “sleep 2 first.” Use the current REST dispatch response workflow_run_id, record it, then read back workflow/event/request guards before cancellation. If you must reconcile an ambiguous old-style dispatch, use the request ID plus workflow/ref/time bounds—not list position.

3. BROKEN — retrying POST blindly after a timeout

# BROKEN pseudo-logic
until gh api -X POST repos/$GH_REPO/actions/workflows/$WORKFLOW/dispatches --input payload.json; do
  sleep 1
  # Repeats the mutation without proving the previous attempt had no side effect.
done

A client timeout does not prove GitHub rejected the request. The first POST may have created a run whose response never reached the client. Generate a request ID before POST, persist it, and reconcile matching workflow runs before sending another dispatch. For cancel/rerun, re-read exact run state before deciding whether another mutation is valid.

4. BROKEN — page 1 is mistaken for the complete inventory

REST collections are paginated; GraphQL connections are cursor-based. Missing a run on page 2 can produce false “not found” logic and trigger an unnecessary dispatch. Conversely, using a full inventory when you already have an exact run ID wastes rate budget and enlarges the failure surface.

# BROKEN for a complete inventory: only the default first page.
gh api repos/{owner}/{repo}/actions/runs

# REPAIR: explicit traversal when completeness is actually required.
gh api --paginate 'repos/{owner}/{repo}/actions/runs?per_page=100'   --jq '.workflow_runs[].id'

5. BROKEN — debugging HTTP by leaking credentials

# DO NOT DO THIS with a real token:
# curl -v -H "Authorization: Bearer $TOKEN" https://api.github.com/...

# Safer: let gh manage authentication and select non-sensitive response data.
gh api -i   -H 'Accept: application/vnd.github+json'   -H 'X-GitHub-Api-Version: 2026-03-10'   rate_limit

Verbose HTTP traces can copy bearer credentials into terminal history, CI logs or support bundles. Preserve endpoint, method, API version, safe response headers/status and request correlation—not the Authorization header. If a credential may have leaked, revoke/rotate it; masking or deleting one log is not sufficient containment.

6. BROKEN — a broad human PAT becomes the platform control identity

A classic PAT with broad repo scope may make a failing command “work,” but it hides which repository capability was actually required and couples automation to a person. For a lab, an authenticated human session is acceptable. For long-lived platform automation, move to a GitHub App or fine-grained credential and grant only the documented repository permissions.

A 403 is therefore not an instruction to grant admin. Inspect endpoint docs and current auth type. Separate “token is valid” from “token has Actions write on this repository” and from “organization policy allows this automation.”

7. BROKEN — accepted control request is reported as workflow success

Current dispatch returns 200 with the new run ID, cancel returns 202 Accepted, and rerun returns 201 Created. Those are control-plane responses. The run can later fail, remain queued, be cancelled, or encounter an environment/provider failure. Always poll/webhook the exact run to terminal state and inspect conclusion plus required downstream/external health.

State separation: dispatch creation, run queueing, job start, check conclusion, artifact publication, deployment approval, provider rollout and external health are separate states. Do not turn one HTTP status into a universal “success.”

8. Specific-job rerun fails with 404: inspect the job identifier

GitHub CLI documents a subtle identity trap: the number shown after /jobs/ in a browser URL is not necessarily the database job ID expected by gh run rerun --job. Query gh run view RUN_ID --json jobs --jq ".jobs[] | {name,databaseId}" and use databaseId. A 404 here is usually resource identity, not runner capacity or application failure.

9. Interpret API failures by layer

Evidence Likely layer Next check
401 Authentication Expired/invalid credential; do not print it.
403 Authorization/policy/rate limit Endpoint permission, org policy, x-ratelimit headers.
404 Repository/workflow/run/job identity or visibility Exact owner/repo/path/ID/databaseId and access.
409 on cancel Current run state conflict Read run status/conclusion; do not force-cancel automatically.
422 Request schema/ref/input validation Workflow_dispatch support, ref, declared input keys/types.
429 / retry-after Rate limiting Honor server timing; reduce polling/concurrency.
5xx / transport timeout after POST Ambiguous mutation Reconcile request/run state before retry.

10. Force-cancel is an escalation, not the default fix

GitHub provides a force-cancel endpoint for runs that do not respond to ordinary cancellation, including cases where workflow conditions keep work alive. It still requires Actions write and exact run identity. Preserve logs/external state, attempt ordinary cancellation, wait a bounded period, then escalate only for the same proven lab resource. Do not teach force-cancel as the normal first command.

11. Rate-limit and GraphQL failures are control-plane failures

REST primary/secondary limits and GraphQL point/node limits can make a controller fail even when workflows are healthy. Honor retry-after and reset headers. For GraphQL, inspect both HTTP status and payload errors; reduce query depth/fields or paginate instead of assuming a partial payload is complete.

The least-destructive correction is usually to improve the controller scheduler/query, not to rerun workflows. Re-running application jobs because the API observer hit a rate limit only consumes more capacity and obscures the real cause.

12. Recovery rules

  • Preserve the original run/attempt and controller ledger before any rerun/cancel/delete.
  • Resolve exact resource identity again immediately before a write operation.
  • Never broaden token scope until endpoint documentation and authorization intent are reviewed.
  • Reconcile ambiguous POST outcomes before retrying; request IDs are operational safety keys.
  • Use a new run for source/input fixes; use rerun for same-source reproduction/diagnostics.
  • Verify external side effects separately before repeating any job that can mutate a provider or deployment target.

13. Lesson summary

Most API-control incidents are not “GitHub Actions bugs”; they are identity, transport, pagination, permission, rate-limit or asynchronous-state mistakes. Preserve evidence, identify the causal layer, and make the controller narrower rather than more privileged.

Next lesson

Checkpoint Lab — GitHub CLI, REST/GraphQL APIs, Workflow Dispatch, and Automation Control

Continue with the next lesson to build on the current concepts, evidence, security boundaries, and operational practices.

Knowledge check

A cancel request returned 202. Is the run definitely stopped?

Why is sleeping before gh run list --limit 1 not a fix for the latest-run race?

A specific-job rerun returns 404 using the browser URL job number. What should you inspect?

A POST timed out. What is the least-destructive next step?

A GraphQL response is HTTP 200 but contains errors. Is it complete success?

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.