Chapter 31Lesson 03~220 minutes

GitHub CLI, REST/GraphQL APIs, Workflow Dispatch, and Automation Control: Configuration, Design Patterns, and Trade-Offs

Choose gh CLI, REST, GraphQL, resource identifiers, polling/webhooks, rerun scope, and automation identity according to reliability, privilege, auditability, and failure-recovery requirements.

Design choicesGitHub AppPaginationIdempotencyWebhooks

Learning objectives

  • Choose gh CLI, REST or GraphQL according to the operation and evidence contract rather than tool preference.
  • Choose workflow path/name versus workflow ID with explicit lifecycle assumptions.
  • Choose bounded polling versus webhooks according to scale, latency and reliability needs.
  • Choose rerun-all, failed-job rerun or specific-job rerun without hiding the original attempt.
  • Choose human authentication, fine-grained PAT or GitHub App identity according to lifetime and privilege boundaries.

1. Design starts with a contract, not a favorite client

Once control becomes reusable platform automation, the interface choice affects correctness. The same logical operation can be exposed through a friendly gh command, a versioned REST endpoint or a GraphQL query. Ask what resource identity you need, whether the operation mutates state, how pagination and retries behave, and what evidence an operator must retain.

A golden control contract should accept explicit repository/workflow/run identifiers, declare the minimum permissions it needs, return durable correlation IDs, and refuse ambiguous state. The underlying client can then change without weakening the contract.

2. gh CLI versus REST versus GraphQL

Choice Best fit Trade-off / evidence
gh workflow/run Human or simple script workflows Excellent UX; some commands hide HTTP details, so record IDs/JSON fields explicitly.
gh api + REST Exact Actions control endpoints and status/header semantics Verbose but explicit method/version/resource/response contract.
gh api graphql Selective connected reads and cursor traversal Flexible schema; must handle GraphQL errors, point limits and cursors.
Direct SDK/client Long-lived service integration Can centralize retries/auth/telemetry; must still obey API version and exact-resource guards.

The common production pattern is mixed: REST for dispatch/cancel/rerun, GraphQL or REST for shaped inventories, and gh for operator tooling. Do not force GraphQL into a mutation merely because the service already uses GraphQL elsewhere.

3. Workflow file name versus workflow ID

A workflow path is readable in code review and accepted by many API/CLI operations. A workflow numeric ID is unambiguous for the current resource and survives a display-name change. Neither is immortal: paths can be renamed, while delete/recreate can yield a new ID. Resolve path → ID at controller startup and record both.

For versioned platform automation, the resolved workflow ID belongs in the execution ledger, while the repository path belongs in configuration and review. If the ID no longer resolves to the expected path/state, stop instead of following the new resource implicitly.

4. Polling versus webhook

Polling is easy for a single disposable run: GET one exact run every few seconds, stop at a deadline, honor rate-limit signals, and never issue mutations inside the poll loop. At platform scale, polling every repository/run wastes budget and increases detection latency variance.

As of 2026-09-10, gh run watch is convenient for interactive observation but does not support fine-grained PAT authentication because the required checks-read capability is not currently grantable that way. The mandatory controller therefore polls the exact REST run resource, which also makes its identity, interval and deadline explicit.

Webhooks invert the model: GitHub notifies your service about workflow-run lifecycle events, and the service verifies repository/run identity before updating its state machine. Webhooks add delivery validation, deduplication, retry and endpoint-availability responsibilities. A robust service often uses webhooks for normal operation plus occasional exact GET reconciliation.

5. Rerun all versus failed jobs versus one job

Rerun scope Use when Risk / proof
Entire run You need a full same-source reproduction Repeats successful side effects unless jobs are idempotent; preserve attempt 1.
Failed jobs The graph supports retrying failed jobs/dependents Smaller blast radius; still same run ID/source with new attempt.
Specific job databaseId A single job and dependents are the intended diagnostic scope Requires exact job database ID, not browser URL number; dependencies may rerun.
New run Source/workflow/input must change Correct proof for a code/config fix; new run ID rather than a rerun.

Rerun is a recovery mechanism, not a substitute for diagnosis. If an external deployment step may have partially completed, inspect provider state before rerunning any job that can repeat side effects.

6. Human token versus GitHub App/automation identity

A human gh session is ideal for a one-off lab and interactive administration because authorization follows an identifiable person. A fine-grained PAT can narrow repository access but remains a user credential with rotation/offboarding concerns. A GitHub App installation token is preferred for long-lived organization automation because installation scope, permissions and expiry fit machine identity better.

Inside GitHub Actions, GITHUB_TOKEN is short-lived and repository-scoped. It is often the right choice for same-repository API work, provided the workflow declares only the permissions required. It is not automatically a replacement for a cross-repository control-plane GitHub App.

7. Permission map: separate read inventory from control mutations

Operation Fine-grained repository permission Mutation?
List/get workflows Actions: read No
List/get runs/log metadata Actions: read No
Dispatch workflow Actions: write Yes
Cancel/force-cancel run Actions: write Yes
Rerun workflow/failed jobs/job Actions: write Yes
Read public metadata unauthenticated Endpoint-dependent No; lower rate budget

Do not grant repository administration merely because the controller needs Actions write. Permission denied is useful evidence: identify the missing capability and decide whether the operation is truly authorized before expanding scope.

8. Mutation design: intent key + exact resource + reconcile-before-retry

A safe dispatch controller generates a request ID before POST, includes it in a validated workflow input/run name, persists the intent locally, then records the returned run ID. This gives the controller a deterministic reconciliation key if the transport fails. Blindly reissuing a POST can create duplicate runs or duplicate external side effects.

Cancel and rerun need the same discipline. Read the exact run immediately before mutation, compare workflow/request/source guards, write the intent to the ledger, issue one mutation, then read final state. For force-cancel, require a separate escalation path because GitHub documents it as a fallback when ordinary cancellation does not respond.

9. Pagination and rate budget are part of the algorithm

For REST, follow Link headers or let gh api --paginate traverse pages. For GraphQL, use first/last within current bounds and cursor pageInfo. If the algorithm says “enumerate every run,” page traversal is a correctness requirement; if it says “inspect this run ID,” a direct resource GET is both cheaper and safer.

Rate limits should change the scheduler, not the authorization model. Slow down, use conditional requests/webhooks where appropriate, and honor retry-after. Never respond to rate pressure by caching a stale “latest run” pointer and mutating it later.

10. Human-readable command output versus machine control ledger

Terminal output helps operators understand what is happening, but a controller should persist machine-readable fields: repository, workflow ID/path, run ID/attempt, event, SHA, request ID, status/conclusion, HTTP operation/status, API version and timestamps. That record is more reliable than scraping colored CLI text.

Treat the ledger as evidence, not as a secret dump. Never store access tokens, Authorization headers, OIDC tokens or whole event contexts. When external systems are involved, store approved correlation IDs and final health state separately from the GitHub run.

11. Worked scenario: choose an architecture

Requirement Choice Reason / observable proof
One developer dispatches a lab and watches it gh workflow run + exact URL/ID read-back Fast UX; verify resulting run metadata explicitly.
Platform service dispatches hundreds of repos GitHub App + REST dispatch + webhook/reconciliation Machine identity, exact returned run IDs, less polling.
Dashboard needs selected workflow/run fields GraphQL cursor query Fetch only required connected fields; record cursors/errors/rate cost.
Incident reruns one failed job REST/gh rerun by job databaseId Narrow scope; preserve old attempt and external side-effect state.
Workflow code was fixed after failure New workflow run, not rerun Rerun would use original source identity and cannot prove the fix.

12. Plan/platform boundaries

The mandatory chapter path requires only a disposable repository with GitHub Actions and GitHub CLI/API access. Enterprise audit streaming, organization policy, SSO-enforced app approval and centralized webhook infrastructure can strengthen a platform controller, but they are optional extensions. GitHub Enterprise Server may expose different API versions/features; document that platform explicitly before copying GitHub.com examples.

13. Lesson summary

Good control-plane design chooses the narrowest interface and identity that can prove the intended transition. Resolve stable-enough resource IDs, separate read and write permissions, prefer event-driven completion at scale, make retries reconcilable, and treat rerun scope as a risk decision rather than a convenience flag.

Next lesson

GitHub CLI, REST/GraphQL APIs, Workflow Dispatch, and Automation Control: Diagnostics, Failure Modes, and Production Practices

Continue with the next lesson to build on the current concepts, evidence, security boundaries, and operational practices.

Knowledge check

Why resolve workflow path to ID at controller startup?

When is polling appropriate?

Why prefer a GitHub App for long-lived organization automation?

A workflow file changed after a failure. Should you rerun the old run?

Does Actions write permission imply repository administration permission is needed?

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.