Chapter 12Lesson 04~155 minutes

GitHub CLI, gh Authentication, Repository Operations, and Scripting Workflows: Diagnostics, Failure Modes, Security, and Performance

CLI failures are often context failures disguised as command failures. A script can be syntactically correct yet act on the wrong repository, wait forever for input, parse decorative output, or authenticate successfully with insufficient authorization. Diagnose the boundary before changing anything.

DiagnosticsExit codesLeast privilegeSupply-chain risk

Learning objectives

  • Use an evidence-first diagnostic sequence to locate failures in host, repository, authentication, authorization, rendering, or extension/alias layers.
  • Repair a wrong-repository script without relying on the current working directory.
  • Prevent automation from hanging on prompts or pagers and replace brittle decorated-text parsing with JSON.
  • Distinguish authentication failure from insufficient permission and choose a least-privilege correction.
  • Inventory aliases/extensions before trusting command behavior and treat third-party extensions as supply-chain dependencies.
Safety boundary: diagnostics are read-only unless a step is explicitly marked disposable. Never broaden a credential, expose a token, bypass repository policy, or delete valuable resources as a troubleshooting shortcut.

1. Diagnostic sequence: preserve context before correcting it

When a CLI job fails, do not immediately re-run it from another directory, switch accounts, broaden a token, or disable policy. Preserve enough evidence to explain the original failure.

  1. Preserve evidence: command line with secrets redacted, timestamp, exit code, stderr, gh --version, intended host/repository, and safe structured output.
  2. Identify scope: local Git clone/remote, GitHub host, repository owner/name, hosted object number/ref, workflow/job, and credential source category (stored credential versus environment token).
  3. Inspect context: active auth, repository identity, viewer permission, alias/extension inventory, prompt/pager config, and API/status response.
  4. Choose the least destructive correction: fix -R/GH_HOST, missing field/permission, or parser contract before changing repository policy or credential breadth.
  5. Verify independently: re-read the exact resource from GitHub and confirm that no unrelated repository/object changed.

2. Failure: the script runs against the wrong repository or host

A copied script contains gh issue list with no -R. An operator runs it inside a fork clone. The command succeeds—against the fork. This is more dangerous than a loud error because the output looks legitimate.

# Preserve evidence from the current context
git remote -v
gh repo view --json nameWithOwner,url,viewerPermission
gh auth status --active --hostname github.com

# Repair: make the intended resource an input
TARGET_REPO="OWNER/EXPECTED-REPO"
gh issue list -R "$TARGET_REPO" --state open --json number,title,url

If your script supports multiple hosts, make host an input too and fail closed when it is not in the supported set. Do not “repair” a wrong-target run by copying data from one repository to another until you first prove whether any mutation occurred.

3. Failure: automation hangs on a prompt, editor, browser, or pager

A command that was tested interactively can block forever in a scheduler because a required value was omitted and the CLI is waiting for a prompt. Long rendered output can also invoke a pager. CI reliability requires eliminating those hidden human dependencies.

# POSIX/Git Bash CI preamble
export GH_PROMPT_DISABLED=1
export GH_PAGER=cat
export NO_COLOR=1

# Supply inputs explicitly; do not open browser/editor
gh issue create -R OWNER/REPO \
  --title "explicit title" \
  --body "explicit body"
PowerShell: set $env:GH_PROMPT_DISABLED = "1" and $env:NO_COLOR = "1". Pager configuration is environment-specific; for CI, prefer machine-readable JSON output that does not depend on terminal paging, and configure the runner image deliberately.

If the job is already hung, capture process/log state before killing it. Then identify which missing input caused interaction rather than adding an arbitrary timeout and calling the problem solved.

4. Intentionally broken example: parsing decorated terminal text

This brittle script assumes the first whitespace-delimited token from terminal output is always the Issue number:

# BROKEN: human rendering is not a parser contract
first_issue="$(gh issue list -R "$REPO" | head -n 1 | awk '{print $1}')"
printf 'issue=%s\n' "$first_issue"

It can fail when output format, terminal width, color, fields, locale-like presentation, or command behavior changes. Preserve the raw output as evidence, then repair by asking for the field directly:

first_issue="$(gh issue list -R "$REPO" --state open --limit 1 \
  --json number --jq '.[0].number // empty')"

if [[ -z "$first_issue" ]]; then
  echo "no open issue" >&2
  exit 0
fi
printf 'issue=%s\n' "$first_issue"

The repair removes presentation parsing and defines the empty-list case. That is a reliability improvement, not merely shorter syntax.

5. Failure: token is too broad or too narrow

A token can authenticate successfully while lacking the permission needed for one endpoint. Conversely, broadening a token until a command works can turn a small diagnostic problem into a large blast-radius problem.

Observation Likely class Least-destructive next step
Exit code 4 / login required No usable authentication context. Fix host/credential injection; do not change repository permissions.
HTTP 401 / bad credentials Credential invalid/expired/revoked. Replace/re-authenticate; do not retry endlessly.
HTTP 403 / “resource not accessible…” Authenticated but token/role/policy blocks operation, or rate limit/policy may apply. Read response headers/body; inspect viewer permission and token/job permission; add only the required permission.
HTTP 404 on private resource Absent resource or intentionally hidden unauthorized resource. Verify owner/name and access out of band; do not infer existence.
Credential leak response: if a real token is printed in CI output, terminal capture, ticket, or screenshot, revoke/rotate it first. Deleting the log or editing Git history does not make the already exposed credential safe.

6. A subtle trap: JSON auth status is not an exit-code health gate

GitHub CLI currently documents that gh auth status exits 1 when an account has authentication problems. However, when --json is used, it normally exits zero despite authentication issues unless there is a fatal error. A script that says “JSON command returned zero, therefore authentication is healthy” is wrong.

# Health gate: use normal status semantics
if ! gh auth status --active --hostname github.com >/dev/null 2>&1; then
  echo "authentication unhealthy" >&2
  exit 65
fi

# JSON is useful for inspection/reporting, but do not infer health solely from its zero exit
gh auth status --json hosts

7. Failure: alias or extension changes the operator’s expectation

A runbook says “run our gh report command,” but the machine has a personal alias and two extensions. The correct diagnostic move is inventory, not immediate reinstall or token broadening.

gh --version
gh config list
gh alias list
gh extension list
type -a gh 2>/dev/null || true

Core commands cannot be overridden by extensions, but extension commands execute third-party code and aliases can invoke shell expressions. In a controlled runner image, maintain an allowlist of required extensions/versions or, preferably, avoid them for critical governance if core commands/API already suffice.

8. Rate limits and performance are causal constraints, not generic warnings

A report that performs one API call per Issue can be slow and consume unnecessary rate limit. Ask first-class commands for all fields you need in one structured request when possible, use API pagination rather than repeated page-number loops, and cache only reads whose staleness is acceptable.

On rate-limit responses, respect GitHub’s documented headers and retry timing. Do not parallelize mutations merely to “go faster.” Throughput is not reliability if you lose ordering, idempotence, or auditability.

9. Controlled failure rehearsal

Use your Chapter 12 disposable repository. First run the report from its clone and capture the successful repository identity. Then cd into a different Git clone and rerun a deliberately broken version that omits -R. Observe the wrong nameWithOwner without mutating anything. Repair it by restoring explicit REPO input and re-run. This demonstrates a real failure mode with read-only evidence and zero cleanup risk.

10. Lesson summary

CLI troubleshooting begins by preserving the original execution context. Wrong repository/host, prompt/pager hangs, decorative-text parsing, insufficient permissions, leaked credentials, and unreviewed alias/extension behavior require different corrections. Broader credentials, policy bypass, repeated retries, or destructive cleanup are not generic fixes.

Knowledge check

A command succeeds but returned data from the fork rather than upstream. Which diagnostic step catches this earliest?

A CI job waits forever after a CLI upgrade. What should you inspect before adding a longer timeout?

An API call returns 403. Should the script immediately request an admin token?

A real token appears in a debug log. What is the first remediation action?

Why inspect gh alias list and gh extension list during strange behavior?

Next lesson

Next: Checkpoint Lab — GitHub CLI, gh Authentication, Repository Operations, and Scripting Workflows

Further reading — current official GitHub sources

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.