REST API, GraphQL API, gh api, Pagination, Rate Limits, and Automation Clients: Diagnostics, Failure Modes, Security, and Performance
Diagnose incomplete pagination, duplicate mutations, API-version drift, misleading 404/403 responses, secondary rate limits, over-scoped credentials, excessive GraphQL selection, and unsafe retries with evidence-first troubleshooting.
Learning objectives
- Preserve request/status/error/page/rate evidence before changing anything.
- Repair first-page-only inventory without hiding the original partial result.
- Prevent duplicate creates after timeouts or retries.
- Separate version mismatch, authorization, primary exhaustion, and secondary throttling.
- Reduce security/performance risk from excess scope, fields, concurrency, and logging.
1. Diagnostic sequence: preserve → scope → inspect → correct → verify
- Preserve: method/endpoint or GraphQL operation, safe variables, timestamp, GitHub request ID, HTTP status, sanitized response, REST API version, rate headers, page/cursor state.
- Scope: github.com/GHES host, authenticated actor/App, owner/repository, token repository selection/permissions.
- Inspect: endpoint docs/schema, version/media headers, pagination, rate evidence, policy/audit data where available.
- Correct: fix traversal/header/field selection/concurrency or exact permission before broadening access.
- Verify: independently re-read IDs/counts/state—not merely exit code.
2. Failure: HTTP 200, business result incomplete
# BROKEN: valid response, incomplete collection.
status, data, meta = request("GET",
f"https://api.github.com/repos/{repo}/issues?state=all&per_page=2")
for issue in data:
process(issue)
# BUG: meta["link"] may contain rel="next".
Nothing has to throw. Preserve the Link header and
repair traversal until no next relation remains. Verify against a
known fixture count or a second documented surface.
3. Failure: a retry “recovers” transport and corrupts state
A create reaches GitHub and succeeds, but the client loses the response. A generic retry sends the same POST and creates a duplicate. A timeout is ambiguous—not proof of failure.
Attempt 1: POST create marker -> response lost
Server: issue #6 created
Retry layer: POST same create again
Server: issue #7 created
Transport: recovered
Business state: wrong
Repair with a durable desired-state key and read-before-create. Where an endpoint documents idempotency/precondition semantics, use them; otherwise do not invent them.
4. Failure: version/media mismatch
set +e
gh api -i -H "X-GitHub-Api-Version: 1900-01-01" "repos/$REPO" >broken.out 2>broken.err
rc=$?
set -e
printf 'exit_code=%s
' "$rc"
sed -n '1,30p' broken.err
# Repair with a supported explicit version.
gh api -i -H "X-GitHub-Api-Version: 2026-03-10" "repos/$REPO" --silent
Do not “repair” an unsupported version by deleting the header and accepting the implicit legacy default. Treat version selection as an owned dependency with migration tests.
5. Failure: throttling misclassified as permission failure
| Evidence | Interpretation | Correction |
|---|---|---|
403/429 + x-ratelimit-remaining: 0 |
Primary budget exhausted. |
Wait until x-ratelimit-reset; reduce request
volume.
|
403/429 + rate message + Retry-After |
Secondary throttling. | Honor Retry-After; reduce concurrency/mutation pressure; bounded backoff. |
| 404 private resource | Absence or insufficient visibility. | Verify actor, host, repo selection, documented permission. |
GraphQL HTTP 200 + errors |
Partial/query/rate/resource failure possible. | Inspect errors before consuming any returned data. |
Do not intentionally hammer GitHub to reproduce secondary limits. Use official limits/fixtures and preserve real production headers when an incident happens.
6. Failure: over-scoped identity and sensitive field selection
An inventory that needs repository names should not run under organization-admin scope or select collaborator emails/security data. Broader token scope and broader query scope both enlarge blast radius. If access fails, identify the exact documented permission first rather than granting administrator access.
7. GraphQL query size, cost, and partial results
GitHub currently requires first/last
between 1 and 100 on connections and enforces point budgets,
node/resource limits, and timeouts. Deeply nested queries can be
expensive or partial. Repair by selecting fewer fields, reducing
page sizes/depth, splitting large queries, and filtering earlier.
Use response headers for ongoing rate observability where possible.
Querying rateLimit can provide cost with the response,
but polling rateLimit itself needlessly consumes resources.
8. Intentionally broken examples: interpret before repair
# A: first page only
gh api -i -H "X-GitHub-Api-Version: 2026-03-10" "repos/$REPO/issues?state=all&per_page=2"
# A repair
gh api --paginate -H "X-GitHub-Api-Version: 2026-03-10" "repos/$REPO/issues?state=all&per_page=2" --jq '.[] | .number'
# B: GraphQL schema error
set +e
gh api graphql -f query='query { viewer { definitelyNotAField } }' >graphql-broken.json 2>graphql-broken.err
rc=$?
set -e
printf 'graphql_exit=%s
' "$rc"
sed -n '1,40p' graphql-broken.err
Preserve the original failure output. The learning goal is not a clean terminal; it is a causal explanation of why the first state was unsafe/incomplete.
9. API incident runbook
- Freeze mutative jobs if duplicate/destructive behavior is possible.
- Capture request IDs, timestamps, operation, sanitized variables, status/errors, pages/cursors, actor identity, and version.
- Reduce concurrency and honor Retry-After/reset before changing credentials/policy.
- For ambiguous creates, read desired state before retrying.
- If a token leaked, revoke/rotate first and assess accessible scope.
- After repair, replay read-only verification and compare IDs/counts/policy state.
Knowledge check
A script exits 0 after exactly 100 repositories. Why should that trigger suspicion?
One hundred is a common page maximum. Exit status proves transport success, not complete traversal.
A create timed out. What evidence should precede another create?
A read/search showing the durable desired-state marker is absent, or documented endpoint idempotency/precondition evidence.
Why is a secondary limit not fixed by a broader token?
Secondary limits protect service reliability/request patterns. Authorization scope is a separate concern; broader access may only increase risk.
What is wrong with fixing a 404 by granting admin immediately?
The cause could be wrong host/path or a narrow missing permission. Diagnose exact ownership/identity/repository selection first.
GraphQL returns partial data plus errors. Can an enforcement job mutate from the partial set?
Not safely by default. The job must prove the input set is complete enough for the policy before any mutation.
Summary
API failures often preserve valid transport while violating business correctness: partial pages, ambiguous creates, concealed private resources, throttling, or partial GraphQL data. Diagnose from evidence, make the least-privilege correction, and verify hosted state independently.
Official references
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.