Chapter 27Lesson 04~190 minutes

GitHub Apps, OAuth Apps, Webhooks, Checks API, and Event-Driven Integrations: Diagnostics, Failure Modes, Security, and Performance

Diagnose forged deliveries, duplicate mutations, excessive App privilege, leaked signing material, token expiry, out-of-order events, and failed-delivery recovery without hiding the cause.

DiagnosticsReplayLeast privilegeOrderingRecovery

Learning objectives

  • Apply an evidence-first diagnostic sequence for integration failures.
  • Repair unsigned/spoofed and duplicate delivery handling.
  • Reduce over-broad App permissions and respond correctly to key/secret exposure.
  • Classify 401/403/404/422 and token-expiry failures.
  • Engineer fast ingress, recovery/redelivery, and out-of-order state reconciliation.

1. Diagnostic sequence: preserve evidence before changing permissions or retrying

Use one sequence across webhook, App-auth, and check failures: preserve evidence → identify account/repository/installation/event/ref scope → inspect App permissions, repository selection, delivery log, response status, API request ID, and current resource state → choose the least destructive correction → independently verify. This prevents “fixes” such as granting organization-wide write or deleting a delivery record merely because an integration is failing.

2. Failure: receiver trusts unsigned or spoofed payloads

# BROKEN PATTERN — do not use
raw = request_body
payload = json.loads(raw)
process_issue(payload)              # mutation happens first
# no HMAC validation
# no X-GitHub-Delivery uniqueness check
return 200

This is not merely missing hardening. Anyone who can reach the endpoint can construct JSON that looks like a GitHub event and trigger the mutation. The repair changes ordering, not only syntax:

# CORRECT ORDER (pseudocode)
raw = read_exact_request_bytes()
verify_hmac_sha256(raw, signature_header)   # reject on mismatch
validate_event_and_action(headers, raw)
if delivery_id_already_committed(x_github_delivery):
    return 200                               # idempotent duplicate
persist_delivery_claim_atomically()
enqueue_or_reconcile_current_state()
return 202

In the Chapter 27 fixture, python send_fixture.py --bad-signature must return 401, must not insert a delivery row, and must not emit the derived-action log. If it does, preserve the request headers and code version, then stop the consumer before accepting more traffic.

3. Failure: redelivery runs the mutation twice

A naive in-memory seen=set() can still duplicate work after a restart or across multiple workers. The durable fix is an atomic uniqueness constraint keyed by X-GitHub-Delivery, with mutation/queue state recorded transactionally enough for your service's consistency model. GitHub uses the same delivery GUID when you request a redelivery, so duplicate suppression should survive intentional recovery.

Do not “fix” duplicates by discarding all redeliveries. If the first attempt failed before the durable derived action committed, your service needs a state machine such as received → validated → queued → applied so a redelivery can resume safely rather than either duplicating or losing work.

4. Failure: App requests organization-wide write it does not need

Symptoms may be subtle: installation works, but a compromise now spans repositories or organization resources unrelated to the service. Treat permission expansion like code deployment. Record the business requirement, endpoint permission needed, installation scope, reviewer, and rollback path. Prefer selected repositories and read permissions until a concrete write endpoint requires more.

Security-sensitive change: Adding App permissions can require installer/organization approval and broadens blast radius. Never use “grant everything, then debug downward” as a troubleshooting technique.

5. Failure: private key or webhook secret is committed/logged

The correct response mirrors Chapter 24: revoke/rotate first. For an exposed App private key, create/activate a replacement through the official App settings, update the service secret store, then delete/revoke the old key and inspect suspicious App/JWT/installation activity. For a leaked webhook secret, rotate the secret and update the receiver. Only after credential invalidation should you clean source/history/log exposure.

Never print a JWT or installation access token for diagnostics. Log token type, installation ID, permission set, repository IDs, issue time/expiry time, API status, and request ID—never the bearer value.

6. Failure: integration assumes event order or immediate consistency

GitHub documents that webhook deliveries can arrive out of event order and may be delayed. A service that sees issue.edited and then a late issue.opened must not overwrite newer derived state with older data. Use event timestamps as evidence and, for consequential mutations, fetch the authoritative current resource before applying the transition.

Similarly, a webhook can arrive before an eventually consistent downstream search/index surface reflects the new object. Use the direct resource endpoint when possible and bounded retry only for reads with a documented reason. Do not turn temporary inconsistency into repeated creates.

7. Failure: expired token and blind retry

Installation tokens expire after one hour. If an API call returns 401 and the token is near/past expires_at, obtain a fresh installation token and retry a safe read once. A 403 should send you to authorization diagnosis, not a refresh loop. A 404 on a private resource can also be permission-related. Record safe response metadata and GitHub request IDs.

Signal Likely class Least-destructive next step
401 Bad Credentials Expired/revoked/invalid token Refresh/re-authenticate once; verify installation still exists; do not log token.
403 Forbidden Permission, repository selection, policy, rate/abuse control Inspect permission/install scope and rate headers; honor Retry-After when present.
404 Not Found Absent resource or concealed inaccessible private resource Verify owner/repo/host and installation access before declaring absent.
422 Validation failed Bad payload/state/endpoint constraint Preserve response body, fix request; do not retry unchanged mutation.

8. Failure: webhook endpoint is slow or unavailable

GitHub.com expects a 2XX response within ten seconds. A long synchronous analyzer can turn successful business logic into a failed delivery. Authenticate/dedupe quickly, enqueue durable work, respond, then perform slower API work. GitHub does not automatically redeliver failed webhooks; build monitoring that inspects failed deliveries and triggers controlled redelivery while the current three-day delivery record is available.

Because a redelivery itself is a mutation trigger, the redelivery worker must be rate-controlled, permissioned, and observable. Never run an unbounded “redeliver all failures” loop.

9. Reliability and performance are causally linked to event volume

Subscribe only to needed events/actions. Queue depth, processing latency, webhook failure rate, signature failures, duplicate rate, API 401/403/429 counts, installation-token refresh failures, and check-run update latency are useful service-level signals. A sudden event surge may cause delivery throttling or API secondary-limit pressure. Scale workers separately from ingress acknowledgements and apply backpressure before expanding API concurrency.

10. Operations deliberately excluded from the mandatory lab

Repository transfer/deletion, force-updating branches/tags, policy bypass, generating real App private keys/tokens, package deletion, runner registration, real secret handling, organization-wide privilege changes, and history rewriting are not required. When a real App key must be created for the optional extension, do it through the official flow, store it outside source/logs, scope the App to the disposable repository, and revoke it during cleanup.

Knowledge check

A webhook payload looks exactly like an issues event but has no signature. What is the correct interpretation?

Why is an in-memory delivery-ID set insufficient for production deduplication?

The App receives 403 after a new endpoint is added. What should you inspect before increasing permissions?

A private key appears in CI logs. What comes before deleting the log/history?

Two webhook events arrive in reverse business order. How should consequential automation determine current truth?

Summary

Integration failures are usually trust, identity, delivery, or state-machine failures—not just HTTP errors. Verify before parse, persist idempotency, treat permissions as production policy, rotate leaked credentials first, reconcile current state instead of trusting arrival order, distinguish 401/403/404/422, and engineer fast acknowledged ingress plus observable recovery.

Next lesson

Checkpoint Lab — GitHub Apps, OAuth Apps, Webhooks, Checks API, and Event-Driven Integrations

Official references

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.