Chapter 02Lesson 04~115 minutes

Account Security, Two-Factor Authentication, SSH Keys, Tokens, and Credential Hygiene: Diagnostics, Failure Modes, Security, and Performance

Diagnose authentication failures without leaking credentials or widening permissions, and distinguish invalid credentials, insufficient authorization, wrong identity, SSO policy, and misleading command exit codes.

DiagnosticsSSOIncident responseFailure analysis

Learning objectives

  • Use a preserve-evidence → scope → authenticate → authorize → policy → verify diagnostic sequence.
  • Interpret SSH, GitHub CLI, and REST failure signals without printing or regenerating credentials prematurely.
  • Diagnose wrong SSH identity/host alias, insufficient token permission, missing repository selection, and SSO authorization.
  • Respond to credential exposure by revoking/rotating before attempting history or log cleanup.
  • Recognize the documented successful SSH-test exit code and avoid false failure automation.
  • Distinguish authentication/authorization failures from rate-limit, network, repository-policy, and billing symptoms.
Availability: The mandatory diagnostics use local/Free-compatible commands and synthetic failure fixtures. Enterprise SSO examples are optional fixtures unless the learner already belongs to an SSO-enabled Enterprise Cloud organization.

1. Diagnostic sequence: preserve evidence before changing credentials

  1. Preserve non-secret evidence. Record timestamp, host, repository coordinate, command/endpoint, HTTP status, request ID, relevant headers, and exact error text. Never record the token/private key.
  2. Identify scope. Which account, organization, repository, ref, workflow, or API resource is failing?
  3. Prove authentication. Which principal does the key/token/gh session resolve to?
  4. Prove authorization. Does the principal have the required repository role and does the credential have the endpoint permission/scope?
  5. Check policy. SSO, PAT policy, IP allow list, repository rules, event context, or enterprise policy may narrow access.
  6. Apply the least-destructive correction. Fix identity/permission/authorization only; do not switch immediately to admin/broad scopes.
  7. Verify independently. Repeat the smallest read-only/functional request and record the expected principal/resource/result.
Do not “debug” by printing secrets. A token value rarely explains why authorization failed, but exposing it creates a second incident.

2. API status codes: classify before retrying

Signal Likely class Next evidence
401 Unauthorized Invalid/missing/expired authentication credential. Check which credential source the client used; reauthenticate safely; do not repeatedly hammer invalid credentials.
403 Forbidden Authenticated but blocked by permission/policy/rate limit, or temporary invalid-login protection. Inspect response message/headers, endpoint required permissions, SSO/policy, and rate-limit headers.
404 Not Found on private resource Resource may not exist or caller may not be authorized to know it exists. Verify owner/repo/path and token repository selection/permission with a known-authorized identity.
422 Unprocessable Entity Request shape/state validation failure. Inspect endpoint schema and response body; changing credentials is usually irrelevant.

GitHub’s REST authentication documentation explicitly notes that insufficient permissions can yield 403 or 404. Therefore “404 proves the repo name is wrong” is not a safe conclusion for private resources.

gh api -i /rate_limit

Use headers only to confirm rate-limit causality when a 403 suggests it. Rate limits are not a generic explanation for every authorization failure.

3. Failure mode: token owner/repository/permission does not match the request

Suppose a fine-grained PAT is valid and belongs to learner-example, but it was created for resource owner learner-example and selected only repository auth-lab-a. A request to private auth-lab-b should be denied even if the user themselves can access both repositories through the web UI.

A second failure is endpoint permission. GitHub’s REST fine-grained-permission reference publishes the required permission and the API may include an X-Accepted-GitHub-Permissions response header. Use that evidence to request one missing read permission instead of granting broad write/admin authority.

Authorization equation: effective access = user’s underlying access ∩ token resource owner/repository selection ∩ token permission ∩ organization/enterprise policy.

4. Failure mode: SSH reaches the right host but the wrong GitHub account

Multiple GitHub accounts often use SSH host aliases. A configuration can be syntactically valid and still offer the key for the wrong account. The first question is not “does SSH work?” but “which principal did it authenticate?”

Inspect resolved OpenSSH configuration without sending a credential

ssh -G github-work | grep -E '^(hostname|user|identityfile) '

Then test the alias with ssh -T github-work or force an exact key using -i PATH_TO_KEY -o IdentitiesOnly=yes. The GitHub greeting names the resolved account. If that account is wrong, fix the host alias/identity mapping. Do not add the same private key to additional accounts as a shortcut.

Privacy note: verbose SSH output can contain usernames, paths, host details, and key fingerprints. Capture only the lines required for diagnosis; never publish private filesystem paths or private key contents.

5. Intentionally broken example: treating ssh -T as a normal success/fail command

This shell check looks reasonable but is wrong for GitHub’s SSH test:

if ssh -T git@github.com; then
  echo "GitHub authentication OK"
else
  echo "GitHub authentication FAILED"
fi

GitHub intentionally does not provide shell access and documents that ssh -T git@github.com exits with code 1 even after printing a successful authentication greeting. Therefore the script enters the else branch on a healthy connection.

Repair: use ssh -T interactively to confirm the username/message, and use an actual Git operation such as git ls-remote for automated functional transport verification. Do not redefine all exit code 1 results as success; that would hide real SSH failures.

6. Failure mode: credential copied into a URL, history, log, screenshot, or source file

If a bearer token or private key has been exposed, treat the credential as compromised even if the repository/file was later deleted. The first technical action is revoke or rotate. The second is to identify where the value propagated. Only after invalidation should you clean remote URLs, shell history, logs, artifacts, caches, commits, or mirrors.

Example: token discovered in a Git remote URL

git remote -v
git remote set-url origin https://github.com/OWNER/REPO.git
git remote -v

These commands remove the credential from the current repository configuration, but they do not revoke it and they do not erase copies from shell history, logs, backups, or screenshots. Revocation/rotation happens first in GitHub. If the secret was committed, use the dedicated history-remediation guidance later; rewriting history alone cannot make a leaked credential safe.

Incident ordering: revoke/rotate → preserve/assess evidence → remove live references → remediate history/logs/caches → verify old credential fails and replacement has only required access.

7. Failure mode: 2FA and recovery factors are unavailable

Do not reproduce an account lockout as a lab. GitHub’s recovery documentation is explicit that if all 2FA credentials and recovery methods are lost, GitHub Support cannot restore the account. A real incident response therefore begins by inventorying the recovery methods still available—recovery codes, passkey, security key, verified device, or other documented recovery paths—without invalidating any remaining factor.

For organization-owned production repositories, account recovery also has a governance dimension: important repositories should not depend on one person’s personal account being recoverable. Organization ownership, multiple owners/admins, and offboarding/recovery runbooks reduce that single-person dependency; those governance controls are developed in later chapters.

8. Failure mode: valid PAT/SSH key lacks Enterprise Cloud SSO authorization

Synthetic fixture: a developer can authenticate to GitHub and access a public repository, but an enterprise organization resource returns an SSO-related denial. The PAT has not necessarily expired and the SSH key is not necessarily broken. The organization has a linked external identity and requires the credential to be authorized for SSO.

For a classic PAT, GitHub documents a Configure SSO authorization flow after creation. Fine-grained PAT authorization occurs during token creation/resource-owner selection. User SSH keys can also require SSO authorization. If an SSH key’s SSO authorization is revoked by an organization, GitHub documents that the same key cannot be reauthorized; create a new key instead.

Optional enterprise fixture only: do not create/alter SSO policy for this course. If you have an enterprise account, follow your organization’s identity policy and current Enterprise Cloud documentation.

9. Failure mode: gh uses a different account or credential source than expected

GitHub CLI can know multiple accounts/hosts, and environment variables can override stored credentials. Before changing anything, inspect:

gh auth status --json hosts
      printf 'GH_HOST=%s
' "${GH_HOST:-<unset>}"
      test -n "${GH_TOKEN:-}" && echo 'GH_TOKEN is set (value intentionally hidden)' || echo 'GH_TOKEN is unset'

On PowerShell, inspect only whether $env:GH_TOKEN is set—never print its value. GitHub CLI documentation states that GH_TOKEN/GITHUB_TOKEN take precedence for github.com and ghe.com subdomains. A CI shell that unexpectedly inherits GH_TOKEN can therefore behave differently from the developer’s stored gh login.

If the active account is wrong, use gh auth switch deliberately. If an environment credential is unintended, remove it from the process/CI configuration rather than deleting unrelated stored accounts.

10. Rate limits, retries, and reliability belong only after classification

Authenticated API use generally has a higher rate limit than unauthenticated use, but authentication failures should not trigger blind retries. GitHub documents temporary blocking after repeated invalid-credential attempts. Repeatedly retrying a stale token can therefore turn a local configuration error into a broader temporary authentication problem.

For automation, classify status first. Retry only transient/retryable conditions with bounded backoff; do not retry authorization denials until permissions/policy are corrected. Preserve request IDs and safe response metadata so platform support/admins can correlate failures without ever receiving the secret token.

11. Diagnostic matrix

Symptom Evidence first Least-destructive correction
SSH greeting names wrong user Exact-key test + ssh -G alias resolution. Fix host alias/IdentityFile or active key selection.
API 401 Credential source, expiry/revocation, gh status. Reauthenticate/rotate the intended credential; do not add permissions.
API 403/404 private repo Repo coordinate, underlying user access, PAT resource/repo selection, endpoint permission, SSO/policy. Add only missing repository/permission/authorization or use correct identity.
Git HTTPS repeatedly prompts Credential helper configuration and active account. Repair GCM/gh integration; remove stale cached credential intentionally.
Enterprise org only fails SSO session/credential authorization and org policy. Authorize according to SSO policy; do not widen token scope.
Token found in URL/log Evidence of exposure and token identity. Revoke/rotate first, then remove copies and verify old credential fails.

12. Disposable diagnostic lab with synthetic evidence

Do not create intentionally leaked real credentials. Instead, save the following synthetic response as api-failure.txt and classify it:

HTTP/2 404
x-github-request-id: SYNTHETIC:1234
x-accepted-github-permissions: contents=read
{
  "message": "Not Found"
}

Scenario context: the target repository is private; the user can open it in the browser; a temporary fine-grained PAT selected the correct repository but was created with Issues: read and no Contents permission. The evidence suggests the token is valid enough to reach authorization, but the endpoint requires Contents: read. The correct change is one missing read permission—not Administration: write, not a classic PAT with broad repo scope.

Now write the verification you would perform after correction: repeat the same GET, expect 200, compare the returned file metadata to the known repository, then revoke the temporary token.

13. Verification checklist

  • You can distinguish 401, 403, and authorization-obscuring 404 responses without assuming one cause.
  • You verify the resolved SSH/GitHub account before changing repository permissions.
  • You know GitHub’s documented ssh -T success exit behavior and use a real Git operation for automated functional tests.
  • A leaked credential triggers revocation/rotation before cleanup or history rewrite.
  • You inspect gh environment precedence without printing GH_TOKEN.
  • SSO failures are labeled Enterprise Cloud/policy-specific and are not “fixed” with broader scopes.
  • Retry/rate-limit analysis occurs only after authentication/authorization classification.

Knowledge check

A private-repository API request returns 404, but the repository definitely exists. Name four authorization checks.

Why is repeatedly retrying an invalid token a bad diagnostic strategy?

A token was committed and then immediately removed in the next commit. What is the first security action?

SSH works with github-personal but not github-work. What should you inspect before creating a new key?

An enterprise organization rejects a valid classic PAT until “Configure SSO” is completed. Authentication or authorization failure?

14. Summary

Secure troubleshooting preserves evidence and narrows the failing layer. Authentication identifies the principal; authorization combines user/app access, credential scope/permission, repository selection, organization policy, and SSO. HTTP status, SSH identity, gh account state, and endpoint permission headers are evidence—not reasons to print a secret.

When security is involved, the least-destructive fix also means the least-privilege fix. Do not widen access to make an error disappear, and do not perform history cleanup before revoking a leaked credential.

Next lesson

Prove the complete authentication matrix

Lesson 05 combines web account protection, SSH, HTTPS credential handling, GitHub CLI/API identity, an intentionally under-privileged fine-grained PAT request, least-privilege correction, and full cleanup evidence.

Authoritative references

 Authenticating to the REST API
 REST API versions
 Permissions required for fine-grained PATs
 Testing your SSH connection
 gh auth status
 GitHub CLI environment variables
 Token expiration and revocation
 Recovering an account after losing 2FA credentials
 About authentication with single sign-on
 Authorizing a PAT for SSO
 Authorizing an SSH key for SSO

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.