Chapter 30Lesson 04~220 minutes

Workflow Logs, Step Summaries, Debug Logging, Annotations, and Observability: Diagnostics, Failure Modes, and Production Practices

Diagnose secret/context leakage, stale or overwritten first-failure evidence, unsafe rendered input and artifact-identity ambiguity by preserving exact run/attempt state before correction.

DiagnosticsFirst failureMaskingUntrusted contentRecovery

Learning objectives

  • Apply an evidence-first diagnostic sequence without destroying the original run or attempt.
  • Diagnose why printing whole contexts and leaving debug enabled can become data-exfiltration paths.
  • Repair unsafe summary rendering without hiding the original cause.
  • Detect when a rerun or log message falsely implies artifact/source identity.
  • Use the least destructive correction and rerun only the smallest equivalent scope.

1. Evidence-first diagnostic sequence

  1. Preserve run ID/attempt and first-failure evidence. Save run metadata, attempt logs and relevant artifacts before cleanup or rerun.
  2. Confirm event/ref/SHA and workflow revision. Do not diagnose a different commit.
  3. Confirm evaluated inputs/conditions/permissions. Record allowed fields, never dump tokens or complete contexts.
  4. Inspect job graph/queue and runner assignment. Determine whether the failure occurred before application code.
  5. Confirm runner/image/toolchain. Record versions that materially affect the step.
  6. Inspect the exact failing process. Use normal logs first; enable debug only if the question remains unanswered.
  7. Inspect outputs/artifacts/caches and deployment/external state. A green upload is not proof of correct bytes.
  8. Apply the least destructive correction. Change one causal layer and preserve the failed evidence.
  9. Rerun the smallest equivalent scope. Same-source rerun for transient/debug questions; new run for source/input fixes.

2. Failure mode: printing entire contexts

BROKEN — do not use in production. The following pattern turns every field GitHub provides into an output surface and makes review nearly impossible. Event payloads contain user-controlled data; other contexts can contain paths, identities or credential-adjacent information. Even if GitHub masks known secrets, a context dump violates least disclosure.

# BROKEN — do not use in production
- name: Dump everything
  run: |
    echo '${{ toJSON(github) }}'
    echo '${{ toJSON(runner) }}'
    echo '${{ toJSON(secrets) }}'

Repair by defining the actual question. If you need source identity, print only event name, ref, SHA, run ID and attempt through environment variables. If you need a request body for forensic work, serialize an allowlisted schema to a restricted artifact after reviewing its sensitivity.

3. Failure mode: debug left enabled around credentials

BROKEN — do not use in production. A repository-wide debug flag remains enabled while later jobs perform registry/cloud authentication. A third-party CLI now emits verbose request metadata. Nothing guarantees every credential-shaped value is registered for masking.

Preserve the affected run and determine exactly what was emitted. Revoke/rotate any credential that may have been exposed; masking after the fact cannot retract downloaded logs. Then disable broad debug and reproduce only the failing safe step in a disposable run using on-demand debug. Never “prove safety” by searching only for the literal token—derived headers, signed URLs and provider responses can also be sensitive.

4. Failure mode: rerun overwrites the mental model of the first failure

A team keeps clicking rerun until a job turns green, then links only to the run’s latest attempt. That hides whether the original failure was deterministic, whether external state changed, and which attempt emitted an artifact. The run ID alone is ambiguous once multiple attempts exist.

Repair by recording RUN_ID + ATTEMPT everywhere. Download attempt 1 first. Compare attempt metadata and artifacts by exact IDs/names. If a code fix was made, start a new run from the corrected SHA instead of calling the old run “fixed.”

5. Failure mode: unescaped untrusted content in a summary

BROKEN — do not use in production. A pull-request title or test name is appended directly as Markdown. The text can create links/images/headings or misleading content in the run summary. Passing the string safely to Bash does not solve rendered-content trust.

# BROKEN — untrusted rendered content
- name: Summary
  env:
    PR_TITLE: ${{ github.event.pull_request.title }}
  run: echo "- failing title: $PR_TITLE" >> "$GITHUB_STEP_SUMMARY"

Repair by showing a content hash/length, or pass the value to a renderer that enforces a length bound and HTML-escapes it into a preformatted code block. If the value could be a secret/private payload, do not render it at all. Preserve the original unsafe run before changing the renderer so incident reviewers know which summary may have been exposed.

6. Failure mode: trusting a log line for artifact identity

A deployment log says using app:latest, and the operator assumes the image matches the CI artifact. This is an observability failure: a mutable tag plus text output does not prove bytes. The repair belongs to artifact/provenance state—record and verify the immutable digest, then include that digest in the human summary as a projection.

Likewise, a log saying “coverage uploaded” does not prove which artifact ID was stored. Query the run artifact metadata or action output and retain the artifact ID/digest when it affects a later decision.

7. Failure mode: annotation and continue-on-error create a false green

BROKEN — do not use in production. A test step has continue-on-error: true, emits an error annotation, and the job ends successfully. Reviewers see mixed signals; automation sees green.

# BROKEN — evidence/failure semantics disagree
- name: Tests
  continue-on-error: true
  run: |
    ./test.sh || echo "::error::tests failed"

Repair by capture-and-gate: preserve the command exit code, create the report/annotation, upload evidence under always(), then explicitly return the original failure. Reserve continue-on-error for intentionally non-gating experiments whose neutralized result is clearly modeled and aggregated.

8. Failure mode: event data becomes workflow-command or shell source

Never build a run: script by directly interpolating an issue title, branch name or PR body. An attacker may inject shell syntax before logging even begins. Pass values through env, quote them, and do not feed raw lines into workflow-command parsing. If raw non-secret output is absolutely required, use GitHub’s stop-command mechanism with a random token and a bounded scope—but derived metadata or an artifact is usually simpler and safer.

9. API/log download failures are another causal layer

A 302 from the log-download REST endpoint is expected: GitHub returns a short-lived redirect URL. An authentication failure, missing Actions read permission, expired redirect URL or rate limit is an API/access problem, not evidence that the workflow itself failed. With gh run view, note that the CLI may fall back to per-job log requests when zip-to-job association is incomplete; very large/missing log sets can therefore behave differently from a simple UI view.

10. Self-hosted causality: worker failure versus application failure

If a self-hosted job disappears, check host Runner_*/Worker_* diagnostics and connectivity before changing test code. If a worker log shows runner process termination or network loss, rerunning application tests is not the least destructive correction. Fix the host/service boundary, preserve diagnostics externally, then rerun the exact source.

11. Causal triage map

Evidence Likely layer Next check
Workflow never created event/filter/workflow selection event + workflow revision/path
Job pending/no runner queue/runner policy/capacity labels, group access, runner online state
401/403 API token/permission/API evaluated permissions + endpoint requirement
Step rc != 0 action/script/tool exact command + toolchain + safe logs
Artifact missing/wrong ID artifact/dataflow upload step, action output, run/attempt
Deployment green, target unhealthy external target provider/Kubernetes health evidence
Summary misleading rendering/observability source record + renderer + untrusted input policy

12. Least-destructive recovery

  • Preserve attempt 1 and any artifact IDs before rerun.
  • Turn debug on only for the smallest failing scope, then turn it back off.
  • Rotate credentials only when exposure is plausible; do not delete the run as a first response.
  • Fix the exact renderer/dataflow/permission layer instead of broadening tokens or bypassing policy.
  • Use a new run for a source/configuration fix; use a rerun for same-source diagnostic/transient validation.

13. Lesson summary

Observability failures often come from mixing layers: rendered text is mistaken for machine evidence, an annotation for a failure, a rerun for a new source, or masking for access control. Preserve exact attempt evidence first, identify the owning layer, then change the smallest thing that restores both diagnosability and confidentiality.

Next lesson

Checkpoint Lab — Workflow Logs, Step Summaries, Debug Logging, Annotations, and Observability

Continue with the next lesson to build on the current concepts, evidence, security boundaries, and operational practices.

Knowledge check

A rerun turns green after attempt 1 failed, but no code changed. What must you preserve before calling it transient?

Why is toJSON(secrets) unacceptable even if masking exists?

A summary shows an attacker-controlled Markdown image. Which two safety layers were confused?

What should you do if a cloud credential may have appeared in debug logs?

Why can a log-download 302 be normal?

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.