Chapter 30Lesson 01~205 minutes

Workflow Logs, Step Summaries, Debug Logging, Annotations, and Observability: Core Concepts and Mental Model

Model GitHub Actions observability as correlated run/job/attempt metadata, bounded logs, annotations, step summaries and retained evidence without turning diagnostics into a data-leak channel.

ObservabilityLogsSummariesAnnotationsEvidence

Learning objectives

  • Explain how stdout/stderr, workflow commands, annotations, summaries and API metadata become distinct observability surfaces.
  • Track run ID, run attempt, job identity, step name, timestamps and exit codes without collapsing them into one green/red signal.
  • Use logs for diagnosis while keeping artifacts or external telemetry for machine-readable evidence and longer-lived correlation.
  • Explain on-demand step/runner debug logging and the additional confidentiality risk of verbose diagnostics.
  • Inspect hosted and self-hosted diagnostic state read-only before changing debug settings or deleting evidence.

1. The practical problem: a failed workflow is not a diagnosis

Chapter 29 turned workflows into a delivery platform. At that scale, “the job failed” is not enough. You need to know which event and source revision ran, which attempt produced the evidence, which job and step failed, what the runner saw, and whether a later rerun changed the evidence or merely repeated the same source. Observability is the discipline of preserving those identities and exposing enough structured signals to explain behavior.

The dangerous opposite is “log everything.” Actions logs, summaries and annotations are visible to repository readers according to repository permissions. Debug modes can expose more runtime detail. Event data can contain attacker-controlled strings. A good observability design therefore asks two questions together: Can an operator prove what happened? and Can the evidence be produced without leaking or rendering data that should never have been surfaced?.

2. Mental model: process output becomes several evidence surfaces

A runner launches actions and shell processes. Their normal stdout/stderr becomes step logs. Lines that intentionally use the workflow-command protocol can create groups, debug lines, masks and annotations. A step can also append Markdown to its own GITHUB_STEP_SUMMARY file. GitHub combines step and job conclusions with run metadata that can be inspected in the UI, with gh, or through the Actions API.

Observability causality
flowchart TD
  A[Event + workflow revision + source SHA] --> B[Run ID + attempt]
  B --> C[Job queued and runner assigned]
  C --> D[Action / shell process]
  D --> E[stdout + stderr]
  D --> F[workflow commands]
  D --> G[GITHUB_STEP_SUMMARY]
  E --> H[Step logs]
  F --> I[Groups / debug / annotations / masks]
  G --> J[Job summary]
  H --> K[Run metadata + API]
  I --> K
  J --> K
  K --> L[Retained evidence / external telemetry]

The arrows are causal, not decorative. An error annotation is not the same state as a failed command. A pretty summary is not proof of artifact identity. A rerun is a new attempt attached to the same run identity and original source/ref. External telemetry is yet another system whose correlation identifiers must be carried deliberately.

3. Define the state before you instrument it

Layer State to record Why it matters
Event/revision event, ref, github.sha, workflow ref Proves which source/configuration the run was intended to evaluate.
Run run ID, run number, run_attempt, triggering/original actor Separates the original execution from diagnostic reruns.
Job/runner job name/ID, runner OS/arch/name, image/tool versions Explains execution environment and queue/runtime differences.
Step/process step name, command, start/end, exit code Identifies the causal failing operation.
Logs stdout/stderr, group boundaries, debug state Human diagnostic stream; potentially sensitive and retention-bound.
Annotations notice/warning/error location + message UI navigation and signal; not automatically the same as step failure.
Summary bounded Markdown produced per step Human run overview, not a canonical machine evidence store.
Artifact/telemetry evidence JSON/digest, artifact ID, external trace/correlation ID Machine-readable/longer-lived proof and cross-system correlation.
Governance retention, access, redaction rules, evidence owner Defines who can see evidence and how long it remains available.

4. Logs are an execution transcript, not a data lake

Logs should answer a bounded diagnostic question: which step ran, what non-sensitive state did it observe, and why did it return its exit code? Prefer stable step names and explicit messages. Group verbose but safe detail under fixed group labels so operators can scan the run without hiding causality.

echo "::group::Toolchain evidence"
printf 'sha=%s\n' "$GITHUB_SHA"
printf 'run=%s attempt=%s\n' "$GITHUB_RUN_ID" "$GITHUB_RUN_ATTEMPT"
python --version
echo "::endgroup::"

Do not place event-derived text directly into shell source. Pass it through env: and quote it. If raw untrusted text does not materially help diagnosis, log only a length, content hash or safely normalized identifier and preserve the original in a controlled artifact only when policy allows.

5. Workflow commands: annotations, groups, debug and masking

GitHub recognizes workflow commands such as group, endgroup, debug, notice, warning, error and add-mask. Use fixed metadata fields or correctly escaped values. A line like ::error ...::message creates an error annotation, but an annotation alone is not a substitute for the command returning a failing exit code. The toolkit’s setFailed concept combines an error signal with a failed process result.

Trust rule: never print an entire github, secrets, OIDC, environment or provider-response context just to “see what is there.” Masking is defense in depth, not authorization. Design a bounded schema of fields that are allowed to leave the process.

6. Step summaries are human views with hard boundaries

Each step receives a unique path in GITHUB_STEP_SUMMARY. Appending GitHub-flavored Markdown produces a summary for that step; after the step completes, later steps cannot edit the already-uploaded summary. As verified September 10, 2026, each step summary is limited to 1 MiB and a maximum of 20 step summaries are displayed per job. Summary-upload errors create an annotation but do not by themselves fail the step or job.

Summaries automatically mask registered secrets that are accidentally included, but that is not permission to write sensitive data. For attacker-controlled or merely untrusted text, prefer derived metadata. If a bounded snippet must be displayed, escape it before rendering and keep it inside a code/preformatted context rather than allowing Markdown links/images/HTML to become active UI.

import hashlib, html, os
raw = os.environ.get("RAW_SAMPLE", "")[:160]
safe = html.escape(raw, quote=True)
with open(os.environ["GITHUB_STEP_SUMMARY"], "a", encoding="utf-8") as f:
    f.write("### Bounded sample\n")
    f.write(f"- length shown: `{len(raw)}`\n")
    f.write(f"- sha256: `{hashlib.sha256(raw.encode()).hexdigest()}`\n")
    f.write(f"<pre><code>{safe}</code></pre>\n")

7. Debug logging is an on-demand incident tool

Step debug logging increases job-log verbosity. Runner diagnostic logging adds runner/worker diagnostic files to the downloadable log archive. GitHub currently supports repository secret or variable flags ACTIONS_STEP_DEBUG=true and ACTIONS_RUNNER_DEBUG=true; if both a secret and variable of the same name exist, the secret wins. The runner.debug context can tell workflow logic whether debug logging is active.

A safer operational pattern is to enable debug only for the exact disposable rerun you are diagnosing. gh run rerun RUN_ID --failed --debug enables diagnostic logging for that rerun. Preserve attempt 1 before doing this, and review the debug archive before sharing it outside the repository.

8. Run versus attempt: rerun evidence must not overwrite the first failure

A rerun keeps the original GITHUB_SHA and GITHUB_REF and uses the privileges of the actor who originally triggered the run. It increments the run attempt rather than creating evidence for a new source revision. GitHub currently permits up to 50 reruns per workflow run and allows reruns for up to 30 days after the initial run.

That leads to an important diagnosis rule: if the correction is a repository code or workflow change, a rerun of the old run cannot prove the new commit. Create a new run from the fixed revision and correlate it with the old run. Use a rerun only when you deliberately want to repeat the same source—for example to gather debug detail or re-test a genuinely transient dependency.

9. Retention and deletion are governance state

GitHub stores Actions artifacts and logs for 90 days by default. Current GitHub.com repository settings allow public repositories to use 1–90 days and private/internal repositories to use 1–400 days, subject to organization/enterprise maximums. Retention changes apply to newly created evidence, not retroactively to existing logs/artifacts.

Deleting a run is disruptive evidence destruction: it also removes associated artifacts and is required if a completed step summary contains sensitive content that must be removed. Therefore preserve first-failure evidence and obtain approval before deletion during an incident.

10. Hosted diagnostics and self-hosted _diag are different surfaces

For a GitHub-hosted debug rerun, runner diagnostics are included in the run’s downloaded log archive under runner-diagnostic-logs. A self-hosted runner also maintains application/job diagnostics in its installation _diag directory: Runner_* files describe the runner service/application and Worker_* files correlate to job execution. Those host files are outside GitHub’s normal job log stream.

For ephemeral self-hosted runners or ARC runner pods, forward diagnostic logs externally before teardown if they are needed for incident response. Do not “fix” connectivity by disabling TLS verification; correct trust-store or network configuration instead.

11. Read-only inspection first

# Local source identity
git rev-parse HEAD
git status --short

# Exact run metadata; no mutation
gh run list --limit 10 --json databaseId,attempt,event,headSha,status,conclusion,url

# Exact run + attempt view
gh run view RUN_ID --attempt ATTEMPT --json attempt,conclusion,createdAt,headSha,jobs,url

# Read-only API request. gh handles the auth header; do not use curl -v with tokens.
gh api \
  -H "Accept: application/vnd.github+json" \
  -H "X-GitHub-Api-Version: 2026-03-10" \
  repos/OWNER/REPO/actions/runs/RUN_ID

Capture the run ID and attempt before opening individual logs. If artifact identity matters, record its artifact ID/digest separately; a log line saying “uploaded build.zip” is not proof that the intended bytes were stored.

12. Lesson summary

Observability in Actions is a correlation problem across source revision, run, attempt, job, runner, process, human-facing signals and retained machine evidence. Produce bounded diagnostics, keep untrusted and sensitive data out of rendered/logged surfaces, enable verbose modes only when needed, and never let a rerun erase the meaning of the original failure.

Next lesson

Workflow Logs, Step Summaries, Debug Logging, Annotations, and Observability: Guided Hands-On Workflow

Continue with the next lesson to build on the current concepts, evidence, security boundaries, and operational practices.

Knowledge check

Does an ::error annotation automatically prove the step failed?

Why record both run ID and run attempt?

What is safer than printing an entire event object?

Can a rerun prove a fix committed after the original run?

Where do self-hosted Runner_ and Worker_ logs live?

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.