Workflow Logs, Step Summaries, Debug Logging, Annotations, and Observability: Core Concepts and Mental Model
Model GitHub Actions observability as correlated run/job/attempt metadata, bounded logs, annotations, step summaries and retained evidence without turning diagnostics into a data-leak channel.
Learning objectives
- Explain how stdout/stderr, workflow commands, annotations, summaries and API metadata become distinct observability surfaces.
- Track run ID, run attempt, job identity, step name, timestamps and exit codes without collapsing them into one green/red signal.
- Use logs for diagnosis while keeping artifacts or external telemetry for machine-readable evidence and longer-lived correlation.
- Explain on-demand step/runner debug logging and the additional confidentiality risk of verbose diagnostics.
- Inspect hosted and self-hosted diagnostic state read-only before changing debug settings or deleting evidence.
1. The practical problem: a failed workflow is not a diagnosis
Chapter 29 turned workflows into a delivery platform. At that scale, “the job failed” is not enough. You need to know which event and source revision ran, which attempt produced the evidence, which job and step failed, what the runner saw, and whether a later rerun changed the evidence or merely repeated the same source. Observability is the discipline of preserving those identities and exposing enough structured signals to explain behavior.
The dangerous opposite is “log everything.” Actions logs, summaries and annotations are visible to repository readers according to repository permissions. Debug modes can expose more runtime detail. Event data can contain attacker-controlled strings. A good observability design therefore asks two questions together: Can an operator prove what happened? and Can the evidence be produced without leaking or rendering data that should never have been surfaced?.
2. Mental model: process output becomes several evidence surfaces
A runner launches actions and shell processes. Their normal
stdout/stderr becomes step logs. Lines that intentionally use the
workflow-command protocol can create groups, debug lines, masks and
annotations. A step can also append Markdown to its own
GITHUB_STEP_SUMMARY file. GitHub combines step and job
conclusions with run metadata that can be inspected in the UI, with
gh, or through the Actions API.
flowchart TD A[Event + workflow revision + source SHA] --> B[Run ID + attempt] B --> C[Job queued and runner assigned] C --> D[Action / shell process] D --> E[stdout + stderr] D --> F[workflow commands] D --> G[GITHUB_STEP_SUMMARY] E --> H[Step logs] F --> I[Groups / debug / annotations / masks] G --> J[Job summary] H --> K[Run metadata + API] I --> K J --> K K --> L[Retained evidence / external telemetry]
The arrows are causal, not decorative. An error annotation is not the same state as a failed command. A pretty summary is not proof of artifact identity. A rerun is a new attempt attached to the same run identity and original source/ref. External telemetry is yet another system whose correlation identifiers must be carried deliberately.
3. Define the state before you instrument it
| Layer | State to record | Why it matters |
|---|---|---|
| Event/revision | event, ref, github.sha, workflow ref |
Proves which source/configuration the run was intended to evaluate. |
| Run |
run ID, run number, run_attempt,
triggering/original actor
|
Separates the original execution from diagnostic reruns. |
| Job/runner | job name/ID, runner OS/arch/name, image/tool versions | Explains execution environment and queue/runtime differences. |
| Step/process | step name, command, start/end, exit code | Identifies the causal failing operation. |
| Logs | stdout/stderr, group boundaries, debug state | Human diagnostic stream; potentially sensitive and retention-bound. |
| Annotations | notice/warning/error location + message | UI navigation and signal; not automatically the same as step failure. |
| Summary | bounded Markdown produced per step | Human run overview, not a canonical machine evidence store. |
| Artifact/telemetry | evidence JSON/digest, artifact ID, external trace/correlation ID | Machine-readable/longer-lived proof and cross-system correlation. |
| Governance | retention, access, redaction rules, evidence owner | Defines who can see evidence and how long it remains available. |
4. Logs are an execution transcript, not a data lake
Logs should answer a bounded diagnostic question: which step ran, what non-sensitive state did it observe, and why did it return its exit code? Prefer stable step names and explicit messages. Group verbose but safe detail under fixed group labels so operators can scan the run without hiding causality.
echo "::group::Toolchain evidence"
printf 'sha=%s\n' "$GITHUB_SHA"
printf 'run=%s attempt=%s\n' "$GITHUB_RUN_ID" "$GITHUB_RUN_ATTEMPT"
python --version
echo "::endgroup::"
Do not place event-derived text directly into shell source. Pass it
through env: and quote it. If raw untrusted text does
not materially help diagnosis, log only a length, content hash or
safely normalized identifier and preserve the original in a
controlled artifact only when policy allows.
5. Workflow commands: annotations, groups, debug and masking
GitHub recognizes workflow commands such as group,
endgroup, debug, notice,
warning, error and add-mask.
Use fixed metadata fields or correctly escaped values. A line like
::error ...::message creates an error annotation, but
an annotation alone is not a substitute for the command returning a
failing exit code. The toolkit’s setFailed concept
combines an error signal with a failed process result.
Trust rule: never print an entire
github, secrets, OIDC, environment or
provider-response context just to “see what is there.” Masking is
defense in depth, not authorization. Design a bounded schema of
fields that are allowed to leave the process.
6. Step summaries are human views with hard boundaries
Each step receives a unique path in
GITHUB_STEP_SUMMARY. Appending GitHub-flavored Markdown
produces a summary for that step; after the step completes, later
steps cannot edit the already-uploaded summary. As verified
September 10, 2026, each step summary is limited to 1 MiB and a
maximum of 20 step summaries are displayed per job. Summary-upload
errors create an annotation but do not by themselves fail the step
or job.
Summaries automatically mask registered secrets that are accidentally included, but that is not permission to write sensitive data. For attacker-controlled or merely untrusted text, prefer derived metadata. If a bounded snippet must be displayed, escape it before rendering and keep it inside a code/preformatted context rather than allowing Markdown links/images/HTML to become active UI.
import hashlib, html, os
raw = os.environ.get("RAW_SAMPLE", "")[:160]
safe = html.escape(raw, quote=True)
with open(os.environ["GITHUB_STEP_SUMMARY"], "a", encoding="utf-8") as f:
f.write("### Bounded sample\n")
f.write(f"- length shown: `{len(raw)}`\n")
f.write(f"- sha256: `{hashlib.sha256(raw.encode()).hexdigest()}`\n")
f.write(f"<pre><code>{safe}</code></pre>\n")
7. Debug logging is an on-demand incident tool
Step debug logging increases job-log verbosity. Runner diagnostic
logging adds runner/worker diagnostic files to the downloadable log
archive. GitHub currently supports repository secret or variable
flags ACTIONS_STEP_DEBUG=true and
ACTIONS_RUNNER_DEBUG=true; if both a secret and
variable of the same name exist, the secret wins. The
runner.debug context can tell workflow logic whether
debug logging is active.
A safer operational pattern is to enable debug only for the exact
disposable rerun you are diagnosing.
gh run rerun RUN_ID --failed --debug enables diagnostic
logging for that rerun. Preserve attempt 1 before doing this, and
review the debug archive before sharing it outside the repository.
8. Run versus attempt: rerun evidence must not overwrite the first failure
A rerun keeps the original GITHUB_SHA and
GITHUB_REF and uses the privileges of the actor who
originally triggered the run. It increments the run attempt rather
than creating evidence for a new source revision. GitHub currently
permits up to 50 reruns per workflow run and allows reruns for up to
30 days after the initial run.
That leads to an important diagnosis rule: if the correction is a repository code or workflow change, a rerun of the old run cannot prove the new commit. Create a new run from the fixed revision and correlate it with the old run. Use a rerun only when you deliberately want to repeat the same source—for example to gather debug detail or re-test a genuinely transient dependency.
9. Retention and deletion are governance state
GitHub stores Actions artifacts and logs for 90 days by default. Current GitHub.com repository settings allow public repositories to use 1–90 days and private/internal repositories to use 1–400 days, subject to organization/enterprise maximums. Retention changes apply to newly created evidence, not retroactively to existing logs/artifacts.
Deleting a run is disruptive evidence destruction: it also removes associated artifacts and is required if a completed step summary contains sensitive content that must be removed. Therefore preserve first-failure evidence and obtain approval before deletion during an incident.
10. Hosted diagnostics and self-hosted _diag are
different surfaces
For a GitHub-hosted debug rerun, runner diagnostics are included in
the run’s downloaded log archive under
runner-diagnostic-logs. A self-hosted runner also
maintains application/job diagnostics in its installation
_diag directory: Runner_* files describe
the runner service/application and Worker_* files
correlate to job execution. Those host files are outside GitHub’s
normal job log stream.
For ephemeral self-hosted runners or ARC runner pods, forward diagnostic logs externally before teardown if they are needed for incident response. Do not “fix” connectivity by disabling TLS verification; correct trust-store or network configuration instead.
11. Read-only inspection first
# Local source identity
git rev-parse HEAD
git status --short
# Exact run metadata; no mutation
gh run list --limit 10 --json databaseId,attempt,event,headSha,status,conclusion,url
# Exact run + attempt view
gh run view RUN_ID --attempt ATTEMPT --json attempt,conclusion,createdAt,headSha,jobs,url
# Read-only API request. gh handles the auth header; do not use curl -v with tokens.
gh api \
-H "Accept: application/vnd.github+json" \
-H "X-GitHub-Api-Version: 2026-03-10" \
repos/OWNER/REPO/actions/runs/RUN_ID
Capture the run ID and attempt before opening individual logs. If artifact identity matters, record its artifact ID/digest separately; a log line saying “uploaded build.zip” is not proof that the intended bytes were stored.
12. Lesson summary
Observability in Actions is a correlation problem across source revision, run, attempt, job, runner, process, human-facing signals and retained machine evidence. Produce bounded diagnostics, keep untrusted and sensitive data out of rendered/logged surfaces, enable verbose modes only when needed, and never let a rerun erase the meaning of the original failure.
Knowledge check
Does an ::error annotation automatically prove the
step failed?
No. It creates an error annotation; the process/step still needs a failing exit status (or equivalent setFailed behavior) when the condition should fail the job.
Why record both run ID and run attempt?
A rerun reuses the run identity/source but increments the attempt. Without both, later diagnostics can be confused with the first failure.
What is safer than printing an entire event object?
Select an allowlisted set of non-sensitive fields and log derived metadata such as length/hash for untrusted strings.
Can a rerun prove a fix committed after the original run?
No. Reruns use the original GITHUB_SHA/GITHUB_REF. A source fix needs a new run from the corrected revision.
Where do self-hosted Runner_ and Worker_ logs live?
In the runner installation _diag directory, separate from ordinary workflow step logs.
Official references and version notes
- Workflow commands — Current commands, annotations, log groups, masking and GITHUB_STEP_SUMMARY behavior/limits.
- Enable debug logging — ACTIONS_STEP_DEBUG, ACTIONS_RUNNER_DEBUG and runner-diagnostic log behavior.
- Re-run workflows and jobs — Rerun identity, attempt behavior, debug reruns and privilege semantics.
- Monitor workflows — Workflow graph, logs, job timing and troubleshooting entry points.
- REST: workflow runs — Run metadata and exact run-attempt log download endpoints.
- REST: workflow jobs — Job metadata and job-log endpoints.
- Variables reference — GITHUB_STEP_SUMMARY, run/attempt and workflow metadata variables.
- Repository Actions settings — Artifact/log retention and repository-level Actions settings.
- Self-hosted runner troubleshooting — Runner_/Worker_ diagnostic files, connectivity checks and external preservation guidance.
- GitHub CLI: gh run rerun — Current --debug/--failed/--job rerun controls.
- GitHub CLI: gh run view — Attempt-aware run inspection and JSON/log views.
- actions/checkout v7.0.1 — Full-SHA pin used by executable examples.
- actions/setup-python v7.0.0 — Full-SHA pin for Python 3.13 used in the labs.
- actions/upload-artifact v7.0.1 — Full-SHA pin for retained machine-readable evidence.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.