Chapter 30Lesson 03~210 minutes

Workflow Logs, Step Summaries, Debug Logging, Annotations, and Observability: Configuration, Design Patterns, and Trade-Offs

Choose logs, summaries, annotations, artifacts, debug modes and runner diagnostics according to audience, integrity, retention, cost and trust boundaries.

Design choicesRetentionTelemetryTrade-offsIntegrity

Learning objectives

  • Choose between logs, annotations, summaries and artifacts based on audience and evidence integrity.
  • Decide when on-demand debug is justified and when verbosity creates more risk than diagnostic value.
  • Separate workflow logs from self-hosted runner diagnostics and external telemetry.
  • Design summary content so convenience never becomes the canonical release/artifact identity.
  • Account for retention, visibility, cost and cleanup when designing an observability contract.

1. Observability surfaces solve different problems

A production workflow should not have one “logging strategy.” It has several evidence surfaces with different consumers and trust properties. Human operators need concise causal clues. Automation needs stable machine fields. Security teams need access-controlled retained evidence. Platform owners need aggregate metrics. Choosing the wrong surface usually creates either noise or data exposure.

2. Human-readable log versus machine artifact

Question Log Artifact / telemetry
Primary consumer human during diagnosis machine or later forensic analysis
Shape chronological text structured JSON/JUnit/SARIF/custom schema
Integrity identity run/job/step context, but text can be ambiguous record artifact ID/digest/schema + run/attempt/SHA
Size strategy bounded diagnostic excerpts complete approved machine evidence
Retention repository log policy artifact policy or external retention policy
Sensitive data avoid entirely still avoid or explicitly protect with separate access policy

If a deployment needs an exact image digest, write the digest to a job output/evidence artifact and verify it. A line in a log saying “deploying :latest” is not a trustworthy artifact identity.

3. Annotation versus failure

Annotations are optimized for navigation: they can point a notice/warning/error at a file/line and give reviewers a concise explanation. A process exit code determines whether a step actually failed. Keep those semantics aligned but separate. For example, emit one bounded error annotation pointing at the validation source, preserve a detailed report artifact, and return non-zero through the authoritative gate.

Do not flood every test assertion into annotations. Large annotation sets become UI noise and can hit service/display limits that evolve over time. Use a summary and artifact for aggregate/detail, and reserve annotations for the most actionable findings.

4. Always-on debug versus on-demand debug

Always-on debug creates extra logs, longer incident review and a larger chance that tools reveal headers, paths or environment detail. On-demand debug narrows the risk and storage to a specific investigation. Prefer an ordinary first attempt, preserve it, then use gh run rerun RUN_ID --failed --debug when the normal log cannot answer the causal question.

Debug design Benefit Cost/risk Default
Always on maximum immediate detail noise, storage, accidental sensitive detail Avoid
Repository variable convenient repeated troubleshooting easy to forget enabled Short, controlled window only
Repository secret same flag with restricted value visibility secret value wins over variable; still broad across eligible runs Rarely needed
Per-rerun debug exact incident scope, preserves attempt relation requires a rerun and repository permission Preferred diagnostic path

5. Runner logs versus workflow logs

Workflow logs describe the workflow process. Self-hosted Runner_* logs describe registration/service coordination; Worker_* logs describe job execution on that host. A network disconnect, runner update or service crash may never be explained by application stdout. Conversely, a failed unit test usually does not require host diagnostics.

Ephemeral runners and ARC pods disappear by design. If runner-level diagnostics are part of your incident requirement, external forwarding is an operational prerequisite, not a troubleshooting afterthought. Keep external telemetry identifiers linked to run ID, attempt and job identity.

6. Summary convenience versus evidence integrity

A summary is a rendered operator view. It is excellent for a table of test counts, links to artifacts and a bounded “what failed” explanation. It is poor as the only record of a release digest, approval decision or provider response because summary content is generated by workflow code and rendered for humans.

The integrity pattern is machine evidence first, summary projection second. Produce evidence.json, compute/record its digest or artifact identity, then render a small subset into the summary. A reader can use the summary quickly and still trace back to the retained source evidence.

7. Rendered untrusted content needs a display policy

Issue titles, branch names, pull-request text, test names and tool output can contain adversarial strings. Shell safety and rendering safety are different. Passing a value through env protects shell-source construction; it does not make Markdown safe. If content must be shown, apply length bounds and escape it into a code/preformatted representation. Often the safer design is to show only a hash/identifier and keep the original in an access-controlled artifact.

Do not rely on secret masking to sanitize arbitrary content. Masking targets registered secret values. It does not decide whether a URL, image, stack trace, source fragment or user input is appropriate to publish.

8. Retention, visibility and cost are part of the design

Current GitHub.com defaults retain logs/artifacts for 90 days. Public repositories can configure 1–90 days; private/internal repositories can configure 1–400 days, subject to higher-level limits. Increasing retention improves forensic reach but increases storage and exposure duration. External telemetry has its own billing, residency and deletion model.

Choose retention from an evidence requirement—release audit, incident response, CI tuning—not from “keep everything forever.” A summary or log may need shorter retention than a signed release artifact or regulated external audit record.

9. Worked design scenarios

Scenario Recommended surface Why
One failing lint location error annotation + non-zero step fast code navigation, authoritative failure
Test counts + links step summary small human overview
10 MB JUnit/coverage bundle artifact machine evidence, downloadable by exact run
Self-hosted runner disconnect _diag Runner_/Worker_ + run metadata host/service layer owns the failure
Temporary unexplained action failure debug rerun after preserving attempt 1 same source with increased diagnostics
Artifact released to consumers artifact/package digest + provenance log text alone cannot identify bytes
Fleet-wide platform SLO external metrics/telemetry aggregate across repositories/runs

10. Permissions and external telemetry

Reading run metadata from inside a workflow needs only the Actions read boundary where the API supports it. Shipping telemetry externally introduces a new credential/network boundary. Prefer OIDC or narrowly scoped credentials where supported, transmit only an allowlisted schema, and keep provider tokens out of summaries/logs. A telemetry sink being “observability” does not justify write-all or unrestricted egress.

11. Rollback means turning verbosity down without losing evidence

If a debug change creates noise or risk, disable only the debug flag or stop the telemetry exporter; do not delete the failed run. If a summary renderer is unsafe, stop rendering raw content and preserve the machine evidence. If an external telemetry schema is wrong, version the schema and keep correlation IDs so old and new records remain interpretable.

12. Lesson summary

Use logs for bounded chronology, annotations for actionable navigation, summaries for human projection, artifacts for machine evidence, runner diagnostics for host failures and external telemetry for aggregate operations. The system is auditable only when those surfaces share run/attempt/SHA correlation without sharing unnecessary sensitive content.

Next lesson

Workflow Logs, Step Summaries, Debug Logging, Annotations, and Observability: Diagnostics, Failure Modes, and Production Practices

Continue with the next lesson to build on the current concepts, evidence, security boundaries, and operational practices.

Knowledge check

When should an annotation be preferred over a large log dump?

Why is always-on debug a poor default?

What does “machine evidence first, summary projection second” mean?

A self-hosted runner loses communication before a test starts. Which evidence layer comes first?

Why is a release digest in a log insufficient?

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.