Chapter 30Lesson 05~270 minutes

Checkpoint Lab — Workflow Logs, Step Summaries, Debug Logging, Annotations, and Observability

Instrument a failing workflow, preserve attempt 1, perform an on-demand debug rerun, correct the invocation, verify a new successful run, and produce a run-to-evidence correlation packet.

CheckpointRun attemptsDebugEvidence packetCorrelation

Learning objectives

  • Predict run/attempt/evidence state before executing a failing observability workflow.
  • Preserve attempt 1 before any diagnostic rerun or cleanup.
  • Enable debug for an exact rerun and prove source identity did not change.
  • Correct the invocation using a new run and correlate failure, debug attempt and success without leaking the demo input.
  • Deliver an evidence packet that separates human summary, logs, artifact identity and runner diagnostics.

1. Checkpoint mission and success criteria

Build the disposable gha-observability-checkpoint repository. Your first run must fail intentionally while still producing a safe summary, one actionable annotation and a machine-readable evidence artifact. Preserve attempt 1. Then rerun the failed job with debug logging to create attempt 2 of the same run. Finally correct the dispatch input and create a new successful run. Your evidence packet must make those three executions distinguishable without containing the raw demo input outside the explicitly escaped summary snippet.

Checkpoint guard: use synthetic text and a disposable repository only. Never insert a real secret to test masking. Do not delete the failed run during the exercise. Debug reruns are diagnostic and can contain more runtime detail.

2. Current assumptions and preflight

Assumption Checkpoint value / verification
GitHub behavior GitHub.com, verified 2026-09-10
Runner ubuntu-24.04; record runner.os/runner.arch and tool versions
Checkout v7.0.1 → 3d3c42e5aac5ba805825da76410c181273ba90b1
Setup Python v7.0.0 → 5fda3b95a4ea91299a34e894583c3862153e4b97; Python 3.13 requested
Evidence upload v7.0.1 → 043fb46d1a93c77aae656e7c1c64a875d1fc6a0a; 7-day lab artifact retention
API version 2026-03-10 for explicit REST examples
Debug on-demand gh run rerun --debug; no persistent debug flag required
Secrets/cloud none; sample text is synthetic and non-sensitive
gh auth status
gh --version
gh repo view --json nameWithOwner,url,visibility
git rev-parse HEAD
git status --short

3. Predict state changes before running

Write your predictions into checkpoint-predictions.md before dispatch. At minimum predict these states and later mark each as verified or disproved:

  • Prediction A: the first dispatch creates a new run ID with attempt 1, checks out its exact headSha, uploads an evidence artifact, then concludes failure with exit code 17 at the final gate.
  • Prediction B: gh run rerun RUN_ID --failed --debug keeps the same run ID/SHA/ref but creates attempt 2 and additional debug/runner diagnostic evidence; it still fails because the original mode=fail input is unchanged.
  • Prediction C: a new dispatch with mode=pass creates a different run ID at attempt 1 and succeeds while preserving the old failed run/attempts.
  • Prediction D: the raw sample does not appear in ordinary workflow logs; only bounded derived metadata and the escaped summary representation are intentionally visible.

4. Exact checkpoint workflow

Use the same reviewed workflow from Lesson 2. Repetition is intentional here: the checkpoint assesses whether you can explain every evidence surface and its state transition, not whether you can invent new YAML.

name: Observability lab
on:
  workflow_dispatch:
    inputs:
      mode:
        description: "Synthetic result to exercise"
        type: choice
        required: true
        options: [fail, pass]
      sample:
        description: "Untrusted demo text; never treat as shell source"
        type: string
        required: false
        default: "hello <observer>"

permissions: {}

jobs:
  observe:
    name: Observe synthetic check
    runs-on: ubuntu-24.04
    permissions:
      contents: read
      actions: read
    steps:
      - name: Checkout exact event revision
        uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
        with:
          persist-credentials: false

      - name: Set up pinned Python
        uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0
        with:
          python-version: "3.13"

      - name: Capture bounded run evidence
        id: evidence
        shell: bash
        env:
          LAB_MODE: ${{ inputs.mode }}
          RAW_SAMPLE: ${{ inputs.sample }}
        run: |
          set -euo pipefail
          python - <<'PY'
          import hashlib, json, os, platform
          raw = os.environ.get("RAW_SAMPLE", "")
          data = {
            "event": os.environ["GITHUB_EVENT_NAME"],
            "sha": os.environ["GITHUB_SHA"],
            "ref": os.environ["GITHUB_REF"],
            "run_id": os.environ["GITHUB_RUN_ID"],
            "run_attempt": os.environ["GITHUB_RUN_ATTEMPT"],
            "mode": os.environ["LAB_MODE"],
            "sample_length": len(raw),
            "sample_sha256": hashlib.sha256(raw.encode()).hexdigest(),
            "python": platform.python_version(),
          }
          with open("evidence.json", "w", encoding="utf-8") as f:
              json.dump(data, f, indent=2, sort_keys=True)
          PY
          printf '::group::Bounded run identity\n'
          printf 'sha=%s\n' "$GITHUB_SHA"
          printf 'run=%s attempt=%s mode=%s\n' "$GITHUB_RUN_ID" "$GITHUB_RUN_ATTEMPT" "$LAB_MODE"
          python --version
          printf '::endgroup::\n'
          printf '::notice title=Observability lab::Bounded run evidence captured; raw sample is not logged.\n'

      - name: Write safe human summary
        shell: bash
        env:
          RAW_SAMPLE: ${{ inputs.sample }}
        run: |
          set -euo pipefail
          python - <<'PY'
          import hashlib, html, json, os
          data = json.load(open("evidence.json", encoding="utf-8"))
          raw = os.environ.get("RAW_SAMPLE", "")[:160]
          safe = html.escape(raw, quote=True)
          with open(os.environ["GITHUB_STEP_SUMMARY"], "a", encoding="utf-8") as f:
              f.write("## Observability lab\n\n")
              f.write(f"- source SHA: `{data['sha']}`\n")
              f.write(f"- run / attempt: `{data['run_id']} / {data['run_attempt']}`\n")
              f.write(f"- synthetic mode: `{data['mode']}`\n")
              f.write(f"- sample sha256: `{data['sample_sha256']}`\n\n")
              f.write("Bounded, HTML-escaped sample (never a secret):\n\n")
              f.write(f"<pre><code>{safe}</code></pre>\n")
          PY

      - name: Run synthetic validation and preserve exit code
        id: check
        shell: bash
        env:
          LAB_MODE: ${{ inputs.mode }}
        run: |
          set -u
          set +e
          python - <<'PY' >check.log 2>&1
          import os, sys
          mode = os.environ["LAB_MODE"]
          print(f"synthetic_check mode={mode}")
          if mode == "fail":
              print("expected synthetic validation failure")
              sys.exit(17)
          print("synthetic validation passed")
          PY
          rc=$?
          set -e
          printf 'exit_code=%s\n' "$rc" >> "$GITHUB_OUTPUT"
          printf 'check_exit_code=%s\n' "$rc"

      - name: Annotate synthetic failure
        if: ${{ steps.check.outputs.exit_code != '0' }}
        shell: bash
        run: |
          echo "::error file=.github/workflows/observability-lab.yml,title=Synthetic validation::The disposable check failed; inspect the preserved evidence artifact."

      - name: Record debug state without credentials
        shell: bash
        env:
          IS_DEBUG: ${{ runner.debug }}
        run: |
          printf 'runner_debug=%s\n' "${IS_DEBUG:-0}" >> evidence.json.debug
          echo "::debug::Debug is enabled; only bounded non-secret metadata is emitted."

      - name: Upload machine-readable evidence
        if: ${{ always() }}
        uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
        with:
          name: observability-evidence-${{ github.run_id }}-${{ github.run_attempt }}
          path: |
            evidence.json
            evidence.json.debug
            check.log
          retention-days: 7

      - name: Final result gate
        if: ${{ always() }}
        shell: bash
        env:
          CHECK_RC: ${{ steps.check.outputs.exit_code }}
        run: |
          set -euo pipefail
          if [[ "${CHECK_RC:-99}" != "0" ]]; then
            echo "synthetic validation failed with exit code ${CHECK_RC:-unknown}" >&2
            exit "${CHECK_RC:-1}"
          fi

5. Run A — intentional failure

gh workflow run observability-lab.yml \
  -f mode=fail \
  -f 'sample=<img src=x onerror=alert(1)> [demo](https://example.invalid/)'

gh run list --workflow observability-lab.yml --limit 5 \
  --json databaseId,attempt,event,headSha,status,conclusion,url

FAIL_RUN_ID=REPLACE_WITH_EXACT_RUN_ID
gh run watch "$FAIL_RUN_ID" --exit-status || true

Verify the failed job, the fixed error annotation, the summary and the evidence artifact. Record the exact source SHA and artifact identity before rerun. The HTML-like sample is intentionally synthetic; it should appear only as escaped code in the summary and never become active HTML.

6. Preserve attempt 1 as immutable incident evidence

ROOT="checkpoint-evidence/run-$FAIL_RUN_ID"
mkdir -p "$ROOT/attempt-1"

gh run view "$FAIL_RUN_ID" --attempt 1 \
  --json attempt,conclusion,createdAt,headSha,jobs,url \
  > "$ROOT/attempt-1/run.json"

gh run view "$FAIL_RUN_ID" --attempt 1 --log \
  > "$ROOT/attempt-1/workflow.log"

gh run download "$FAIL_RUN_ID" \
  -p 'observability-evidence-*' \
  -D "$ROOT/attempt-1/artifacts"

# Record one API response without exposing auth headers.
gh api \
  -H 'Accept: application/vnd.github+json' \
  -H 'X-GitHub-Api-Version: 2026-03-10' \
  "repos/{owner}/{repo}/actions/runs/$FAIL_RUN_ID" \
  > "$ROOT/attempt-1/api-run.json"

sha256sum "$ROOT/attempt-1"/* 2>/dev/null || true

Do not rewrite these files after the rerun. If your host supports it, make the folder read-only or copy it to a separate lab archive. The goal is evidentiary continuity, not cryptographic non-repudiation.

7. Run A, attempt 2 — same source with debug

gh run rerun "$FAIL_RUN_ID" --failed --debug
gh run watch "$FAIL_RUN_ID" --exit-status || true

mkdir -p "$ROOT/attempt-2"
gh run view "$FAIL_RUN_ID" --attempt 2 \
  --json attempt,conclusion,createdAt,headSha,jobs,url \
  > "$ROOT/attempt-2/run.json"
gh run view "$FAIL_RUN_ID" --attempt 2 --log \
  > "$ROOT/attempt-2/workflow.log"

Your comparison must show: same run ID, same headSha, attempt 2 instead of 1, and the same synthetic failure. Note in the evidence packet that a full debug log archive can also contain runner-diagnostic-logs. Review it for sensitive content before copying or sharing.

8. Run B — corrected invocation

gh workflow run observability-lab.yml \
  -f mode=pass \
  -f 'sample=<img src=x onerror=alert(1)> [demo](https://example.invalid/)'

gh run list --workflow observability-lab.yml --limit 5 \
  --json databaseId,attempt,headSha,status,conclusion,url

PASS_RUN_ID=REPLACE_WITH_EXACT_SUCCESS_RUN_ID
gh run watch "$PASS_RUN_ID" --exit-status

gh run view "$PASS_RUN_ID" \
  --json attempt,conclusion,createdAt,headSha,jobs,url \
  > "checkpoint-evidence/run-$PASS_RUN_ID.json"

Run B should have its own run ID and attempt 1. If it is green but the artifact is missing, the checkpoint is not complete: job success and retained evidence are distinct acceptance conditions.

9. Build the run-to-evidence map

Execution Identity to record Expected conclusion Evidence
Run A / attempt 1 FAIL_RUN_ID + attempt 1 + SHA failure normal log, summary, annotation, artifact ID/digest/name
Run A / attempt 2 same FAIL_RUN_ID + attempt 2 + same SHA failure debug log + runner diagnostic availability + new attempt artifact
Run B / attempt 1 PASS_RUN_ID + attempt 1 + corrected invocation/SHA success summary + success artifact + final gate rc 0

Add timestamps, job names/database IDs, runner OS/arch, action SHAs and gh --version. If you use external telemetry, add only its opaque correlation/trace ID and approved status—not an entire provider payload.

10. Required evidence packet

  • checkpoint-predictions.md with predictions and independent verification.
  • Run A attempt 1 metadata, complete safe workflow log, artifact metadata/files and first-failure annotation description.
  • Run A attempt 2 metadata/log plus a note that debug was enabled and source identity remained unchanged.
  • Run B metadata and successful evidence artifact.
  • Exact workflow/action manifest: workflow SHA/ref, full action SHAs, runner label and requested Python version.
  • A table of step names and authoritative exit codes/conclusions.
  • A summary-safety note: what untrusted content was derived, escaped, bounded or withheld.
  • Retention/access assumptions and any external telemetry limitation.
  • A one-paragraph residual-risk note: GitHub-rendered/logged evidence is useful operational proof but not a substitute for artifact provenance/signature or provider-side audit state.

11. Verification checklist

  • Run A attempt 1 and attempt 2 have the same run ID and source SHA.
  • Attempt 1 evidence was captured before attempt 2 was created.
  • Attempt 2 used debug and did not magically repair a deterministic input failure.
  • Run B has a different run ID and a success conclusion.
  • No whole contexts, tokens, secrets, Authorization headers or OIDC tokens appear in the retained output.
  • The synthetic sample is bounded/escaped in the summary and absent from ordinary logs.
  • Failure annotation and non-zero gate are both present; the annotation is not the only failure mechanism.
  • Evidence artifacts are named by run ID/attempt and retained for a bounded lab duration.

12. Faithful local simulation

If GitHub Actions is temporarily unavailable, simulate the evidence contract locally. Set synthetic GITHUB_SHA, GITHUB_RUN_ID and GITHUB_RUN_ATTEMPT values, run the Python evidence/escaping snippets, execute the synthetic validator for fail/pass modes, and archive each result under separate attempt directories. Mark the result explicitly as a simulation: local files cannot reproduce GitHub annotations, run summaries, hosted runner diagnostics, run-attempt APIs or GitHub retention policy.

13. Cleanup and rollback

  • Do not delete Run A until the evidence review is complete.
  • Remove any persistent debug repository variables you created; the checkpoint does not require them.
  • Delete only the exact disposable repository after confirming evidence is no longer required.
  • If accidental sensitive data entered a completed summary, treat it as an incident; deleting the entire workflow run is required to remove uploaded summaries, and exposed credentials should be revoked/rotated rather than trusting deletion alone.

14. What Chapter 30 adds — and the bridge to Chapter 31

Chapter 30 adds an observability operating model: exact run/attempt correlation, bounded logs, safe summaries, actionable annotations, on-demand debug and runner-level evidence, all without turning diagnostics into uncontrolled output. Chapter 31 builds on these identifiers to control workflows programmatically with GitHub CLI, REST and GraphQL APIs—where exact resource guards, pagination, idempotency and asynchronous state become the next operational challenge.

Next lesson

GitHub CLI, REST/GraphQL APIs, Workflow Dispatch, and Automation Control: Core Concepts and Mental Model

Continue with the next lesson to build on the current concepts, evidence, security boundaries, and operational practices.

Knowledge check

Why did the checkpoint require a debug rerun that still fails?

Run B is green, but there is no evidence artifact. Is the checkpoint complete?

What proves attempt 2 did not run a source fix?

Why is the synthetic HTML-like input safe to use here?

A completed summary accidentally contains a credential. Is adding a mask in a later step enough?

Official references and version notes

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.