Chapter 23Lesson 05~220 minutes

Checkpoint Lab — CodeQL, Code Scanning, SARIF, Custom Queries, Autofix, and Security Gates

The checkpoint integrates analysis, SARIF identity, triage, remediation evidence, and merge governance. You will make one synthetic finding appear, prove which analysis created it, remove it with a clean rerun, and write a security-gate policy that treats new risk differently from inherited backlog.

Checkpoint labTriageFixed stateGate policyChapter 24 bridge

Learning objectives

  • Operate a two-analysis checkpoint: managed CodeQL plus a synthetic SARIF tool with explicit revision/category identity.
  • Predict, create, triage, remove, and independently verify one training finding without dismissing it.
  • Exercise a disposable code-scanning gate and prove whether the PR is blocked by evidence rather than workflow folklore.
  • Write a production security-gate policy covering new versus existing findings, thresholds, dismissal authority, exception expiry, and missing analysis.
  • Design an optional custom-query fixture and explain when it belongs in advanced setup.
Checkpoint assumptions: GitHub.com public disposable repository, GitHub Free, repository admin permission, Actions enabled, standard hosted runners. If the repository from Lesson 2 still exists, reuse it; otherwise recreate the same safe source/workflow. No paid Code Security or Copilot subscription is required.
Prediction requirement: write your predictions before each mutation. At minimum predict (1) which exact tool/category/commit will own the synthetic alert, (2) whether the target PR will be blocked at High-or-higher, and (3) whether a clean rerun should produce fixed rather than dismissed state.

1. Preflight: prove repository, analysis, and policy state

REPO=OWNER/code-scanning-evidence-lab
DEFAULT_BRANCH=$(gh repo view "$REPO" --json defaultBranchRef --jq .defaultBranchRef.name)

gh repo view "$REPO" --json nameWithOwner,visibility,viewerPermission,defaultBranchRef,isArchived
gh api -H "X-GitHub-Api-Version: 2026-03-10" "repos/$REPO/code-scanning/default-setup" || true
gh api --paginate -H "X-GitHub-Api-Version: 2026-03-10" \
  "repos/$REPO/code-scanning/analyses?per_page=100" \
  --jq '.[] | {id,tool:.tool.name,ref,commit_sha,category}'
gh api --paginate -H "X-GitHub-Api-Version: 2026-03-10" \
  "repos/$REPO/code-scanning/alerts?per_page=100" \
  --jq '.[] | {number,state,tool:.tool.name,rule:.rule.id,security:.rule.security_severity_level,fixed_at,dismissed_reason}'

If the repository is archived from Lesson 2, unarchive it only if you intentionally want to repeat the live checkpoint; otherwise create a fresh disposable repository. Never perform security-policy mutations on a production repository for course completion.

2. Ensure the checkpoint fixture is committed

The source remains harmless; the synthetic result is the training signal. If needed, recreate src/app.js and the exact pinned workflow below. The two action SHAs are immutable references resolved at chapter generation time.

name: Academy SARIF training
on:
  workflow_dispatch:
    inputs:
      result_set:
        description: Synthetic analysis result
        required: true
        type: choice
        options:
          - finding
          - clean
        default: finding
  pull_request:
    branches:
      - gate-target
permissions:
  contents: read
  security-events: write
jobs:
  sarif:
    runs-on: ubuntu-latest
    steps:
      - name: Checkout exact workflow revision
        uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
        with:
          persist-credentials: false
      - name: Choose training result
        id: choose
        env:
          MANUAL_MODE: ${{ inputs.result_set || 'auto' }}
        shell: bash
        run: |
          mode="$MANUAL_MODE"
          if [[ "$mode" == "auto" ]]; then
            if [[ -f training/finding.flag ]]; then mode="finding"; else mode="clean"; fi
          fi
          echo "mode=$mode" >> "$GITHUB_OUTPUT"
      - name: Generate valid SARIF fixture
        env:
          RESULT_MODE: ${{ steps.choose.outputs.mode }}
        shell: bash
        run: |
          python - <<'PYFIX'
          import json, os
          finding = {
            "ruleId": "ACADEMY/TRAIN001",
            "level": "warning",
            "message": {"text": "Synthetic training finding on harmless code; not a vulnerability claim."},
            "partialFingerprints": {"primaryLocationLineHash": "academy-train-001-v1"},
            "locations": [{"physicalLocation": {
              "artifactLocation": {"uri": "src/app.js"},
              "region": {"startLine": 1, "startColumn": 1}
            }}]
          }
          doc = {
            "$schema": "https://json.schemastore.org/sarif-2.1.0.json",
            "version": "2.1.0",
            "runs": [{
              "tool": {"driver": {
                "name": "Academy Training Static Analyzer",
                "semanticVersion": "1.0.0",
                "rules": [{
                  "id": "ACADEMY/TRAIN001",
                  "shortDescription": {"text": "Training-only security finding"},
                  "fullDescription": {"text": "A synthetic result used to teach SARIF identity and gating."},
                  "properties": {
                    "tags": ["security", "training"],
                    "precision": "very-high",
                    "security-severity": "8.0"
                  }
                }]
              }},
              "results": [finding] if os.environ["RESULT_MODE"] == "finding" else []
            }]
          }
          with open("results.sarif", "w", encoding="utf-8") as f:
            json.dump(doc, f, indent=2)
          print("Synthetic result mode:", os.environ["RESULT_MODE"])
          PYFIX
      - name: Upload SARIF
        uses: github/codeql-action/upload-sarif@ff2f1c621b7f889edc0d3c761ac2e6a3f8cdb0dd # v4.37.7 line
        with:
          sarif_file: results.sarif
          category: academy-training

Commit and push the fixture, record the resulting Git SHA, and confirm CodeQL default setup has analyzed JavaScript. This establishes the baseline “managed analyzer + external SARIF analyzer” model.

3. Predict and create one finding

Prediction: manual finding mode will create/open ACADEMY/TRAIN001 in the academy-training category for the workflow run’s head SHA. It should be classified High because SARIF security score 8.0 maps to High.

gh workflow run academy-sarif.yml --repo "$REPO" --ref "$DEFAULT_BRANCH" -f result_set=finding
RUN_ID=$(gh run list --repo "$REPO" --workflow academy-sarif.yml --event workflow_dispatch \
  --limit 1 --json databaseId --jq '.[0].databaseId')
gh run watch "$RUN_ID" --repo "$REPO" --exit-status
RUN_SHA=$(gh run view "$RUN_ID" --repo "$REPO" --json headSha --jq .headSha)

ALERT_JSON=$(gh api --paginate -H "X-GitHub-Api-Version: 2026-03-10" \
  "repos/$REPO/code-scanning/alerts?state=open&per_page=100" \
  --jq '[.[] | select(.tool.name=="Academy Training Static Analyzer" and .rule.id=="ACADEMY/TRAIN001")][0]')
printf '%s\n' "$ALERT_JSON" | jq '{number,state,rule:.rule.id,security:.rule.security_severity_level,sha:.most_recent_instance.commit_sha,category:.most_recent_instance.category}'
printf 'Workflow head SHA: %s\n' "$RUN_SHA"

Independently compare the workflow head SHA with the alert instance commit SHA. Open the alert and read its tool, rule, message, location, severity, and history. The message explicitly states that the finding is synthetic; if it does not, you are looking at the wrong alert.

4. Triage without convenience dismissal

Write a one-paragraph triage record: “TRAIN001 is a training-only result produced by Academy Training Static Analyzer 1.0.0. It points at harmless src/app.js and is intentionally synthetic. It exists to test analysis identity/gating. The correct resolution is a clean rerun of the same analyzer/category, not dismissal.”

This record demonstrates the distinction between “not a real vulnerability” and “dismiss this alert.” For production false positives, the durable record must include technical evidence showing why the data flow is infeasible or safely constrained.

5. Predict and test the PR gate

Create gate-target if it does not already exist. In the repository ruleset, target only that branch and require Academy Training Static Analyzer with Security alerts High or higher. Prediction: a same-repository PR carrying training/finding.flag will be blocked after SARIF processing.

git fetch origin
if ! git show-ref --verify --quiet refs/remotes/origin/gate-target; then
  git switch -c gate-target "origin/$DEFAULT_BRANCH"
  git push -u origin gate-target
fi
git switch -c lab/checkpoint-finding origin/gate-target
mkdir -p training
printf 'training-only\n' > training/finding.flag
git add training/finding.flag
git commit -m "Checkpoint: produce synthetic finding"
git push -u origin HEAD
PR=$(gh pr create --repo "$REPO" --base gate-target --head lab/checkpoint-finding \
  --title "Checkpoint: synthetic code scanning finding" \
  --body "Training-only PR; expected to be blocked by High-or-higher code-scanning merge protection." | sed -n 's#.*/pull/##p')
echo "PR=$PR"
gh pr checks "$PR" --repo "$REPO" || true
gh pr view "$PR" --repo "$REPO" --json number,headRefOid,mergeStateStatus,statusCheckRollup

Inspect the merge box/ruleset result. If the analyzer run is still processing, “waiting for required analysis” is distinct from a High finding. If the workflow failed, diagnose the workflow before concluding that the rule blocked the PR for security reasons.

6. Remove the finding and prove the state transition

Remove the marker file and push a second commit. The same PR ref is reanalyzed with the same tool/category but an empty result set. Prediction: the training alert becomes fixed for that ref and the code-scanning gate is satisfied once processing completes.

git rm training/finding.flag
git commit -m "Checkpoint: clean synthetic analysis"
git push
FIX_SHA=$(git rev-parse HEAD)

gh pr checks "$PR" --repo "$REPO" --watch || true

gh api --paginate -H "X-GitHub-Api-Version: 2026-03-10" \
  "repos/$REPO/code-scanning/alerts?per_page=100" \
  --jq '.[] | select(.tool.name=="Academy Training Static Analyzer" and .rule.id=="ACADEMY/TRAIN001") | {number,state,fixed_at,dismissed_reason,sha:.most_recent_instance.commit_sha,category:.most_recent_instance.category}'
printf 'Fixed-branch SHA: %s\n' "$FIX_SHA"

Do not require the alert number to change. The important lifecycle evidence is that the logical alert/analysis stream no longer has an active finding for the new revision and that no dismissal reason was manufactured. Close the PR after verification; merging is not required.

7. Write the production security-gate policy

Create docs/code-scanning-policy.md in the disposable repository or your notes. It must cover the following decisions in complete sentences, with named owners/expiry concepts rather than “security team decides.”

Policy field Minimum checkpoint decision
Coverage Required languages/build mode/tool status; missing analysis fails closed for protected production branches.
New vs existing findings New High/Critical security findings block; inherited backlog receives owners and remediation SLA rather than blanket dismissal.
Ordinary severity Define whether error/warning quality findings block separately from security severity.
Dismissal authority Named role/team may dismiss only with durable technical rationale; convenience/green-dashboard dismissals prohibited.
Exception expiry Every temporary exception has owner, ticket/reference, compensating control, and explicit review/expiry date.
Autofix Suggestion requires human review, relevant tests, and rerun of the same analysis; no bulk acceptance.
SARIF identity Stable tool GUID/name/category policy per logical analysis slice; no timestamp categories.
Action/query versioning Executable Actions/query packs use reviewed immutable/versioned references and controlled updates.
Failure behavior Differentiate scanner failure, no supported language, permission denial, and clean zero-result analysis.

A gate policy is production operating code even when expressed in prose. It defines which failures stop delivery and who can override them, so it needs change review and auditability.

8. Optional custom-query fixture: define the interface before writing QL

Do not make a custom CodeQL query mandatory for a beginner checkpoint. Instead, design one safely. Suppose the fictional application exposes dangerousOperation() and requires every caller to pass through authorize(). Write a query contract before implementation: target language, source/sink/guard model, expected true-positive fixtures, expected negative fixtures, rule ID, security severity, precision, owner, and pack version.

Optional query contract
Rule ID: academy/js/missing-authorization
Language: JavaScript/TypeScript
Purpose: flag calls to fictional dangerousOperation() not dominated by fictional authorize()
Fixtures: 2 positive, 2 negative
Precision target: high
Security severity: 8.0
Owner: Platform Security
Distribution: versioned CodeQL query pack
Execution: advanced setup only after query tests pass

Only then scaffold a CodeQL query pack and run it against a disposable database. The course does not pretend that a few lines of untested QL are suitable for organization-wide merge blocking.

9. Final verification and cleanup

  • Record the CodeQL default-setup state and at least one CodeQL analysis commit SHA.
  • Record the synthetic SARIF tool name, category, action SHA, alert number, High classification, and fixed state.
  • Record the PR head SHA before and after removing the marker and the observed ruleset/merge state.
  • Confirm no alert was dismissed and no valuable branch/history was rewritten.
  • Close the training PR, remove/disable the disposable ruleset, and archive the repository after reviewing the evidence.
gh pr close "$PR" --repo "$REPO" --comment "Checkpoint complete; training repository retained only for evidence." || true
gh api --paginate -H "X-GitHub-Api-Version: 2026-03-10" \
  "repos/$REPO/code-scanning/alerts?per_page=100" \
  --jq '.[] | {number,state,tool:.tool.name,rule:.rule.id,fixed_at,dismissed_reason}'

# Remove/disable the chapter23-training-gate in Settings → Rules → Rulesets, then:
gh repo archive "$REPO" --yes
Cleanup is not “erase the security history.” Do not delete analyses, rewrite Git history, or bulk-dismiss alerts to make the lab disappear. Archive the disposable repository after removing the training-only enforcement rule.

Knowledge check

The SARIF alert says High, but the source line is harmless. Is GitHub wrong?

Why does the checkpoint resolve the result with a clean rerun rather than a dismissal?

A PR has no code-scanning result because the workflow failed. Should the policy treat that as clean?

Why must new findings and inherited backlog be governed differently?

Who should be allowed to dismiss a production alert?

What should happen before a custom query becomes a required organization-wide gate?

10. Production operating model and Chapter 24 bridge

Chapter 23 adds a structured static-analysis control plane to your GitHub operating model: managed or explicit analysis configuration, reproducible tool/query identity, SARIF interoperability, evidence-preserving triage, remediation verification, and merge policy that distinguishes severe new risk from inherited debt. The critical habit is to prove what was analyzed before interpreting what was not found.

Chapter 24 changes the security object entirely: instead of code-flow vulnerabilities, you will govern credential/token material through secret scanning, push protection, custom patterns, bypass controls, and incident response. The response priority also changes—when a real credential is exposed, revoke/rotate it first; scanning/history cleanup comes afterward.

Next chapter

Secret Scanning, Push Protection, Custom Patterns, Bypass Controls, and Incident Response: Concepts, Architecture, and Mental Model

Official references

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.