Checkpoint Lab — CodeQL, Code Scanning, SARIF, Custom Queries, Autofix, and Security Gates
The checkpoint integrates analysis, SARIF identity, triage, remediation evidence, and merge governance. You will make one synthetic finding appear, prove which analysis created it, remove it with a clean rerun, and write a security-gate policy that treats new risk differently from inherited backlog.
Learning objectives
- Operate a two-analysis checkpoint: managed CodeQL plus a synthetic SARIF tool with explicit revision/category identity.
- Predict, create, triage, remove, and independently verify one training finding without dismissing it.
- Exercise a disposable code-scanning gate and prove whether the PR is blocked by evidence rather than workflow folklore.
- Write a production security-gate policy covering new versus existing findings, thresholds, dismissal authority, exception expiry, and missing analysis.
- Design an optional custom-query fixture and explain when it belongs in advanced setup.
1. Preflight: prove repository, analysis, and policy state
REPO=OWNER/code-scanning-evidence-lab
DEFAULT_BRANCH=$(gh repo view "$REPO" --json defaultBranchRef --jq .defaultBranchRef.name)
gh repo view "$REPO" --json nameWithOwner,visibility,viewerPermission,defaultBranchRef,isArchived
gh api -H "X-GitHub-Api-Version: 2026-03-10" "repos/$REPO/code-scanning/default-setup" || true
gh api --paginate -H "X-GitHub-Api-Version: 2026-03-10" \
"repos/$REPO/code-scanning/analyses?per_page=100" \
--jq '.[] | {id,tool:.tool.name,ref,commit_sha,category}'
gh api --paginate -H "X-GitHub-Api-Version: 2026-03-10" \
"repos/$REPO/code-scanning/alerts?per_page=100" \
--jq '.[] | {number,state,tool:.tool.name,rule:.rule.id,security:.rule.security_severity_level,fixed_at,dismissed_reason}'
If the repository is archived from Lesson 2, unarchive it only if you intentionally want to repeat the live checkpoint; otherwise create a fresh disposable repository. Never perform security-policy mutations on a production repository for course completion.
2. Ensure the checkpoint fixture is committed
The source remains harmless; the synthetic result is the training
signal. If needed, recreate src/app.js and the exact
pinned workflow below. The two action SHAs are immutable references
resolved at chapter generation time.
name: Academy SARIF training
on:
workflow_dispatch:
inputs:
result_set:
description: Synthetic analysis result
required: true
type: choice
options:
- finding
- clean
default: finding
pull_request:
branches:
- gate-target
permissions:
contents: read
security-events: write
jobs:
sarif:
runs-on: ubuntu-latest
steps:
- name: Checkout exact workflow revision
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
- name: Choose training result
id: choose
env:
MANUAL_MODE: ${{ inputs.result_set || 'auto' }}
shell: bash
run: |
mode="$MANUAL_MODE"
if [[ "$mode" == "auto" ]]; then
if [[ -f training/finding.flag ]]; then mode="finding"; else mode="clean"; fi
fi
echo "mode=$mode" >> "$GITHUB_OUTPUT"
- name: Generate valid SARIF fixture
env:
RESULT_MODE: ${{ steps.choose.outputs.mode }}
shell: bash
run: |
python - <<'PYFIX'
import json, os
finding = {
"ruleId": "ACADEMY/TRAIN001",
"level": "warning",
"message": {"text": "Synthetic training finding on harmless code; not a vulnerability claim."},
"partialFingerprints": {"primaryLocationLineHash": "academy-train-001-v1"},
"locations": [{"physicalLocation": {
"artifactLocation": {"uri": "src/app.js"},
"region": {"startLine": 1, "startColumn": 1}
}}]
}
doc = {
"$schema": "https://json.schemastore.org/sarif-2.1.0.json",
"version": "2.1.0",
"runs": [{
"tool": {"driver": {
"name": "Academy Training Static Analyzer",
"semanticVersion": "1.0.0",
"rules": [{
"id": "ACADEMY/TRAIN001",
"shortDescription": {"text": "Training-only security finding"},
"fullDescription": {"text": "A synthetic result used to teach SARIF identity and gating."},
"properties": {
"tags": ["security", "training"],
"precision": "very-high",
"security-severity": "8.0"
}
}]
}},
"results": [finding] if os.environ["RESULT_MODE"] == "finding" else []
}]
}
with open("results.sarif", "w", encoding="utf-8") as f:
json.dump(doc, f, indent=2)
print("Synthetic result mode:", os.environ["RESULT_MODE"])
PYFIX
- name: Upload SARIF
uses: github/codeql-action/upload-sarif@ff2f1c621b7f889edc0d3c761ac2e6a3f8cdb0dd # v4.37.7 line
with:
sarif_file: results.sarif
category: academy-training
Commit and push the fixture, record the resulting Git SHA, and confirm CodeQL default setup has analyzed JavaScript. This establishes the baseline “managed analyzer + external SARIF analyzer” model.
3. Predict and create one finding
Prediction: manual finding mode will create/open
ACADEMY/TRAIN001 in the
academy-training category for the workflow run’s head
SHA. It should be classified High because SARIF security score 8.0
maps to High.
gh workflow run academy-sarif.yml --repo "$REPO" --ref "$DEFAULT_BRANCH" -f result_set=finding
RUN_ID=$(gh run list --repo "$REPO" --workflow academy-sarif.yml --event workflow_dispatch \
--limit 1 --json databaseId --jq '.[0].databaseId')
gh run watch "$RUN_ID" --repo "$REPO" --exit-status
RUN_SHA=$(gh run view "$RUN_ID" --repo "$REPO" --json headSha --jq .headSha)
ALERT_JSON=$(gh api --paginate -H "X-GitHub-Api-Version: 2026-03-10" \
"repos/$REPO/code-scanning/alerts?state=open&per_page=100" \
--jq '[.[] | select(.tool.name=="Academy Training Static Analyzer" and .rule.id=="ACADEMY/TRAIN001")][0]')
printf '%s\n' "$ALERT_JSON" | jq '{number,state,rule:.rule.id,security:.rule.security_severity_level,sha:.most_recent_instance.commit_sha,category:.most_recent_instance.category}'
printf 'Workflow head SHA: %s\n' "$RUN_SHA"
Independently compare the workflow head SHA with the alert instance commit SHA. Open the alert and read its tool, rule, message, location, severity, and history. The message explicitly states that the finding is synthetic; if it does not, you are looking at the wrong alert.
4. Triage without convenience dismissal
Write a one-paragraph triage record: “TRAIN001 is a training-only
result produced by Academy Training Static Analyzer 1.0.0. It points
at harmless src/app.js and is intentionally synthetic.
It exists to test analysis identity/gating. The correct resolution
is a clean rerun of the same analyzer/category, not dismissal.”
This record demonstrates the distinction between “not a real vulnerability” and “dismiss this alert.” For production false positives, the durable record must include technical evidence showing why the data flow is infeasible or safely constrained.
5. Predict and test the PR gate
Create gate-target if it does not already exist. In the
repository ruleset, target only that branch and require
Academy Training Static Analyzer with Security
alerts High or higher. Prediction: a
same-repository PR carrying training/finding.flag will
be blocked after SARIF processing.
git fetch origin
if ! git show-ref --verify --quiet refs/remotes/origin/gate-target; then
git switch -c gate-target "origin/$DEFAULT_BRANCH"
git push -u origin gate-target
fi
git switch -c lab/checkpoint-finding origin/gate-target
mkdir -p training
printf 'training-only\n' > training/finding.flag
git add training/finding.flag
git commit -m "Checkpoint: produce synthetic finding"
git push -u origin HEAD
PR=$(gh pr create --repo "$REPO" --base gate-target --head lab/checkpoint-finding \
--title "Checkpoint: synthetic code scanning finding" \
--body "Training-only PR; expected to be blocked by High-or-higher code-scanning merge protection." | sed -n 's#.*/pull/##p')
echo "PR=$PR"
gh pr checks "$PR" --repo "$REPO" || true
gh pr view "$PR" --repo "$REPO" --json number,headRefOid,mergeStateStatus,statusCheckRollup
Inspect the merge box/ruleset result. If the analyzer run is still processing, “waiting for required analysis” is distinct from a High finding. If the workflow failed, diagnose the workflow before concluding that the rule blocked the PR for security reasons.
6. Remove the finding and prove the state transition
Remove the marker file and push a second commit. The same PR ref is reanalyzed with the same tool/category but an empty result set. Prediction: the training alert becomes fixed for that ref and the code-scanning gate is satisfied once processing completes.
git rm training/finding.flag
git commit -m "Checkpoint: clean synthetic analysis"
git push
FIX_SHA=$(git rev-parse HEAD)
gh pr checks "$PR" --repo "$REPO" --watch || true
gh api --paginate -H "X-GitHub-Api-Version: 2026-03-10" \
"repos/$REPO/code-scanning/alerts?per_page=100" \
--jq '.[] | select(.tool.name=="Academy Training Static Analyzer" and .rule.id=="ACADEMY/TRAIN001") | {number,state,fixed_at,dismissed_reason,sha:.most_recent_instance.commit_sha,category:.most_recent_instance.category}'
printf 'Fixed-branch SHA: %s\n' "$FIX_SHA"
Do not require the alert number to change. The important lifecycle evidence is that the logical alert/analysis stream no longer has an active finding for the new revision and that no dismissal reason was manufactured. Close the PR after verification; merging is not required.
7. Write the production security-gate policy
Create docs/code-scanning-policy.md in the disposable
repository or your notes. It must cover the following decisions in
complete sentences, with named owners/expiry concepts rather than
“security team decides.”
| Policy field | Minimum checkpoint decision |
|---|---|
| Coverage | Required languages/build mode/tool status; missing analysis fails closed for protected production branches. |
| New vs existing findings | New High/Critical security findings block; inherited backlog receives owners and remediation SLA rather than blanket dismissal. |
| Ordinary severity | Define whether error/warning quality findings block separately from security severity. |
| Dismissal authority | Named role/team may dismiss only with durable technical rationale; convenience/green-dashboard dismissals prohibited. |
| Exception expiry | Every temporary exception has owner, ticket/reference, compensating control, and explicit review/expiry date. |
| Autofix | Suggestion requires human review, relevant tests, and rerun of the same analysis; no bulk acceptance. |
| SARIF identity | Stable tool GUID/name/category policy per logical analysis slice; no timestamp categories. |
| Action/query versioning | Executable Actions/query packs use reviewed immutable/versioned references and controlled updates. |
| Failure behavior | Differentiate scanner failure, no supported language, permission denial, and clean zero-result analysis. |
A gate policy is production operating code even when expressed in prose. It defines which failures stop delivery and who can override them, so it needs change review and auditability.
8. Optional custom-query fixture: define the interface before writing QL
Do not make a custom CodeQL query mandatory for a beginner
checkpoint. Instead, design one safely. Suppose the fictional
application exposes dangerousOperation() and requires
every caller to pass through authorize(). Write a query
contract before implementation: target language, source/sink/guard
model, expected true-positive fixtures, expected negative fixtures,
rule ID, security severity, precision, owner, and pack version.
Optional query contract
Rule ID: academy/js/missing-authorization
Language: JavaScript/TypeScript
Purpose: flag calls to fictional dangerousOperation() not dominated by fictional authorize()
Fixtures: 2 positive, 2 negative
Precision target: high
Security severity: 8.0
Owner: Platform Security
Distribution: versioned CodeQL query pack
Execution: advanced setup only after query tests pass
Only then scaffold a CodeQL query pack and run it against a disposable database. The course does not pretend that a few lines of untested QL are suitable for organization-wide merge blocking.
9. Final verification and cleanup
- Record the CodeQL default-setup state and at least one CodeQL analysis commit SHA.
- Record the synthetic SARIF tool name, category, action SHA, alert number, High classification, and fixed state.
- Record the PR head SHA before and after removing the marker and the observed ruleset/merge state.
- Confirm no alert was dismissed and no valuable branch/history was rewritten.
- Close the training PR, remove/disable the disposable ruleset, and archive the repository after reviewing the evidence.
gh pr close "$PR" --repo "$REPO" --comment "Checkpoint complete; training repository retained only for evidence." || true
gh api --paginate -H "X-GitHub-Api-Version: 2026-03-10" \
"repos/$REPO/code-scanning/alerts?per_page=100" \
--jq '.[] | {number,state,tool:.tool.name,rule:.rule.id,fixed_at,dismissed_reason}'
# Remove/disable the chapter23-training-gate in Settings → Rules → Rulesets, then:
gh repo archive "$REPO" --yes
Knowledge check
The SARIF alert says High, but the source line is harmless. Is GitHub wrong?
No. GitHub is faithfully rendering the synthetic analyzer’s declared security score. Severity metadata is evidence from the tool, not an independent truth oracle.
Why does the checkpoint resolve the result with a clean rerun rather than a dismissal?
It demonstrates analysis lifecycle: the same tool/category no longer reports the result. Dismissal would demonstrate a human triage decision instead of remediation evidence.
A PR has no code-scanning result because the workflow failed. Should the policy treat that as clean?
No. Missing/failed analysis must be distinguished from a successful zero-result analysis; protected branches should fail closed according to policy.
Why must new findings and inherited backlog be governed differently?
Otherwise a legacy backlog can deadlock unrelated delivery or incentivize mass dismissal. New severe findings can be gated while backlog has explicit owners and SLA.
Who should be allowed to dismiss a production alert?
A policy-defined role/team with technical evidence and an auditable reason. The authority should not be granted merely to whoever needs the current PR merged.
What should happen before a custom query becomes a required organization-wide gate?
Define its contract, test true/false fixtures, version and review the query pack, measure precision/coverage, establish ownership, and roll it out in observation mode first.
10. Production operating model and Chapter 24 bridge
Chapter 23 adds a structured static-analysis control plane to your GitHub operating model: managed or explicit analysis configuration, reproducible tool/query identity, SARIF interoperability, evidence-preserving triage, remediation verification, and merge policy that distinguishes severe new risk from inherited debt. The critical habit is to prove what was analyzed before interpreting what was not found.
Chapter 24 changes the security object entirely: instead of code-flow vulnerabilities, you will govern credential/token material through secret scanning, push protection, custom patterns, bypass controls, and incident response. The response priority also changes—when a real credential is exposed, revoke/rotate it first; scanning/history cleanup comes afterward.
Official references
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.