Chapter 23Lesson 04~175 minutes

CodeQL, Code Scanning, SARIF, Custom Queries, Autofix, and Security Gates: Diagnostics, Failure Modes, Security, and Performance

Code scanning can fail quietly through incomplete extraction, noisy rules, bad SARIF identity, stale baselines, or policy that optimizes for a green dashboard instead of useful evidence. This lesson engineers those failures and repairs the causal boundary instead of suppressing the symptom.

Build coverageSARIF categoriesDismissalsBaselinesDiagnostics

Learning objectives

  • Diagnose “green but incomplete” analyses by inspecting language/build/database coverage before changing gates.
  • Explain SARIF category replacement and fragmentation and repair the analysis identity rather than deleting evidence.
  • Distinguish fixed, dismissed, missing, and not-analyzed states and preserve dismissal rationale.
  • Prevent legacy findings and over-broad admission scripts from becoming permanent delivery deadlocks.
  • Review autofix, Actions permissions, logs, API errors, and rate/cost behavior only where they causally affect scanning.
Diagnostic order: preserve evidence → identify repository/account/ref/workflow/analysis/category scope → inspect feature state, permissions, tool status, logs and REST objects → choose the least destructive correction → rerun and verify. Do not start by deleting an analysis, dismissing alerts, or loosening the ruleset.

1. Failure: analysis succeeded but expected code was not represented

The most dangerous scanning failure is a successful run with incomplete coverage. Typical causes include an unsupported language, a compiled-language build mode that omits generated code, a manual build that never invokes one target, Kotlin introduced into a Java no-build configuration, or a path configuration that excludes security-sensitive code.

Preserve the run and tool-status evidence. Compare repository language/build expectations with the analysis database and logs. A “success” conclusion only means the configured workflow completed; it does not certify that every production compilation unit entered the database.

REPO=OWNER/REPO
# Analysis records identify tool, commit, category and environment.
gh api --paginate -H "X-GitHub-Api-Version: 2026-03-10" \
  "repos/$REPO/code-scanning/analyses?per_page=100" \
  --jq '.[] | {id,ref,commit_sha,tool:.tool.name,category,environment,created_at}'

# CodeQL database inventory is another coverage clue.
gh api -H "X-GitHub-Api-Version: 2026-03-10" \
  "repos/$REPO/code-scanning/codeql/databases" \
  --jq '.[] | {language,id,content_type,commit_oid,updated_at}' || true

Repair the extraction/build mode first. Only after the intended code is represented should you decide whether the query suite or gate needs changing.

2. Failure: SARIF categories overwrite or fragment evidence

Suppose a monorepo scanner uploads backend and frontend results using the same tool/category for the same commit. A later upload can replace the earlier result set. The opposite error is embedding a timestamp in every category, which creates a fresh logical analysis identity every run and prevents clean lifecycle tracking.

Current behavior: if a second SARIF file for the same commit uses the same tool/category, earlier results for that identity are overwritten. If a single workflow run attempts multiple uploads with the same tool/category, GitHub detects the misconfiguration and the run fails.

3. Intentionally broken example: duplicate category in one run

This fragment is intentionally wrong. Both uploads claim to be the same analysis identity inside one workflow run.

permissions:
  contents: read
  security-events: write
steps:
  - name: Upload backend results
    uses: github/codeql-action/upload-sarif@ff2f1c621b7f889edc0d3c761ac2e6a3f8cdb0dd
    with:
      sarif_file: backend.sarif
      category: monorepo-scan
  - name: Upload frontend results
    uses: github/codeql-action/upload-sarif@ff2f1c621b7f889edc0d3c761ac2e6a3f8cdb0dd
    with:
      sarif_file: frontend.sarif
      category: monorepo-scan  # WRONG: duplicate identity in one run

Interpret the failed workflow as an identity collision, not “SARIF is broken.” Preserve the failed logs. Repair by consolidating the result sets if they are one logical analysis, or use stable categories such as monorepo-backend and monorepo-frontend if they are independent slices. Never use random run IDs as categories merely to get past the error.

4. Failure: alert dismissed for convenience

A dismissal removes an alert from the active count but does not change source. GitHub records the dismissal reason and optional comment. If the team uses “false positive” to unblock a merge without proving why the modeled path is impossible, the dashboard improves while the engineering evidence degrades.

Preserve the alert JSON, code location, query help/data flow, current commit, and dismissal history. Reopen the alert if the prior decision is unsupported; fix the code or document a durable reason tied to architecture/tests. For temporary exceptions, your governance system should record owner and expiry even if the GitHub dismissal object itself does not enforce that expiry policy.

5. Failure: a gate turns legacy debt into a delivery deadlock

GitHub’s native code-scanning merge-protection rule evaluates required tools and configured thresholds with PR-specific limitations. Teams sometimes add a broader custom status script that simply asks “are there any open High alerts in the repository?” and fails every PR. That script silently changes the policy from “do not introduce severe risk” to “nobody can merge until all history is remediated.”

# BROKEN POLICY SHAPE — read-only illustration only.
# It ignores PR relevance and would make legacy backlog block unrelated changes.
COUNT=$(gh api --paginate -H "X-GitHub-Api-Version: 2026-03-10" \
  "repos/$REPO/code-scanning/alerts?state=open&severity=high&per_page=100" \
  --jq 'length' | awk '{s+=$1} END{print s+0}')
if [ "$COUNT" -gt 0 ]; then
  echo "Policy failure: any historical High alert blocks this PR"
  exit 1
fi

Repair the policy semantics: distinguish alerts introduced by the proposed diff from inherited backlog, define severity/precision thresholds, assign backlog ownership/SLA, and fail closed when required analysis is missing. Do not dismiss legacy alerts to satisfy a badly designed gate.

6. Failure: autofix silences the query but changes behavior

A generated fix can escape the data flow that triggered one query while introducing a functional regression, weakening authorization, changing error handling, or leaving a broader design flaw. Preserve the original alert and proposed diff. Review it against the threat model and application tests, then rerun the same analysis after the patch. For custom queries or third-party tools, do not assume agentic validation can prove the alert is resolved.

Autofix creation/commit operations are security-sensitive source mutations. The mandatory course path remains review-only; learners are never instructed to bulk-accept autofixes or apply them to valuable repositories.

7. Failure: scanner permissions or event context are wrong

SARIF upload from Actions requires security-events: write. A workflow can also fail because policy disables Actions, a fork PR cannot receive the same write authority, or a private repository lacks Code Security. Treat a 403 as evidence about authorization/feature state—not as proof that the SARIF is invalid.

# Read-only diagnostic requests with current API version.
gh api -i -H "X-GitHub-Api-Version: 2026-03-10" \
  "repos/$REPO/code-scanning/default-setup" || true
gh api -i -H "X-GitHub-Api-Version: 2026-03-10" \
  "repos/$REPO/code-scanning/alerts?per_page=1" || true

# Preserve failed workflow logs instead of immediately rerunning with broader permissions.
gh run view RUN_ID --repo "$REPO" --log-failed

Repair only the permission the operation actually requires. Do not grant broad repository write authority to a scanner because one upload step was denied.

8. Reliability, performance, and cost where they matter

Code scanning consumes compute and can lengthen PR feedback loops. Advanced manual builds may be substantially more expensive than no-build analysis. Security-extended/custom suites add query execution and triage load. External SARIF upload is rate-limited; the REST SARIF upload endpoint currently documents a 1,000-requests-per-hour limit per user/app installation.

Optimize after proving coverage: avoid redundant matrix dimensions, use the correct build mode, schedule deeper scans where appropriate, and measure queue/run time. A faster scan that silently excludes generated security-critical code is not an optimization. Public repositories can use code scanning without purchasing Code Security, but Actions runner consumption and agentic autofix billing still have their own product rules.

9. Diagnostic runbook

Step Evidence Question
1 Preserve Run URL/logs, alert JSON, analysis ID, commit/ref, category What exactly failed or disappeared?
2 Scope Repo visibility/plan, tool, language, build mode, event, branch Are we investigating the correct product/resource?
3 Authorize Feature state, Actions policy, security-events permission Is execution/upload allowed?
4 Coverage CodeQL database/tool status/build logs Was expected code actually analyzed?
5 Identity Tool GUID/name, category, fingerprints, analysis key Did result sets replace or fragment each other?
6 Policy Ruleset threshold, PR diff, required tool Is the gate expressing the intended risk rule?
7 Repair Smallest build/config/category/policy change Can we correct cause without deleting evidence?
8 Verify New analysis + alert state + PR policy state Did the same evidence chain become healthy?

Knowledge check

A CodeQL run succeeded but generated source was never built. Is the repository “clean”?

Two monorepo slices upload the same tool/category in one workflow and the run fails. What is the root cause?

What should happen first when an alert seems like a false positive?

Why can a “block if any repository High alert exists” script be harmful?

What does a 403 during SARIF upload tell you?

Summary

The reliable diagnostic unit is the whole evidence chain: revision, coverage, query/rule, SARIF identity, alert lifecycle, and policy. Fixing the nearest red message without that chain often produces false green state.

The checkpoint now asks you to operate that chain end-to-end and turn it into a durable gate policy.

Next lesson

Checkpoint Lab — CodeQL, Code Scanning, SARIF, Custom Queries, Autofix, and Security Gates

Official references

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.