CodeQL, Code Scanning, SARIF, Custom Queries, Autofix, and Security Gates: Diagnostics, Failure Modes, Security, and Performance
Code scanning can fail quietly through incomplete extraction, noisy rules, bad SARIF identity, stale baselines, or policy that optimizes for a green dashboard instead of useful evidence. This lesson engineers those failures and repairs the causal boundary instead of suppressing the symptom.
Learning objectives
- Diagnose “green but incomplete” analyses by inspecting language/build/database coverage before changing gates.
- Explain SARIF category replacement and fragmentation and repair the analysis identity rather than deleting evidence.
- Distinguish fixed, dismissed, missing, and not-analyzed states and preserve dismissal rationale.
- Prevent legacy findings and over-broad admission scripts from becoming permanent delivery deadlocks.
- Review autofix, Actions permissions, logs, API errors, and rate/cost behavior only where they causally affect scanning.
1. Failure: analysis succeeded but expected code was not represented
The most dangerous scanning failure is a successful run with incomplete coverage. Typical causes include an unsupported language, a compiled-language build mode that omits generated code, a manual build that never invokes one target, Kotlin introduced into a Java no-build configuration, or a path configuration that excludes security-sensitive code.
Preserve the run and tool-status evidence. Compare repository language/build expectations with the analysis database and logs. A “success” conclusion only means the configured workflow completed; it does not certify that every production compilation unit entered the database.
REPO=OWNER/REPO
# Analysis records identify tool, commit, category and environment.
gh api --paginate -H "X-GitHub-Api-Version: 2026-03-10" \
"repos/$REPO/code-scanning/analyses?per_page=100" \
--jq '.[] | {id,ref,commit_sha,tool:.tool.name,category,environment,created_at}'
# CodeQL database inventory is another coverage clue.
gh api -H "X-GitHub-Api-Version: 2026-03-10" \
"repos/$REPO/code-scanning/codeql/databases" \
--jq '.[] | {language,id,content_type,commit_oid,updated_at}' || true
Repair the extraction/build mode first. Only after the intended code is represented should you decide whether the query suite or gate needs changing.
2. Failure: SARIF categories overwrite or fragment evidence
Suppose a monorepo scanner uploads backend and frontend results using the same tool/category for the same commit. A later upload can replace the earlier result set. The opposite error is embedding a timestamp in every category, which creates a fresh logical analysis identity every run and prevents clean lifecycle tracking.
3. Intentionally broken example: duplicate category in one run
This fragment is intentionally wrong. Both uploads claim to be the same analysis identity inside one workflow run.
permissions:
contents: read
security-events: write
steps:
- name: Upload backend results
uses: github/codeql-action/upload-sarif@ff2f1c621b7f889edc0d3c761ac2e6a3f8cdb0dd
with:
sarif_file: backend.sarif
category: monorepo-scan
- name: Upload frontend results
uses: github/codeql-action/upload-sarif@ff2f1c621b7f889edc0d3c761ac2e6a3f8cdb0dd
with:
sarif_file: frontend.sarif
category: monorepo-scan # WRONG: duplicate identity in one run
Interpret the failed workflow as an identity collision, not
“SARIF is broken.” Preserve the failed logs. Repair by consolidating
the result sets if they are one logical analysis, or use stable
categories such as monorepo-backend and
monorepo-frontend if they are independent slices. Never
use random run IDs as categories merely to get past the error.
4. Failure: alert dismissed for convenience
A dismissal removes an alert from the active count but does not change source. GitHub records the dismissal reason and optional comment. If the team uses “false positive” to unblock a merge without proving why the modeled path is impossible, the dashboard improves while the engineering evidence degrades.
Preserve the alert JSON, code location, query help/data flow, current commit, and dismissal history. Reopen the alert if the prior decision is unsupported; fix the code or document a durable reason tied to architecture/tests. For temporary exceptions, your governance system should record owner and expiry even if the GitHub dismissal object itself does not enforce that expiry policy.
5. Failure: a gate turns legacy debt into a delivery deadlock
GitHub’s native code-scanning merge-protection rule evaluates required tools and configured thresholds with PR-specific limitations. Teams sometimes add a broader custom status script that simply asks “are there any open High alerts in the repository?” and fails every PR. That script silently changes the policy from “do not introduce severe risk” to “nobody can merge until all history is remediated.”
# BROKEN POLICY SHAPE — read-only illustration only.
# It ignores PR relevance and would make legacy backlog block unrelated changes.
COUNT=$(gh api --paginate -H "X-GitHub-Api-Version: 2026-03-10" \
"repos/$REPO/code-scanning/alerts?state=open&severity=high&per_page=100" \
--jq 'length' | awk '{s+=$1} END{print s+0}')
if [ "$COUNT" -gt 0 ]; then
echo "Policy failure: any historical High alert blocks this PR"
exit 1
fi
Repair the policy semantics: distinguish alerts introduced by the proposed diff from inherited backlog, define severity/precision thresholds, assign backlog ownership/SLA, and fail closed when required analysis is missing. Do not dismiss legacy alerts to satisfy a badly designed gate.
6. Failure: autofix silences the query but changes behavior
A generated fix can escape the data flow that triggered one query while introducing a functional regression, weakening authorization, changing error handling, or leaving a broader design flaw. Preserve the original alert and proposed diff. Review it against the threat model and application tests, then rerun the same analysis after the patch. For custom queries or third-party tools, do not assume agentic validation can prove the alert is resolved.
Autofix creation/commit operations are security-sensitive source mutations. The mandatory course path remains review-only; learners are never instructed to bulk-accept autofixes or apply them to valuable repositories.
7. Failure: scanner permissions or event context are wrong
SARIF upload from Actions requires
security-events: write. A workflow can also fail
because policy disables Actions, a fork PR cannot receive the same
write authority, or a private repository lacks Code Security. Treat
a 403 as evidence about authorization/feature state—not as proof
that the SARIF is invalid.
# Read-only diagnostic requests with current API version.
gh api -i -H "X-GitHub-Api-Version: 2026-03-10" \
"repos/$REPO/code-scanning/default-setup" || true
gh api -i -H "X-GitHub-Api-Version: 2026-03-10" \
"repos/$REPO/code-scanning/alerts?per_page=1" || true
# Preserve failed workflow logs instead of immediately rerunning with broader permissions.
gh run view RUN_ID --repo "$REPO" --log-failed
Repair only the permission the operation actually requires. Do not grant broad repository write authority to a scanner because one upload step was denied.
8. Reliability, performance, and cost where they matter
Code scanning consumes compute and can lengthen PR feedback loops. Advanced manual builds may be substantially more expensive than no-build analysis. Security-extended/custom suites add query execution and triage load. External SARIF upload is rate-limited; the REST SARIF upload endpoint currently documents a 1,000-requests-per-hour limit per user/app installation.
Optimize after proving coverage: avoid redundant matrix dimensions, use the correct build mode, schedule deeper scans where appropriate, and measure queue/run time. A faster scan that silently excludes generated security-critical code is not an optimization. Public repositories can use code scanning without purchasing Code Security, but Actions runner consumption and agentic autofix billing still have their own product rules.
9. Diagnostic runbook
| Step | Evidence | Question |
|---|---|---|
| 1 Preserve | Run URL/logs, alert JSON, analysis ID, commit/ref, category | What exactly failed or disappeared? |
| 2 Scope | Repo visibility/plan, tool, language, build mode, event, branch | Are we investigating the correct product/resource? |
| 3 Authorize |
Feature state, Actions policy,
security-events permission
|
Is execution/upload allowed? |
| 4 Coverage | CodeQL database/tool status/build logs | Was expected code actually analyzed? |
| 5 Identity | Tool GUID/name, category, fingerprints, analysis key | Did result sets replace or fragment each other? |
| 6 Policy | Ruleset threshold, PR diff, required tool | Is the gate expressing the intended risk rule? |
| 7 Repair | Smallest build/config/category/policy change | Can we correct cause without deleting evidence? |
| 8 Verify | New analysis + alert state + PR policy state | Did the same evidence chain become healthy? |
Knowledge check
A CodeQL run succeeded but generated source was never built. Is the repository “clean”?
No. The analysis may be incomplete. Fix database/build coverage before interpreting the absence of alerts.
Two monorepo slices upload the same tool/category in one workflow and the run fails. What is the root cause?
An analysis-identity collision. Consolidate the results or give the independent slices stable unique categories.
What should happen first when an alert seems like a false positive?
Preserve the finding/data flow and prove why the path is infeasible or harmless. Dismissal is a recorded governance decision, not a convenience button.
Why can a “block if any repository High alert exists” script be harmful?
It conflates inherited backlog with new-change risk and can deadlock unrelated delivery, encouraging suppression instead of remediation.
What does a 403 during SARIF upload tell you?
It points to authorization, event-policy, repository feature/entitlement, or repository state. It does not by itself prove malformed SARIF.
Summary
The reliable diagnostic unit is the whole evidence chain: revision, coverage, query/rule, SARIF identity, alert lifecycle, and policy. Fixing the nearest red message without that chain often produces false green state.
The checkpoint now asks you to operate that chain end-to-end and turn it into a durable gate policy.
Official references
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.