CodeQL, Code Scanning, SARIF, Custom Queries, Autofix, and Security Gates: Concepts, Architecture, and Mental Model
Dependency policy asks whether the components you consume are acceptable; code scanning asks what risky behavior may exist in the code you ship. This lesson separates the scanner, the analysis database, the query set, the SARIF evidence, the alert record, and the merge policy so a green badge is never mistaken for proof of security.
Learning objectives
- Explain the boundary among CodeQL extraction/databases, query execution, SARIF transport, GitHub code-scanning alerts, and merge policy.
- Choose default or advanced CodeQL setup based on build/language/control needs instead of habit.
- Interpret rule severity, security severity, alert state, location/data-flow evidence, and dismissal history without treating a finding as certainty.
- Explain what Copilot Autofix can and cannot prove, and why security gates must account for false positives and false negatives.
1. The problem: “scanner passed” is not a security model
Chapter 22 governed dependencies. Chapter 23 governs static-analysis evidence about your own code. The common mistake is to collapse many independent questions into one status: “CodeQL is green.” A trustworthy operating model asks what revision was analyzed, what languages and generated code were represented, which queries ran, which tool/category produced each result, how alerts were triaged, and what policy consumed that evidence.
Static analysis is deliberately incomplete. A tool can miss a vulnerability because the language, build, framework model, query, path, or runtime behavior is outside its model. It can also report a path that is technically possible but impossible in your application. The goal is therefore reproducible evidence plus accountable triage—not a claim of perfect detection.
2. Mental model: source → database/result → alert → decision
flowchart TD
C["Commit / PR revision"] -->|extract or analyze| D["CodeQL database or third-party analyzer"]
Q["Query suite / rules"] -->|executes against| D
D -->|findings encoded as SARIF| S["SARIF analysis identity"]
S -->|upload| A["GitHub code scanning alerts"]
A -->|triage + fix/dismiss| T["Alert state + audit evidence"]
A -->|thresholds| G["Ruleset code-scanning gate"]
G -->|admission decision| M["Merge allowed or blocked"]
The first arrow binds analysis to a Git revision. CodeQL creates a language-specific database representing the code it extracted; a third-party tool may produce SARIF directly instead. Query suites or rules define what patterns are searched. SARIF gives GitHub a structured interchange record. GitHub then creates alert objects, while a separate ruleset may turn selected alert thresholds into a merge decision. No arrow says “secure”: each arrow has a scope and failure mode.
3. CodeQL: database, query, query suite, and coverage
CodeQL models source code as data that queries can inspect. A CodeQL database represents one language for one analyzed source/build context. A query asks a security or quality question over that representation. A query suite selects a group of queries.
| Term | What it means | What can go wrong |
|---|---|---|
| Database | Extracted representation of a particular language/revision/build context. | Files that never enter extraction cannot be found by later queries. |
| Query | A program that searches the database for one pattern/data-flow condition. | A query can be inapplicable, noisy, or blind to framework-specific semantics. |
default suite |
High-precision built-in queries intended for routine code scanning. | Lower-noise does not mean complete coverage. |
security-extended |
Default plus additional, somewhat lower-precision security queries. | More findings can increase review cost and false positives. |
| Custom query / pack | Organization/project-specific query logic distributed in a CodeQL pack. | Requires ownership, tests, versioning, and advanced setup for custom suites. |
For compiled languages, database creation can depend on build mode.
Current CodeQL supports none, autobuild,
and manual in language-dependent combinations. “No
build” can be convenient but may miss generated code or infer
dependencies differently from the production build. The correct
question is not “did the workflow succeed?” but “did the database
represent the code we expected?”
4. Default setup versus advanced setup
Default setup is GitHub-managed configuration. GitHub detects CodeQL-supported languages, selects a build mode, runs on relevant branch/PR events and a schedule, and lets you choose built-in query suites without maintaining a workflow file. This is the recommended starting point for most eligible repositories.
Advanced setup places the CodeQL workflow/configuration under your control. Use it when you need manual build commands, non-default trigger logic, custom query packs/suites, private-registry preparation, special runner topology, or other settings that default setup cannot express. Greater control also means greater maintenance responsibility.
5. SARIF is the evidence envelope, not the analyzer
SARIF (Static Analysis Results Interchange Format) is a structured JSON format for static-analysis results. GitHub code scanning can ingest CodeQL results automatically or accept SARIF generated by compatible third-party tools. Important fields identify the tool, rule, message, source location, severity/security metadata, and fingerprints used to track a result across analyses.
Analysis identity matters. GitHub expects one result set per analysis identity for a commit. If the same tool/category is uploaded again, the new analysis represents the current result set for that identity; findings that disappear can be closed as fixed. If you need multiple independent analyses for one tool/commit, give them distinct categories. Reusing a category accidentally can replace earlier results; inventing a new category on every run fragments one logical stream into many unrelated streams.
6. Alert state, ordinary severity, and security severity
A code-scanning alert is GitHub’s persisted
representation of a rule result. The alert connects a rule to one or
more instances on refs/commits. Rule severity (for
example error/warning/note) and
security severity are different dimensions. SARIF
security rules can provide a numeric
security-severity score; GitHub maps 7.0–8.9 to High
and values over 9.0 to Critical.
An alert may remain open, become fixed because subsequent analysis no longer reports it, or be dismissed with an explicit reason/comment. Dismissal is a governance action, not a source-code fix. GitHub records the dismissal context; production policy should require a durable technical reason and an owner rather than “we needed the merge.”
7. Autofix is a proposed patch, not security proof
Copilot Autofix can generate a targeted suggestion for supported code-scanning alerts. It is currently available to public GitHub.com repositories that use CodeQL without requiring a separate GitHub Copilot subscription. Agentic autofix is a different, public-preview path that may use a cloud agent session and carries separate availability/billing considerations.
Treat an autofix like any other code change: review the data-flow reasoning, tests, behavior, compatibility, and whether the patch addresses the design cause or merely silences the query. A suggestion that makes one query disappear is evidence of one analysis result—not proof that the vulnerability class is impossible.
8. Security gates: policy consumes alerts
A repository ruleset can add Require code scanning results. For each required tool, the administrator chooses ordinary alert thresholds and security-alert thresholds such as Critical, High or higher, Medium or higher, or All. This gate is separate from required status checks and does not itself enable scanning.
Current merge protection also has limits: it does not apply to merge queue groups or Dependabot pull requests analyzed by default setup, and alert locations must exist in the pull-request diff for the protection to apply. That is why “turn on a gate” is not a complete policy. You still need ownership, baseline strategy, exception handling, and coverage monitoring.
9. Read-only inspection before any mutation
Before enabling code scanning in a real repository, inspect visibility, permissions, default-setup state, existing analyses, and alerts. A missing feature, an empty successful result, and a permission failure are different states.
REPO=OWNER/REPO
gh repo view "$REPO" --json nameWithOwner,visibility,viewerPermission,defaultBranchRef
# Current default-setup configuration (may report not configured on a fresh repo).
gh api -H "X-GitHub-Api-Version: 2026-03-10" \
"repos/$REPO/code-scanning/default-setup" || true
# Existing alert inventory; [] means the request succeeded and there are no matching alerts.
gh api --paginate \
-H "X-GitHub-Api-Version: 2026-03-10" \
"repos/$REPO/code-scanning/alerts?state=open&per_page=100" \
--jq '.[] | {number,state,rule:.rule.id,severity:.rule.severity,security:.rule.security_severity_level,tool:.tool.name,path:.most_recent_instance.location.path}' || true
The UI equivalent is the repository’s Security/code-scanning surface plus the CodeQL tool-status/configuration view. Use UI for human investigation and REST/CLI for repeatable evidence; do not scrape HTML when a documented object exists.
Knowledge check
What does a successful CodeQL workflow prove?
It proves that the configured analysis completed for its represented languages/build/query scope. It does not prove that every relevant file was modeled or that no vulnerability exists.
Why is SARIF category part of security evidence?
It distinguishes logical analysis streams. Reusing one category can replace prior results, while arbitrary new categories can fragment a single stream and confuse lifecycle tracking.
When should you prefer advanced setup?
When default setup cannot express required build commands, trigger/runner configuration, custom query packs/suites, registry preparation, or other repository-specific analysis needs.
Does dismissing an alert change the source code?
No. It changes the GitHub alert state and records a governance decision. A fixed state should instead be demonstrated by a new analysis where the finding is gone.
Why is a code-scanning ruleset not the same as a required status check?
GitHub implements code-scanning merge protection as a distinct ruleset rule that evaluates required tools and alert thresholds; it does not simply consume an arbitrary check name.
Summary
Code scanning is a pipeline: revision → extracted/analyzed representation → queries/rules → SARIF analysis identity → alert lifecycle → policy decision. Each boundary can be inspected independently, which makes false confidence and silent coverage gaps easier to detect.
Next, you will create that pipeline in a disposable public repository and prove a synthetic finding all the way from SARIF to alert state and merge governance.
Official references
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.