Chapter 25Lesson 05~230 minutes

Checkpoint Lab — SAST, Dependency Scanning, Secret Detection, Container Scanning, DAST, IaC Scanning, and Security Gates

Scan a disposable vulnerable sample with two scanner classes, preserve findings, triage one synthetic issue, enforce a transparent gate, prove remediation, and assemble a security evidence packet without production access.

CheckpointSecurity gateTriageRemediation evidenceCleanup

Learning objectives

Checkpoint objectives

  • Scan a disposable vulnerable sample with Semgrep SAST and a synthetic-secret detector.
  • Preserve scanner version, source SHA, raw findings, report checksums, and policy result.
  • Triage the synthetic marker with a narrow, expiring exception while remediating the SAST finding in source.
  • Prove the gate changes from fail to pass for explainable reasons rather than because evidence was deleted.
  • Produce an evidence packet and bridge naturally to Chapter 26 supply-chain/SBOM evidence.

1. Checkpoint scenario

Capture CI_PIPELINE_SOURCE, ref, CI_COMMIT_SHA, pipeline ID, and each scan/gate job ID before triage so the first and second runs remain independently attributable.

Use only the disposable glci-ch25-security-lab project. The initial source has two findings: one Semgrep SAST finding for dynamic eval and one synthetic training marker. The scanner jobs preserve both reports. The gate initially fails. You then remediate the SAST issue and add a narrowly scoped exception for the synthetic marker. A second pipeline must pass while preserving the original failed pipeline and its reports.

Scope guard: no real secrets, real vulnerable production application, production URL, cloud account, external registry write credential, or privileged runner is required. If your environment cannot run the Semgrep image, use the faithful local simulation described below instead of weakening runner isolation.

2. Current assumptions

Item Checkpoint assumption
GitLab tier Free-compatible; no GitLab Dependency Scanning, DAST, or security-policy entitlement is required.
Runner Linux container-capable runner for Semgrep path; otherwise local deterministic SAST simulation.
SAST tool Semgrep OSS 1.176.0 for the documented pinned example.
Secret tool Course-local Python detector v1.0 matching only the synthetic marker namespace.
Evidence retention Seven days for lab artifacts unless your organization requires a different training horizon.
External side effects None. No DAST/network target and no image push are required.

3. Predict state changes before running

Prediction How to verify independently
Two scan jobs create separate raw JSON evidence. Artifact inventory + JSON parse + SHA-256 files.
First gate fails because SAST + synthetic marker are both unexcepted. Gate trace + exact counts from both reports.
Adding an exception does not modify the secret report. Report checksum remains evidence; exception file is separate policy input.
Fixing eval removes only the Semgrep finding. Second Semgrep report has zero matching results; source diff shows safe parser.
Second gate passes because no unexcepted findings remain. Gate trace names accepted exception and shows zero blocking findings.

4. Exception schema for the synthetic marker

{
  "exceptions": [
    {
      "kind": "synthetic_training_marker",
      "path": "src/app.py",
      "classification": "false_positive_training_data",
      "owner": "lab-security-reviewer",
      "rationale": "Marker contains no credential material and exists only for scanner training.",
      "expires": "2026-09-30"
    }
  ]
}

The exception is narrow by kind and path. In a real program, prefer the scanner/GitLab finding identifier when stable and include ticket/approval references according to policy. An expired exception must fail closed or require renewed review.

5. Checkpoint gate with explicit exception handling

# tools/security_gate.py
from pathlib import Path
from datetime import date
import json, sys

semgrep = json.loads(Path("security-reports/semgrep.json").read_text())
secrets = json.loads(Path("security-reports/secrets.json").read_text())
policy = json.loads(Path("security/exceptions.json").read_text()) if Path("security/exceptions.json").exists() else {"exceptions": []}

sast_findings = semgrep.get("results", [])
blocking_secrets = []
for finding in secrets.get("findings", []):
    matched = None
    for exc in policy["exceptions"]:
        if exc.get("kind") == finding.get("kind") and exc.get("path") == finding.get("path"):
            if date.fromisoformat(exc["expires"]) >= date.today():
                matched = exc
                break
    if matched:
        print("EXCEPTED", finding["id"], matched["classification"], matched["expires"])
    else:
        blocking_secrets.append(finding)

print("sast_blocking=", len(sast_findings))
print("secret_blocking=", len(blocking_secrets))
sys.exit(1 if sast_findings or blocking_secrets else 0)

6. Exact pipeline

stages: [scan, gate]

sast_semgrep:
  stage: scan
  image: semgrep/semgrep:1.176.0
  script:
    - mkdir -p security-reports
    - semgrep --version | tee security-reports/semgrep-version.txt
    - semgrep scan --config security/semgrep-rules.yml --json --output security-reports/semgrep.json src/
    - printf 'source=%s pipeline=%s job=%s\n' "$CI_COMMIT_SHA" "$CI_PIPELINE_ID" "$CI_JOB_ID" > security-reports/sast-identity.txt
    - sha256sum security-reports/* > security-reports/sast-SHA256SUMS
  artifacts:
    when: always
    expire_in: 7 days
    paths: [security-reports/]

secret_synthetic:
  stage: scan
  image: python:3.12-alpine
  script:
    - python tools/secret_scan.py
    - printf 'source=%s pipeline=%s job=%s\n' "$CI_COMMIT_SHA" "$CI_PIPELINE_ID" "$CI_JOB_ID" > security-reports/secret-identity.txt
    - sha256sum security-reports/* > security-reports/secret-SHA256SUMS
  artifacts:
    when: always
    expire_in: 7 days
    paths: [security-reports/]

security_gate:
  stage: gate
  image: python:3.12-alpine
  needs:
    - job: sast_semgrep
      artifacts: true
    - job: secret_synthetic
      artifacts: true
  script:
    - python tools/security_gate.py
  artifacts:
    when: always
    expire_in: 7 days
    paths:
      - security/exceptions.json
      - security-reports/

7. First run: preserve the expected failure

Before editing anything, record the failed gate's pipeline/job IDs, CI_COMMIT_SHA, both scanner job IDs, Semgrep version, both raw report checksums, and gate output. Confirm that the Semgrep report contains the synthetic eval finding and that the secret report contains only the training marker hash/location—not a usable credential.

python -m json.tool security-reports/semgrep.json >/dev/null
python -m json.tool security-reports/secrets.json >/dev/null
sha256sum security-reports/*.json
# Record counts without dumping environment variables or credentials.
python - <<'PY2'
import json
print('sast=', len(json.load(open('security-reports/semgrep.json'))['results']))
print('synthetic_secret=', len(json.load(open('security-reports/secrets.json'))['findings']))
PY2

8. Triage and remediate

Classify the synthetic marker as a training false positive and add the narrow exception above. Do not except the SAST finding. Replace dynamic eval with an intentionally bounded parser suitable for the toy problem, for example accepting only decimal integers:

# src/app.py -- repaired training version

def calculate(user_expression: str):
    text = user_expression.strip()
    if not text or not text.lstrip("-").isdigit():
        raise ValueError("only decimal integers are accepted in this lab")
    return int(text)

LAB_MARKER = "LAB_SECRET_EXAMPLE_NOT_CREDENTIAL"

Commit the remediation + exception as reviewable repository state. The new scan should show zero Semgrep matches for the course rule, one synthetic marker, one accepted non-expired exception, and zero blocking findings.

9. Faithful local simulation when Semgrep container execution is unavailable

Do not switch to a privileged runner just to complete the lesson. Replace the Semgrep job with a deterministic Python source check that reports the presence of the literal eval( call into a raw JSON file, using the same gate/evidence contract. Mark the packet as a simulation, not a GitLab SAST result. The learning objective is scanner→report→gate→remediation traceability, not a specific runtime.

10. Required evidence packet

Evidence Required contents
Source/config Source/ref/SHA, pipeline ID, merged pipeline/config note, Semgrep rule file SHA.
Scanner identity Semgrep 1.176.0 output or simulation identity; synthetic-secret scanner v1.0.
Reports Raw Semgrep JSON, secret JSON, checksums, finding counts.
Policy Gate script SHA, exception JSON, exception owner/rationale/expiry.
First failure Original failed gate job ID/trace and counts before repair.
Remediation Source diff removing dynamic eval; second report showing zero SAST matches.
Final outcome Second pipeline/job IDs and passing gate output explaining the accepted synthetic exception.
Assumptions Free/local path, no DAST/prod target, no real secrets, no paid policy UI used.

11. Verification checklist

  • Both reports parse as JSON and have recorded SHA-256 values.
  • Scanner inputs and source SHA are explicit.
  • First failed pipeline remains preserved.
  • No real credential was created or printed.
  • The SAST finding is remediated, not excepted.
  • The synthetic marker exception is narrow and has an expiry.
  • Second gate passes for an explainable policy reason.
  • No production endpoint, registry write, cloud identity, or privileged runner was used.

12. Cleanup / rollback

Delete the disposable project/branch after evidence review or let its seven-day artifacts expire. If you added a temporary runner/tag only for the lab, remove that disposable association according to your runner policy. There are no external resources to destroy in the mandatory path.

Knowledge check

Why must the first failed gate remain preserved after remediation?

Why is the SAST finding fixed while the synthetic secret is excepted?

What proves the second gate did not pass because reports were deleted?

If Semgrep cannot run on the available runner, what should you avoid doing?

What will Chapter 26 add to this operating model?

13. What Chapter 25 adds to the production operating model

You can now treat security scanning as a reproducible evidence chain rather than a dashboard icon: exact input and source → scanner/template/rules identity → raw/standard report → finding → explicit policy/exception → remediation → new evidence. This makes false positives reviewable, paid-tier boundaries honest, and production trust boundaries enforceable.

Next chapter

SBOM Generation, Dependency Lists, Vulnerability Reports, Policy Evaluation, and Supply-Chain Evidence

Extend scanner evidence into software-supply-chain evidence: enumerate components, generate SBOMs, correlate vulnerability data, evaluate policy, and preserve provenance across the build.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-12. Basic GitLab SAST, pipeline Secret Detection, Container Scanning, and IaC scanning are documented for Free/Premium/Ultimate. Current GitLab Dependency Scanning and DAST are Ultimate. Security policy enforcement is Ultimate. GitLab's current Dependency Scanning direction is SBOM-based; the legacy Gemnasium-based implementation was deprecated in GitLab 17.9 and is proposed for removal in GitLab 20.0. Stable security templates are recommended for production workflows; Latest templates change more frequently. AST_ENABLE_MR_PIPELINES="true" is the current recommended switch for supported application-security scans in merge-request pipelines. The checkpoint deliberately uses normal artifacts for raw local scanner JSON. A GitLab security report should be declared only when the producer emits the supported GitLab security-report schema for that category/version.

Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.