Checkpoint Lab — SAST, Dependency Scanning, Secret Detection, Container Scanning, DAST, IaC Scanning, and Security Gates
Scan a disposable vulnerable sample with two scanner classes, preserve findings, triage one synthetic issue, enforce a transparent gate, prove remediation, and assemble a security evidence packet without production access.
Learning objectives
Checkpoint objectives
- Scan a disposable vulnerable sample with Semgrep SAST and a synthetic-secret detector.
- Preserve scanner version, source SHA, raw findings, report checksums, and policy result.
- Triage the synthetic marker with a narrow, expiring exception while remediating the SAST finding in source.
- Prove the gate changes from fail to pass for explainable reasons rather than because evidence was deleted.
- Produce an evidence packet and bridge naturally to Chapter 26 supply-chain/SBOM evidence.
1. Checkpoint scenario
Capture CI_PIPELINE_SOURCE, ref,
CI_COMMIT_SHA, pipeline ID, and each scan/gate job ID
before triage so the first and second runs remain independently
attributable.
Use only the disposable glci-ch25-security-lab project.
The initial source has two findings: one Semgrep SAST finding for
dynamic eval and one synthetic training marker. The
scanner jobs preserve both reports. The gate initially fails. You
then remediate the SAST issue and add a narrowly scoped exception
for the synthetic marker. A second pipeline must pass while
preserving the original failed pipeline and its reports.
2. Current assumptions
| Item | Checkpoint assumption |
|---|---|
| GitLab tier | Free-compatible; no GitLab Dependency Scanning, DAST, or security-policy entitlement is required. |
| Runner | Linux container-capable runner for Semgrep path; otherwise local deterministic SAST simulation. |
| SAST tool | Semgrep OSS 1.176.0 for the documented pinned example. |
| Secret tool | Course-local Python detector v1.0 matching only the synthetic marker namespace. |
| Evidence retention | Seven days for lab artifacts unless your organization requires a different training horizon. |
| External side effects | None. No DAST/network target and no image push are required. |
3. Predict state changes before running
| Prediction | How to verify independently |
|---|---|
| Two scan jobs create separate raw JSON evidence. | Artifact inventory + JSON parse + SHA-256 files. |
| First gate fails because SAST + synthetic marker are both unexcepted. | Gate trace + exact counts from both reports. |
| Adding an exception does not modify the secret report. | Report checksum remains evidence; exception file is separate policy input. |
Fixing eval removes only the Semgrep finding.
|
Second Semgrep report has zero matching results; source diff shows safe parser. |
| Second gate passes because no unexcepted findings remain. | Gate trace names accepted exception and shows zero blocking findings. |
4. Exception schema for the synthetic marker
{
"exceptions": [
{
"kind": "synthetic_training_marker",
"path": "src/app.py",
"classification": "false_positive_training_data",
"owner": "lab-security-reviewer",
"rationale": "Marker contains no credential material and exists only for scanner training.",
"expires": "2026-09-30"
}
]
}
The exception is narrow by kind and path. In a real program, prefer the scanner/GitLab finding identifier when stable and include ticket/approval references according to policy. An expired exception must fail closed or require renewed review.
5. Checkpoint gate with explicit exception handling
# tools/security_gate.py
from pathlib import Path
from datetime import date
import json, sys
semgrep = json.loads(Path("security-reports/semgrep.json").read_text())
secrets = json.loads(Path("security-reports/secrets.json").read_text())
policy = json.loads(Path("security/exceptions.json").read_text()) if Path("security/exceptions.json").exists() else {"exceptions": []}
sast_findings = semgrep.get("results", [])
blocking_secrets = []
for finding in secrets.get("findings", []):
matched = None
for exc in policy["exceptions"]:
if exc.get("kind") == finding.get("kind") and exc.get("path") == finding.get("path"):
if date.fromisoformat(exc["expires"]) >= date.today():
matched = exc
break
if matched:
print("EXCEPTED", finding["id"], matched["classification"], matched["expires"])
else:
blocking_secrets.append(finding)
print("sast_blocking=", len(sast_findings))
print("secret_blocking=", len(blocking_secrets))
sys.exit(1 if sast_findings or blocking_secrets else 0)
6. Exact pipeline
stages: [scan, gate]
sast_semgrep:
stage: scan
image: semgrep/semgrep:1.176.0
script:
- mkdir -p security-reports
- semgrep --version | tee security-reports/semgrep-version.txt
- semgrep scan --config security/semgrep-rules.yml --json --output security-reports/semgrep.json src/
- printf 'source=%s pipeline=%s job=%s\n' "$CI_COMMIT_SHA" "$CI_PIPELINE_ID" "$CI_JOB_ID" > security-reports/sast-identity.txt
- sha256sum security-reports/* > security-reports/sast-SHA256SUMS
artifacts:
when: always
expire_in: 7 days
paths: [security-reports/]
secret_synthetic:
stage: scan
image: python:3.12-alpine
script:
- python tools/secret_scan.py
- printf 'source=%s pipeline=%s job=%s\n' "$CI_COMMIT_SHA" "$CI_PIPELINE_ID" "$CI_JOB_ID" > security-reports/secret-identity.txt
- sha256sum security-reports/* > security-reports/secret-SHA256SUMS
artifacts:
when: always
expire_in: 7 days
paths: [security-reports/]
security_gate:
stage: gate
image: python:3.12-alpine
needs:
- job: sast_semgrep
artifacts: true
- job: secret_synthetic
artifacts: true
script:
- python tools/security_gate.py
artifacts:
when: always
expire_in: 7 days
paths:
- security/exceptions.json
- security-reports/
7. First run: preserve the expected failure
Before editing anything, record the failed gate's pipeline/job IDs,
CI_COMMIT_SHA, both scanner job IDs, Semgrep version,
both raw report checksums, and gate output. Confirm that the Semgrep
report contains the synthetic eval finding and that the
secret report contains only the training marker hash/location—not a
usable credential.
python -m json.tool security-reports/semgrep.json >/dev/null
python -m json.tool security-reports/secrets.json >/dev/null
sha256sum security-reports/*.json
# Record counts without dumping environment variables or credentials.
python - <<'PY2'
import json
print('sast=', len(json.load(open('security-reports/semgrep.json'))['results']))
print('synthetic_secret=', len(json.load(open('security-reports/secrets.json'))['findings']))
PY2
8. Triage and remediate
Classify the synthetic marker as a training false positive and add
the narrow exception above. Do not except the SAST
finding. Replace dynamic eval with an intentionally
bounded parser suitable for the toy problem, for example accepting
only decimal integers:
# src/app.py -- repaired training version
def calculate(user_expression: str):
text = user_expression.strip()
if not text or not text.lstrip("-").isdigit():
raise ValueError("only decimal integers are accepted in this lab")
return int(text)
LAB_MARKER = "LAB_SECRET_EXAMPLE_NOT_CREDENTIAL"
Commit the remediation + exception as reviewable repository state. The new scan should show zero Semgrep matches for the course rule, one synthetic marker, one accepted non-expired exception, and zero blocking findings.
9. Faithful local simulation when Semgrep container execution is unavailable
Do not switch to a privileged runner just to complete the lesson.
Replace the Semgrep job with a deterministic Python source check
that reports the presence of the literal eval( call
into a raw JSON file, using the same gate/evidence contract. Mark
the packet as a simulation, not a GitLab SAST result. The
learning objective is scanner→report→gate→remediation traceability,
not a specific runtime.
10. Required evidence packet
| Evidence | Required contents |
|---|---|
| Source/config | Source/ref/SHA, pipeline ID, merged pipeline/config note, Semgrep rule file SHA. |
| Scanner identity | Semgrep 1.176.0 output or simulation identity; synthetic-secret scanner v1.0. |
| Reports | Raw Semgrep JSON, secret JSON, checksums, finding counts. |
| Policy | Gate script SHA, exception JSON, exception owner/rationale/expiry. |
| First failure | Original failed gate job ID/trace and counts before repair. |
| Remediation | Source diff removing dynamic eval; second report showing zero SAST matches. |
| Final outcome | Second pipeline/job IDs and passing gate output explaining the accepted synthetic exception. |
| Assumptions | Free/local path, no DAST/prod target, no real secrets, no paid policy UI used. |
11. Verification checklist
- Both reports parse as JSON and have recorded SHA-256 values.
- Scanner inputs and source SHA are explicit.
- First failed pipeline remains preserved.
- No real credential was created or printed.
- The SAST finding is remediated, not excepted.
- The synthetic marker exception is narrow and has an expiry.
- Second gate passes for an explainable policy reason.
- No production endpoint, registry write, cloud identity, or privileged runner was used.
12. Cleanup / rollback
Delete the disposable project/branch after evidence review or let its seven-day artifacts expire. If you added a temporary runner/tag only for the lab, remove that disposable association according to your runner policy. There are no external resources to destroy in the mandatory path.
Knowledge check
Why must the first failed gate remain preserved after remediation?
It proves the original findings and policy behavior existed and prevents troubleshooting from rewriting history.
Why is the SAST finding fixed while the synthetic secret is excepted?
The SAST issue represents intentionally unsafe code and has a direct repair. The secret marker is known training data, so a narrow expiring false-positive exception is the correct policy action.
What proves the second gate did not pass because reports were deleted?
The second raw reports/checksums are retained: Semgrep shows zero course-rule findings, the secret report still shows the marker, and the gate records the matching exception.
If Semgrep cannot run on the available runner, what should you avoid doing?
Do not switch to an untrusted privileged runner or weaken isolation. Use the faithful local simulation and label it clearly.
What will Chapter 26 add to this operating model?
SBOM generation, dependency lists, vulnerability-report/policy interpretation, and stronger supply-chain evidence that binds components and advisories to exact builds.
13. What Chapter 25 adds to the production operating model
You can now treat security scanning as a reproducible evidence chain rather than a dashboard icon: exact input and source → scanner/template/rules identity → raw/standard report → finding → explicit policy/exception → remediation → new evidence. This makes false positives reviewable, paid-tier boundaries honest, and production trust boundaries enforceable.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-12. Basic
GitLab SAST, pipeline Secret Detection, Container Scanning, and IaC
scanning are documented for Free/Premium/Ultimate. Current GitLab
Dependency Scanning and DAST are Ultimate. Security policy
enforcement is Ultimate. GitLab's current Dependency Scanning
direction is SBOM-based; the legacy Gemnasium-based implementation
was deprecated in GitLab 17.9 and is proposed for removal in GitLab
20.0. Stable security templates are recommended for production
workflows; Latest templates change more frequently.
AST_ENABLE_MR_PIPELINES="true" is the current
recommended switch for supported application-security scans in
merge-request pipelines. The checkpoint deliberately uses normal
artifacts for raw local scanner JSON. A GitLab security report
should be declared only when the producer emits the supported GitLab
security-report schema for that category/version.
- SAST — official reference.
- SAST analyzers — official reference.
- Secret detection — official reference.
- Secret detection configuration — official reference.
- Container scanning — official reference.
- Dependency scanning — official reference.
- Dependency scanning migration to SBOM — official reference.
- DAST — official reference.
- DAST browser analyzer — official reference.
- IaC scanning — official reference.
- Security configuration — official reference.
- Security policies — official reference.
- Security scanner integration and report schemas — official reference.
- CI/CD YAML syntax — official reference.
Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.