Chapter 33Lesson 04~370 minutes

GitLab Duo and Agentic DevSecOps: AI Features, Security Remediation, Governance, and Usage Controls: Diagnostics, Failure Modes, Security, and Performance

Diagnose unsafe or unavailable AI workflows: plausible but wrong remediation, excessive tool authority, sensitive-data exposure, feature-status drift, runner/governance mismatch, and exhausted or disabled credits.

Failure AnalysisPrompt InjectionPermissionsCreditsDiagnostics

Learning objectives

  • Diagnose AI workflow failures by separating product availability, authorization, governance, execution, model quality, and verification failures.
  • Recognize prompt/context data exposure and excessive tool authority as security incidents or near misses rather than mere UX problems.
  • Handle background Runner governance differently from interactive approvals and avoid policies that assume a human prompt exists.
  • Detect product/version drift such as beta/experimental behavior, removed/renamed features, changed models, or credit-policy changes.
  • Use an evidence-first sequence that preserves session/audit/pipeline artifacts before applying the least destructive correction.
Availability and governance baseline — verified 2026-08-22 against GitLab 19.3. The failure examples are fixtures. Live troubleshooting must use the exact GitLab version and current feature documentation. Agent Platform features, audit reports, tool governance, MCP controls, model defaults, and billing rules can change rapidly; never copy a 19.3-era workaround into a later production environment without revalidation.

1. First identify which layer failed

“Duo failed” is too vague to troubleshoot. An unavailable button can be subscription or namespace configuration. A denied tool call can be correct governance. A flow stuck queued can be runner selection. A flow that executes but creates bad code is a model/verification problem. A rejected push can be normal protected-branch enforcement. Diagnose the layer before changing settings.

Symptom Likely layer to inspect first Wrong first reaction
Feature absent Offering/version/entitlement/namespace/preview Grant Maintainer or enable everything
Tool unavailable/denied Tool governance + role + action category Turn governance off
Flow queued Runner tag/executor/availability + compute limits Regenerate prompt repeatedly
API returns forbidden Execution identity/resource RBAC Give Owner globally
Patch fails tests Proposal/model/context correctness Bypass tests
Credits unavailable/exhausted Usage/billing/namespace state Enable on-demand spend without owner approval
MR cannot merge Protected branch/approval/security/pipeline policy Ask agent to force push/bypass

2. Failure mode: the remediation sounds plausible, so it gets merged

Language-model fluency can produce false confidence. Reuse the unsafe string-prefix path check from Lesson 2. A reviewer reading only the explanation may accept it; the negative sibling-directory test rejects it. The correction is not “use a better prompt.” It is to restore the verification gate and require evidence tied to the security invariant.

BROKEN REVIEW
AI says: "The patch prevents traversal by verifying the resolved path begins with the template directory."
Reviewer: "Looks good." -> MERGE

CORRECT REVIEW
Invariant: resolved child must be structurally inside resolved base.
Evidence: negative tests include ../secret and ../templates-backup/secret.
Result: string-prefix candidate FAILS.
Decision: reject proposal; retain failed evidence.
Never disable a failing test or scanner merely because the AI says the finding is a false positive. False-positive disposition itself needs evidence and the appropriate security role/process.

3. Failure mode: tool authority is broader than task intent

An issue says “update the README example.” The agent has repository read/write, command execution, project settings write, and delete tools on automatic allow. The mismatch is architectural: the capability envelope is wider than the user’s request. Prompt instructions such as “please do not delete anything” are not an authorization boundary.

Requested intent Needed authority Over-broad authority to remove
Edit README on branch Read repo, write one branch, Git commit Project settings, members, releases, deploy, delete
Explain CI failure Read job log/config Repository write, command execution, deployment
Create remediation MR Read finding/source, branch write, MR create Protected branch direct write, project delete, secret management

Correction: narrow tool visibility/governance, narrow GitLab role/resource scope, run in a feature branch, and let protected-resource policies remain authoritative.

4. Failure mode: sensitive code or vulnerability data is sent without policy review

The incident is not solved by deleting the chat message after the fact. First stop further transmission. If a credential was included, revoke/rotate it immediately. Preserve enough audit/session evidence to identify the affected feature/provider and scope, then follow organizational incident/privacy procedures. Review current GitLab data usage documentation for provider, caching, history, usage-logging, and deletion/retention behavior.

AI DATA INCIDENT TRIAGE
1. Stop the session/flow and prevent repeated transmission.
2. If credential/secret exposed -> revoke/rotate/disable FIRST.
3. Preserve sanitized session identifiers, time, feature, model/provider, namespace.
4. Determine exactly what context/tool output was transmitted.
5. Check current history/cache/logging/provider retention controls.
6. Notify security/privacy owner under policy.
7. Remediate context collection / tool / data-classification control.
8. Verify no further prohibited data is sent.

5. Failure mode: an unattended flow assumes Always Ask will pause for a human

Current GitLab tool-governance documentation explicitly notes that Runner/background access cannot use interactive Always Ask in the same way as Web/Local sessions. A policy design that depends on someone clicking Approve while a background flow runs is invalid. The safe patterns are: deny the high-impact tool in Runner context, or create a separate workflow boundary where a human approves an MR/environment/change and deterministic GitLab controls authorize the consequence.

# Conceptual policy intent, not a GitLab import schema.
web:
  edit_feature_branch: always_ask
  delete_release: always_ask
runner:
  edit_feature_branch: always_allow   # only for an explicitly approved flow
  delete_release: always_deny         # no interactive approver exists

6. Failure mode: beta/experimental behavior is documented as a stable contract

GitLab 19.3 contains a mix of GA, beta, and feature-flagged AI capabilities. For example, AI audit-event reporting is documented as beta, and 19.3 introduces beta MCP-server blocking behind a feature flag. A runbook that says “this control always exists and is on” will eventually be wrong.

Production documentation should record the verification date and exact version, link the feature page, state maturity, and describe a fallback. For an audit feature, the fallback may be ordinary GitLab audit events plus CI/MR/session evidence rather than no audit trail at all.

Feature-preview principle: beta/experimental enablement is a governance decision. Do not silently turn on preview features to make a tutorial work.

7. Failure mode: credits/model behavior changes and automation silently degrades

An AI workflow can fail because the namespace has no eligible credits, a usage pool is exhausted, on-demand usage is disabled, a feature changes from pre-release to GA billing, or the selected model/provider becomes unavailable. The recovery plan should preserve core delivery without AI.

Bad dependency Resilient alternative
Release depends on AI summary finishing Release uses deterministic changelog/evidence; AI summary is optional
Pipeline fix requires agent credits Expose logs and deterministic diagnostics; human can repair YAML
Security triage requires AI explanation Scanner finding + source + security runbook remain primary evidence
Merge requires AI review only Human/CODEOWNERS/required CI remain authoritative controls

Credits should have budget alerts/ownership, but do not design a system where exceeding an AI budget halts ordinary Git operations, required CI, or emergency remediation.

8. Failure mode: flow execution environment is misconfigured or unsafe

Current flow execution can involve CI/CD runners. A flow may stay queued if an expected runner/tag is absent, fail if its executor/network cannot reach required services, or become dangerous if you solve those problems by attaching an over-privileged persistent runner. Diagnose runner state the same way as ordinary CI: tags, executor, protected status, network, image, variables, queue, and logs.

SAFE DIAGNOSIS ORDER
AI session -> associated pipeline/job -> runner selection/tags -> executor support
-> network reachability -> image/setup -> job log -> tool/session result

DO NOT "FIX" BY:
- registering an untrusted shared runner on a production host
- mounting Docker socket / host secrets without explicit design
- dumping every environment variable to logs
- disabling protected-runner controls
- giving broad cloud credentials to the flow

9. Intentionally broken governance example

Consider this policy record:

{
  "background_runner": {
    "read_repository": "always_allow",
    "write_file": "always_ask",
    "delete_project": "always_ask"
  },
  "fallback_on_governance_error": "allow_all"
}

Two defects are immediate. First, background Runner execution cannot rely on interactive Always Ask for the human-in-the-loop behavior described for interactive sessions. Second, allow_all on a governance-resolution failure violates fail-closed expectations. A safer intent is to allow only explicitly preapproved background writes and deny destructive tools; if governance resolution fails, expose no tools.

{
  "background_runner": {
    "read_repository": "always_allow",
    "write_file": "always_deny",
    "delete_project": "always_deny"
  },
  "fallback_on_governance_error": "deny_all"
}

10. Evidence-first diagnostic sequence for AI/agent incidents

Use the same discipline as Chapter 32, but add AI-specific scope.

PRESERVE
  session/audit artifact IDs, MR/diff, pipeline/job IDs, tool decisions, usage record
IDENTIFY
  offering/version + namespace/project + actor + feature/maturity + model/provider
  + context/data class + requested tool + execution environment
INSPECT
  entitlement/credits + Duo/Agent settings + role/RBAC + tool governance
  + flow runner + MR/CI/security policy + prompt/context + test/scan evidence
CORRECT
  least destructive: deny/narrow tool, revert proposal, rotate leaked secret,
  restore runner policy, replenish/disable usage only with budget owner approval
VERIFY
  unsafe action did not execute, data exposure stopped, intended workflow passes,
  normal GitLab governance remains enforced, usage/audit evidence is complete

Knowledge check

An agent tool call is denied. Why should you not immediately weaken governance?

A proposed fix passes one happy-path test. What key evidence is still missing?

A background flow needs project deletion but the policy says Always Ask. What is wrong with the design?

What is the first response if a real credential appears in an AI prompt or trace?

Why should AI unavailability not break core delivery?

11. Lesson summary and bridge

  • Diagnose the specific failed layer instead of labeling every problem “Duo failure.”
  • Fluent explanations and successful tool execution do not prove a change is correct or authorized.
  • Prompt instructions cannot substitute for narrow tool/RBAC authority.
  • Background flows need deny/preapprove patterns rather than imaginary interactive approvals.
  • Feature maturity, model/provider, and credits are volatile dependencies; core delivery needs non-AI fallback.
  • AI incidents use the same evidence-first operational discipline as other GitLab incidents, plus context/provider/tool/usage evidence.

Next, the checkpoint combines these lessons into one governed workflow: classify data and tools, review competing remediation proposals, deny an unsafe action, create an approval record, and hand a verified result into the production capstone.

Primary sources and version notes

These lessons were finalized against current official GitLab documentation on 2026-08-22 with GitLab 19.3 as the release baseline. GitLab Duo and the Agent Platform are unusually version-volatile: feature maturity, model selection, credits, availability, governance behavior, UI placement, and data-processing details can change between releases. Re-check the documentation for the exact offering, subscription, add-on, and GitLab version you use. Examples intentionally avoid live credentials, AI calls, protected-resource mutations, and Runner changes. Production troubleshooting must preserve the organization’s security/privacy evidence requirements and current GitLab feature contracts.

Next lesson

Checkpoint Lab

Assemble the complete governance package and independently approve or reject synthetic AI work.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.