Chapter 28Lesson 04~320 minutes

Kubernetes Agent, GitOps Workflows, Cluster Access, Environments, and Deployment Integration: Diagnostics, Failure Modes, Security, and Performance

Diagnose over-broad agent authorization, unsafe production access, stale GitLab deployment records, residual credentials, Kubernetes RBAC denials, connectivity failures, and GitOps drift.

DiagnosticsRBACKASDriftCredential response

Learning objectives

  • Apply a layer-by-layer diagnostic sequence to GitLab-to-Kubernetes failures.
  • Diagnose over-broad ci_access and unsafe production access from untrusted refs/runners.
  • Separate GitLab environment state from actual Kubernetes/Flux convergence.
  • Respond correctly to leaked or residual agent credentials.
  • Interpret Kubernetes RBAC denial, KAS connectivity, and version/configuration failures without destroying evidence.
Availability baseline — verified 2026-08-22 against GitLab 19.3. The GitLab Agent for Kubernetes, agent REST API, agent registration/token management, and the basic CI/CD cluster workflow are available on Free, Premium, and Ultimate across GitLab.com, Self-Managed, and Dedicated. Project/group CI access and environment filters are Free-compatible. CI job impersonation and user impersonation require Premium/Ultimate. User-access integration is currently Beta. The Agent's old built-in pull-based gitops.manifest_projects functionality was removed in GitLab 17.0; current GitOps guidance uses Flux plus the GitLab Agent. Therefore the mandatory chapter path is manifest/identity simulation with no real cluster or paid feature; any live cluster exercise is optional and disposable.

1. Diagnostic sequence: preserve → scope → identity → transport → authorization → observed state

Do not begin by “reinstalling the agent.” Preserve the Git commit SHA, pipeline/job ID, environment, agent ID, configuration revision, runner identity, selected kubecontext name, exact kubectl error, and Flux/Kubernetes status. Then determine which layer is actually failing.

Layer Questions Useful evidence
GitLab scope Is the project/group/environment authorized? config.yaml SHA, project path, environment name
Job/ref/runner trust Should this pipeline be allowed near the cluster? pipeline source, protected status, runner ID/tags
KAS/agent transport Is agent connected and current? agent status/version, agentk logs, KAS/TLS error
Kubernetes identity Which user/group/service account is effective? kubectl auth can-i, audit identity where available
RBAC/policy What verb/resource/namespace is denied? Role/RoleBinding, admission response
Reconciliation Does observed state match desired revision? Flux Ready/revision, deployment image digest

2. Failure: a group authorization grants too much cluster access

Suppose ci_access.groups: company was chosen for convenience. A new project under company/experiments now receives the agent context. Nothing is “broken” technically—the scope is wrong.

# Too broad
ci_access:
  groups:
    - id: company

# Safer: narrow to the deployment project and environment
ci_access:
  projects:
    - id: company/platform/deploy
      environments: [staging]

Preserve which project unexpectedly gained access, then narrow the GitLab authorization. Independently verify the agent service account is not over-privileged. If a broad group is operationally necessary, add environment filters and Kubernetes RBAC/impersonation appropriate to the tier.

3. Failure: an unprotected branch or untrusted runner can reach production

The dangerous symptom may be a successful kubectl get from a feature branch. That is a policy failure even if no destructive operation occurred. Check project/group agent authorization, environment filters, branch protection, CI rules, runner protection/isolation, and Kubernetes-side permissions.

Do not test the blast radius by deploying to production. Use kubectl auth can-i and a disposable namespace/cluster. If a real production credential or agent token was exposed to an untrusted job, revoke/rotate it first.

4. Failure: GitLab says deployed, but the cluster drifted

A GitLab environment record is historical application state, not a live assertion. Diagnose with three identities:

# Sanitized example checks
# 1) GitLab: deployment commit/job/environment from UI/API
# 2) Flux: desired/applied revision
flux get kustomizations -A
# 3) Kubernetes: actual image identity
kubectl -n ch28-lab get deploy ch28-demo   -o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}'

If Flux reports the expected revision but the workload differs, inspect controller events/admission/mutation. If Flux is not Ready, do not “fix” the GitLab environment badge; repair the reconciliation cause.

5. Failure: agent token outlives the project or leaks

An agent token authenticates agentk to GitLab. GitLab returns the secret value only when the token is created, and an agent can have at most two active tokens. If the token leaks, revoke it using the Agent API/UI/glab, create a replacement only if the agent still belongs in service, update the cluster installation, and verify the old token returns revoked/not found behavior.

# Metadata only; do not output token secret
glab api projects/:id/cluster_agents/:agent_id/tokens   | jq '.[] | {id,name,status,last_used_at}'

# Destructive/security action: revoke only a known compromised/lab token
# glab cluster agent token revoke <agent-id> <token-id>

Deleting the GitLab project/agent is not sufficient cluster cleanup. Remove agentk secrets/deployments/namespace resources separately and verify they are gone.

6. Intentionally broken example: Kubernetes says Forbidden

Assume the lab identity may list Pods but not Secrets:

$ kubectl -n ch28-lab get secrets
Error from server (Forbidden): secrets is forbidden: User "gitlab:ci_job:4207" cannot list resource "secrets" in API group "" in the namespace "ch28-lab"

This is a successful security control, not an agent transport failure. Confirm with:

kubectl auth can-i list pods -n ch28-lab
# yes
kubectl auth can-i list secrets -n ch28-lab
# no

Repair only if the workload requirement genuinely needs that permission. Then change the narrow Role, not the agent connection or cluster-admin binding. Preserve the original denial in the evidence packet.

7. Connectivity and version failures belong to KAS/agent transport

Common examples include an agent that is not connected, a self-signed KAS certificate untrusted by the job, or incompatible agent versions. GitLab recommends that agentk match the GitLab major/minor version, with previous/next minor versions also supported. On Self-Managed/Dedicated, KAS endpoint/TLS configuration is part of platform administration, not project CI.

Symptom Likely layer Least-destructive first check
No agent connection agentk ↔ KAS transport Agent status, agentk logs, KAS address/TLS
x509: certificate signed by unknown authority Trust/TLS Correct CA trust; do not disable TLS verification globally
Kubecontext missing in job GitLab authorization/config propagation config SHA, project/environment match, wait propagation
Forbidden from Kubernetes RBAC/admission Effective identity + verb/resource/namespace

8. Failure: treating GitLab Agent as the GitOps reconciler

If a team writes new gitops.manifest_projects configuration expecting GitLab 19.3 to synchronize manifests, the design is obsolete. The built-in pull feature was removed in 17.0. Migrate the desired-state source to Flux, keep the Agent for GitLab connectivity/access/visibility, and record controller revision separately from agent connection status.

9. Reliability and cost where they are causal

Broad project/group authorization increases configuration and audit complexity. Excessive agents per cluster increase lifecycle work; GitLab recommends one agent per cluster in typical multi-tenant designs. Flux reconciliation intervals that are too aggressive can create unnecessary API/Git/registry traffic; intervals that are too slow increase convergence latency. Tune only after measuring controller, Git server, registry, and API behavior—not by disabling security controls to make deploys “faster.”

Knowledge check

A feature-branch job can reach a production agent. What should you inspect before Kubernetes RBAC?

A kubectl command returns Forbidden. Should you reinstall the agent?

What is the first response to a leaked real agent token?

GitLab environment says deployed but Flux is not Ready. Which state wins for cluster convergence?

What should replace legacy agent pull-based manifest synchronization?

10. Summary and next bridge

Most cluster-integration incidents become tractable when you refuse to collapse GitLab scope, runner/ref trust, KAS transport, Kubernetes identity/RBAC, and reconciliation into one “Kubernetes problem.” The checkpoint now combines these layers into one disposable operational proof.

Primary sources and version notes

These lessons were finalized against current official GitLab documentation on 2026-08-22. Kubernetes, Flux, agentk, glab, KAS, and cluster security evolve independently, so re-check current GitLab version/tier, supported agent/Kubernetes versions, CLI syntax, and cluster policy before production use.

Next lesson

Checkpoint Lab

Prove a complete disposable trust path with one allowed operation, one denied operation, cross-system identity evidence, and verified cleanup.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.