Kubernetes Agent, GitOps Workflows, Cluster Access, Environments, and Deployment Integration: Diagnostics, Failure Modes, Security, and Performance
Diagnose over-broad agent authorization, unsafe production access, stale GitLab deployment records, residual credentials, Kubernetes RBAC denials, connectivity failures, and GitOps drift.
Learning objectives
- Apply a layer-by-layer diagnostic sequence to GitLab-to-Kubernetes failures.
- Diagnose over-broad ci_access and unsafe production access from untrusted refs/runners.
- Separate GitLab environment state from actual Kubernetes/Flux convergence.
- Respond correctly to leaked or residual agent credentials.
- Interpret Kubernetes RBAC denial, KAS connectivity, and version/configuration failures without destroying evidence.
gitops.manifest_projects functionality was removed in
GitLab 17.0; current GitOps guidance uses
Flux plus the GitLab Agent. Therefore the mandatory
chapter path is manifest/identity simulation with no real cluster or
paid feature; any live cluster exercise is optional and disposable.
1. Diagnostic sequence: preserve → scope → identity → transport → authorization → observed state
Do not begin by “reinstalling the agent.” Preserve the Git commit
SHA, pipeline/job ID, environment, agent ID, configuration revision,
runner identity, selected kubecontext name, exact
kubectl error, and Flux/Kubernetes status. Then
determine which layer is actually failing.
| Layer | Questions | Useful evidence |
|---|---|---|
| GitLab scope | Is the project/group/environment authorized? |
config.yaml SHA, project path, environment name
|
| Job/ref/runner trust | Should this pipeline be allowed near the cluster? | pipeline source, protected status, runner ID/tags |
| KAS/agent transport | Is agent connected and current? | agent status/version, agentk logs, KAS/TLS error |
| Kubernetes identity | Which user/group/service account is effective? |
kubectl auth can-i, audit identity where
available
|
| RBAC/policy | What verb/resource/namespace is denied? | Role/RoleBinding, admission response |
| Reconciliation | Does observed state match desired revision? | Flux Ready/revision, deployment image digest |
2. Failure: a group authorization grants too much cluster access
Suppose ci_access.groups: company was chosen for
convenience. A new project under
company/experiments now receives the agent context.
Nothing is “broken” technically—the scope is wrong.
# Too broad
ci_access:
groups:
- id: company
# Safer: narrow to the deployment project and environment
ci_access:
projects:
- id: company/platform/deploy
environments: [staging]
Preserve which project unexpectedly gained access, then narrow the GitLab authorization. Independently verify the agent service account is not over-privileged. If a broad group is operationally necessary, add environment filters and Kubernetes RBAC/impersonation appropriate to the tier.
3. Failure: an unprotected branch or untrusted runner can reach production
The dangerous symptom may be a successful
kubectl get from a feature branch. That is a policy
failure even if no destructive operation occurred. Check
project/group agent authorization, environment filters, branch
protection, CI rules, runner protection/isolation, and
Kubernetes-side permissions.
kubectl auth can-i and a disposable
namespace/cluster. If a real production credential or agent token
was exposed to an untrusted job, revoke/rotate it first.
4. Failure: GitLab says deployed, but the cluster drifted
A GitLab environment record is historical application state, not a live assertion. Diagnose with three identities:
# Sanitized example checks
# 1) GitLab: deployment commit/job/environment from UI/API
# 2) Flux: desired/applied revision
flux get kustomizations -A
# 3) Kubernetes: actual image identity
kubectl -n ch28-lab get deploy ch28-demo -o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}'
If Flux reports the expected revision but the workload differs, inspect controller events/admission/mutation. If Flux is not Ready, do not “fix” the GitLab environment badge; repair the reconciliation cause.
5. Failure: agent token outlives the project or leaks
An agent token authenticates agentk to GitLab. GitLab returns the secret value only when the token is created, and an agent can have at most two active tokens. If the token leaks, revoke it using the Agent API/UI/glab, create a replacement only if the agent still belongs in service, update the cluster installation, and verify the old token returns revoked/not found behavior.
# Metadata only; do not output token secret
glab api projects/:id/cluster_agents/:agent_id/tokens | jq '.[] | {id,name,status,last_used_at}'
# Destructive/security action: revoke only a known compromised/lab token
# glab cluster agent token revoke <agent-id> <token-id>
Deleting the GitLab project/agent is not sufficient cluster cleanup. Remove agentk secrets/deployments/namespace resources separately and verify they are gone.
6. Intentionally broken example: Kubernetes says Forbidden
Assume the lab identity may list Pods but not Secrets:
$ kubectl -n ch28-lab get secrets
Error from server (Forbidden): secrets is forbidden: User "gitlab:ci_job:4207" cannot list resource "secrets" in API group "" in the namespace "ch28-lab"
This is a successful security control, not an agent transport failure. Confirm with:
kubectl auth can-i list pods -n ch28-lab
# yes
kubectl auth can-i list secrets -n ch28-lab
# no
Repair only if the workload requirement genuinely needs that permission. Then change the narrow Role, not the agent connection or cluster-admin binding. Preserve the original denial in the evidence packet.
7. Connectivity and version failures belong to KAS/agent transport
Common examples include an agent that is not connected, a self-signed KAS certificate untrusted by the job, or incompatible agent versions. GitLab recommends that agentk match the GitLab major/minor version, with previous/next minor versions also supported. On Self-Managed/Dedicated, KAS endpoint/TLS configuration is part of platform administration, not project CI.
| Symptom | Likely layer | Least-destructive first check |
|---|---|---|
| No agent connection | agentk ↔ KAS transport | Agent status, agentk logs, KAS address/TLS |
x509: certificate signed by unknown authority
|
Trust/TLS | Correct CA trust; do not disable TLS verification globally |
| Kubecontext missing in job | GitLab authorization/config propagation | config SHA, project/environment match, wait propagation |
Forbidden from Kubernetes |
RBAC/admission | Effective identity + verb/resource/namespace |
8. Failure: treating GitLab Agent as the GitOps reconciler
If a team writes new
gitops.manifest_projects configuration expecting GitLab
19.3 to synchronize manifests, the design is obsolete. The built-in
pull feature was removed in 17.0. Migrate the desired-state source
to Flux, keep the Agent for GitLab connectivity/access/visibility,
and record controller revision separately from agent connection
status.
9. Reliability and cost where they are causal
Broad project/group authorization increases configuration and audit complexity. Excessive agents per cluster increase lifecycle work; GitLab recommends one agent per cluster in typical multi-tenant designs. Flux reconciliation intervals that are too aggressive can create unnecessary API/Git/registry traffic; intervals that are too slow increase convergence latency. Tune only after measuring controller, Git server, registry, and API behavior—not by disabling security controls to make deploys “faster.”
Knowledge check
A feature-branch job can reach a production agent. What should you inspect before Kubernetes RBAC?
GitLab agent scope/environment filters, CI rules/protected refs, and runner trust. Kubernetes RBAC is another layer, not the only one.
A kubectl command returns Forbidden. Should you reinstall the agent?
No. A Kubernetes Forbidden response proves the transport reached the API; diagnose the effective identity and RBAC/admission policy.
What is the first response to a leaked real agent token?
Revoke/rotate the credential, then clean up leaked copies and verify residual access is gone.
GitLab environment says deployed but Flux is not Ready. Which state wins for cluster convergence?
The controller/cluster evidence. The GitLab environment record is not proof of current convergence.
What should replace legacy agent pull-based manifest synchronization?
Flux-based GitOps, with the GitLab Agent remaining as the integration/access bridge.
10. Summary and next bridge
Most cluster-integration incidents become tractable when you refuse to collapse GitLab scope, runner/ref trust, KAS transport, Kubernetes identity/RBAC, and reconciliation into one “Kubernetes problem.” The checkpoint now combines these layers into one disposable operational proof.
Primary sources and version notes
These lessons were finalized against current official GitLab documentation on 2026-08-22. Kubernetes, Flux, agentk, glab, KAS, and cluster security evolve independently, so re-check current GitLab version/tier, supported agent/Kubernetes versions, CLI syntax, and cluster policy before production use.
- GitLab 19.3 release
- Connecting a Kubernetes cluster with GitLab
- Install the GitLab Agent for Kubernetes
- Using GitLab CI/CD with a Kubernetes cluster
- Grant users Kubernetes access
- Using GitOps with a Kubernetes cluster
- Migrate legacy agent GitOps to Flux
- Managing agent instances and tokens
- Kubernetes Agent REST API
- Kubernetes managed resources and environments
- GitLab deprecations and removals
- GitLab token security overview
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.