Kubernetes Agent, Cluster Access, GitOps-Oriented Delivery, Kubernetes Deployments, and Environment Integration: Diagnostics, Failure Modes, Security, and Performance
Diagnose over-broad cluster access, namespace collisions, wrong contexts, premature green jobs, image drift, deprecated integrations, and rollout failures from preserved evidence.
Learning objectives
- Diagnose Kubernetes delivery failures by separating pipeline compilation, Agent/context selection, RBAC, manifest, runtime, environment and external-health layers.
- Preserve original pipeline/job/rollout evidence before correction.
- Repair over-broad access and mutable-image failures without masking the original cause.
- Recognize deprecated certificate-based and legacy Agent GitOps guidance.
- Retry only the smallest safe scope after proving the failed layer.
1. Evidence-first diagnostic sequence
- Preserve pipeline/job IDs, first failing trace and any GitLab deployment ID.
-
Confirm
CI_PIPELINE_SOURCE, ref,CI_COMMIT_SHAand compiled configuration. - Confirm workflow/job-rule decisions and non-secret effective inputs.
- Inspect job graph, queue, runner/executor/image/tool versions.
-
Inspect Agent context name and
kubectl auth can-i; never expose kubeconfig credentials. - Inspect manifest/image digest and the command/tool/network failure.
- Inspect GitLab environment/deployment record.
- Inspect Kubernetes Deployment/ReplicaSet/Pod/Service and runtime imageID.
- Verify the external endpoint independently.
- Apply the least destructive correction; rerun only the failed safe scope.
2. Failure-layer map
| Symptom | Likely layer | First proof | Do not do |
|---|---|---|---|
| No deploy job exists | Rules / compiled config | Merged YAML and job-rule result | Change cluster RBAC |
| Expected Agent context absent | Agent ci_access / propagation |
Context list plus Agent config revision | Paste a production kubeconfig variable |
| Context exists, API says forbidden | Kubernetes RBAC / impersonation | Exact auth can-i and denial |
Grant cluster-admin |
| Apply succeeds, rollout times out | Workload/runtime | Deployment conditions, Events, Pods | Mark job green because apply returned 0 |
| Rollout is Ready, wrong code runs | Artifact identity/tag drift | Manifest image + runtime imageID + producer digest | Blindly restart Pods |
| GitLab environment looks green, endpoint fails | External/application health | Independent HTTP/service check | Rewrite deployment history |
3. Failure: production cluster credential stored in a variable
Broken pattern: a generic project variable contains a kubeconfig with cluster-admin privileges and every branch can run the deployment job. This combines long-lived credential exposure, broad source trust and broad cluster privilege.
4. Failure: Agent authorizes too many projects
Suppose the Agent config grants a whole group while only one deployer project needs access:
ci_access:
groups:
- id: platform
Before changing it, list which projects depend on that connection and preserve the current Agent config commit. Then narrow to explicit projects/environments:
ci_access:
projects:
- id: platform/deployer
environments:
- staging
- production
Authorization changes can take time to propagate. Verify context availability and denials before retrying.
5. Failure: namespace collision
Two projects both deploy deployment/web into namespace
staging. One pipeline patches the other team’s
workload. The fix is not a blind rename in the running cluster.
Preserve resource UIDs, owner labels, pipeline/job IDs and manifest
revisions; then move to deterministic project-owned namespaces or
stronger naming/ownership conventions plus RBAC.
kubectl -n staging get deployment/web -o json > evidence/web-before.json
kubectl -n staging get rolebinding -o yaml > evidence/rolebindings-before.yaml
kubectl -n staging get events --sort-by=.lastTimestamp > evidence/events-before.txt
6. Intentionally broken example: green job before rollout health
deploy_staging:
stage: deploy
script:
- kubectl config use-context "$KUBE_CONTEXT"
- kubectl apply -f k8s/staging.yaml
environment:
name: staging
This job can succeed as soon as the Kubernetes API accepts the manifest. It does not wait for Pods to become Ready. Preserve the successful job/deployment record, then inspect the workload. Repair by adding an explicit bounded rollout/health check:
script:
- kubectl config use-context "$KUBE_CONTEXT"
- kubectl apply -f k8s/staging.yaml
- kubectl rollout status deployment/example -n app-staging --timeout=180s
- ./verify-staging-health.sh
If rollout fails, do not replace the original job trace. The failed rollout is the causal evidence.
7. Failure: image tag drifts
A manifest uses example/app:stable. Pipeline 410
deploys it successfully. Later the tag moves. A restarted Pod now
pulls different bytes even though no manifest commit changed.
Diagnose by comparing the original registry digest, manifest
reference and current Pod imageID. Correct the manifest
to the verified digest and preserve the old runtime evidence.
kubectl -n app-staging get deployment/example -o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}'
kubectl -n app-staging get pods -l app=example -o jsonpath='{range .items[*]}{.metadata.name}{" "}{.status.containerStatuses[0].imageID}{"\n"}{end}'
cat evidence/approved-image-digest.txt
The safe correction is to promote the already-approved digest, not rebuild the image during troubleshooting.
8. Failure: correct permissions, wrong cluster context
A job is authorized to two Agents and relies on whichever context is current. The script succeeds against staging while the GitLab environment says production—or vice versa. Make context selection an explicit reviewed variable and record it before mutation:
kubectl config get-contexts -o name
printf 'selected_context=%s\n' "$KUBE_CONTEXT"
kubectl config use-context "$KUBE_CONTEXT"
kubectl config current-context
kubectl get namespace "$TARGET_NAMESPACE" -o jsonpath='{.metadata.uid}{"\n"}'
Do not infer cluster identity only from a namespace name, because the same namespace name can exist in many clusters.
9. Failure: template copied from deprecated certificate-based integration
Symptoms include advice to configure project/group/instance cluster
certificates, a gitlab-deploy context assumption, or
legacy built-in Agent gitops: manifest_projects.
Preserve the source/template version, then migrate to the Agent
CI/CD workflow or Flux. Do not add feature flags to resurrect
deprecated integration in a new design.
10. Runtime evidence without destructive debugging
kubectl -n app-staging get deployment,replicaset,pod,service -o wide
kubectl -n app-staging describe deployment/example
kubectl -n app-staging get events --sort-by=.lastTimestamp
kubectl -n app-staging logs deployment/example --tail=200
kubectl -n app-staging get pod -l app=example -o jsonpath='{range .items[*]}{.metadata.name}{" "}{.status.phase}{" "}{.status.containerStatuses[0].imageID}{"\n"}{end}'
These commands preserve observable state. Avoid deleting Pods merely to see whether they come back; that destroys first-failure evidence and can create side effects before you know the cause.
11. Security-sensitive actions
- Changing Agent project/group authorization is an authorization change: review and record it.
- Changing an Agent service-account Role/ClusterRole changes external privilege: use least privilege and preserve before/after RBAC.
- Rotating a leaked kubeconfig or Agent token is incident-response work: do not post token material into tickets or logs.
- Changing production Flux sources/reconciliation targets can deploy broadly: use reviewed exact revisions.
- Never troubleshoot by disabling TLS verification, granting cluster-admin, bypassing protected environments or using an unreviewed mutable image.
12. Performance and reliability: avoid making the control plane the bottleneck
Large pipelines can overload an API server with many parallel
imperative jobs. Serialize conflicting deployments (Chapter 18),
prefer controller reconciliation for production desired state, bound
kubectl rollout status timeouts, and keep manifests
small enough to review. A queued runner is not a slow cluster;
distinguish job queue time, API request latency, reconciliation time
and application readiness.
13. Least-destructive recovery map
| Failure | Smallest safe correction |
|---|---|
| Agent context absent |
Correct ci_access/environment authorization and
wait for propagation; do not add static production
credentials.
|
| RBAC denied | Add only required resource/verb/namespace permission; rerun the single deployment job after proof. |
| Bad manifest | Commit/review corrected desired state or revert the exact manifest revision; do not patch production ad hoc unless emergency procedure requires it. |
| Rollout failed | Preserve Events/Pod logs, then fix application/config/resource issue; use rollout undo only when the exact previous revision is understood. |
| Image drift | Restore verified digest reference; do not rebuild from source as a “rollback”. |
| External health failed | Keep deployment record; diagnose service/network/dependency layer separately. |
Knowledge check
The Agent context is present but kubectl returns Forbidden. Which layer failed?
Kubernetes authorization. Preserve the Agent/context evidence, then inspect the effective ServiceAccount/impersonation identity and RBAC rather than changing pipeline rules.
Why is adding rollout status to the job important?
kubectl apply only proves the API accepted desired state. rollout status makes controller progress a separate bounded condition before the job reports success.
A restarted Pod runs a different imageID with no manifest commit. What is the most likely cause?
A mutable tag drifted. Compare registry/manifest/runtime digests and move the manifest to the approved immutable digest.
Why should copied certificate-cluster instructions be rejected?
The certificate-based integration has been deprecated since GitLab 14.5. Current designs should use the Agent, and current GitOps should use Flux.
Why preserve a successful GitLab deployment record when external health fails?
It accurately proves what GitLab authorized/ran. External health is a later independent state; rewriting or deleting history would hide the causal boundary.
Version and compatibility note
GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.
Official references and version notes
Documentation verification date: 2026-09-12.
Current GitLab documentation lists the Agent CI/CD workflow as Free,
Premium, and Ultimate on GitLab.com, Self-Managed, and Dedicated.
Authorized CI/CD jobs receive a KUBECONFIG containing
contexts for authorized agent connections; the context identifies
the agent configuration project and agent name.
ci_access can authorize projects/groups and can
restrict access to named environment patterns. The current per-agent
authorization limit is 500 projects and 500 groups. CI-job
impersonation with access_as: ci_job is
Premium/Ultimate; the mandatory lab therefore demonstrates true
Kubernetes namespace-scoped ServiceAccount RBAC locally instead of
pretending that paid impersonation is Free. GitLab recommends Flux
for GitOps; the Agent’s legacy built-in pull-based GitOps
functionality was removed in GitLab 17.0. Certificate-based cluster
integration was deprecated in GitLab 14.5. The Kubernetes dashboard
is currently Beta and available on all tiers. For
environment:kubernetes, dashboard namespace/Flux
resource settings belong under dashboard:; the older
direct namespace/flux_resource_path form
is deprecated. For production GitOps diagnosis, add Flux
Source/Kustomization/HelmRelease conditions and the exact
desired-state revision to the same evidence sequence before changing
reconciliation settings.
- Using GitLab CI/CD with a Kubernetes cluster — official reference.
- GitLab Agent for Kubernetes — official reference.
- Get started connecting a Kubernetes cluster — official reference.
- Install the Agent for Kubernetes — official reference.
- Migrate legacy Agent GitOps to Flux — official reference.
- Dashboard for Kubernetes — official reference.
- CI/CD YAML syntax — environment:kubernetes — official reference.
- Deprecated CI/CD keywords — official reference.
- Migrate from certificate-based Kubernetes integration — official reference.
- GitLab-managed Kubernetes resources — official reference.
- Kubernetes RBAC authorization — official reference.
- Kubernetes Deployments — official reference.
- kind quick start — official reference.
- Flux documentation — official reference.
Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0
Send only Ethereum/ERC-20 compatible assets to this
address.