Chapter 28Lesson 04~210 minutes

Kubernetes Agent, Cluster Access, GitOps-Oriented Delivery, Kubernetes Deployments, and Environment Integration: Diagnostics, Failure Modes, Security, and Performance

Diagnose over-broad cluster access, namespace collisions, wrong contexts, premature green jobs, image drift, deprecated integrations, and rollout failures from preserved evidence.

DiagnosticsWrong contextRBACRolloutImage drift

Learning objectives

  • Diagnose Kubernetes delivery failures by separating pipeline compilation, Agent/context selection, RBAC, manifest, runtime, environment and external-health layers.
  • Preserve original pipeline/job/rollout evidence before correction.
  • Repair over-broad access and mutable-image failures without masking the original cause.
  • Recognize deprecated certificate-based and legacy Agent GitOps guidance.
  • Retry only the smallest safe scope after proving the failed layer.

1. Evidence-first diagnostic sequence

  1. Preserve pipeline/job IDs, first failing trace and any GitLab deployment ID.
  2. Confirm CI_PIPELINE_SOURCE, ref, CI_COMMIT_SHA and compiled configuration.
  3. Confirm workflow/job-rule decisions and non-secret effective inputs.
  4. Inspect job graph, queue, runner/executor/image/tool versions.
  5. Inspect Agent context name and kubectl auth can-i; never expose kubeconfig credentials.
  6. Inspect manifest/image digest and the command/tool/network failure.
  7. Inspect GitLab environment/deployment record.
  8. Inspect Kubernetes Deployment/ReplicaSet/Pod/Service and runtime imageID.
  9. Verify the external endpoint independently.
  10. Apply the least destructive correction; rerun only the failed safe scope.

2. Failure-layer map

Symptom Likely layer First proof Do not do
No deploy job exists Rules / compiled config Merged YAML and job-rule result Change cluster RBAC
Expected Agent context absent Agent ci_access / propagation Context list plus Agent config revision Paste a production kubeconfig variable
Context exists, API says forbidden Kubernetes RBAC / impersonation Exact auth can-i and denial Grant cluster-admin
Apply succeeds, rollout times out Workload/runtime Deployment conditions, Events, Pods Mark job green because apply returned 0
Rollout is Ready, wrong code runs Artifact identity/tag drift Manifest image + runtime imageID + producer digest Blindly restart Pods
GitLab environment looks green, endpoint fails External/application health Independent HTTP/service check Rewrite deployment history

3. Failure: production cluster credential stored in a variable

Broken pattern: a generic project variable contains a kubeconfig with cluster-admin privileges and every branch can run the deployment job. This combines long-lived credential exposure, broad source trust and broad cluster privilege.

Security-sensitive correction: do not print or download the credential for inspection. Restrict the job immediately, inventory where the credential was used, rotate/revoke it through the cluster/provider process, migrate to Agent/GitOps connectivity, and preserve audit evidence. Do not “fix” this by merely masking the variable.

4. Failure: Agent authorizes too many projects

Suppose the Agent config grants a whole group while only one deployer project needs access:

ci_access:
  groups:
    - id: platform

Before changing it, list which projects depend on that connection and preserve the current Agent config commit. Then narrow to explicit projects/environments:

ci_access:
  projects:
    - id: platform/deployer
      environments:
        - staging
        - production

Authorization changes can take time to propagate. Verify context availability and denials before retrying.

5. Failure: namespace collision

Two projects both deploy deployment/web into namespace staging. One pipeline patches the other team’s workload. The fix is not a blind rename in the running cluster. Preserve resource UIDs, owner labels, pipeline/job IDs and manifest revisions; then move to deterministic project-owned namespaces or stronger naming/ownership conventions plus RBAC.

kubectl -n staging get deployment/web -o json > evidence/web-before.json
kubectl -n staging get rolebinding -o yaml > evidence/rolebindings-before.yaml
kubectl -n staging get events --sort-by=.lastTimestamp > evidence/events-before.txt

6. Intentionally broken example: green job before rollout health

deploy_staging:
  stage: deploy
  script:
    - kubectl config use-context "$KUBE_CONTEXT"
    - kubectl apply -f k8s/staging.yaml
  environment:
    name: staging

This job can succeed as soon as the Kubernetes API accepts the manifest. It does not wait for Pods to become Ready. Preserve the successful job/deployment record, then inspect the workload. Repair by adding an explicit bounded rollout/health check:

script:
  - kubectl config use-context "$KUBE_CONTEXT"
  - kubectl apply -f k8s/staging.yaml
  - kubectl rollout status deployment/example -n app-staging --timeout=180s
  - ./verify-staging-health.sh

If rollout fails, do not replace the original job trace. The failed rollout is the causal evidence.

7. Failure: image tag drifts

A manifest uses example/app:stable. Pipeline 410 deploys it successfully. Later the tag moves. A restarted Pod now pulls different bytes even though no manifest commit changed. Diagnose by comparing the original registry digest, manifest reference and current Pod imageID. Correct the manifest to the verified digest and preserve the old runtime evidence.

kubectl -n app-staging get deployment/example   -o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}'
kubectl -n app-staging get pods -l app=example   -o jsonpath='{range .items[*]}{.metadata.name}{"  "}{.status.containerStatuses[0].imageID}{"\n"}{end}'
cat evidence/approved-image-digest.txt

The safe correction is to promote the already-approved digest, not rebuild the image during troubleshooting.

8. Failure: correct permissions, wrong cluster context

A job is authorized to two Agents and relies on whichever context is current. The script succeeds against staging while the GitLab environment says production—or vice versa. Make context selection an explicit reviewed variable and record it before mutation:

kubectl config get-contexts -o name
printf 'selected_context=%s\n' "$KUBE_CONTEXT"
kubectl config use-context "$KUBE_CONTEXT"
kubectl config current-context
kubectl get namespace "$TARGET_NAMESPACE" -o jsonpath='{.metadata.uid}{"\n"}'

Do not infer cluster identity only from a namespace name, because the same namespace name can exist in many clusters.

9. Failure: template copied from deprecated certificate-based integration

Symptoms include advice to configure project/group/instance cluster certificates, a gitlab-deploy context assumption, or legacy built-in Agent gitops: manifest_projects. Preserve the source/template version, then migrate to the Agent CI/CD workflow or Flux. Do not add feature flags to resurrect deprecated integration in a new design.

10. Runtime evidence without destructive debugging

kubectl -n app-staging get deployment,replicaset,pod,service -o wide
kubectl -n app-staging describe deployment/example
kubectl -n app-staging get events --sort-by=.lastTimestamp
kubectl -n app-staging logs deployment/example --tail=200
kubectl -n app-staging get pod -l app=example   -o jsonpath='{range .items[*]}{.metadata.name}{"  "}{.status.phase}{"  "}{.status.containerStatuses[0].imageID}{"\n"}{end}'

These commands preserve observable state. Avoid deleting Pods merely to see whether they come back; that destroys first-failure evidence and can create side effects before you know the cause.

11. Security-sensitive actions

  • Changing Agent project/group authorization is an authorization change: review and record it.
  • Changing an Agent service-account Role/ClusterRole changes external privilege: use least privilege and preserve before/after RBAC.
  • Rotating a leaked kubeconfig or Agent token is incident-response work: do not post token material into tickets or logs.
  • Changing production Flux sources/reconciliation targets can deploy broadly: use reviewed exact revisions.
  • Never troubleshoot by disabling TLS verification, granting cluster-admin, bypassing protected environments or using an unreviewed mutable image.

12. Performance and reliability: avoid making the control plane the bottleneck

Large pipelines can overload an API server with many parallel imperative jobs. Serialize conflicting deployments (Chapter 18), prefer controller reconciliation for production desired state, bound kubectl rollout status timeouts, and keep manifests small enough to review. A queued runner is not a slow cluster; distinguish job queue time, API request latency, reconciliation time and application readiness.

13. Least-destructive recovery map

Failure Smallest safe correction
Agent context absent Correct ci_access/environment authorization and wait for propagation; do not add static production credentials.
RBAC denied Add only required resource/verb/namespace permission; rerun the single deployment job after proof.
Bad manifest Commit/review corrected desired state or revert the exact manifest revision; do not patch production ad hoc unless emergency procedure requires it.
Rollout failed Preserve Events/Pod logs, then fix application/config/resource issue; use rollout undo only when the exact previous revision is understood.
Image drift Restore verified digest reference; do not rebuild from source as a “rollback”.
External health failed Keep deployment record; diagnose service/network/dependency layer separately.

Knowledge check

The Agent context is present but kubectl returns Forbidden. Which layer failed?

Why is adding rollout status to the job important?

A restarted Pod runs a different imageID with no manifest commit. What is the most likely cause?

Why should copied certificate-cluster instructions be rejected?

Why preserve a successful GitLab deployment record when external health fails?

Next lesson

Next: checkpoint

Lesson 5 combines identity, digest, manifest, rollout, environment mapping and exact cleanup into one evidence packet, including an intentionally denied cluster-scoped action.

Version and compatibility note

GitLab and GitLab Runner evolve continuously. Treat version-sensitive YAML, runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations as assumptions to verify against the current official GitLab documentation before production use. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for any reproducible lab or incident record.

Official references and version notes

Documentation verification date: 2026-09-12. Current GitLab documentation lists the Agent CI/CD workflow as Free, Premium, and Ultimate on GitLab.com, Self-Managed, and Dedicated. Authorized CI/CD jobs receive a KUBECONFIG containing contexts for authorized agent connections; the context identifies the agent configuration project and agent name. ci_access can authorize projects/groups and can restrict access to named environment patterns. The current per-agent authorization limit is 500 projects and 500 groups. CI-job impersonation with access_as: ci_job is Premium/Ultimate; the mandatory lab therefore demonstrates true Kubernetes namespace-scoped ServiceAccount RBAC locally instead of pretending that paid impersonation is Free. GitLab recommends Flux for GitOps; the Agent’s legacy built-in pull-based GitOps functionality was removed in GitLab 17.0. Certificate-based cluster integration was deprecated in GitLab 14.5. The Kubernetes dashboard is currently Beta and available on all tiers. For environment:kubernetes, dashboard namespace/Flux resource settings belong under dashboard:; the older direct namespace/flux_resource_path form is deprecated. For production GitOps diagnosis, add Flux Source/Kustomization/HelmRelease conditions and the exact desired-state revision to the same evidence sequence before changing reconciliation settings.

Current assumptions used in this chapter: Version-sensitive YAML, Runner/executor behavior, APIs, security features, policy controls, tiers, and deprecations must be verified against the exact GitLab, GitLab Runner, tool, and external-system versions used in production. Preserve the exact project/ref/SHA, compiled configuration, pipeline/job IDs, runner/tool versions, and external target evidence used for reproducible work.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.