Chapter 27Lesson 04~235 minutes

Containers, Kubernetes, Ephemeral Injectors, and Infrastructure Automation: Diagnostics, Failure Modes, and Production Practices

An ephemeral injector can disappear with the evidence needed to explain why it failed. Diagnostics therefore begin before cleanup. Capture image/input identity, container/Pod spec, exit/OOM/restart state, JTL/jmeter.log, target events and generator resource evidence before rebuilding the image, changing the manifest or deleting the failed object.

Floating tagsSecret layersLost artifactslocalhost namespacesPermissions/scaling

Learning objectives

  • Detect community-image provenance assumptions and mutable tags.
  • Prevent credentials/secrets from entering image layers/history.
  • Diagnose JTL loss caused by ephemeral filesystem cleanup.
  • Interpret container localhost/network DNS failures.
  • Control Kubernetes Job parallelism/retry/deadline load multiplication.
  • Repair non-root artifact-directory permission problems without hiding the cause.

1. Preserve first-failure evidence

All runnable diagnosis remains disposable/private. Preserve first-failure JTL, matching jmeter.log, Docker/Kubernetes inspect/status/logs, image ID/digests/history, input hashes/mounts, resource limits/OOM state and target events before correction. Do not delete the failed container/Pod/Job until needed evidence is externalized. Do not add workload.

2. Diagnostic sequence

Container/Kubernetes failure diagnosis

Containerization changes how the injector is packaged, started, networked, limited, and cleaned up; it does not change JMeter thread semantics. The image/runtime lifecycle and the test-plan lifecycle must therefore be measured and evidenced separately.

flowchart TD
E[Preserve JTL + jmeter.log + image/mount/runtime + target evidence] --> V[Confirm JMeter / Java / plugin / Docker-kubectl-kind versions]
V --> C[Confirm image ID + JMX/data/properties/CLI + authorized target]
C --> S[Validate tree scope / resolved properties / injector-count load math]
S --> P[Inspect HTTP session / target DNS / mounted data state]
P --> G[Inspect JVM + container CPU/memory/PID/filesystem/network state]
G --> T[Inspect SUT events/telemetry]
T --> O[Inspect Job/Pod/container exit/OOM/restart/CI state]
O --> F[Least-destructive correction]
F --> R[Small controlled rerun]

3. Failure mode: assuming an unofficial image is Apache-maintained

An image named someone/jmeter can be useful, but the namespace/name does not establish Apache provenance. Inspect its source/Dockerfile, base image, JMeter download verification, plugin contents, update process and digest; run jmeter -v/java -version.

Repair: adopt the verified custom build or an organization-vetted community image pinned by digest. Document the maintainer accurately.

4. Failure mode: floating latest

Yesterday's latest and today's latest can resolve to different content. A regression then mixes application change with injector/runtime change.

Repair: use version-specific tags plus resolved digest/image ID in evidence; rebuild/upgrade deliberately and re-baseline.

5. Failure mode: baking a secret into the image

Broken Dockerfile:

ARG API_TOKEN
ENV API_TOKEN=$API_TOKEN
COPY secret.properties /opt/jmeter/bin/user.properties

Build args/environment/layers/history/caches can leak credential material. Never put a reusable secret into JMX, image layers or Git.

Repair: remove it from build context/Dockerfile, rotate the real credential if exposure occurred, rebuild cleanly, and inject the secret at runtime through an approved secret mechanism with redacted JTL/logs. Mandatory lab uses no credentials at all.

6. Failure mode: JTL written only inside the ephemeral container

Broken run writes -l /work/results.jtl with no artifact mount; CI then removes the container. Exit code is known, but raw samples are gone.

Repair: preserve the failed container if still present and docker cp the files out; then change future runs to /artifacts bind/PVC collection. Do not delete the only surviving container before copying evidence.

7. Intentionally broken example: wrong localhost

Run the JMeter container on p27-net but pass:

-Jtarget.host=127.0.0.1
-Jtarget.port=8027

Expected evidence:

  • JTL samples fail with connection errors because nothing listens on 8027 inside the JMeter container;
  • jmeter.log records the connection failure;
  • p27-fixture target event count remains zero;
  • Docker network inspection shows both containers attached, so the topology exists but target identity is wrong.

Repair: preserve failed JTL/log/inspect evidence, change only -Jtarget.host=p27-fixture, rerun the same 1×5 smoke before the full 5×20 test. Do not increase timeouts/retries.

8. Failure mode: unrestricted Kubernetes scaling/retries

A Job is changed from parallelism 2 to 20 while each Pod still runs 100 samples: configured total becomes 2,000 samples. If backoffLimit allows retries, failures can add more execution attempts.

Repair: explicit target authorization, total-load equation, conservative parallelism/completions, backoffLimit: 0 for this teaching workload, active deadline, resource limits and target abort/count gate.

9. Failure mode: non-root bind-mount permission denied

The JMeter image runs UID 10001. On a Linux host, /artifacts bind mount is owned by another user and denies writes. JMeter may fail before producing JTL.

Diagnosis:

  • container exit/log error;
  • docker inspect confirms mount is writable in Docker metadata;
  • host filesystem ownership/mode denies UID 10001.

Repair one dedicated result directory: pre-create/chown it intentionally, or on local Linux run the container with the current host UID/GID when compatible. Do not switch the whole pattern to root or chmod 777 as a universal fix. Docker Desktop filesystem semantics differ; record the platform.

10. Failure mode: container OOM mistaken for JMeter/test failure

If memory limit is below effective Java/native requirement, Docker may report OOMKilled=true or Kubernetes Pod reason OOMKilled. JTL can be truncated. This is generator infrastructure saturation, not a target regression.

Repair from Chapter 26 evidence: reduce plan allocation/threads or assign measured memory/heap headroom. Keep -Xmx below the container limit.

11. Causal symptom table

Symptom Container/orchestrator cause Target cause to distinguish Evidence
JTL connect errors, target zero events wrong namespace/DNS/localhost target service down JTL/log + Docker/K8s DNS/network + target logs.
JTL disappears after run ephemeral filesystem/cleanup no samples generated container/Pod history + missing mount + target count.
Same commit behaves differently next day floating image tag/base SUT regression resolved image IDs/digests + input hashes.
Pod OOMKilled, partial JTL resource limit/heap mismatch server aborting connections Pod/container state + JVM/resource evidence + target events.
More target requests than JMX math extra Pods/Job retries application retry behavior Job status/pod attempts + injector IDs + target counts.
Permission denied on JTL non-root host-mount ownership JMeter sampler failure container log/exit + mount metadata + host permissions.

12. Security-sensitive boundaries

Image build contexts/layers, runtime env vars, Docker socket, registry credentials, Kubernetes Secrets/service accounts, CI secrets, mounted files, target credentials, recorder certs/RMI keys and privileged/root settings are sensitive. Do not mount the host Docker socket into a JMeter container; it grants broad daemon/host control and is unnecessary here.

13. Troubleshooting shortcuts to reject

  • Do not add blanket retries or arbitrary long sleeps.
  • Do not assign giant heaps to outrun a memory quota.
  • Do not mass-disable evidence/listeners without measuring their cost.
  • Do not use global property hacks to hide namespace/config mistakes.
  • Do not disable TLS/RMI verification.
  • Do not test container fixes against production/public targets.
  • Do not increase Pod parallelism while validity is unresolved.
  • Do not delete failed containers/Pods/JTL/logs before evidence collection.

Knowledge check

Target has zero events while container JTL says connection refused to 127.0.0.1. What is the likely layer?

Why is latest unsafe for a regression baseline?

What should happen after discovering a real secret was baked into an image?

Why is OOMKilled not proof the server failed?

Why can a Kubernetes retry violate load safety?

Next lesson

Checkpoint: portable Docker run and optional local Job

Lesson 5 freezes the provenance/input/artifact contract, verifies native and Docker counts, then designs a two-Pod kind layout with bounded resources and artifact collection.

Official references and version notes

Version and compatibility note

Version-sensitive statements were rechecked against current primary documentation on 2026-09-05. The course baseline remains Apache JMeter 5.6.3, requiring Java 8+; this chapter uses Eclipse Temurin 17.0.20+8 JRE (Jammy) as the version-specific Java base and records its resolved registry digest locally. The JMeter layer is built from Apache's binary tarball and verifies Apache's published SHA-512: 5978a1a35edb5a7d428e270564ff49d2b1b257a65e17a759d259a9283fc17093e522fe46f474a043864aea6910683486340706d745fcdf3db1505fd71e689083. Apache JMeter does not need a third-party plugin for the mandatory lab. Docker bind mounts are writable by default, so the JMX/data/config inputs are explicitly readonly; only the artifact mount is writable. The optional Kubernetes path uses current batch/v1 Job semantics with restartPolicy: Never, backoffLimit: 0, activeDeadlineSeconds, resource requests/limits, and no cloud requirement. The local cluster example assumes kind v0.32.0; a current minikube installation is an optional equivalent.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0 Send only Ethereum/ERC-20 compatible assets to this address.