Deploying Grid with Containers, Kubernetes, and Cloud Infrastructure: Diagnostics, Failure Modes, and Production Practices
Container failures are often misdiagnosed as Selenium failures because the symptom appears during a WebDriver command. This lesson practices layer-by-layer diagnosis: preserve first evidence, prove image/session/network/resource state, then repair the smallest failing boundary.
Learning objectives
- Diagnose browser-container inability to reach an AUT without hiding it with retries.
- Recognize image/version skew and resource-pressure symptoms.
- Separate orchestrator readiness from Grid readiness and browser/AUT readiness.
- Prevent artifact, video, and secret handling from becoming an operational incident.
- Apply the chapter diagnostic sequence to one intentionally broken topology.
1. Diagnostic sequence for containerized Grid
- Preserve the first failing test evidence and container/Grid logs.
- Record Selenium client, Grid image tag, returned browser capabilities, Docker/Compose/Kubernetes versions.
- Confirm target URL, synthetic data, and environment selection.
- Confirm session ID, current URL, browser context, and Grid readiness.
- Check DNS/TCP reachability from the browser runtime to the AUT.
- Check AUT/browser/network evidence.
- Check Grid queue, Node health, container/pod restarts, CPU/memory/shared memory, disk, and artifact growth.
- Apply the least destructive correction and rerun the smallest scenario.
2. Broken example: host/localhost confusion
This variant is intentionally wrong:
The following example makes the Broken example: host/localhost confusion behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
from selenium import webdriver
from selenium.webdriver.common.by import By
with webdriver.Remote("http://127.0.0.1:4444", options=webdriver.ChromeOptions()) as driver:
# BROKEN: localhost is the browser container itself.
driver.get("http://127.0.0.1:8000/")
assert driver.find_element(By.CSS_SELECTOR, "[data-testid=status]").text == "container-grid-ok"
The session can be created successfully, proving the host-to-Grid
path works. Navigation then fails or lands nowhere useful because
the browser resolves loopback inside its own container. The correct
repair in the Compose lab is http://aut:8000/. Preserve
the navigation exception and Grid/container logs before changing
anything.
3. Prove network state from the right namespace
The following example makes the Prove network state from the right namespace behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
docker compose ps
docker network inspect selenium-ch20-lab
# Inspect DNS/network configuration without installing ad-hoc tools into the browser image:
docker inspect selenium-ch20-lab-selenium-1 --format '{{json .NetworkSettings.Networks}}' 2>/dev/null || true
docker compose logs --no-color aut | tail -n 40
Do not mutate a production browser container by installing curl/debug packages during an incident. Prefer orchestration metadata, logs, a disposable debug container on the same network, or application-side access logs.
4. Incompatible image/version assumptions
A Grid image tag is a bundle, not just a browser label. If CI uses a different tag than local development, compare the actual image ID, Grid status, and returned capabilities. Do not “fix” version skew by manually replacing a driver binary inside an official container. Upgrade/downgrade the image as a unit unless the official image contract explicitly supports your customization.
docker compose images
docker inspect --format '{{.Config.Image}} {{.Image}}' selenium-ch20-lab-selenium-1 2>/dev/null || true
curl -fsS http://127.0.0.1:4444/status
5. Shared memory, CPU, memory, disk, and session overcommit
Browser renderer crashes under container pressure can produce
WebDriver disconnections, session deletions, or blank-page symptoms.
The official browser images recommend shm_size: 2gb.
That is not permission to run unlimited sessions. A host with
insufficient CPU/memory/disk will still fail.
docker stats --no-stream
docker system df
docker compose ps
Avoid --privileged, blanket
--no-sandbox, and unlimited resources as
troubleshooting shortcuts. They change the security model and can
hide the real capacity problem.
6. Artifact growth and secret leakage
Videos, downloads, page source, browser logs, and URLs may contain tokens or PII. Per-session names prevent collisions, but they do not make content safe. Bound collection, redact before persistence when possible, restrict access, and expire artifacts. Never pass cloud/API secrets as plain environment variables if your platform exposes them in logs or inspection output; use your CI/Kubernetes secret mechanism and avoid echoing values.
7. Kubernetes readiness is not browser readiness
A ready pod means its readiness probe passed. It does not prove a matching browser slot is free, a session can be created, the browser can reach the AUT, or the AUT is ready for the next interaction. Treat these as different state machines and correlate them by timestamps/session ID.
| Signal | What it proves | What it does not prove |
|---|---|---|
| Pod Ready | Configured readiness probe succeeded | Matching browser session can start |
Grid /status ready |
Grid can accept work under its readiness model | AUT is reachable from browser |
| Session created | Grid placed a browser session | Page/app state is ready |
| AUT app-ready signal | Application condition is ready | Grid resources are healthy globally |
8. Production runbook boundaries
For a real incident, attach: image tags/digests, Grid status, session ID/capabilities, relevant container/pod restart/resource state, browser/AUT error evidence, sanitized network endpoints, and CI run/attempt ID. Do not delete failed artifacts just because a retry later passes. Do not restart the entire cluster before preserving state unless safety/availability requires it.
Knowledge check
A session is created, but the browser cannot load
127.0.0.1:8000. Which layer failed?
Browser-container to AUT networking/addressing, not Grid session routing.
Why is --privileged a bad first response to browser
crashes?
It weakens isolation/security and can hide resource/configuration problems without identifying the cause.
A Kubernetes pod is Ready but WebDriver session creation times out. Is that contradictory?
No. Pod readiness and Grid capacity/session scheduling are different states.
Why can unlimited video recording destabilize Grid?
Video consumes additional CPU and storage; unbounded retention can fill disks and amplify contention.
A retry passes after a Node restart. Should the first failure bundle be deleted?
No. Preserve the first failure; the retry is an additional observation, not proof that the original incident was meaningless.
Official references and current-version notes
- SeleniumHQ/docker-selenium — official images, Compose, Dynamic Grid, troubleshooting
- Docker Selenium 4.47.0-20260808 release
- Official Selenium Grid Helm chart
- Selenium Grid documentation
- Grid CLI/configuration options
These lessons pin Selenium Python and Grid concepts to
4.47.0, Docker Selenium image tag
4.47.0-20260808, and Helm chart 0.58.0.
The nightly images track Selenium 4.48.0-SNAPSHOT and
are intentionally excluded from the mandatory path.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.