Chapter 20Lesson 04~235 minutes

Deploying Grid with Containers, Kubernetes, and Cloud Infrastructure: Diagnostics, Failure Modes, and Production Practices

Container failures are often misdiagnosed as Selenium failures because the symptom appears during a WebDriver command. This lesson practices layer-by-layer diagnosis: preserve first evidence, prove image/session/network/resource state, then repair the smallest failing boundary.

DiagnosticsNetworkingResourcesSecretsReadiness

Learning objectives

  • Diagnose browser-container inability to reach an AUT without hiding it with retries.
  • Recognize image/version skew and resource-pressure symptoms.
  • Separate orchestrator readiness from Grid readiness and browser/AUT readiness.
  • Prevent artifact, video, and secret handling from becoming an operational incident.
  • Apply the chapter diagnostic sequence to one intentionally broken topology.

1. Diagnostic sequence for containerized Grid

  1. Preserve the first failing test evidence and container/Grid logs.
  2. Record Selenium client, Grid image tag, returned browser capabilities, Docker/Compose/Kubernetes versions.
  3. Confirm target URL, synthetic data, and environment selection.
  4. Confirm session ID, current URL, browser context, and Grid readiness.
  5. Check DNS/TCP reachability from the browser runtime to the AUT.
  6. Check AUT/browser/network evidence.
  7. Check Grid queue, Node health, container/pod restarts, CPU/memory/shared memory, disk, and artifact growth.
  8. Apply the least destructive correction and rerun the smallest scenario.

2. Broken example: host/localhost confusion

This variant is intentionally wrong:

The following example makes the Broken example: host/localhost confusion behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.

from selenium import webdriver
from selenium.webdriver.common.by import By

with webdriver.Remote("http://127.0.0.1:4444", options=webdriver.ChromeOptions()) as driver:
    # BROKEN: localhost is the browser container itself.
    driver.get("http://127.0.0.1:8000/")
    assert driver.find_element(By.CSS_SELECTOR, "[data-testid=status]").text == "container-grid-ok"

The session can be created successfully, proving the host-to-Grid path works. Navigation then fails or lands nowhere useful because the browser resolves loopback inside its own container. The correct repair in the Compose lab is http://aut:8000/. Preserve the navigation exception and Grid/container logs before changing anything.

3. Prove network state from the right namespace

The following example makes the Prove network state from the right namespace behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.

docker compose ps
docker network inspect selenium-ch20-lab
# Inspect DNS/network configuration without installing ad-hoc tools into the browser image:
docker inspect selenium-ch20-lab-selenium-1 --format '{{json .NetworkSettings.Networks}}' 2>/dev/null || true
docker compose logs --no-color aut | tail -n 40

Do not mutate a production browser container by installing curl/debug packages during an incident. Prefer orchestration metadata, logs, a disposable debug container on the same network, or application-side access logs.

4. Incompatible image/version assumptions

A Grid image tag is a bundle, not just a browser label. If CI uses a different tag than local development, compare the actual image ID, Grid status, and returned capabilities. Do not “fix” version skew by manually replacing a driver binary inside an official container. Upgrade/downgrade the image as a unit unless the official image contract explicitly supports your customization.

docker compose images
docker inspect --format '{{.Config.Image}} {{.Image}}' selenium-ch20-lab-selenium-1 2>/dev/null || true
curl -fsS http://127.0.0.1:4444/status

5. Shared memory, CPU, memory, disk, and session overcommit

Browser renderer crashes under container pressure can produce WebDriver disconnections, session deletions, or blank-page symptoms. The official browser images recommend shm_size: 2gb. That is not permission to run unlimited sessions. A host with insufficient CPU/memory/disk will still fail.

docker stats --no-stream
docker system df
docker compose ps
Do not jump to privileged mode

Avoid --privileged, blanket --no-sandbox, and unlimited resources as troubleshooting shortcuts. They change the security model and can hide the real capacity problem.

6. Artifact growth and secret leakage

Videos, downloads, page source, browser logs, and URLs may contain tokens or PII. Per-session names prevent collisions, but they do not make content safe. Bound collection, redact before persistence when possible, restrict access, and expire artifacts. Never pass cloud/API secrets as plain environment variables if your platform exposes them in logs or inspection output; use your CI/Kubernetes secret mechanism and avoid echoing values.

7. Kubernetes readiness is not browser readiness

A ready pod means its readiness probe passed. It does not prove a matching browser slot is free, a session can be created, the browser can reach the AUT, or the AUT is ready for the next interaction. Treat these as different state machines and correlate them by timestamps/session ID.

Signal What it proves What it does not prove
Pod Ready Configured readiness probe succeeded Matching browser session can start
Grid /status ready Grid can accept work under its readiness model AUT is reachable from browser
Session created Grid placed a browser session Page/app state is ready
AUT app-ready signal Application condition is ready Grid resources are healthy globally

8. Production runbook boundaries

For a real incident, attach: image tags/digests, Grid status, session ID/capabilities, relevant container/pod restart/resource state, browser/AUT error evidence, sanitized network endpoints, and CI run/attempt ID. Do not delete failed artifacts just because a retry later passes. Do not restart the entire cluster before preserving state unless safety/availability requires it.

Knowledge check

A session is created, but the browser cannot load 127.0.0.1:8000. Which layer failed?

Why is --privileged a bad first response to browser crashes?

A Kubernetes pod is Ready but WebDriver session creation times out. Is that contradictory?

Why can unlimited video recording destabilize Grid?

A retry passes after a Node restart. Should the first failure bundle be deleted?

Next lesson

Checkpoint Lab — Deploying Grid with Containers, Kubernetes, and Cloud Infrastructure

Continue with Checkpoint Lab — Deploying Grid with Containers, Kubernetes, and Cloud Infrastructure. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

Official references and current-version notes

Version baseline — August 2026

These lessons pin Selenium Python and Grid concepts to 4.47.0, Docker Selenium image tag 4.47.0-20260808, and Helm chart 0.58.0. The nightly images track Selenium 4.48.0-SNAPSHOT and are intentionally excluded from the mandatory path.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.