Checkpoint Lab — Deploying Grid with Containers, Kubernetes, and Cloud Infrastructure
The checkpoint combines deployment, reachability, evidence, failure injection, and cleanup. You will operate a disposable Compose Grid, prove a browser-container session can reach a synthetic AUT, intentionally break the network contract, diagnose it from correlated state, restore service, and leave no shared browser state behind.
Learning objectives
- Generate a complete local lab from one helper script.
- Predict container/network/session changes before executing them.
- Capture a minimal evidence packet from a successful remote session.
- Inject a network misconfiguration and diagnose it without retries or sleeps.
- Prove project-scoped teardown and document what would change on Kubernetes/cloud.
1. Preflight and safety contract
- Docker Engine/Desktop and Docker Compose v2 are installed.
- Port
4444is free on the host. - The lab uses only synthetic static content.
-
Grid port 4444 is published to
127.0.0.1, not all interfaces. - No Docker socket mount, privileged mode, TLS bypass, personal browser profile, or production target is used.
-
Expected Selenium image:
selenium/standalone-chrome:4.47.0-20260808.
2. Generate the disposable lab
Run this helper from a scratch directory. It creates only the checkpoint folder and does not touch Docker until you explicitly run Compose.
from pathlib import Path
root = Path("selenium-ch20-checkpoint")
(root / "fixture").mkdir(parents=True, exist_ok=True)
(root / "evidence").mkdir(exist_ok=True)
(root / "fixture" / "index.html").write_text("""<!doctype html>
<html lang="en"><head><meta charset="utf-8"><title>Chapter 20 checkpoint</title></head>
<body><main><h1>Checkpoint AUT</h1><p data-testid="status">container-grid-ok</p></main></body></html>
""", encoding="utf-8")
(root / "compose.yaml").write_text("""services:
selenium:
image: selenium/standalone-chrome:4.47.0-20260808
shm_size: 2gb
ports:
- "127.0.0.1:4444:4444"
depends_on: [aut]
networks: [lab]
aut:
image: python:3.13-alpine
working_dir: /site
command: ["python", "-m", "http.server", "8000", "--bind", "0.0.0.0"]
volumes:
- ./fixture:/site:ro
networks: [lab]
networks:
lab:
name: selenium-ch20-checkpoint
""", encoding="utf-8")
(root / "run.py").write_text("""from pathlib import Path
import json
from selenium import webdriver
from selenium.webdriver.common.by import By
GRID = "http://127.0.0.1:4444"
AUT = "http://aut:8000/"
out = Path("evidence")
out.mkdir(exist_ok=True)
options = webdriver.ChromeOptions()
driver = webdriver.Remote(command_executor=GRID, options=options)
try:
driver.get(AUT)
status = driver.find_element(By.CSS_SELECTOR, "[data-testid='status']").text
assert status == "container-grid-ok"
record = {
"session_id": driver.session_id,
"browserName": driver.capabilities.get("browserName"),
"browserVersion": driver.capabilities.get("browserVersion"),
"platformName": driver.capabilities.get("platformName"),
"url": driver.current_url,
"title": driver.title,
"status": status,
}
(out / "session.json").write_text(json.dumps(record, indent=2), encoding="utf-8")
driver.save_screenshot(str(out / "viewport.png"))
finally:
driver.quit()
""", encoding="utf-8")
print(root.resolve())
The following example makes the Generate the disposable lab behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
python make_lab.py
cd selenium-ch20-checkpoint
python -m venv .venv
# macOS/Linux: source .venv/bin/activate
# Windows PowerShell: .venv\Scripts\Activate.ps1
python -m pip install "selenium==4.47.0"
3. Predict changes before execution
Write predictions before running anything. At minimum:
-
After
docker compose up -d, two containers and one named project network should exist; no WebDriver session exists yet. -
After
python run.py, one temporary browser session should be created and quit; the AUT should log an HTTP GET; two evidence files should remain on the host. -
Changing
AUTtohttp://127.0.0.1:8000/should leave Grid session creation successful but break browser-to-AUT navigation.
4. Start, verify, and execute the healthy path
The following example makes the Start, verify, and execute the healthy path behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
docker compose pull
docker compose up -d
docker compose ps
# Wait until /status says ready using the bounded preflight from Lesson 2.
python run.py
cat evidence/session.json
docker compose logs --no-color aut | tail -n 30
Verification: the JSON has a session ID plus returned browser
version; the URL is http://aut:8000/; status is
container-grid-ok; the AUT access log shows a request
from the Compose network; and the browser session no longer appears
after quit().
5. Inject one network error without changing test data or Grid version
Edit only the AUT constant in run.py:
The following example makes the Inject one network error without changing test data or Grid version behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
AUT = "http://127.0.0.1:8000/" # intentionally wrong inside the browser container
Run the smallest scenario once. Preserve the exception and current
Grid/container logs. Do not add a retry, longer page-load timeout,
JavaScript navigation, --network host, or privileged
mode. The failure is expected and controlled.
6. Trace the failed request through infrastructure state
The following example makes the Trace the failed request through infrastructure state behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
curl -fsS http://127.0.0.1:4444/status
docker compose ps
docker network inspect selenium-ch20-checkpoint
docker compose logs --no-color selenium | tail -n 100
docker compose logs --no-color aut | tail -n 50
Expected diagnosis: Grid is ready and can create a session, but the
AUT sees no request for the failing run because browser loopback
points at the browser container. Restore
http://aut:8000/, rerun once, and show that the AUT
access log and session evidence return.
7. Optional resource-misconfiguration variant
If your machine has enough spare capacity and you want a second
exercise, reduce shm_size to an unrealistically small
value in this disposable lab and observe whether browser stability
changes. Do not force a crash on a shared CI host, do not overcommit
multiple sessions, and do not treat one machine’s threshold as
universal. The preferred mandatory injection remains the
deterministic network error above.
8. Required evidence packet and conclusions
Keep: session.json, one viewport screenshot from the
healthy path, the exact Compose/image configuration, sanitized Grid
status excerpt, relevant Selenium/AUT log excerpts for healthy and
failed runs, and a short conclusion:
| Observation | Layer | Conclusion |
|---|---|---|
| Grid ready + session created | Grid/container runtime | Host-to-Grid path works |
| AUT access log on healthy run | Browser network → AUT | Service DNS/path works |
| No AUT request when URL is browser loopback | Browser-container network | Wrong address namespace, not locator timing |
Healthy rerun after restoring aut:8000 |
Controlled correction | Network contract repaired without masking failure |
9. Prove teardown leaves no shared lab state
The following example makes the Prove teardown leaves no shared lab state behavior concrete. Read it with the stated assumptions, then compare its observable output or state changes with the explanation that follows.
docker compose down --remove-orphans
# Verify no lab containers are running:
docker compose ps
# The explicitly named lab network should be gone after Compose teardown:
docker network inspect selenium-ch20-checkpoint 2>/dev/null && echo "unexpected: network remains" || true
cd ..
rm -rf selenium-ch20-checkpoint
Do not run broad Docker pruning. The checkpoint creates no named data volume and persists browser state nowhere outside the disposable container. On Kubernetes the equivalent proof would include namespace/release deletion, pod/PVC/secret review, and any external artifact lifecycle; on a hosted cloud it would include session termination and provider artifact-retention controls.
10. What Chapter 20 adds to the operating model
You can now distinguish Selenium test semantics from browser-infrastructure deployment semantics, and you can prove whether a failure belongs to Grid routing, container networking, resource pressure, orchestration, or the AUT. Chapter 21 builds on this infrastructure foundation to study parallel execution, isolation, concurrency, and test sharding—where Grid capacity and test independence meet.
Knowledge check
In the checkpoint, why is the wrong loopback URL a better failure injection than random network latency?
It is deterministic, reversible, local, and isolates one clear network-namespace error.
What proves that the failure is not a Grid session-allocation problem?
Grid is ready and a RemoteWebDriver session is successfully created; the failure occurs during browser navigation to the wrong address.
Why does the evidence packet keep the Compose file/image tag?
Infrastructure identity is part of reproducibility and lets an operator correlate browser/Grid behavior with the exact deployed runtime.
Should the cleanup remove every unused Docker image on the developer machine?
No. Cleanup must be scoped to lab containers/network/files; broad pruning can destroy unrelated state.
How would the same diagnosis change on Kubernetes?
You would additionally inspect pod/service DNS, NetworkPolicy, readiness/restarts/resources, chart/image versions, and cluster routing, while preserving the same session/AUT evidence model.
Official references and current-version notes
- SeleniumHQ/docker-selenium — official images, Compose, Dynamic Grid, troubleshooting
- Docker Selenium 4.47.0-20260808 release
- Official Selenium Grid Helm chart
- Selenium Grid documentation
- Grid CLI/configuration options
These lessons pin Selenium Python and Grid concepts to
4.47.0, Docker Selenium image tag
4.47.0-20260808, and Helm chart 0.58.0.
The nightly images track Selenium 4.48.0-SNAPSHOT and
are intentionally excluded from the mandatory path.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.