Chapter 19Lesson 04~235 minutes

Selenium Grid Architecture, Roles, Routing, and Session Distribution: Diagnostics, Failure Modes, and Production Practices

Grid failures often look similar from the test runner: a session does not start, a session disappears, or commands time out. The correction depends on which control-plane state is wrong. This lesson applies one diagnostic sequence to realistic incidents so “restart Grid” is not the default response.

Grid diagnosticsFailure modesSecurityCapacityIncidents

Learning objectives

  • Diagnose unavailable stereotypes, queue timeouts, Node registration failures, and stale/disconnected Nodes.
  • Separate Grid control-plane incidents from AUT/test synchronization problems.
  • Preserve first-failure queue/node/session evidence before destructive recovery.
  • Recognize public exposure and resource overcommit as production risks.
  • Repair one intentionally broken Grid configuration by changing the smallest responsible layer.

1. Use one diagnostic sequence every time

  1. Preserve first-failure evidence.
  2. Confirm Selenium binding/browser/driver/Grid versions.
  3. Confirm target environment and synthetic test data.
  4. Inspect requested capabilities, session ID, and current browsing context if a session exists.
  5. Inspect locator/synchronization/AUT evidence for in-session failures.
  6. Inspect Grid /status, queue, GraphQL, Node health, and logs for remote execution failures.
  7. Check host CPU/memory/network/ports only after the routing evidence points there.
  8. Apply the least destructive correction and rerun the smallest controlled request.

2. Failure taxonomy by observable state

The following table organizes the key choices and evidence for Failure taxonomy by observable state. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.

Symptom Likely layer Evidence Least destructive next step
Request sits in queue; no matching stereotype capability/Distributor model requested caps + node stereotypes + queue correct request or add matching lane
Node absent from status registration/Event Bus/network Node logs, Hub/Router address, ports 4442/4443 + node port fix reachability/registration config
Queue grows; all matching slots occupied capacity slot/session occupancy, queue time, host metrics free/add measured capacity
Session ID exists but commands fail to Node Node/network/session ownership Session Map/GraphQL, node URI/status, node logs restore route/node or fail session cleanly
Grid healthy; test fails after click AUT/test code browser DOM/log/network evidence diagnose app/locator/wait; do not tune Grid
Parallel tests corrupt same account test isolation test IDs/data/account collisions fix fixtures/data ownership; Router cannot solve it

3. Failure mode: Grid exposed publicly

This is a security incident, not a convenience configuration. Selenium warns that public Grid access can expose internal applications/files or permit arbitrary browser/binary execution. Preserve access/network logs, stop accepting untrusted traffic using firewall/private-network controls, rotate any affected credentials, and rebuild disposable lab Nodes if compromise is plausible.

Never “test” exposure on production

Do not scan public endpoints or demonstrate abuse. This course uses loopback/private Grid only.

4. Failure mode: unavailable stereotype/capability

An impossible browser version is a controlled way to reproduce this. The queue accepts the request because it is syntactically valid, but the Distributor cannot match it to a free slot. Do not convert the requested version into “latest” silently: the requested capability is part of test meaning.

from selenium import webdriver
from selenium.common.exceptions import WebDriverException

opts = webdriver.ChromeOptions()
opts.browser_version = "999.0-NOT-INSTALLED"
try:
    webdriver.Remote("http://127.0.0.1:4444", options=opts)
except WebDriverException as exc:
    print("creation failed as expected")
    print(str(exc)[:500])

Repair by choosing a supported browser/version or provisioning the intended stereotype. Preserve the first request payload and queue/Node evidence.

5. Failure mode: Node does not register

A Node registers through the Event Bus and the Distributor confirms the Node over HTTP. In Hub/Node mode, default Hub Event Bus ports are 4442/4443 and the Node service port must also be mutually reachable. A wrong --hub address, firewall rule, hostname advertised from inside a container, or registration period expiry can keep the Node out of the Grid model.

# On the Hub/control-plane host
curl http://127.0.0.1:4444/status

# On the Node host, use the exact pinned server help before changing flags
java -jar selenium-server-4.47.0.jar node --help

Do not “fix” registration by disabling security controls or exposing all ports publicly. Correct private reachability and the advertised Grid/Event Bus addresses.

6. Failure mode: session queue timeout

A queue timeout is the consequence of a request waiting beyond configured policy. The cause can be no matching stereotype, no free slot, or unhealthy Nodes. Increasing --session-request-timeout only changes how long the client waits; it does not create capacity or correct capabilities.

Observation Meaning
matching slot exists but occupied capacity/queueing
no matching stereotype configuration/matrix gap
Node status UNAVAILABLE health/registration/resource problem
AUT test runtime long occupied slots may be a downstream test/runtime issue
queue timeout only at peak capacity and CI concurrency need joint analysis

7. Failure mode: stale or disconnected Node

Nodes send heartbeat events to the Distributor. Current Node options include --heartbeat-period and health-related controls. If a Node disappears mid-session, Session Map still describes ownership until cleanup/reconciliation, but routing cannot make a dead browser live again. Preserve session ID, node URI, last known status, Node logs, and infrastructure events before restarting anything.

8. Failure mode: port/firewall mismatch

Standalone hides most internal port wiring. Hub/Node and distributed modes require intended bidirectional reachability. A browser session request that never reaches a registered Node is fundamentally different from an AUT network request inside the browser. Map the failing TCP/HTTP path explicitly before touching test waits.

Grid control-plane networking and browser-to-AUT networking are separate paths

The following diagram visualizes the relationships described in Failure mode: port/firewall mismatch. Read the nodes in sequence and use the arrows to connect the conceptual state changes to the explanation around the diagram.

flowchart LR
 Runner[Test runner] -->|4444 HTTP| Router[Hub/Router]
 Node[Node host] <-->|registration HTTP + Event Bus| Router
 Node -->|WebDriver remote end| Browser[Browser]
 Browser -->|private AUT URL| AUT[AUT]

9. Failure mode: session overcommit

If queue time appears low but browsers crash, become unresponsive, or session creation slows under concurrency, inspect CPU, memory, browser process count, container shared memory, and evidence I/O. Raising --max-sessions is not a performance optimization if the host cannot sustain those browsers.

Measure one variable at a time

Record host resources, Grid/browser versions, slot occupancy, queue time, session startup time, test runtime, and failure rate before and after a capacity change.

10. Intentionally broken example: wrong responsibility

Suppose a suite adds a 30-second explicit wait because RemoteWebDriver session creation times out. That wait runs only after a driver exists, so it cannot repair Grid session scheduling. The first failure is before test DOM interaction.

# Broken diagnosis: this code can never run if session creation itself fails.
driver = webdriver.Remote("http://127.0.0.1:4444", options=options)
WebDriverWait(driver, 30).until(lambda d: d.title != "")

Repair: inspect requested capabilities, queue, Node stereotypes/health, and session-request timeout. Keep application waits scoped to application conditions after a session exists.

11. Router load balancing does not create test isolation

Two sessions can be routed perfectly to different Nodes and still mutate the same synthetic account or write to the same evidence filename. Grid placement solves browser capacity, not shared state. Reuse Chapters 11, 14, and 15: per-worker profiles/files/data, deterministic cleanup, first-failure evidence, and explicit concurrency ownership remain mandatory.

12. Minimal production runbook

  1. Freeze retries and preserve the first failing attempt.
  2. Record versions and exact requested capabilities.
  3. Query /status, GraphQL placement/queue, and relevant Node status.
  4. Correlate Grid logs using session/test ID and timestamps.
  5. Check private network/ports and host capacity only where evidence points.
  6. Fix one layer: capability, capacity, registration, node health, or test/AUT.
  7. Rerun one controlled request, then restore normal concurrency.

13. Lesson summary

  • Grid incidents must be classified by observable routing/health/capacity state.
  • Queue timeouts are symptoms, not root causes.
  • Node registration needs Event Bus and HTTP reachability in non-Standalone topologies.
  • A dead Node cannot be repaired by Router routing or application waits.
  • Public Grid exposure and overcommit are operational failures.
  • Grid does not replace test-data and evidence isolation.

Knowledge check

A session request times out before a driver object exists. Should you increase an explicit DOM wait?

What does a growing queue with matching but occupied slots indicate?

Why can a Node be missing even though its process is running?

What is the risk of exposing a Grid Router publicly?

Why can perfectly routed sessions still corrupt one another?

Next lesson

Checkpoint Lab — Selenium Grid Architecture, Roles, Routing, and Session Distribution

Continue with Checkpoint Lab — Selenium Grid Architecture, Roles, Routing, and Session Distribution. It builds directly on the state, evidence, and operating assumptions established here, so carry those constraints forward rather than treating the next page as an isolated topic.

Official references and version notes

Version and compatibility note

Version-sensitive behavior was rechecked against Selenium primary documentation on 2026-08-28. Mandatory examples pin Selenium Python and Selenium Server/Grid 4.47.0, Python 3.10+, and Java 11+ for the server path. Local Standalone/Hub/Node labs remain on loopback/private networking and use Selenium Manager on the Grid Node only through the current documented --selenium-manager true option. Queue timeouts are deliberately shortened only for disposable failure exercises. Grid 4 roles are taught as Router, New Session Queue, Distributor, Session Map, Event Bus, Nodes, slots/stereotypes; Standalone packages rather than replaces those responsibilities. External state backends are optional advanced architecture and must be revalidated against the pinned server/API before production use. Paid browser clouds, public Grid endpoints, enterprise identity/proxy infrastructure, and managed Kubernetes/cloud are not required.

Keep the academy open

Support free, practical DevOps education.

Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.

Ethereum / ERC-20
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this address.