Selenium Grid Architecture, Roles, Routing, and Session Distribution: Core Concepts and Mental Model
A local WebDriver session hides most routing because the test runner, driver, and browser are on one machine. Grid introduces a scheduler and routing fabric between the client and browser. The important beginner shift is to stop thinking “Grid is a remote browser” and instead trace how one session request becomes a queue entry, a slot match, a browser process on a Node, and then a stable session-to-Node route.
Learning objectives
- Trace a new session from client to Router, queue, Distributor, matching slot, and Node.
- Explain how Session Map routes commands after session creation.
- Define Event Bus, Node heartbeat, slot, stereotype, capacity, and session ownership.
- Inspect Grid state before creating or changing sessions.
- Explain why Standalone packages the same logical roles rather than replacing the architecture.
1. Why distributed browser execution needs explicit routing
With local webdriver.Chrome(), one process starts one
driver and one browser. With Grid, many clients may ask for
different browsers at the same time, while browser capacity lives on
one or many Node hosts. Grid therefore needs a front door, a waiting
area, a scheduler, a map of running sessions, health information,
and a transport path to the chosen Node.
Grid schedules and routes WebDriver sessions. It does not create test-data isolation, assertion meaning, application readiness, or safe parallel test design for you.
2. The six logical Grid roles
The following diagram visualizes the relationships described in The six logical Grid roles. Read the nodes in sequence and use the arrows to connect the conceptual state changes to the explanation around the diagram.
flowchart TD C[Test client] -->|POST /session| R[Router] R --> Q[New Session Queue] Q --> D[Distributor] D -->|match request to free slot| N[Node / browser slot] N -->|session created| D D --> M[Session Map] M -->|session id → node URI| R R -->|later session commands| N E[Event Bus] <-->|registration / internal events| D E <-->|registration / heartbeat events| N
The Router is the external entry point. New-session requests go to the New Session Queue. The Distributor maintains the Grid model and consumes queue requests when a matching free slot exists. The Node owns the browser process and WebDriver session. Once creation succeeds, the Session Map records the session ID to Node URI relationship. The Event Bus carries asynchronous internal events such as Node registration. Later commands use the Router and Session Map to reach the owning Node without asking the Distributor to schedule the session again.
3. Role-by-role responsibilities
The following table organizes the key choices and evidence for Role-by-role responsibilities. Use it together with the surrounding prose so the rows serve as a comparison aid rather than standalone rules.
| Role | Primary responsibility | Observable evidence | What it does not do |
|---|---|---|---|
| Router | External entry point; sends new sessions to queue and existing-session commands to owning Node | HTTP response, Grid URL, routed session command | Does not choose application test data or assertion policy |
| New Session Queue | Holds pending session requests until matching capacity becomes available or request times out | queue size/request payloads, timeout behavior | Does not launch browsers itself |
| Distributor | Tracks Nodes/slots and matches pending requests to suitable free slots | Node model, matching/creation logs, GraphQL placement | Does not route every later command |
| Session Map | Maps session ID to Node address | session → node URI relationship | Does not mean browser cookies/storage or AUT state |
| Event Bus | Internal asynchronous message path for Grid components | registration/heartbeat-related logs and traces | Not the WebDriver BiDi event channel from Chapter 18 |
| Node | Runs browser sessions in slots and executes commands | Node status, slot state, browser capabilities, session lifecycle | Does not interpret test assertions |
4. Slots, stereotypes, and max sessions
A slot is a place where one browser session may
run. A slot has a stereotype: the minimal
capabilities it can match, such as
{"browserName":"firefox"}. A Node may advertise many
possible browser slots while still enforcing a lower
max-sessions concurrency limit. Therefore “number of
stereotypes/slots shown” and “sessions this host can safely run at
once” are not interchangeable metrics.
{
"slot": {
"id": "slot-uuid",
"stereotype": {"browserName": "firefox"},
"session": null
}
}
Current Grid defaults use processor count as a basis for max sessions, but Selenium explicitly warns that overcommitting can reduce stability. Browser memory, CPU, container shared memory, evidence I/O, AUT capacity, and CI host limits all matter.
5. New-session lifecycle versus running-session routing
A session has two different routing phases. During creation, capabilities must be matched against available slots. After creation, the session already has an owner. This distinction is why a full Grid can scale scheduling and command routing independently.
| Phase | Path | Key state | Typical failure |
|---|---|---|---|
| Creation | client → Router → queue → Distributor → Node slot | requested capabilities, queue deadline, free slots, Node health | no matching stereotype, all slots busy, queue timeout, Node creation failure |
| Running | client → Router → Session Map → owning Node | session ID → Node URI | stale/disconnected Node, invalid session ID, network route failure |
| Teardown | client → Router → Node; Grid releases map/slot | browser exit, slot becomes free, Session Map entry removed | orphan browser, delayed cleanup, client disappears |
6. Read-only inspection before mutation
Before opening a session, prove what Grid thinks exists. Current
Grid exposes /status and GraphQL. The status payload
includes registered Nodes, availability, sessions, and slots.
GraphQL can answer targeted questions such as session count, queue
size, node URIs, stereotypes, and session placement.
# Grid health / capacity
curl http://127.0.0.1:4444/status
# Targeted capacity query
curl -s -X POST -H "Content-Type: application/json" --data '{"query":"{ grid { maxSession sessionCount sessionQueueSize } nodesInfo { nodes { id uri status slotCount sessionCount stereotypes } } }"}' http://127.0.0.1:4444/graphql
Do not memorize one JSON shape forever. Record the Grid/server version and inspect fields defensively because observability surfaces can evolve.
7. Standalone packages roles; it does not erase them
standalone starts a full Grid in one process and on one
machine. That is ideal for learning because the external surface is
simple, but logically the same Router, Queue, Distributor, Session
Map, Event Bus, Node, and slots still exist. A queue timeout in
Standalone is still a queue timeout; a stereotype mismatch is still
a Distributor/slot-matching problem.
java -jar selenium-server-4.47.0.jar standalone --selenium-manager true --port 4444
8. State stores and trust boundaries
Keep Grid infrastructure state separate from browser/AUT state. The
Session Map stores routing metadata, not cookies. A Node owns a
browser process, but the test runner may own evidence files. A
remote browser may not be able to reach the same
127.0.0.1 URL as the runner. Grid logs can contain URLs
and capability data, and the Router/Node endpoints can control
browsers—so they are operationally sensitive.
| State | Owner | Example | Isolation requirement |
|---|---|---|---|
| Grid model | Distributor | Node health and slots | private infrastructure control plane |
| Routing map | Session Map | session ID → Node URI | consistent with running sessions |
| Browser state | Node/browser | cookies, storage, windows | one session/test ownership policy |
| AUT state | application | synthetic account/order state | independent test-data isolation |
| Evidence | runner/CI workspace | screenshots, status JSON, Grid logs | per test/attempt paths and retention |
| Credentials | secret store/config boundary | Grid auth or proxy secrets | never hard-coded in lesson artifacts |
9. Event Bus versus WebDriver BiDi
The names can be confusing after Chapter 18. Grid Event Bus is internal infrastructure communication among Grid components. WebDriver BiDi is a browser automation protocol/event channel associated with a browser session. A BiDi console event and a Node registration event are different event systems with different producers, consumers, schemas, and security boundaries.
10. DevOps connection: browser capacity becomes schedulable infrastructure
Once browser execution is centralized, capacity, health, routing, queue delay, session ownership, and version drift become operational concerns. A useful CI incident record therefore includes client/binding version, Grid version, requested capabilities, returned capabilities, session ID, Node placement, queue evidence when applicable, and browser/AUT evidence. “Remote test failed” is not enough to identify the failing layer.
11. Common wrong mental models
- “Router load-balances my tests.” The Distributor matches new sessions; after creation the Session Map routes commands to the owning Node.
- “A slot is one CPU.” A slot is a capability-matching place; safe concurrency depends on measured resources.
- “Grid makes parallel tests isolated.” Grid supplies browser capacity; your suite still owns data/account/profile/evidence isolation.
- “Standalone has no queue or distributor.” Those logical roles are packaged into one process.
- “Event Bus is BiDi.” They solve different infrastructure/browser event problems.
12. Lesson summary
- Grid 4 is a session-routing system, not merely a remote browser launcher.
- New sessions flow through Router, Queue, Distributor, matching slot, and Node.
- Session Map records ownership so subsequent commands route directly to the correct Node.
- Event Bus coordinates Grid components; it is not WebDriver BiDi.
- Slots/stereotypes describe matchable capacity; max safe sessions is an operational measurement.
- Standalone preserves the logical architecture in one process.
Knowledge check
Which component decides where a new browser session should run?
The Distributor, using its model of registered Nodes and free slots/stereotypes while consuming matching requests from the New Session Queue.
Why does the Router not need the Distributor for every command after session creation?
The Session Map records the session ID to Node relationship, allowing the Router to forward existing-session commands to the owning Node.
Is the Grid Event Bus the same thing as WebDriver BiDi?
No. Grid Event Bus is internal Grid-component messaging; WebDriver BiDi is a browser automation event/control protocol tied to a browser session.
A Node advertises Chrome and Firefox slots. Does that prove it can safely run all advertised slots simultaneously?
No. Slot stereotypes describe matchable session types. Safe concurrent session count is constrained separately and must be measured against CPU, memory, browser processes, shared memory, AUT load, and evidence overhead.
Why can a test still leak state even when every session is routed correctly?
Grid routes browsers, but test data, accounts, downloads, profiles, assertions, and cleanup remain responsibilities of the test architecture.
Official references and version notes
- Selenium 4.47 release notes — current stable client and Grid baseline.
- Selenium downloads — Python and Selenium Server/Grid 4.47.0 stable releases.
- Selenium Grid — purpose and current Grid documentation entry point.
- Grid components — Router, Distributor, Session Map, New Session Queue, Node, and Event Bus responsibilities.
- Grid architecture — slots, stereotypes, sessions, and logical relationships.
- Grid getting started — Standalone, Hub/Node, Distributed roles, ports, Java/browser prerequisites, and Selenium Manager option.
- Grid CLI options — current Node/session queue/capacity/heartbeat/BiDi/managed-download options.
- Grid TOML configuration — reviewable configuration examples and Router authentication.
- Grid endpoints — status, Node, session, and New Session Queue endpoints.
- Grid GraphQL support — session placement, Node, queue, and capacity queries.
- External datastore — JDBC/Redis-backed Session Map patterns.
- Grid configuration help — use the pinned server JAR help as implementation-grounded configuration truth.
Version-sensitive behavior was rechecked against Selenium primary
documentation on 2026-08-28. Mandatory examples pin Selenium
Python and Selenium Server/Grid 4.47.0, Python 3.10+, and Java 11+
for the server path. Local Standalone/Hub/Node labs remain on
loopback/private networking and use Selenium Manager on the Grid
Node only through the current documented
--selenium-manager true option. Queue timeouts are
deliberately shortened only for disposable failure exercises. Grid
4 roles are taught as Router, New Session Queue, Distributor,
Session Map, Event Bus, Nodes, slots/stereotypes; Standalone
packages rather than replaces those responsibilities. External
state backends are optional advanced architecture and must be
revalidated against the pinned server/API before production use.
Paid browser clouds, public Grid endpoints, enterprise
identity/proxy infrastructure, and managed Kubernetes/cloud are
not required.
Keep the academy open
Support free, practical DevOps education.
Every lesson is designed to remain readable in a browser, downloadable from GitHub, and usable without a paid learning platform. Contributions help expand and maintain the curriculum.
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0Send only Ethereum/ERC-20 compatible assets to this
address.