Resource library

QA How-To

How to Fix Selenium Grid "Could not start a new session"

Fix Selenium Grid could not start a new session by checking Router access, Node registration, capability matching, queue capacity, and browser Node logs.

22 min read | 3,267 words

TL;DR

Check Router reachability, then query Grid GraphQL for registered Nodes, matching stereotypes, occupied slots, and queue size. If a compatible slot is free, inspect the Node log for the browser startup failure and verify with a real RemoteWebDriver smoke test.

Key Takeaways

  • Probe the Router from the same network as the test runner before changing browser settings.
  • Use GraphQL to compare Node stereotypes, active sessions, and queued requests.
  • A healthy Router does not guarantee a registered, matching, free browser slot.
  • Read Node logs when a matching free slot still cannot launch a browser.
  • In Compose, use the Hub service name from test containers and wait for Node registration.
  • Release every WebDriver session in a finally block or test fixture teardown.

If you need to fix Selenium Grid could not start a new session, the error appears when a RemoteWebDriver request reaches a Grid that cannot connect to, match, or launch a browser session. Read the nested cause and inspect Grid capacity before changing the test itself.

org.openqa.selenium.SessionNotCreatedException: Could not start a new session. Possible causes are invalid address of the remote server or browser start-up failure.

That wording is deliberately broad. A wrong Router URL, an unregistered Node, incompatible capabilities, a full session queue, and a browser crash can all surface at session creation. Selenium has recorded this exact Java message in a Grid issue. The commands below use a small Python client so you can repeat one request while examining the Grid. For a full setup, see the Selenium Python framework guide.

TL;DR

Probe the Router from the same machine or container that runs your test, then inspect slots and queue state:

curl -fsS http://127.0.0.1:4444/status | python3 -m json.tool
curl -fsS -H 'Content-Type: application/json' -d '{"query":"{ grid { nodeCount maxSession sessionCount sessionQueueSize } nodesInfo { nodes { uri status maxSession sessionCount stereotypes } } }"}' http://127.0.0.1:4444/graphql | python3 -m json.tool

If the first command cannot connect, fix the URL or network. If nodeCount is zero, restore Node registration. If no stereotype matches your requested browser, correct the options or add that Node type. If sessionCount equals usable capacity or the queue grows, reduce concurrency or add Nodes. If a free matching slot exists, inspect the Node log for the browser or driver failure. The Grid GraphQL reference documents these fields.

What the Error Actually Means

Selenium Grid's Router receives a W3C new-session request and places it in the New Session Queue. The Distributor looks for an available slot on a registered Node whose stereotype matches the requested capabilities. The Node then starts the browser driver and browser. Failure at any of these boundaries can leave the client without a session ID. A successful /status response only establishes that the Router answers; it does not prove a matching free slot or successful browser launch. The official Grid architecture describes these components.

Capture the exception class and the last Caused by: line, not just the top-level phrase. Connection refused suggests the client cannot reach the Router. A queue timeout points to capacity or matching. A Chrome startup error points to the Node's browser process. Do not diagnose all three by reinstalling the client library. Also distinguish a failed new session from a session that was created and later lost: the latter has a session ID and calls for a different investigation.

For consistent probes, install the Python Selenium package already approved by your project and save this as grid_smoke.py. Match the package version to your repository's dependency policy; no new package pin is required here.

import os
from selenium import webdriver

url = os.environ.get("GRID_URL", "http://127.0.0.1:4444")
options = webdriver.ChromeOptions()
options.add_argument("--headless=new")
driver = webdriver.Remote(command_executor=url, options=options)
try:
    driver.get("https://example.com/")
    print("session:", driver.session_id)
    print("title:", driver.title)
finally:
    driver.quit()

Run python3 grid_smoke.py after each change. It tests session creation, navigation, and cleanup. If your Grid contains Firefox Nodes only, change ChromeOptions() to FirefoxOptions() and remove the Chrome-specific argument. Do not request a browser the Grid cannot supply.

Root-Cause Decision Table

Symptom Root cause Fix
curl /status fails from the test runner Wrong Router address or blocked network Use the Router's reachable host and port
Grid responds, nodeCount is zero Node cannot register or is unavailable Repair Event Bus configuration and Node connectivity
Nodes appear but no stereotype matches Browser, platform, or version constraint mismatch Request supported capabilities or provision the requested Node
Queue grows while all matching slots are busy More concurrent requests than capacity Limit workers or add browser Nodes
A free matching slot exists but launch fails Browser or driver startup failure Read Node logs, repair image/resources, and retry
Local works but CI times out on first request CI startup race or wrong service hostname Gate on Grid readiness and probe inside the CI network
Docker test uses localhost for another service Container loopback mismatch Use the Compose service name and internal port
Sessions linger after tests finish Missing quit() or stuck clients Close drivers reliably and inspect active sessions

Treat a row as a hypothesis. Run its verification command and keep the resulting Router response, GraphQL snapshot, and Node log together. This separates queue pressure from browser startup failure even when both produce the same top-level Java message.

1. Fix Selenium Grid Could Not Start a New Session Caused by the Wrong Router URL

A remote client must call the Router it can actually reach. On a laptop running a published Docker port, http://127.0.0.1:4444 is reasonable. Inside a separate Compose service, that address refers to the test container itself. In a CI job, a service alias or forwarded port may differ from the laptop address. A stale /wd/hub suffix from a Grid 3 example can also obscure what endpoint your deployment exposes; start with the Grid 4 Router root shown in Selenium's Remote WebDriver documentation.

Set the URL explicitly in the client environment. From the same shell that starts tests, inspect the response and run the smoke script:

export GRID_URL=http://127.0.0.1:4444
curl -i --max-time 5 "$GRID_URL/status"
python3 grid_smoke.py

The first command should return HTTP data containing a value object. If it reports Connection refused, confirm the Grid process or published port; if it hangs, inspect routing, DNS, and firewall rules. If it returns an HTML login page, you may be calling a proxy rather than Grid. The second command should print a session ID and Example Domain as the title. For a remote deployment, replace only GRID_URL with the reachable Router address. Never copy the Node's internal address into the client: the Router is the public WebDriver endpoint. Keep credentials and any reverse proxy settings in your deployment's normal configuration rather than embedding them in article examples.

2. Restore Node Registration and Event Bus Connectivity

A Router may report ready while no browser Node has registered. In Hub and Node mode, the Node connects through the Event Bus. For the official Docker images, SE_EVENT_BUS_HOST must identify the Hub from the Node network, and the Hub's event ports must be reachable. A Node that repeatedly restarts or announces an unreachable URI may appear briefly and then disappear. Check actual registration before attempting to solve a nonexistent browser mismatch.

For a local Docker network, use the same matching image tag for Hub and Node. Replace <your-selenium-image-tag> with a tag matching the Selenium image version you have chosen; do not mix arbitrary tags. These commands are a minimal disposable setup:

docker network create selenium-grid
docker run -d --name selenium-hub --network selenium-grid -p 4444:4444 selenium/hub:<your-selenium-image-tag>
docker run -d --name selenium-chrome --network selenium-grid -e SE_EVENT_BUS_HOST=selenium-hub --shm-size=2g selenium/node-chrome:<your-selenium-image-tag>

If those names already exist, use your existing deployment and inspect it instead of creating duplicates. The Hub and Node need a shared Docker network; publishing 4444 lets a host client reach the Router. The Node's browser traffic stays on the Docker network. Verify registration and inspect both logs:

curl -fsS -H 'Content-Type: application/json' -d '{"query":"{ grid { nodeCount } nodesInfo { nodes { uri status stereotypes } } }"}' http://127.0.0.1:4444/graphql | python3 -m json.tool
docker logs selenium-hub --tail 80
docker logs selenium-chrome --tail 80

Look for at least one Node with status UP and a Chrome stereotype. If it is absent, check the Node's Event Bus host, container network membership, and repeated registration errors. Grid's Docker setup guide covers the deployment choices. In a distributed installation, repair the named Distributor, Queue, and Event Bus endpoints according to your actual topology rather than assuming Hub mode.

3. Match Browser Capabilities to an Available Slot

A Node can be healthy yet incapable of satisfying your request. browserName, browserVersion, and platformName are selection constraints. Requesting Firefox against Chrome-only Nodes, a precise version no Node advertises, or a platform name copied from another OS can leave the request queued. Extra vendor capabilities can also affect matching. First inspect the Node stereotypes, then remove unnecessary constraints from the client.

Run this query before editing test code:

curl -fsS -H 'Content-Type: application/json' -d '{"query":"{ nodesInfo { nodes { status stereotypes slotCount sessionCount } } }"}' http://127.0.0.1:4444/graphql | python3 -m json.tool

For the Chrome smoke client defined earlier, the request is intentionally simple: ChromeOptions() supplies browserName: chrome, while --headless=new is a browser launch argument. Do not add an invented browserVersion to force a match. If the query lists only Firefox, use this complete alternative as grid_firefox_smoke.py:

import os
from selenium import webdriver

options = webdriver.FirefoxOptions()
options.add_argument("-headless")
driver = webdriver.Remote(
    command_executor=os.environ.get("GRID_URL", "http://127.0.0.1:4444"),
    options=options,
)
try:
    driver.get("https://example.com/")
    print(driver.session_id, driver.title)
finally:
    driver.quit()

Verify with python3 grid_firefox_smoke.py; a session ID proves a Firefox slot matched and launched. If the product genuinely requires a particular browser or platform, add an appropriate Node rather than weakening the test requirement. When a Node advertises the expected stereotype yet matching still fails, capture the exact requested capabilities from the exception and compare each field with the GraphQL response. The session queue monitoring guide extends this comparison to sustained load.

4. Relieve Queue Pressure Without Hiding a Capacity Problem

Grid queues new-session requests until a matching slot becomes available or the request times out. The official Docker browser Node defaults to one concurrent session per container. Ten parallel test workers pointed at one Node do not create ten browser slots. They create contention, especially if tests retain drivers after failures. Measure occupied sessions and queued requests before raising a timeout.

curl -fsS -H 'Content-Type: application/json' -d '{"query":"{ grid { maxSession sessionCount sessionQueueSize } nodesInfo { nodes { status maxSession sessionCount } } }"}' http://127.0.0.1:4444/graphql | python3 -m json.tool

If the only compatible Node is full, reduce test concurrency to its available capacity for the next diagnostic run. For example, with pytest-xdist already installed, pytest -n 1 requests one worker. A plain pytest run is also serial. Then run python3 grid_smoke.py with no other browser users; a pass under low load supports a capacity diagnosis. If the suite needs sustained parallelism, scale the Node service instead of cramming many Chrome instances into one container. In Compose, docker compose up -d --scale chrome=3 scales a service named chrome whose configuration supports replicas. Verify the GraphQL nodeCount and maxSession increased, then rerun the intended worker count. The parallel test sharding guide helps plan workload against slot capacity.

SE_SESSION_REQUEST_TIMEOUT changes how long a request may wait in the queue, measured in seconds for the official images. Increase it only for expected bursts that clear within a measured interval. It does not add browser capacity or repair mismatched capabilities. Similarly, SE_NODE_MAX_SESSIONS and SE_NODE_OVERRIDE_MAX_SESSIONS can raise sessions per Node, but the Docker Selenium guidance recommends measuring CPU and memory first. More slots on an overloaded container can make browser startup less reliable.

5. Repair Browser Startup on a Matching Node

If a compatible Node has a free slot but session creation still fails, the Distributor has done its selection job and the Node or browser launch is the next target. Look below the client wrapper for Chrome failed to start, driver errors, sandbox problems, out-of-memory kills, or an image architecture mismatch. These are different failures, so preserve the first Node-side error line rather than replacing every setting at once.

In the Docker setup from section 2, inspect the Node and its resource state:

docker logs selenium-chrome --tail 150
docker inspect selenium-chrome --format '{{.State.Status}} {{.State.OOMKilled}}'
docker stats --no-stream selenium-chrome

A true OOMKilled value or a container that restarts during browser launch points to a resource limit. The official browser image examples use --shm-size=2g as a practical shared-memory setting, while its documentation says the right amount depends on workload. Do not assume 2 GB fixes a CPU-starved host. Compare memory and CPU while reproducing one session, then assign resources or reduce concurrent browsers. The Node image already includes its supported browser and driver; an unreviewed manual driver download inside that image can introduce a mismatch. For custom images, confirm the installed browser and driver pairing using the image's build documentation.

After a resource change, run python3 grid_smoke.py twice, one run after another. Both should print distinct session IDs, and GraphQL sessionCount should return to zero after each quit(). If the first succeeds and the second fails, inspect leftover browser processes, disk pressure, and Node cleanup behavior. The top-level Java message alone cannot prove a driver-version mismatch; the Node log must support that diagnosis.

6. Fix Selenium Grid Could Not Start a New Session in Docker Compose and CI

Compose introduces two addresses: the host-published Router address and the internal service address. A test container must use http://hub:4444 if the Hub service is named hub; localhost in that container is its own loopback. CI also needs to wait for Hub readiness and Node registration, not merely for the Hub container process to exist. A health check on /status can succeed before the browser Node appears, so gate on both endpoint reachability and nodeCount.

Save this as compose.yaml only in an isolated example project or adapt its service names to your existing file. Replace the placeholder image tag in both services with the same valid tag matched to your installed Selenium version. The tests service assumes your project contains grid_smoke.py and the Selenium Python package is available in the test image; install it through your project's locked dependencies before running the service.

services:
  hub:
    image: selenium/hub:<your-selenium-image-tag>
    ports:
      - "4444:4444"
  chrome:
    image: selenium/node-chrome:<your-selenium-image-tag>
    shm_size: "2gb"
    environment:
      SE_EVENT_BUS_HOST: hub
    depends_on:
      - hub
  tests:
    image: python:<your-python-image-tag>
    working_dir: /work
    volumes:
      - .:/work
    environment:
      GRID_URL: http://hub:4444
    depends_on:
      - hub
      - chrome
    command: sh -c 'python3 -m pip install -r requirements.txt && python3 grid_smoke.py'

Place your approved Selenium dependency in requirements.txt with the version your project uses. depends_on orders container startup; it does not prove the Node is registered. Start the services, then query Grid from the network or run the smoke test after a readiness gate:

docker compose up -d hub chrome
curl -fsS http://127.0.0.1:4444/status | python3 -m json.tool
curl -fsS -H 'Content-Type: application/json' -d '{"query":"{ grid { nodeCount } }"}' http://127.0.0.1:4444/graphql | python3 -m json.tool
docker compose run --rm tests

Run the last command only after nodeCount is at least one. In CI, implement a bounded poll of the same GraphQL field before launching the test service and fail with Hub and Node logs if the bound is exceeded. A ready: true Router response without a Node does not authorize test startup. For a broader setup, see Docker Compose test environments and Selenium Grid on Kubernetes. Kubernetes uses Service DNS and pod readiness instead of Compose service names, but the client-to-Router and Node-registration checks remain the same.

7. Release Sessions That Occupy Slots After Test Failures

A suite can create a session successfully and still prevent the next one from starting. A test that throws before calling quit() leaves its Node slot occupied until cleanup or session timeout. This pattern often looks like a Grid that works for the first few cases and then fails only under a full suite. Check active sessions with GraphQL and identify whether the count falls when tests complete.

curl -fsS -H 'Content-Type: application/json' -d '{"query":"{ grid { sessionCount } sessionsInfo { sessions { id nodeUri startTime } } }"}' http://127.0.0.1:4444/graphql | python3 -m json.tool

The earlier smoke scripts put quit() in finally, which runs after navigation or assertion exceptions. For pytest, a yield fixture gives the same ownership rule without requiring every test to remember teardown. Save this as conftest.py in a Python test project:

import os
import pytest
from selenium import webdriver

@pytest.fixture
def driver():
    options = webdriver.ChromeOptions()
    options.add_argument("--headless=new")
    browser = webdriver.Remote(
        command_executor=os.environ.get("GRID_URL", "http://127.0.0.1:4444"),
        options=options,
    )
    try:
        yield browser
    finally:
        browser.quit()

Verify with a test that deliberately fails after the fixture is created, then query sessionCount again. For example, save test_cleanup.py containing def test_cleanup(driver): assert driver.session_id; assert False and run pytest -q test_cleanup.py; the test failure is intentional, while the Grid count should return to its prior value. If it does not, inspect client termination, Node logs, and sessions that predate your run. Avoid deleting live sessions owned by other jobs. For long-running suites, periodically compare Grid sessions with active test workers and alert on a count that remains high after jobs finish.

How to Verify the Fix

Use a layered check from the actual test environment. First, curl -fsS "$GRID_URL/status" must reach the Router. Second, GraphQL must show an UP Node and a stereotype matching your browser. Third, python3 grid_smoke.py must print a nonempty session ID and the expected page title. Fourth, after quit(), GraphQL sessionCount should return to its baseline. These checks distinguish connectivity, matching, browser launch, and cleanup.

Repeat the smoke request several times serially, then at the concurrency your suite actually uses. A single passing request cannot prove that a ten-worker CI run has enough slots. Record maxSession, sessionCount, and sessionQueueSize during the parallel test. If the queue rises only during bursts but clears before the configured request timeout, capacity may be acceptable; if it stays high or sessions fail, add compatible Nodes or lower worker count. A CI rerun is meaningful only when it uses the same container network, image tags, environment, and browser options as the failed run.

Keep the full exception chain when the smoke still fails. A different exception after the fix can show that you passed one boundary and exposed the next one. For example, changing a bad URL may turn Connection refused into a Node browser launch error. That is progress, not proof that the URL change was ineffective. Compare the current failure's first Node log line and queue snapshot with the previous run before making another change.

Prevent It From Coming Back

Make Grid readiness an explicit CI precondition: Router reachable, required Node count present, and matching browser stereotype advertised. Keep Hub and Node images on the same selected tag and store that tag centrally in deployment configuration. Preserve Node logs as CI artifacts when a new-session request fails. A simple GraphQL dashboard or periodic query can reveal a growing queue, draining Nodes, and persistent occupied slots before the suite times out.

Set test worker count from measured usable slots, not from the number of CPU cores on the test runner. The browser processes run on Nodes, possibly on other machines. Give every driver a teardown path, and monitor whether sessions remain after the job ends. When scaling, add Nodes with the browsers the suite requests rather than only increasing a global timeout. If the Grid is internet-accessible, apply the authentication and network controls your organization requires; Grid itself is an execution service, so its reachability should be deliberate.

Interview Questions and Answers

Q: Why can Grid /status be healthy while new sessions fail?

The Router can answer while no suitable Node slot exists. I query GraphQL for Node status, stereotypes, occupancy, and queue size, then inspect Node logs if a free matching slot still cannot launch.

Q: What separates a capability mismatch from exhausted capacity?

A mismatch means no advertised stereotype satisfies the requested browser, version, or platform, even when slots are free. Exhaustion means a matching Node exists but its sessions occupy the available slots. I compare the request with Node stereotypes and sessionCount.

Q: Why does localhost often fail in Compose?

Each container owns its own loopback interface. A test container must call the Hub by its service DNS name and internal port, while a host process can use the published host port.

Q: When should SE_SESSION_REQUEST_TIMEOUT change?

Only when a measured, expected queue burst clears shortly after the current limit. A longer wait cannot create a missing browser Node, free a leaked session, or repair a browser crash.

Q: What evidence supports a browser resource failure?

I look for the first browser launch error in the Node log, container OOM state, and CPU or memory pressure during a one-session reproduction. The generic client exception by itself is insufficient.

Q: How do you prove cleanup works?

I force a test failure after session creation and check that fixture teardown still calls quit(). The Grid session count should return to its pretest value, accounting for any other jobs using the same Grid.

Common Mistakes

  • Assuming HTTP 200 from /status means Chrome can launch. Check matching Node slots and make one real session request.
  • Increasing a client test timeout to solve a Router connection error. Inspect reachability from the client network.
  • Pinning browserVersion to a value no Node advertises. Use the actual Grid stereotype as evidence.
  • Raising sessions per Node without checking CPU, memory, and shared memory. Add Nodes when one container is saturated.
  • Using host localhost inside the test container. Use the Hub service name.
  • Treating depends_on as a Node-registration check. Query GraphQL before starting tests.
  • Catching a test error without running driver.quit(). Use finally or a fixture teardown.
  • Reading only the top-level SessionNotCreatedException. Preserve the nested cause and the Node log.

Conclusion

To fix Selenium Grid "Could not start a new session", find the first boundary that fails: Router reachability, Node registration, capability matching, free capacity, browser launch, or session cleanup. Run the smoke request from the failing environment, make one targeted change, and verify both a new session and its teardown. Keep Grid state and Node logs beside CI failures so the next incident starts with evidence.

Interview Questions and Answers

How would you triage a Selenium Grid new-session failure?

I first reproduce from the same environment as the test runner and probe the Router `/status` endpoint. I query GraphQL for Node status, stereotypes, active sessions, and queue size. If a matching slot is free, I inspect the Node's first browser launch error and make one targeted change.

What does the Distributor do during session creation?

It obtains queued new-session requests and selects an available registered Node slot whose stereotype matches the requested capabilities. A healthy Router does not mean the Distributor can find such a slot. Queue and Node metrics show where selection is blocked.

How would you diagnose a full Grid?

I compare active sessions with matching Node capacity and watch the session queue during the failing workload. Then I run one serial smoke request to separate load from basic launch health. I lower worker count or add compatible Nodes based on measured demand.

Why is Grid status alone an incomplete CI readiness gate?

The Router may respond before any browser Node has registered. I gate on both Router reachability and the required Node count, then attempt a real smoke session. A successful request proves the requested browser can actually launch.

What is a Selenium Grid capability stereotype?

It describes what a Node slot can satisfy, such as browser name, version, and platform. I compare it with the requested W3C capabilities instead of guessing from the container name. An overly specific version constraint can exclude otherwise healthy Nodes.

How would you handle session leaks in a parallel suite?

I use teardown that calls `quit()` even after assertions fail and compare GraphQL session counts before and after a forced failure. I preserve active session IDs when investigating other jobs on the same Grid. I do not delete unknown sessions just to make a test pass.

When would you increase per-Node concurrency?

Only after measuring CPU, memory, shared memory, and browser stability with the current slot count. The official Docker images default to one session per browser Node. Scaling Node replicas is often easier to reason about than overloading a single container.

Frequently Asked Questions

What does Selenium Grid Could not start a new session mean?

The remote client did not receive a new browser session ID. The cause may be Router connectivity, missing or incompatible Nodes, a full queue, or a browser launch error. Use the nested exception and Grid state to identify the failing boundary.

Can Grid be ready but unable to create sessions?

Yes. Router readiness shows the HTTP endpoint responds, while Node registration, capability matching, and browser startup are separate checks. Query GraphQL and attempt one remote smoke session.

How do I see whether Selenium Grid has free slots?

Query `/graphql` for `grid { maxSession sessionCount sessionQueueSize }` and inspect `nodesInfo` for each Node's status and sessions. Compare available capacity only among Nodes matching the requested browser.

Why does Selenium Grid fail only in Docker Compose?

The test container may be calling its own `localhost` instead of the Hub service. Set the remote URL to the Hub service name and internal port, then probe it inside the test container. Also wait until the browser Node registers.

Should I increase SE_SESSION_REQUEST_TIMEOUT?

Only if requests wait behind valid occupied slots and the queue usually clears after the existing limit. It is measured in seconds in the official Docker images. It cannot solve missing Nodes, bad capabilities, or crashed browsers.

Why do later tests fail after the first Grid test passes?

Earlier tests may leave sessions open when their error paths skip `quit()`. Query active sessions after each run and put cleanup in `finally` or a yield fixture. Node resource exhaustion can show a similar pattern, so also inspect Node logs.

How do I tell a browser crash from a capability mismatch?

A mismatch leaves no suitable stereotype for the requested capabilities. A browser crash can occur after a matching free slot is selected and normally leaves a Node-side startup error. Compare GraphQL stereotypes with the request, then read the Node log.

Related Guides