QA How-To
How to Fix Playwright Tests That Pass Locally but Fail in CI
Fix Playwright tests that pass locally but fail in CI: check browser installs, server readiness, environment variables, timing, shared state, and screenshots.
21 min read | 3,591 words
TL;DR
Use a clean install, install Playwright browsers with dependencies, and rerun one failing spec with CI settings and one worker. Probe the app URL from the runner, inspect the first failure trace, and fix the specific installation, startup, configuration, timing, state, or screenshot difference it reveals.
Key Takeaways
- Classify whether failure happens at installation, server startup, navigation, assertion, or screenshot comparison.
- Install browser binaries and Linux dependencies after the locked npm install, or match the Docker image to Playwright.
- Probe the exact app URL from the CI runner and disable local server reuse in CI.
- Replace fixed sleeps with awaited locators and assertions tied to observable state.
- Use one worker to diagnose races, then isolate mutable backend data before scaling parallelism.
- Generate and review screenshot baselines in the same environment as CI.
- Keep first-failure traces and reports so retries do not hide the cause.
If you search for "fix Playwright tests pass locally but fail in CI," the problem usually appears after a clean CI checkout launches the same suite that passed on your laptop. The last line may be a timeout, a missing browser executable, a failed navigation, or a screenshot mismatch. Treat that line as a symptom and compare the two environments before changing test code.
Timeout of 30000ms exceeded.
That is a real Playwright test-timeout message, not a universal message for every CI failure. Your CI log may instead say browserType.launch: Executable doesn't exist at ... or Error: Timed out waiting 60000ms from config.webServer. Preserve the exact error, project name, failed assertion, and runner log. The path and timeout value in your own run identify which part failed.
TL;DR
Run npm ci, npx playwright install --with-deps, and npx playwright test in a clean Linux environment. In playwright.config.ts, set workers: process.env.CI ? 1 : undefined, disable web server reuse in CI, and retain a trace from the first failing attempt. Probe the app URL from the runner. If the test still fails, classify it with the table below instead of adding sleeps or blindly increasing every timeout.
npm ci
npx playwright install --with-deps
CI=1 npx playwright test --workers=1
These commands assume a Node project with @playwright/test in its lockfile and a Linux CI agent where browser dependencies can be installed. In the official Playwright CI guide, the same install sequence is the baseline. If npm ci fails, repair the dependency or lockfile error before diagnosing browser behavior. If only one spec fails, rerun that file with CI=1 npx playwright test tests/example.spec.ts --workers=1, substituting its actual path.
What the Error Actually Means
"Passes locally, fails in CI" is not one Playwright error. It describes a difference in inputs to the test: operating system, Node dependencies, browser binaries, app build, service availability, network access, environment variables, time zone, viewport, CPU budget, or shared backend state. Playwright can wait for actionable elements and retry web assertions, but it cannot supply an absent secret or make two workers safely edit the same account.
First locate the boundary. A failure before any test name usually means installation, browser launch, or webServer startup. A failure at page.goto() points toward base URL, DNS, TLS, or app readiness. A failure at a locator assertion points toward page state, data, permissions, or timing. A screenshot mismatch can mean a real UI regression, but also a baseline generated on another OS or with different fonts. An intermittent failure that disappears with one worker suggests resource contention or shared state; it does not prove either one without a controlled rerun.
Use npx playwright test --list to confirm what CI collected. Record the failing project with npx playwright test --list --project=chromium only if your configuration defines a project named chromium. Then read the first meaningful error in the job log, not just the summary. A timeout in a test and a webServer.timeout have separate controls, as explained by Playwright's timeout documentation.
Root-Cause Decision Table
| Symptom in CI | Likely root cause | First fix to test |
|---|---|---|
Executable doesn't exist or launch dependency error |
Browser binary or Linux libraries absent, possibly image/package mismatch | Install browsers with dependencies after npm ci; match Docker image to package |
config.webServer times out or page.goto() cannot connect |
App did not start, wrong host or port, or unhealthy dependency | Start the app deterministically and probe its exact URL |
| Test fails only on clean checkout or on pull requests | Missing env value, fixture, authentication file, or build artifact | Declare required inputs and generate state in the job |
| Locator or assertion times out only on slower runner | Test observes a transitional state or uses a fixed sleep | Wait for a meaningful locator, response, or URL |
Failure disappears at --workers=1 |
Shared backend data or CPU contention | Give tests unique data and limit workers deliberately |
toHaveScreenshot() differs on Linux |
Baseline environment, fonts, animation, or data differ | Generate baseline in matching environment and stabilize page |
| Local and CI execute different tests or browser projects | Script, config, filter, or dependency drift | Run the exact CI command from a clean checkout |
Treat each row as a hypothesis. Make one change, run its verification command, and compare the next failure. A passing retry alone is weak evidence because retries start a fresh worker and may encounter different data. Keep the first-run trace when the cause is still unknown.
1. Fix Playwright Tests Pass Locally but Fail in CI When Browsers Are Missing
The package in node_modules is not the browser executable. npm ci installs the locked JavaScript packages; npx playwright install --with-deps installs the browsers and, on supported Linux systems, their operating-system dependencies. A developer laptop may have browser files from a previous run while a fresh worker has none. The launch error often includes a cache path and tells you to run the install command. Read which project is failing before installing only Chromium when the suite also uses Firefox or WebKit.
Put installation after dependency installation in the job. This GitHub Actions example uses the action major versions shown in the current official CI example; check that page when updating your workflow. It assumes a root package-lock.json and tests invoked from the repository root.
name: Playwright tests
on: [push, pull_request]
jobs:
e2e:
runs-on: ubuntu-latest
timeout-minutes: 60
steps:
- uses: actions/checkout@v6
- uses: actions/setup-node@v6
with:
node-version: lts/*
cache: npm
- run: npm ci
- run: npx playwright install --with-deps
- run: npx playwright test
env:
CI: 'true'
- uses: actions/upload-artifact@v5
if: ${{ !cancelled() }}
with:
name: playwright-report
path: playwright-report/
if-no-files-found: ignore
Verify the install in the job with npx playwright --version followed by npx playwright test --list. Then run one actual smoke spec, for example npx playwright test tests/smoke.spec.ts, using a file that exists in your repository. Listing tests does not launch browsers; the smoke run does. If launch still fails, run DEBUG=pw:browser npx playwright test tests/smoke.spec.ts and inspect the missing executable or library in the browser logs. The browser installation guide covers that narrower failure.
For Docker, use mcr.microsoft.com/playwright:v<your-playwright-version>-noble and replace the placeholder with the version printed by npx playwright --version from the locked project. The official Docker guide warns that an image/package mismatch can make executables undiscoverable. The image supplies browsers and system libraries, but your job still needs npm ci to install the project's package. The Docker for Playwright guide explains image setup in more depth.
2. Fix Playwright Tests Pass Locally but Fail in CI When the App Is Not Ready
Your laptop may have a dev server already listening. A clean CI worker starts with no such process, so a successful local page.goto('/') says little about the CI startup path. Use Playwright's webServer to start the app and wait for an HTTP response. For a Vite application whose dev script accepts CLI arguments, choose a fixed host and port and make the readiness URL match baseURL. Replace the script if your project uses another server.
// playwright.config.ts
import { defineConfig, devices } from '@playwright/test';
const origin = 'http://127.0.0.1:4173';
export default defineConfig({
testDir: './tests',
projects: [{ name: 'chromium', use: { ...devices['Desktop Chrome'] } }],
use: { baseURL: origin, trace: 'retain-on-failure' },
workers: process.env.CI ? 1 : undefined,
webServer: {
command: 'npm run dev -- --host 127.0.0.1 --port 4173 --strictPort',
url: `${origin}/`,
reuseExistingServer: !process.env.CI,
timeout: 120_000,
stdout: 'pipe',
stderr: 'pipe',
},
});
--strictPort makes a conflict visible rather than letting Vite silently select another port. reuseExistingServer: !process.env.CI lets a developer use a local process but makes CI own its server. The startup timeout applies to webServer; it does not extend a slow test assertion. If your app's / returns a failure until an API starts, configure a readiness path that correctly represents essential service availability. For a dedicated server failure, see the Playwright webServer configuration guide.
Verify without Playwright first: in one terminal run npm run dev -- --host 127.0.0.1 --port 4173 --strictPort; in another run curl -i --max-time 5 http://127.0.0.1:4173/. Stop that manual server, then run CI=1 npx playwright test tests/smoke.spec.ts --project=chromium. If the app starts locally but not in CI, inspect piped startup output for a missing package, compile failure, or required env value. If the test process is in a different Docker container, 127.0.0.1 refers to the test container; use the app service name and internal port instead. Probe from inside the test container with docker compose exec tests node -e 'fetch("http://app:4173/").then(r => console.log(r.status)).catch(e => { console.error(e); process.exit(1); })', assuming services named tests and app. The Docker Compose test environments guide covers that network boundary.
3. Supply the Configuration and Authentication CI Actually Needs
A shell profile, ignored .env file, or locally generated playwright/.auth/user.json will not appear in a fresh checkout. If CI logs show a missing file, an unauthorized response, or navigation to an empty base URL, list the inputs that the test needs. Do not print secret values to logs. Make the test fail early on absent configuration, and create authentication state during the job through a setup project or a login fixture. Never commit a real storage-state file: it may contain session cookies.
Here is a small check for a required app origin. The variable is safe to name in logs; its value should be provided by the CI job or environment. It replaces hard-coded origins when staging and local deployments differ.
// playwright.config.ts
import { defineConfig } from '@playwright/test';
const baseURL = process.env.E2E_BASE_URL;
if (!baseURL) throw new Error('E2E_BASE_URL is required');
new URL(baseURL);
export default defineConfig({
testDir: './tests',
use: { baseURL, trace: 'retain-on-failure' },
});
Verify config loading with E2E_BASE_URL=http://127.0.0.1:4173 npx playwright test --list; remove the variable and confirm the intended early error appears. Then set the actual CI URL in the job's env map and probe it with curl -i --max-time 5 "$E2E_BASE_URL" before tests. A 401 response may be expected for a protected app, but its auth flow must be explicit. If login state is generated in setup, make the setup wait for the final authenticated page before saving storageState, as described in Playwright's authentication guide. If the setup project fails, diagnose it first rather than allowing dependent tests to use a stale local auth file. The Playwright global setup guide shows the setup-project pattern.
Pull requests from forks often have a different secret policy than internal branches. Make the workflow's branch and secret assumptions visible. A test that needs credentials should report a missing required input or be deliberately scoped to a trusted environment. Hard-coding a token into a fixture makes the test pass once and creates a security problem.
4. Replace Timing Guesses with Observable Conditions
A local machine can render so quickly that expect(await locator.isVisible()).toBe(true) happens after the element appears. On a slower CI worker, the same immediate check may run during loading. page.waitForTimeout(2000) only changes the race window; it does not prove the page is ready. Use a locator action followed by an awaited Playwright assertion, which retries until the expected state or its assertion timeout. Prefer a stable user-facing role or label over a generated CSS class.
This standalone test uses the public Playwright site so the APIs can be pasted without a custom app. In your suite, replace the URL and accessible names with your application's real contract.
// tests/readiness.spec.ts
import { test, expect } from '@playwright/test';
test('documentation landing page is ready', async ({ page }) => {
await page.goto('https://playwright.dev/');
await expect(page.getByRole('heading', { name: /Playwright enables reliable web automation/ })).toBeVisible();
});
For an application action, wait for the result the user sees: await page.getByRole('button', { name: 'Save' }).click(); await expect(page.getByRole('status')).toHaveText('Saved');. That pair assumes your app exposes a Save button and a status message. If a response is the contract, register page.waitForResponse() before the click, ideally with Promise.all, and check its status. Avoid networkidle as a blanket readiness signal when an app keeps analytics, streaming, or polling requests open.
Verify this section with CI=1 npx playwright test tests/readiness.spec.ts --workers=1. A public-site check needs network access; for an isolated CI network use your own app's stable route instead. Compare the failing call log with the locator's accessible name and the page state in the trace. When the test itself exceeds its budget, adjust only the appropriate timeout after measuring the real operation; the Playwright test timeout guide separates test and assertion limits. The official best-practices guide documents locator auto-waiting and web assertions.
5. Remove Shared-State Races and Resource Contention
Playwright gives each test an isolated browser context, but two workers can still modify the same backend record or account. CI may run more workers than your laptop or may have less memory per worker. Start diagnosis with CI=1 npx playwright test --workers=1 --repeat-each=3. If a failure disappears, inspect both shared data and resource usage. One worker is a sensible CI baseline from the official CI guide, but do not treat it as a permanent substitute for data isolation if you later scale out.
Use a unique record identifier per test. The following spec assumes your app exposes an API endpoint POST /api/notes and returns a 201 JSON response with a title field. Replace those application-specific details with your real endpoint; request.post() and testInfo.testId are Playwright APIs. The example illustrates a concrete isolation contract, rather than a magic random delay.
// tests/notes.spec.ts
import { test, expect } from '@playwright/test';
test('creates a uniquely named note', async ({ request }, testInfo) => {
const title = `ci-note-${testInfo.testId}`;
const response = await request.post('/api/notes', { data: { title } });
expect(response.status()).toBe(201);
const note = await response.json();
expect(note.title).toBe(title);
});
This API spec requires use.baseURL in playwright.config.ts, a running app, and an endpoint matching the stated contract. If your backend needs authentication, create a distinct test account per worker or use an isolated tenant. A shared account may be safe for read-only tests but unsafe when tests change profile settings, carts, or permissions. Avoid deriving uniqueness only from workerIndex, because a failed worker can be replaced; testInfo.testId identifies the test, and you can add a CI run identifier when two jobs target the same backend. The Playwright parallelism guide explains worker isolation and shared backend data.
Verify by running the same suite at --workers=1 and then at --workers=2, with a clean backend each time. If both pass with unique records, raise worker count gradually while watching runner memory and service limits. If the one-worker run still fails, return to its trace; parallelism was not sufficient to explain the problem. The flaky test CI guide covers suite-level tracking.
6. Match Visual Baselines to the CI Environment
Screenshot comparison is sensitive to browser build, OS, fonts, viewport, animations, and dynamic content. A macOS baseline can differ from Linux even when the page is functionally identical. Playwright's visual comparison guide recommends generating baselines in the same environment used for comparison. First inspect the actual, expected, and difference images in the report. A mismatch in one label after a product change calls for review; text edges differing everywhere suggest rendering environment drift.
Use an explicit viewport and stabilize known animation and moving regions. This example compares a public page only as an API demonstration; for a real baseline, choose a page you control, seed its data, and generate the expected image in the CI-equivalent environment.
// tests/visual.spec.ts
import { test, expect } from '@playwright/test';
test.use({ viewport: { width: 1280, height: 720 } });
test('landing screenshot', async ({ page }) => {
await page.goto('https://playwright.dev/');
await expect(page.locator('body')).toHaveScreenshot('landing.png', {
animations: 'disabled',
});
});
Run npx playwright test tests/visual.spec.ts --update-snapshots in the same Linux image or runner type as CI, inspect the generated snapshot, and commit it only after confirming the UI is correct. Then verify with CI=1 npx playwright test tests/visual.spec.ts. A public site may change independently, so an owned page with deterministic data is a better lasting test. Do not set a huge pixel tolerance to hide an unexamined regression. If the first run only reports a missing snapshot, generate and review the baseline before comparing. The Playwright screenshot comparison guide goes deeper on fonts and masking.
7. Reproduce the Exact CI Command and Capture the First Failure
Local runs often use --headed, one browser, a selected file, or an already built app, while CI uses all configured projects and a clean install. Write down the CI command and run it with CI=1 from a clean checkout. npx playwright test --list reveals what the runner sees, including projects and test titles. A wrong testDir, .gitignore interaction, test.only, or shell filter can make local and CI suites different without a browser problem. Use forbidOnly: !!process.env.CI to make an accidental focused test fail clearly.
Configure a report and retain the first failed attempt's trace. trace: 'retain-on-failure' records an initial failure even when retries are zero; on-first-retry records the retry instead. For a one-off investigation, this config keeps failure evidence and constrains worker count. Do not stack unexplained retries on top of a failing test.
// playwright.config.ts
import { defineConfig } from '@playwright/test';
export default defineConfig({
testDir: './tests',
forbidOnly: !!process.env.CI,
retries: 0,
workers: process.env.CI ? 1 : undefined,
reporter: [['list'], ['html', { open: 'never' }]],
use: { trace: 'retain-on-failure' },
});
Verify with CI=1 npx playwright test --list, then CI=1 npx playwright test. After a failed run, use npx playwright show-report locally on the downloaded report or npx playwright show-trace path/to/trace.zip for an extracted trace. The trace viewer guide shows actions, snapshots, network events, and call logs. Upload playwright-report/ and test-results/ as CI artifacts if your workflow otherwise deletes them. Treat traces as potentially sensitive: URLs, DOM snapshots, and request data can contain test credentials. The GitHub Actions for Playwright guide covers report collection.
How to Verify the Fix
Start with the smallest failing spec on the same OS, browser project, and command as CI. Record the original failure category, then run the section's probe. A browser install fix is verified by launching a real browser, not by listing tests. A server fix is verified by curl from the test environment and by a browser navigation. A state-isolation fix is verified at both one and multiple workers. A visual fix is verified by reviewing the baseline and running the comparison again.
After the focused run passes, run the complete suite with CI=1 npx playwright test --workers=1. If that passes, restore the intended worker count and run again. Compare all configured projects rather than assuming Chromium proves WebKit or Firefox. A green rerun with no code or environment change is evidence of intermittence, not proof of repair. Inspect the first-run trace and repeat the scenario under the condition that triggered it.
A useful local reproduction is a clean dependency install followed by the exact CI command. If you use Docker, replace <your-playwright-version> with the version from the project's lockfile before running the container; do not pull an arbitrary image tag. Keep the app and test process network relationship the same as CI. When you cannot reproduce locally, preserve CI artifacts and add a small diagnostic command in the job, such as node --version, npx playwright --version, or a curl of the configured base URL. Avoid logging secrets or entire environment dumps.
Prevent It From Coming Back
Keep package-lock.json committed and use npm ci in CI. Install the browsers required by configured projects, or pin a Playwright Docker image that matches the package version. Make service startup deterministic with one origin, a health probe, and visible stderr. Generate auth state during the job, isolate mutable backend data, and use web assertions tied to observable behavior. Review screenshot baselines in a matching environment when dependencies or browser builds change.
Retain traces and the HTML report on failed jobs. Run a short smoke test before a large matrix so browser launch, server readiness, and login failures appear early. Measure how often retries rescue tests; recurring rescued tests still need root-cause work. Review changes to the CI image, browser package, fonts, Node runtime, environment variables, and service topology together, since any one can invalidate a previously stable comparison. The CI troubleshooting interview guide offers a useful checklist for explaining these boundaries to a team.
Interview Questions and Answers
Q: What is the first distinction you make when a Playwright test fails only in CI?
I locate whether failure occurs before test execution, during navigation, or at an assertion. That tells me whether to inspect installation and web server setup, network and configuration, or page state. I preserve the first failure log and trace before changing timeouts.
Q: Why can npm ci succeed while a browser launch fails?
The npm package and browser executables are separate installations. A clean Linux job may also lack browser system libraries. I run npx playwright install --with-deps or use a version-matched Playwright image, then launch a smoke spec.
Q: How do you distinguish slow UI from shared-state interference?
I rerun the failing spec with one worker and inspect the trace's page state and call log. If the result changes with worker count, I check shared accounts, records, and resource pressure. If one worker still sees a late element, I wait for the actual UI condition.
Q: Why might a screenshot pass on macOS and fail on Linux?
Rendering differs by browser build, fonts, and OS. I inspect the diff and generate baselines where CI compares them. I also fix viewport, animations, and dynamic data before considering tolerance.
Q: When would you raise a timeout?
Only after finding the precise phase and measuring a legitimate operation that needs longer. webServer.timeout, test timeout, and expect timeout are separate. A missing dependency, wrong URL, or fixed-sleep race will not be repaired by a larger number.
Q: What do retries tell you?
A passing retry proves the test can pass under a changed attempt, not that the first failure was harmless. I capture a trace of the first attempt or first retry, inspect worker and backend state, and count flaky outcomes as work to fix.
Common Mistakes
- Installing JavaScript packages but omitting browser binaries or Linux libraries.
- Copying a Docker image tag that differs from the locked Playwright package.
- Reusing a local dev server while CI has no equivalent process.
- Raising every timeout before probing the exact failed URL or locator.
- Using
page.waitForTimeout()or an immediateisVisible()check for asynchronously rendered UI. - Sharing one mutable account or record across parallel workers and CI jobs.
- Updating screenshot baselines on a different OS without reviewing the diff.
- Relying on a green retry while discarding the first-run trace.
- Running a selected local project and assuming it matches the full CI matrix.
Conclusion
To fix Playwright tests pass locally but fail in CI, make the environments comparable and test the first broken boundary: browser installation, app readiness, required configuration, UI synchronization, shared state, or visual rendering. Verify each proposed fix with a focused command, then run the full CI-shaped suite. Keep the failure artifacts so the next divergence has evidence rather than guesswork.
Interview Questions and Answers
How would you triage a Playwright test that fails only in CI?
I classify the first failure by phase: install, browser launch, server setup, navigation, assertion, or screenshot. I compare the local and CI command, OS, browser package, app URL, and required inputs. Then I run one focused verification command for the suspected boundary and preserve the trace.
Why is npm ci not enough to run Playwright on a Linux runner?
npm ci installs the locked Node packages, while Playwright browser executables and Linux libraries are separate. I run npx playwright install --with-deps after npm ci, or use a Docker image matched to the installed package. I verify with a browser-launching smoke spec.
How can webServer reuse hide a local versus CI difference?
A local listener may already satisfy Playwright readiness, so the configured command never has to start. CI begins without that listener and reveals a broken command or missing input. I set reuseExistingServer to !process.env.CI and probe the configured URL.
How do you separate a timeout caused by slow UI from a bad locator?
I inspect the trace snapshot and Playwright call log at the failed assertion. If the target eventually appears, I use an awaited web assertion tied to the expected state. If it never appears, I check accessible name, navigation, auth, and application errors before raising a timeout.
What does a one-worker rerun prove?
It provides a controlled comparison. If the failure disappears, I inspect shared backend records and CPU or memory contention. I still rerun with unique data and the intended worker count before claiming the race is fixed.
Why can screenshot baselines differ across environments?
Browser build, OS fonts, viewport, animation, and dynamic content affect pixels. I generate baselines in a CI-equivalent environment and inspect expected, actual, and diff images. A changed product UI requires review before updating the baseline.
Which trace setting would you use with no retries?
I use trace: 'retain-on-failure' so the first failing attempt is available. With retries enabled, on-first-retry captures the retry, which may be useful but can miss the initial page state. I choose based on which attempt I need to inspect.
Frequently Asked Questions
Why do Playwright tests pass locally but fail in GitHub Actions?
GitHub Actions starts from a clean Linux checkout, which may lack browser binaries, system libraries, local environment files, or a prestarted app. Compare the first failing step and run the same install and test commands locally in a clean environment. Probe the app URL from the job.
Does npm ci install Playwright browsers?
No. It installs the JavaScript dependencies from the lockfile. Run npx playwright install --with-deps on supported Linux runners, or use a Playwright image whose tag matches the package version.
Should I increase the Playwright timeout for CI failures?
Only if the log identifies a legitimately slow operation and you have measured it. Test, expect, and webServer timeouts govern different phases. A missing browser, wrong base URL, or absent secret needs a different fix.
How do I debug a test that passes on a retry in CI?
Retain a trace of the first failed run, inspect its call log and page snapshot, then rerun under the same worker and data conditions. A retry uses a fresh attempt and may avoid a race without removing it. Track repeated rescued tests as flaky failures.
Why do screenshot tests fail on CI but pass on my Mac?
Linux, macOS, browser builds, fonts, and viewport settings can render different pixels. Inspect the diff and regenerate the baseline in the same environment CI uses. Stabilize animation and dynamic content before accepting a new baseline.
What does --workers=1 tell me about a CI-only failure?
If a failure disappears with one worker, shared backend data or resource contention becomes more likely. It is not a complete diagnosis. Give each test unique records or accounts, then verify again with the intended parallelism.
What should I upload after a failed Playwright CI run?
Keep the HTML report and test-results artifacts, including traces and screenshot diffs. They show the first failing action and page state. Handle them as sensitive test artifacts because they may contain URLs or application data.
Related Guides
- How to Run tests in headed mode in Playwright (2026)
- How to Debug a failing test in VS Code in Playwright (2026)
- How to Fix "Cannot find module '@playwright/test'" in Playwright
- How to Fix "The Cypress binary is missing" in CI
- How to Fix Appium "No Chromedriver found that can automate Chrome" in a WebView
- How to Fix Playwright "Element is not attached to the DOM"