Resource library

QA How-To

How to Fix Playwright toHaveScreenshot "Screenshot comparison failed"

Fix Playwright toHaveScreenshot screenshot comparison failed errors by reading image diffs, stabilizing captures, and matching CI baselines across browsers.

19 min read | 3,044 words

TL;DR

Open the expected, actual, and diff images, then separate dimension changes from same-sized pixel changes. Fix unstable layout, content, fonts, or environment; update the baseline only for reviewed intentional UI changes. Re-run the focused test without update mode in the environment where it failed.

Key Takeaways

  • Inspect expected, actual, and diff images and check dimensions before adjusting comparison options.
  • Update a baseline only after confirming an intentional UI change and reviewing its new image.
  • Assert the final loaded state before capture because screenshot settling can occur on a skeleton.
  • Use deterministic data and narrow masks for volatile content that is outside the visual requirement.
  • Match browser, operating system, fonts, viewport, scale, and Playwright version across baseline and CI runs.
  • Apply a small measured pixel budget only for known rendering noise, then verify a real regression still fails.

To fix Playwright toHaveScreenshot screenshot comparison failed errors, first open the expected, actual, and diff images from the failed test. The error appears when a captured page or element does not match its stored baseline, including cases where its dimensions differ. Decide whether the UI changed intentionally, the capture is unstable, or the baseline was made in another rendering environment before changing any comparison setting.

Error: Screenshot comparison failed:

That line is the real Playwright error prefix. The following lines vary by failure: they may report different pixel counts, different image dimensions, and paths labeled Expected, Received, and Diff. Keep those artifacts. A changed button label should fail; a clock or unpinned font should not make every build noisy.

TL;DR

  1. Open the HTML report with npx playwright show-report and compare the expected, actual, and diff images. Read the dimensions before examining pixel counts.
  2. If the design change is intended, review the new image and run the focused test with --update-snapshots; commit its baseline. If it is unintended, fix the UI.
  3. If only volatile content differs, make test data deterministic or mask the exact region. If the images have different dimensions, stabilize layout, viewport, fonts, and environment.
  4. Capture baselines in the same browser, operating system, fonts, viewport, and headless mode used by CI. Use a Playwright Docker image whose version matches the installed package when containerizing the suite.

Start with the Playwright toHaveScreenshot guide if the matcher itself is new to you. This article concentrates on why an existing comparison fails and how to prove each repair.

What the Error Actually Means

await expect(page).toHaveScreenshot('name.png') asks Playwright Test to capture a stable screenshot and compare it with a committed reference image. The matcher waits until two consecutive screenshots match before comparing the last capture with the baseline. This settling behavior helps, but it cannot make a live clock, random avatar, late font, changing ad, or inconsistent server response deterministic.

Read the error in this order. Expected points to the baseline. Received is what the browser rendered now. Diff highlights changed pixels. If Playwright says Expected an image 646px by 169px, received 646px by 168px., the image sizes differ. A pixel difference budget is not a remedy for a height mismatch. If the sizes match and the error reports different pixels, the diff image tells you where the change occurred. A compact colored area usually suggests a real component change; scattered text-edge pixels often suggest rendering or font drift. These are diagnostic patterns, not a reason to approve an image without review.

A missing baseline is a separate case. Playwright's documented first-run message begins Error: A snapshot doesn't exist at, then writes an actual image for you to inspect. Store an approved baseline in version control. Never make the CI job update snapshots automatically just to turn a failure green. For examples of naming and scoping captures, see Playwright toHaveScreenshot examples.

Root-Cause Decision Table

Symptom Likely root cause Fix First verification
A snapshot doesn't exist at Baseline absent or wrong snapshot path Generate, review, and commit the correct baseline Run the same test again without update mode
Large, coherent changed region Intended redesign or product regression Approve a new baseline or repair the UI Review expected, actual, and diff side by side
Expected an image ... received ... Size, viewport, font, or content-height drift Stabilize dimensions and rendering Compare image sizes in report
Date, avatar, ad, or count changes Nondeterministic content Fix data, mask a small region, or apply screenshot CSS Repeat the focused test
Spinner or skeleton appears Capture starts before real readiness Assert a meaningful loaded state first Run the test repeatedly
Thin halos around text Browser, OS, font, scale, or rasterization drift Match the rendering environment; use a small measured tolerance only if needed Compare local and CI settings
Passes locally, fails in CI Different image, browsers, fonts, or headless mode Regenerate baseline in the same CI-like environment Run in pinned container

The table is a routing tool. Diagnose a size difference before tuning pixels; inspect a visible design change before calling it noise. The snapshot testing guide explains where visual assertions fit alongside functional checks.

1. Fix Playwright toHaveScreenshot Screenshot Comparison Failed When the Baseline Is Missing or Stale

Create one small, reproducible case first. In a Playwright Test project with @playwright/test installed, save this as tests/visual.spec.ts. It uses page.setContent, so it needs no application server. The named screenshot makes review easier than an automatically generated ordinal name.

import { test, expect } from '@playwright/test';

test('invoice card visual baseline', async ({ page }) => {
  await page.setViewportSize({ width: 800, height: 600 });
  await page.setContent(`
    <main style="font: 16px Arial; width: 320px; padding: 24px;
                 border: 1px solid #333; background: white">
      <h1>Invoice ready</h1>
      <p>Order QA-1042</p>
      <strong>Total: $42.00</strong>
    </main>
  `);
  const card = page.locator('main');
  await expect(card).toHaveScreenshot('invoice-card.png');
});

Verify and generate the first baseline deliberately:

npx playwright test tests/visual.spec.ts --project=chromium --update-snapshots
npx playwright test tests/visual.spec.ts --project=chromium

If your project names differ, replace chromium with the project name from npx playwright test --list. The second run must pass without update mode. Inspect the generated image in the test's snapshot directory and commit it. For a real UI revision, review the received image against the product change, then run the same focused update command. If the new screenshot is wrong, fix the product or test setup instead of refreshing the expected file. A baseline update is an approval of visible behavior, not a generic error fix.

Snapshot naming can also change when a test title, project name, test path, or snapshotPathTemplate changes. A moved test may look like a missing baseline even though the old PNG is still present. Check the Expected path before creating a duplicate. Restore or migrate the expected asset intentionally, and verify that the final committed path is the one the runner reads.

2. Fix Playwright toHaveScreenshot Screenshot Comparison Failed When Dimensions Differ

When expected and actual sizes differ, start with layout. An element screenshot inherits the element's bounding box. New text wrapping, an unloaded font, expanded validation message, or a one-pixel border change can change height. A page screenshot also depends on viewport, and fullPage: true depends on total document height. Changing maxDiffPixels cannot bring differently sized images into alignment.

For a controlled layout, set a viewport and assert the content that determines its size before capture. Replace the test in tests/visual.spec.ts with this standalone case if you are investigating dimensions:

import { test, expect } from '@playwright/test';

test('stable receipt dimensions', async ({ page }) => {
  await page.setViewportSize({ width: 800, height: 600 });
  await page.setContent(`
    <style>html, body { margin: 0 } .receipt { box-sizing: border-box;
      width: 320px; height: 180px; padding: 16px; background: white;
      font: 16px Arial; border: 1px solid black }</style>
    <section class="receipt" aria-label="Receipt">
      <h1>Payment complete</h1><p>Reference QA-1042</p>
    </section>
  `);
  const receipt = page.getByRole('region', { name: 'Receipt' });
  await expect(receipt).toHaveCSS('width', '320px');
  await expect(receipt).toHaveCSS('height', '180px');
  await expect(receipt).toHaveScreenshot('receipt.png');
});
npx playwright test tests/visual.spec.ts --project=chromium --update-snapshots
npx playwright test tests/visual.spec.ts --project=chromium --repeat-each=5

The explicit height is appropriate only when the product has a fixed-height component. Do not force a variable-height article or accessible text area into a fixed box merely to satisfy a snapshot. For variable content, seed the exact copy and wait for its final state. Choose an element screenshot when the feature is a component; it excludes unrelated page chrome. If a full-page capture is required, check that lazy-loaded sections, banners, and sticky elements reach the same state on every run. Use visual regression masking techniques for volatile areas, not as a substitute for diagnosing actual layout drift.

3. Stabilize Dynamic Text and Images Before Comparing Pixels

Timestamps, randomized examples, rotating banners, user avatars, animation frames, and live counters are common sources of false differences. Prefer deterministic fixtures when the data is part of the feature under test. Mask only when the pixels are irrelevant to the assertion and cannot reasonably be fixed. A mask covers the locator's box, so an element that appears or disappears can still change layout or the mask's size.

This self-contained case freezes one dynamic timestamp at the source, then masks an irrelevant rotating token. Both operations are visible in the test and can be reviewed:

import { test, expect } from '@playwright/test';

test('account card ignores rotating token', async ({ page }) => {
  await page.setContent(`
    <article aria-label="Account card" style="width: 320px; padding: 20px;
      background: white; color: black; border: 1px solid black">
      <h2>Account ready</h2>
      <p data-testid="created">Created: 2026-10-05</p>
      <p data-testid="token">Session token: ${Math.random()}</p>
    </article>
  `);
  const card = page.getByRole('article', { name: 'Account card' });
  await expect(page.getByTestId('created')).toHaveText('Created: 2026-10-05');
  await expect(card).toHaveScreenshot('account-card.png', {
    mask: [page.getByTestId('token')],
    maskColor: '#FF00FF',
  });
});
npx playwright test tests/visual.spec.ts --project=chromium --update-snapshots
npx playwright test tests/visual.spec.ts --project=chromium --repeat-each=10

A real application should normally inject fixed dates and API fixture values rather than hardcode product behavior inside a test. If the token itself matters, assert its format or business behavior separately. Avoid a page-wide mask that hides the component you intended to check. Playwright also offers stylePath for screenshot-only CSS; use it for consistent treatment of known volatile visual details such as a blinking cursor or embedded widget. Keep the stylesheet narrow and in the same repository as the test. Review the resulting baseline to confirm you still test the meaningful UI.

4. Wait for the Final UI State Instead of Taking an Early Screenshot

A stable pair of consecutive screenshots is not proof that the correct screen loaded. A skeleton can be perfectly stable for two captures, and a network response can arrive afterward. Add a functional assertion for the state you mean to compare. It gives a clearer failure when a backend or route is broken and prevents a snapshot from being taken on a loading placeholder.

This test simulates an asynchronous card without relying on a live server. The assertion waits for the final heading and for the loading text to disappear:

import { test, expect } from '@playwright/test';

test('captures the loaded profile', async ({ page }) => {
  await page.setContent(`
    <section aria-label="Profile" style="width: 300px; padding: 20px">
      <p id="loading">Loading profile...</p>
    </section>
    <script>
      setTimeout(() => {
        document.querySelector('[aria-label="Profile"]').innerHTML =
          '<h1>Ada Rivera</h1><p>QA engineer</p>';
      }, 150);
    </script>
  `);
  await expect(page.getByRole('heading', { name: 'Ada Rivera' })).toBeVisible();
  await expect(page.getByText('Loading profile...')).toHaveCount(0);
  await expect(page.getByRole('region', { name: 'Profile' }))
    .toHaveScreenshot('loaded-profile.png');
});
npx playwright test tests/visual.spec.ts --project=chromium --update-snapshots
npx playwright test tests/visual.spec.ts --project=chromium --repeat-each=10

The short timer is only a local demonstration. In a product test, the final heading or a ready indicator is the meaningful gate. Do not replace it with page.waitForTimeout(2000): a fixed delay can be longer than needed on one runner and still too short on another. If a screenshot assertion itself times out while trying to settle, inspect the changing region and the page's pending work before raising its timeout option. The Playwright trace on retry guide helps preserve evidence for a CI-only loading race.

5. Match Fonts, Browser, Viewport, and Scale Across Machines

Visual baselines are environment-specific. Text antialiasing and wrapping can differ with operating system fonts, installed browser binaries, hardware rendering, device scale, or headed versus headless execution. Even if the CSS is identical, a font substitution can alter glyph widths and shift every line. Playwright's visual comparison documentation recommends capturing reference images in the same environment used for comparison.

Make the browser project explicit. This complete playwright.config.ts gives every test a known viewport, locale, color scheme, and device scale. It also saves failure artifacts and scopes snapshots by project using Playwright's default project-aware naming:

import { defineConfig } from '@playwright/test';

export default defineConfig({
  testDir: './tests',
  use: {
    viewport: { width: 800, height: 600 },
    deviceScaleFactor: 1,
    locale: 'en-US',
    colorScheme: 'light',
    trace: 'retain-on-failure',
    screenshot: 'only-on-failure',
  },
  projects: [{ name: 'chromium', use: { browserName: 'chromium' } }],
});
npx playwright install chromium
npx playwright test tests/visual.spec.ts --project=chromium

Verify the font used in your product is installed and loaded before approving a baseline. Web fonts should come from a stable asset and should not silently fall back in an offline CI run. Playwright waits for fonts during screenshot capture, but it cannot download a missing font or make two host font libraries render identically. For cross-platform work, maintain separate baselines per project or capture and compare in one consistent container. Do not commit a macOS baseline and expect it to be pixel-identical in Linux solely because both say Chromium.

The screenshot option scale: 'css' is the default and records one image pixel per CSS pixel. scale: 'device' follows device pixels. Keep the setting consistent when generating and comparing. Likewise, fullPage, animations, and caret affect the captured image. Use the same options on every run and regenerate a baseline only after reviewing the effect of an intentional option change.

6. Use Pixel Tolerance Only for Measured Rendering Noise

Playwright exposes threshold, maxDiffPixels, and maxDiffPixelRatio. They solve different problems. threshold is the allowed perceived color difference for an individual pixel; the other two allow an amount or ratio of different pixels across the image. A large maxDiffPixelRatio can hide a broken icon or missing small component, especially on a full-page screenshot. None of these options should be used to wave through a size mismatch or a visibly changed layout.

For a small, repeatable text-edge difference, use a narrow budget on the affected assertion. The following test is complete and demonstrates a local option; the numbers are examples to calibrate from your actual diff, not recommended defaults:

import { test, expect } from '@playwright/test';

test('small label visual comparison', async ({ page }) => {
  await page.setContent(`
    <span aria-label="Status label" style="display: inline-block;
      padding: 8px; font: 16px Arial; color: #222; background: #eee">
      Ready
    </span>
  `);
  await expect(page.locator('[aria-label="Status label"]')).toHaveScreenshot('status-label.png', {
    maxDiffPixels: 12,
    threshold: 0.2,
  });
});
npx playwright test tests/visual.spec.ts --project=chromium --update-snapshots
npx playwright test tests/visual.spec.ts --project=chromium --repeat-each=10

Start with no custom budget. If a stable, environment-matched image still has a small benign difference, inspect exactly which pixels differ and select the smallest allowance that covers that known variation. Test one deliberate visual mutation, such as removing the label background, and confirm the assertion still fails. A tolerance is useful only if it filters noise while retaining the regressions your team cares about. Element-scoped captures usually need less tolerance than full-page images because they exclude unrelated motion and content.

7. Reproduce CI and Docker Screenshot Failures in the Same Environment

When the test passes on a laptop and fails in CI, compare the lockfile, Playwright package version, browser project, container image, fonts, locale, and display mode. Do not assume that identical test code means identical pixels. First run a focused test with one worker in CI so parallel traffic does not confuse the diagnosis. Then inspect the uploaded HTML report and diff, especially the dimension line.

A minimal CI sequence for a non-containerized Linux runner is:

npm ci
npx playwright install --with-deps chromium
npx playwright test tests/visual.spec.ts --project=chromium --workers=1

Use the package version from npm ls @playwright/test or your lockfile when selecting a Docker tag. Replace the placeholder below with that exact installed version before running it. The image contains browser binaries and system dependencies; your project still needs npm ci inside the container. Generate an approved baseline there, then run the same command without update mode:

docker run --rm --ipc=host -v "$PWD:/work" -w /work   mcr.microsoft.com/playwright:v<your-playwright-version>-noble   sh -lc 'npm ci && npx playwright test tests/visual.spec.ts --project=chromium --workers=1'

The placeholder is intentionally not a release recommendation. A mismatched image and npm package can fail to find the expected browser executable, and even a runnable mismatched browser may render differently. Use the Docker for Playwright guide for container setup and the visual regression in CI guide for artifact retention and baseline review. If a containerized baseline passes in the same container locally but fails in CI, inspect fonts installed by the application, environment-driven content, and server responses before changing tolerances.

How to Verify the Fix

Verify the specific failure, not just a fresh baseline. Run the focused test without --update-snapshots, repeat it, and run it in the original failing environment. For the examples above, use these commands after keeping the relevant standalone test in tests/visual.spec.ts:

npx playwright test tests/visual.spec.ts --project=chromium
npx playwright test tests/visual.spec.ts --project=chromium --repeat-each=10
npx playwright show-report

A pass is stronger when the report's expected and actual images match for the reason you intended. If the original issue was a 1-pixel height change, verify equal dimensions. If it was dynamic content, repeat enough times for the volatile value to change and confirm the mask or fixed fixture works. If it was CI-only, rerun the same job image and settings; a local pass alone does not close that case.

Inspect the Git diff of any baseline PNG or WebP before merging. A new snapshot must show the approved UI, correct copy, expected data, and no hidden loading overlay. Keep the diff artifact from the original failure long enough to explain the change in review. When a deliberate product regression is introduced for a check, confirm the assertion fails again. This guards against an overbroad mask or tolerance that made the test insensitive.

Prevent It From Coming Back

Give visual tests ownership of their rendering inputs. Pin the Playwright package with the lockfile, use a matching browser image in CI, and generate baselines there. Specify viewport, locale, color scheme, and device scale in configuration. Keep fonts as reliable application assets or install the expected system fonts in the runner. Seed test data, fix clocks where needed, and avoid network content that changes between runs.

Prefer screenshots of a component or meaningful region over a full page when the requirement is local. That keeps unrelated banners, cookie prompts, and footer changes out of the assertion. Keep functional assertions alongside the image check: a visual baseline can be approved accidentally, but an explicit heading or status assertion states the behavior in text. When a feature genuinely changes, have a reviewer inspect both the code diff and the updated image. Preserve the expected, actual, and diff files in CI so the next failure has evidence instead of a bare error line.

Document why each mask or tolerance exists. A later teammate should know which pixels are intentionally excluded and what still must fail. Revisit visual baselines when browser or OS upgrades change rendering. If many snapshots change at once, treat that as an environment migration and review a representative sample plus any product-sensitive screens before accepting the batch.

Interview Questions and Answers

Q: What does Screenshot comparison failed mean in Playwright?

The captured image does not match the stored reference image. I first determine whether the dimensions differ or whether the same-sized images have changed pixels, then inspect the expected, actual, and diff artifacts.

Q: Why can a screenshot test pass locally but fail in CI?

The browser, operating system, fonts, viewport, scale, headless mode, or content may differ. I reproduce it in the CI image and align those inputs before adding tolerance.

Q: Does toHaveScreenshot wait for the page to finish loading?

It waits for two consecutive screenshots to be the same, but that can happen while a stable skeleton is displayed. I assert the final business state before the screenshot.

Q: Will maxDiffPixels fix images with different widths or heights?

No. A size mismatch needs a layout or environment fix. I inspect wrapping, fonts, viewport, element bounds, and fullPage behavior.

Q: When should a baseline be updated?

Only after a deliberate UI change has been reviewed against the received and diff images. I update the focused snapshot, commit it, and verify a normal run passes.

Q: How would you handle a changing timestamp in a visual test?

I prefer a fixed clock or deterministic fixture if the timestamp is part of the tested UI. If it is irrelevant to the visual requirement, I mask that specific region and test the timestamp behavior separately.

Common Mistakes

  • Running --update-snapshots across the whole suite without reviewing each changed image.
  • Increasing maxDiffPixelRatio until a real missing element no longer fails.
  • Trying pixel tolerance on an image dimension mismatch.
  • Taking the screenshot before the actual content replaces a loading state.
  • Masking a parent container so the feature under test disappears from coverage.
  • Mixing local macOS baselines with Linux CI captures.
  • Upgrading the Playwright package while keeping an old Docker browser image.
  • Using a full-page screenshot for a tiny component that could be captured directly.
  • Treating a first-run missing snapshot as a visual regression without checking the expected path.
  • Committing snapshots generated with different viewport or display settings from the CI project.

Conclusion

To fix Playwright toHaveScreenshot "Screenshot comparison failed", inspect the image artifacts and classify the mismatch before acting. Restore or approve a baseline for intentional change, stabilize content and layout for flaky captures, and align the rendering environment for CI differences. Run the focused test repeatedly without update mode, then verify it in the environment that originally failed.

Interview Questions and Answers

How would you triage a Playwright screenshot comparison failure?

I inspect the expected, actual, and diff artifacts, then check whether their dimensions agree. For equal sizes, I look at the shape and location of changed pixels. I classify the cause as product change, unstable content, capture timing, or rendering environment before changing the test.

What is the distinction between threshold and maxDiffPixels?

`threshold` controls how different the color of one pixel may be before it counts as changed. `maxDiffPixels` allows a count of changed pixels across the image. Neither option repairs a width or height mismatch, and both need deliberate calibration.

Why is automatic baseline updating in CI dangerous?

It can accept a broken UI as the new expectation without review. A missing button or error screen may become the approved snapshot. I generate updates in a controlled run and require image review before committing them.

How do you make a visual test deterministic?

I set viewport, locale, color scheme, and device scale, then make data and fonts stable. I assert that the final UI has loaded before capture. For content outside the visual requirement, I use a narrow mask and keep a separate functional check.

How would you debug a screenshot whose height changes by one pixel?

I treat it as a layout issue rather than raising a pixel budget. I compare the element box, font metrics, wrapping, borders, viewport, and headed or headless mode. I capture and compare in the same browser environment as the baseline.

When would you use an element screenshot instead of a page screenshot?

I use an element capture when the requirement concerns one card, dialog, or control. It excludes unrelated banners, navigation, and live page content, making the failure easier to diagnose. I still assert the component is in its final state first.

What evidence proves a visual regression fix is stable?

The focused test passes repeatedly without snapshot update mode and passes in the originally failing CI image. I inspect the expected and actual artifacts, verify dimensions, and introduce a deliberate visual change to confirm the test can still fail.

Frequently Asked Questions

How do I fix Playwright toHaveScreenshot Screenshot comparison failed?

Open the expected, actual, and diff images in the Playwright report. If the UI change is intended, review and update the focused baseline; otherwise stabilize the page, test data, fonts, or environment that caused the mismatch.

Why does Playwright report different image dimensions?

A viewport, element bounding box, font, text wrapping, or full-page height changed between the baseline and current capture. Fix the layout or environment difference first; pixel tolerance cannot correct a size mismatch.

How do I update a Playwright screenshot baseline?

Run the focused test with `npx playwright test tests/visual.spec.ts --project=chromium --update-snapshots`, adjusting the path and project to your suite. Review the new image, commit the approved baseline, and rerun without the flag.

Why does toHaveScreenshot fail only in CI?

CI may use a different browser build, operating system, fonts, device scale, locale, or headless mode. Reproduce the assertion in the same CI image and compare its actual image with the baseline before changing thresholds.

Does maxDiffPixels help with a one-pixel height difference?

No. A height mismatch is a dimension problem, even if it is only one pixel. Inspect element bounds, font loading, borders, wrapping, viewport, and screenshot options.

How do I ignore a dynamic timestamp in a Playwright screenshot?

Prefer a fixed clock or deterministic fixture when the timestamp is part of the tested behavior. If its pixels are irrelevant to the visual requirement, pass a narrow locator in the `mask` option and assert the timestamp separately.

Does toHaveScreenshot wait for application readiness?

It waits for two consecutive screenshots to match, which can also happen on a stable loading screen. Assert an explicit final heading, status, or loaded element before the screenshot assertion.

Related Guides