QA How-To
How to Rerun Failed Tests in pytest with pytest-rerunfailures
Learn to rerun failed tests in pytest with pytest-rerunfailures, targeted markers, error filters, retry reports, CI exit policies, and last-failed cache.
20 min read | 3,351 words
TL;DR
Install pytest-rerunfailures and run python -m pytest --reruns 1 to retry each eligible failure once in the same session. Mark individual tests with @pytest.mark.flaky(reruns=1), restrict transient errors with --only-rerun, and inspect rerun outcomes before deciding whether CI should pass.
Key Takeaways
- Use --reruns N for up to N additional attempts per failing test.
- Use @pytest.mark.flaky to scope retries to a reviewed test.
- Filter by exception and prove both matching and nonmatching cases.
- Use --rerun-show-tracebacks to retain evidence from recovered tests.
- Use --fail-on-flaky when CI must reject a test that passed only on rerun.
- Use --lf for failures cached from a previous invocation, not same-run retries.
Rerun failed tests in pytest during the same test session: install pytest-rerunfailures and run python -m pytest --reruns 1. The number is the maximum extra attempts per failing test, so a test can run twice in total. For an individual test, use @pytest.mark.flaky(reruns=1). A recovered test can still reveal a defect, so inspect the rerun report instead of treating the final green result as the whole story.
This tutorial builds a tiny suite with a predictable first-attempt failure. You will see a rerun succeed, restrict retries by exception type, expose a permanent assertion failure, and make a deliberate CI decision about recovered tests. All example commands run from the root of a small Python project. The examples do not depend on random values, a live API, or a fixed sleep.
Pytest also has --lf, which runs failures remembered from an earlier invocation. That solves a different problem from retrying a failure inside the current invocation. The distinction matters when you are debugging locally or deciding what a CI job should do.
What You Will Build
- A local pytest environment with a known pytest and pytest-rerunfailures pair.
- A deterministic test that fails on its first attempt and passes on its second.
- A targeted
flakymarker and exception filters for retryable errors. - Report commands that show the failed attempt and its eventual outcome.
- A separate, permanent failure to demonstrate the limit of retries.
- A practical policy for using retries in CI without erasing failure evidence.
The runnable tests are intentionally small. Their counters simulate a transient result within one Python process; they are teaching devices, not a recommended pattern for production tests. In a real suite, the transient event might be a service timeout, an eventually consistent search index, or an intermittent browser startup. The assertion should still express the product contract.
Prerequisites
Use Python 3.10 or newer. For this reproducible walkthrough, install pytest 9.1.1 and pytest-rerunfailures 16.7, the releases listed on pytest on PyPI and pytest-rerunfailures on PyPI at publication. The plugin documents pytest 8.2 or newer as its minimum. If your project already has a locked dependency set, match the versions in that lockfile and verify the plugin options with python -m pytest --help before copying commands. Do not silently upgrade a shared environment to follow a tutorial.
Start in an empty directory or an isolated scratch project. You need permission to create tests/, pytest.ini, and a virtual environment there. On macOS or Linux, use the shell commands below. On Windows, activate the environment with .venv\Scripts\Activate.ps1 in PowerShell and use py -m pytest if python is not on your path.
python3 --version
python3 -m venv .venv
source .venv/bin/activate
python -m pip install "pytest==9.1.1" "pytest-rerunfailures==16.7"
python -m pytest --version
python -m pip show pytest-rerunfailures
Verify: The first command reports Python 3.10 or newer. The final two commands identify the installed pytest and plugin versions. python -m pytest --help should list --reruns, --reruns-delay, --only-rerun, and --fail-on-flaky. Using python -m pytest throughout keeps the runner tied to the interpreter where you installed the plugin. If you are new to test discovery and assertions, read the pytest tutorial for beginners before continuing.
Step 1: Create Tests That Expose a Retry
Create tests/test_retry_demo.py with the complete code below. Keep this file in one process for the first experiments; do not add xdist or parallel workers yet.
# tests/test_retry_demo.py
from itertools import count
gateway_attempts = count(1)
socket_attempts = count(1)
def test_transient_gateway():
attempt = next(gateway_attempts)
assert attempt >= 2, f"gateway still unavailable on attempt {attempt}"
def test_temporary_socket():
attempt = next(socket_attempts)
if attempt == 1:
raise OSError("temporary socket reset")
assert attempt == 2
def test_stable_invoice():
assert 19 + 4 == 23
The module-level counters advance only while this test module remains loaded. Each fresh pytest command starts a new process and starts them at one again. That gives every command in this guide an observable first failure. The stable test acts as a control: it should never receive a rerun. In a production test, do not make success depend on a global attempt count, because parallel execution and test order would change the result.
Run the gateway test without plugin options:
python -m pytest -q tests/test_retry_demo.py::test_transient_gateway
Verify: The command exits with a failure and displays gateway still unavailable on attempt 1. There is one failed test and no RERUN outcome. This baseline proves that the assertion really fails before you let the plugin intervene. If it passes here, check that you copied the file exactly and selected the right node ID.
Step 2: Rerun Failed Tests in pytest During One Invocation
Run the same node with one allowed rerun and ask pytest to show rerun information:
python -m pytest -q tests/test_retry_demo.py::test_transient_gateway --reruns 1 -rR
Verify: The output includes a RERUN event followed by a final pass, and the process exits successfully. You should see one rerun, not two. The initial execution is attempt one; --reruns 1 permits attempt two. On every fresh invocation, the counter resets, so repeating this exact command repeats the demonstration.
Now run the whole small module with the same setting:
python -m pytest -q tests/test_retry_demo.py --reruns 1 -rR
Verify: Both transient tests recover and test_stable_invoice passes on its first attempt. There are three final passes and two rerun events. This shows the scope of a global flag: each failed test can consume its own retry allowance. A broad --reruns 3 does not mean three additional attempts for the entire suite. If a thousand tests fail once, that setting may schedule up to a thousand extra executions before any second retries are considered.
Use the smallest allowance supported by the failure mode. One extra attempt often distinguishes an intermittent result from a stable bug without making a broken build wait through many identical failures. A retry is not a repair: the reason for the first failure remains relevant even when the final status is green. The guide to reducing flaky tests in CI covers root-cause work beyond this runner feature.
Step 3: Add Delay Only When the Failure Needs Time
A short delay can make sense when a dependency has a documented recovery window. Run the gateway example with a half-second wait between attempts:
python -m pytest -q tests/test_retry_demo.py::test_transient_gateway --reruns 1 --reruns-delay 0.5 -rR
Verify: The final result still passes after one rerun. The command takes roughly half a second longer than the no-delay version, subject to normal machine overhead. The value is seconds and accepts a fraction. A delay is applied between attempts; it does not change the assertion or make the test wait inside its first run.
The demo counter does not actually need recovery time, so the delay is purely illustrative. In an API suite, a delay is justified only if you know what can change between attempts. An eventually consistent read may benefit from a bounded interval; a deterministic schema mismatch will not. For true eventual consistency, an explicit polling assertion with a deadline often communicates the expected behavior better than replaying an entire test with all its setup.
The current plugin also supports --reruns-delay-backoff-factor. For example, --reruns 3 --reruns-delay 1 --reruns-delay-backoff-factor 2 waits 1, 2, then 4 seconds before successive retries, according to the plugin's documented formula. That can reduce repeated pressure on a recovering service, but it grows the worst-case runtime. Do not add backoff to every suite by default. Put a ceiling on total work and investigate the source of repeated failures.
Step 4: Rerun Failed Tests in pytest With a Marker
Replace tests/test_retry_demo.py with this complete version. It adds a marker to the gateway test and preserves the same names and signatures used in the previous steps.
# tests/test_retry_demo.py
from itertools import count
import pytest
gateway_attempts = count(1)
socket_attempts = count(1)
@pytest.mark.flaky(reruns=1, reruns_delay=0.1)
def test_transient_gateway():
attempt = next(gateway_attempts)
assert attempt >= 2, f"gateway still unavailable on attempt {attempt}"
def test_temporary_socket():
attempt = next(socket_attempts)
if attempt == 1:
raise OSError("temporary socket reset")
assert attempt == 2
def test_stable_invoice():
assert 19 + 4 == 23
Run only the marked test, then run the unmarked socket test:
python -m pytest -q tests/test_retry_demo.py::test_transient_gateway -rR
python -m pytest -q tests/test_retry_demo.py::test_temporary_socket
Verify: The first command passes after one rerun even without --reruns. The second fails on its first OSError, because it has no marker and no global retry flag. This difference is why a marker is useful for a known, bounded exception: reviewers can see exactly which test has a temporary policy.
The marker's reruns count takes priority over a global --reruns value in the default strict mode. The plugin also offers --force-reruns and an append mode, but do not use them casually: they can override a carefully reviewed local limit. Keep the marker beside an issue or comment that explains the underlying cause and a removal criterion. If the test becomes stable after the product fix, remove the marker and confirm an ordinary run stays green. Quarantine is a different operational choice; see flaky test quarantine in CI when a test must temporarily stop blocking a merge.
Step 5: Rerun Only the Errors You Intend to Retry
Use the unmarked socket test to make exception filtering visible. This command reruns its OSError, then passes on the second attempt:
python -m pytest -q tests/test_retry_demo.py::test_temporary_socket --reruns 1 --only-rerun OSError -rR
Verify: Expect one RERUN event and a final pass. Run the same test with a filter that does not match its failure:
python -m pytest -q tests/test_retry_demo.py::test_temporary_socket --reruns 1 --only-rerun AssertionError -rR
Verify: This second command fails once, with no rerun. Its nonzero exit is expected. --only-rerun takes a regular expression, not a Python import path, and can be repeated to allow several patterns. Choose patterns that correspond to actual transient failures in your suite. A broad AssertionError filter can retry ordinary product regressions, because most test assertions raise that type.
You can invert the policy with --rerun-except AssertionError, which blocks retries for matching errors while allowing other failures. Test the filter against a known negative case as well as a positive case. A marker-level only_rerun or rerun_except overrides the command-line filter for that marked test, so the marked gateway example is not suitable for checking a global filter. The plugin's filter rules and precedence are useful when a command appears to ignore an option.
Exception names alone do not prove an error is temporary. A repeated OSError caused by a wrong file path or missing certificate is stable. Record the exception message and environment evidence before granting the test another attempt. A narrow, explainable policy is easier to review than an unconditional retry around every assertion.
Step 6: Read Rerun Evidence and Decide the Exit Policy
A final pass normally gives the pytest process a successful exit even though an earlier attempt failed. Show the earlier attempt's traceback with the plugin's report option:
python -m pytest -q tests/test_retry_demo.py::test_temporary_socket --reruns 1 --only-rerun OSError --rerun-show-tracebacks
Verify: The command passes, but the rerun summary includes the original OSError: temporary socket reset traceback. This option is available in the pinned plugin release. It is useful when ordinary compact output gives only a rerun count and you need to diagnose whether the same failure recurs across jobs.
Now make recovery visible to CI as a failing job:
python -m pytest -q tests/test_retry_demo.py::test_temporary_socket --reruns 1 --only-rerun OSError --fail-on-flaky
Verify: The test recovers after one rerun, yet the command exits with status 7. In a POSIX shell, run echo $? immediately afterward to inspect that status. The plugin documents this special exit code for a test that passes on a rerun. Do not put another command between pytest and the exit-code check.
Choose the gate intentionally. A team investigating instability may enable --fail-on-flaky so a recovered test cannot quietly turn a CI job green. A team that allows bounded recovery should still retain the rerun summary and count rerun events in its test reporting. Attach evidence from the failed attempt when the environment supports it; see publishing test evidence as CI artifacts. Either way, do not confuse final test outcome with clean first-attempt execution.
Step 7: Keep a Permanent Failure Red
Create a separate file for a stable contract bug. Keeping it apart from the transient examples lets every earlier command stay focused.
# tests/test_contract_bug.py
def test_invoice_total_is_wrong():
actual_total = 25
expected_total = 23
assert actual_total == expected_total
Run it with two allowed reruns:
python -m pytest -q tests/test_contract_bug.py --reruns 2 -rR
Verify: The output shows two rerun attempts followed by a final failure. The process exits nonzero. A rerun budget of two means three executions at most, including the initial run. The assertion is unchanged each time, so additional retries waste time and may bury the useful first traceback under repeated copies.
This is also the moment to check setup and teardown cost. The plugin can re-execute a failed function-scoped fixture or setup phase. Broader fixture scopes may persist across attempts, so a rerun is not necessarily an entirely fresh environment. A test that mutates a database, charges a payment method, or sends a message needs idempotent setup and cleanup before retries are safe. Otherwise, attempt two may pass because attempt one changed the world, while the product behavior remains wrong.
Keep known deterministic failures outside an automatic retry policy where practical. Use markers sparingly, a narrow exception filter for the remaining suite, or the plugin's --rerun-exclude-path option for a directory that should never retry. A rerun result cannot replace a defect ticket with a reproducible input, logs, and a fix. For structured failure analysis, read flaky test root-cause analysis with AI as an aid to triage rather than a substitute for evidence.
Step 8: Compare Reruns With pytest's Last-Failed Cache
Pytest's built-in --lf selects tests that failed in a previous command. It does not rerun a failing test inside its current attempt. Demonstrate that distinction with an isolated cache directory so your project-wide pytest cache does not affect the result:
python -m pytest -q tests/test_contract_bug.py -o cache_dir=.pytest_cache_demo --cache-clear
python -m pytest -q tests/ -o cache_dir=.pytest_cache_demo --lf --lfnf none
Verify: The first command fails and records the contract test as a last failure. The second command collects that one remembered failure from tests/ and fails it again; it does not automatically run the other tests. Both nonzero exits are expected. The --lfnf none choice prevents a surprise full-suite run if the cache has no remembered failures. When you finish experimenting, the scratch cache directory can be deleted in your own project.
| Option | When selection happens | Scope | Typical use |
|---|---|---|---|
--reruns 1 |
After a failure within the current invocation | Up to one extra attempt per eligible test | Bound an intermittent failure |
@pytest.mark.flaky(reruns=1) |
After this marked test fails | The marked test | Document a temporary local exception |
--lf |
At collection, using a prior run's cache | Previously failed node IDs | Debug yesterday's red tests locally |
--ff |
At collection, using a prior run's cache | Entire suite, with known failures first | Get early feedback while retaining full coverage |
Verify the comparison: Run python -m pytest --help and find --lf, --ff, and the plugin's --reruns options. They can be combined, but doing so means "select old failures, then retry eligible failures during this new run." On a fresh CI worker, the last-failed cache may be absent. In that case --lf defaults to the full suite unless you set --lfnf none. Do not mistake an empty cache for a clean test history.
Troubleshooting
Problem: unrecognized arguments: --reruns -> fix: Run python -m pip show pytest-rerunfailures and python -m pytest --version through the same active interpreter. Install the plugin in that environment, then inspect python -m pytest --help. An editor, virtual environment, or CI step can point at a different Python than your terminal.
Problem: A marked test never retries -> fix: Check that pytest-rerunfailures loaded, the decorator is @pytest.mark.flaky(reruns=1), and the failure is in a phase the plugin can rerun. Inspect --only-rerun and --rerun-except filters, including marker-level overrides. Test the node ID by itself before diagnosing a broad suite.
Problem: Retries run but the test remains red -> fix: Read the first and final tracebacks. If the same assertion and input fail every time, fix the product or expectation. If the failure changes, inspect fixture state, data cleanup, clocks, and shared services. Raising the count without a hypothesis increases runtime and can create false confidence.
Problem: The suite passes locally but CI reports exit code 7 -> fix: Look for --fail-on-flaky in the CI command or configuration. The plugin uses 7 when a test passed only after a rerun. Decide whether recovered tests should block that pipeline, then document the policy; do not suppress the exit code without preserving the rerun report.
Problem: --lf runs everything or nothing unexpected -> fix: Inspect whether .pytest_cache exists in the current workspace and whether the first run stored failures. Use --lfnf none to run no tests when there are no known failures. Use an explicit -o cache_dir=... when you need a separate scratch cache for a demonstration.
Problem: A retry changes external data or breaks another test -> fix: Make setup and cleanup independent across attempts. Confirm that repeated writes are idempotent or use fresh test data per attempt. If the operation cannot be safely repeated, do not add a retry to that test until the side effect is controlled.
Interview Questions and Answers
Q: What does --reruns 2 mean? It allows two executions after the initial failure, for at most three attempts per eligible test. I would confirm the count in a small failing example because teams sometimes mistake the flag for total attempts.
Q: How do --lf and pytest-rerunfailures differ? --lf reads pytest's cache at collection and chooses tests that failed in a previous invocation. The plugin reacts to a failure during the current invocation and reruns that test before reporting its final outcome.
Q: How would you retry only a transient network error? I would start with a narrow --only-rerun expression for the observed exception, plus a small rerun limit. Then I would check a negative case to prove that deterministic assertions are not retried and retain the original error in the report.
Q: Why can a retry be unsafe for a mutating test? The first attempt may have already written data or sent a request before the assertion failed. A second attempt can duplicate that action or pass because state changed. I would establish idempotency and per-attempt cleanup before enabling the plugin.
Q: What does a recovered test tell you? It proves the test passed on a later attempt under the same invocation; it does not prove the first failure was harmless. I would record the exception, environment, and frequency, then fix the underlying cause or remove the temporary retry.
Q: When would you use --fail-on-flaky? I would use it when the CI policy treats any recovered failure as a signal requiring investigation. The runner returns exit code 7 for a passed-on-rerun case, so dashboards and scripts must recognize that status separately from ordinary assertion failure.
Q: How do markers interact with global retry settings? In the default strict mode, a marker's rerun count takes precedence over the command-line count, which takes precedence over configuration. I would keep local exceptions visible in code and avoid force or append modes unless the team explicitly needs those semantics.
Common Mistakes
- Treating
--reruns 3as three attempts total instead of one initial run plus up to three extra attempts. - Running
--lfon a fresh CI worker and assuming it remembers failures from another machine. - Applying
--only-rerun AssertionErrorto a whole suite without checking whether it hides ordinary regressions. - Adding a delay without a plausible state change that could occur during that interval.
- Using a global counter, like this teaching example, as the acceptance criterion for a real test.
- Reporting only final passes while dropping the retry count and first failure traceback.
- Retrying tests with irreversible writes before confirming they are safe to repeat.
Where To Go Next
Transfer the small command to one real failing test. Capture its first exception, choose a maximum of one or two extra attempts based on the dependency's behavior, and run a negative case that must remain red. If the failure needs a wait for a known state transition, replace broad reruns with a focused wait inside the test. If instability crosses workers or services, use the test automation CI/CD guide to check environment and pipeline design.
For browser suites built on pytest, Playwright Python fixtures help isolate state before you add retries, while parallel Playwright Python tests with pytest-xdist explain why concurrency changes failure patterns. Once your suite reports both first-attempt and final outcomes, review whether each retry still earns its place. Remove temporary markers after the underlying defect is fixed.
Conclusion
To rerun failed tests in pytest, install pytest-rerunfailures and use --reruns N for a bounded same-run retry or @pytest.mark.flaky(reruns=N) for a named test. Use --lf when you want to select failures from an earlier run. Keep a permanent failure in your validation set, inspect rerun evidence, and choose explicitly whether a recovered failure should keep CI red.
Interview Questions and Answers
How would you introduce retries to an existing pytest suite?
I would first reproduce a specific intermittent failure and capture its original traceback. Then I would add one bounded rerun to that test or failure type, verify a permanent bug still fails, and publish rerun counts in CI. I would track the underlying defect and remove the exception after the fix.
What exactly does pytest-rerunfailures do when a test fails?
It can execute an eligible failed test again within the same pytest invocation, up to the configured rerun count. The initial run is separate from the configured extra attempts. The final outcome determines the normal test result unless --fail-on-flaky changes the process exit policy.
Why is --lf not a replacement for pytest-rerunfailures?
--lf uses cached node IDs to choose what to run on a later command. It does not react to a failure by immediately executing that test again. I use --lf for local debugging after a red run and the plugin only when bounded same-run recovery is justified.
How do you prevent retries from hiding assertion regressions?
I restrict retries to specific tests or exception patterns and run a known deterministic failure as a negative control. I retain failed-attempt tracebacks or rerun summaries in CI. If recovered failures must block the pipeline, I enable --fail-on-flaky.
What is dangerous about retrying tests that write to a database?
The failed attempt may already have committed a write before its assertion failed. A rerun can duplicate the write, pass because state changed, or contaminate another test. I require idempotent operations, isolated data, and reliable cleanup before allowing a retry.
How do marker and command-line retry counts interact?
In strict mode, the per-test flaky marker takes priority over the global --reruns count; command-line settings take priority over configuration defaults. I check for marker-level exception filters too, because they can override global filters. I avoid --force-reruns and append mode unless the changed precedence is intentional.
When would a delay or backoff help?
A delay helps only when a real dependency can recover between attempts, such as a bounded service restart or eventual consistency window. Backoff increases the spacing between later attempts and can reduce pressure on a recovering service. I calculate the worst-case runtime and prefer an explicit wait for a known condition where possible.
Frequently Asked Questions
How do I rerun failed tests in pytest automatically?
Install pytest-rerunfailures, then run python -m pytest --reruns 1. Each eligible failure gets at most one additional execution in the same pytest invocation. Add -rR to display rerun information.
Does --reruns 2 mean two total test attempts?
No. It allows two reruns after the initial execution, for a maximum of three attempts per test. A passing first attempt uses no retries.
How can I rerun only one flaky pytest test?
Decorate that test with @pytest.mark.flaky(reruns=1) after installing the plugin. The marker applies to that test without enabling global retries. Document why it is temporary and when it will be removed.
What is the difference between pytest --lf and --reruns?
--lf selects tests that failed in a previous invocation using pytest's cache. --reruns retries eligible failures during the current invocation. They solve different problems and can be combined when that behavior is intentional.
How do I retry only OSError failures?
Run pytest with --reruns 1 --only-rerun OSError. The filter is a regular expression applied to the failure; check a nonmatching case to confirm other failures remain red. A marker-level filter can override the command-line filter.
Can CI fail when a test passes after a rerun?
Yes. Add --fail-on-flaky with pytest-rerunfailures. The plugin returns exit code 7 for a test that ultimately passes after being rerun, allowing CI to flag recovered instability.
Will reruns repeat failed fixture setup?
The plugin can re-execute a failed function-scoped fixture or setup phase. Broader fixture state may remain, so a rerun is not guaranteed to start from a completely clean environment. Review side effects and cleanup before retrying mutating tests.
Related Guides
- How to Run tests in headed mode in Cypress (2026)
- How to Run tests in headed mode in Playwright (2026)
- How to Run tests in headed mode in Selenium (2026)
- How to Debug a failing test in VS Code in Cypress (2026)
- How to Debug a failing test in VS Code in Playwright (2026)
- How to Debug a failing test in VS Code in Selenium (2026)