QA How-To
pytest parametrize Examples for Data-Driven Tests
Learn pytest parametrize examples with runnable pricing tests, clear case IDs, invalid inputs, stacked decorators, indirect fixtures, and JSON test data.
20 min read | 2,931 words
TL;DR
Use @pytest.mark.parametrize to run one test with several named input rows. Give important rows explicit IDs, use stacked decorators only for meaningful combinations, and reserve indirect fixtures for setup that turns data into a resource.
Key Takeaways
- Use one parametrized test to give each meaningful data row its own result and failure ID.
- Name complex cases with pytest.param IDs so CI failures point to a specific scenario.
- Separate positive totals, rejection behavior, and cross-product invariants into focused tests.
- Use indirect fixtures when a parameter must create a resource during setup.
- Load small committed JSON catalogs at collection time and keep expected values independent.
- Verify collection and execution after each change, then run the complete suite.
Pytest parametrize examples are most useful when each row represents a real behavior and a failure names the exact input that broke it. In this tutorial, you will build a small order-pricing suite with @pytest.mark.parametrize, readable case IDs, boundary cases, Cartesian combinations, indirect fixtures, and JSON-backed data. Every command runs against one consistent implementation, so you can paste the files into a fresh directory and follow along.
The example stays local and deterministic. That makes it easy to inspect collection, compare expected values, and distinguish a broken pricing rule from a test-data mistake. If you are new to the runner itself, read the Pytest tutorial for beginners before adding these patterns to an API or browser suite.
What You Will Build
- A
pricing.pyfunction that calculates an order total with a regional shipping fee and optional discount. - A table of named examples that exercises normal inputs, boundaries, and rejected values.
- A small marked smoke subset, a two-dimensional combination test, and one indirect fixture.
- An external JSON case file that keeps larger sets reviewable while preserving useful test IDs.
- A final command that runs the complete suite and a collection command that exposes every generated case.
The worked contract is deliberately narrow: subtotal arrives as a decimal string, region is US, CA, or LOCAL, and coupon is either absent or SAVE10. A ten percent discount applies to merchandise before shipping; the result has two decimal places. Negative subtotals, unknown regions, and unknown coupons raise ValueError. These rules are enough to make every parameter choice meaningful without a network service or a database.
| Technique | Use it when | Watch for |
|---|---|---|
| One decorator with tuple rows | Several inputs determine one expected outcome | Tuple order drifting from argument names |
pytest.param(..., id=...) |
A case needs a searchable name or row-specific mark | IDs that repeat or disclose sensitive data |
| Stacked decorators | Dimensions genuinely combine independently | A large Cartesian product with redundant coverage |
indirect fixture |
Input must become a resource during setup | Hiding a simple literal behind unnecessary setup |
| JSON case file | Non-programmers review a stable data catalog | Loading mutable or untrusted data at collection |
Prerequisites
Use Python 3.12.7 and pytest 9.1.1 for the exact environment used to check this guide. The examples use only Python's standard library beyond pytest. Pytest 9.1.1 supports Python 3.10 or newer according to its package metadata; if your local Python differs, check compatibility and match the installed version instead of copying an unrelated pin. Use an empty directory so earlier tests cannot change the counts shown here.
python3 --version
python3 -m venv .venv
. .venv/bin/activate
python -m pip install 'pytest==9.1.1'
python -m pytest --version
On Windows PowerShell, activate with .venv\Scripts\Activate.ps1. The last command should report pytest 9.1.1. If your shell does not resolve python3, use the command that starts the supported Python installation on your machine. All later commands assume the virtual environment is active and you are in the project root. The Pytest documentation on parametrizing tests is the upstream reference for decorator behavior.
Step 1: Create a Deterministic Pricing Target
Create pricing.py in the project root. This function is the production code every later test imports. Keep the business rule visible: shipping is added after the discount, and Decimal avoids binary-float surprises such as a displayed 0.30000000000000004. The input type is a string because real prices often cross a text or JSON boundary before validation.
# pricing.py
from decimal import Decimal, InvalidOperation
SHIPPING = {
"US": Decimal("5.00"),
"CA": Decimal("8.00"),
"LOCAL": Decimal("0.00"),
}
def order_total(subtotal: str, region: str, coupon: str | None = None) -> Decimal:
try:
amount = Decimal(subtotal)
except InvalidOperation as exc:
raise ValueError("subtotal must be a decimal number") from exc
if not amount.is_finite() or amount < 0:
raise ValueError("subtotal must be finite and nonnegative")
if region not in SHIPPING:
raise ValueError("unknown region")
if coupon not in (None, "SAVE10"):
raise ValueError("unknown coupon")
discount = Decimal("0.90") if coupon == "SAVE10" else Decimal("1.00")
return (amount * discount + SHIPPING[region]).quantize(Decimal("0.01"))
Run a direct smoke check before adding pytest. This catches indentation, import, and arithmetic mistakes while the file is still small.
python -c 'from pricing import order_total; assert str(order_total("100.00", "US", "SAVE10")) == "95.00"'
Verification: the command exits with status zero and prints nothing. An AssertionError means the implementation or copied expected value differs. A ModuleNotFoundError usually means the command was run outside the directory containing pricing.py. The direct check is a starting point, not the final test suite: it does not show which input row failed once the examples multiply.
Step 2: Write Pytest Parametrize Examples as Named Rows
Create test_pricing.py. One decorator maps each tuple to subtotal, region, coupon, and expected in that exact order. Pytest creates one test item per row during collection. Use Decimal in expected results so an equality failure compares monetary values of the same type. The final row checks that discount precedes shipping: ten percent off a 10.00 subtotal yields 9.00, then Canadian shipping adds 8.00.
# test_pricing.py
from decimal import Decimal
import pytest
from pricing import order_total
@pytest.mark.parametrize(
("subtotal", "region", "coupon", "expected"),
[
pytest.param("0.00", "LOCAL", None, Decimal("0.00"), id="zero-local"),
pytest.param("10.00", "US", None, Decimal("15.00"), id="us-shipping"),
pytest.param("100.00", "US", "SAVE10", Decimal("95.00"), id="us-discount"),
pytest.param("10.00", "CA", "SAVE10", Decimal("17.00"), id="discount-before-ca-shipping"),
],
)
def test_order_total_examples(subtotal, region, coupon, expected):
assert order_total(subtotal, region, coupon) == expected
Run both collection and execution. Collection prints the generated node IDs without executing the function; execution proves the assertions. A readable ID makes a CI failure such as test_order_total_examples[discount-before-ca-shipping] immediately actionable.
python -m pytest --collect-only -q test_pricing.py
python -m pytest -q test_pricing.py
Verification: collection lists four cases, and execution reports 4 passed. If pytest says fixture 'subtotal' not found, compare every name in the decorator with the test function signature. If it reports a row-width error, count the four values in each pytest.param call. Parameter data is bound positionally even though the argument names read like a schema. Keep the table short enough that a reviewer can see each price rule at a glance.
A useful way to review this table is to ask which one-line mutation each row would catch. Removing shipping breaks us-shipping; applying the coupon after shipping breaks discount-before-ca-shipping; treating zero as invalid breaks zero-local. The us-discount row catches a forgotten coupon branch. If two rows would always fail together for the same reason, one may be redundant. If a rule has no row capable of failing when that rule is removed, add a targeted example instead of another arbitrary price.
Step 3: Cover Invalid Inputs and Boundaries Without Ambiguous Failures
Positive rows alone do not prove the rejection rules. Append the following test to test_pricing.py. pytest.raises makes the exception type and message part of the expected behavior. Regex matching is appropriate here because the messages are intentionally stable; when an application treats messages as incidental, assert the exception type and a durable error code instead.
@pytest.mark.parametrize(
("subtotal", "region", "coupon", "message"),
[
pytest.param("-0.01", "US", None, "nonnegative", id="negative-cent"),
pytest.param("not-a-price", "US", None, "decimal number", id="invalid-decimal"),
pytest.param("NaN", "US", None, "finite", id="non-finite"),
pytest.param("10.00", "EU", None, "unknown region", id="unsupported-region"),
pytest.param("10.00", "US", "SAVE50", "unknown coupon", id="unsupported-coupon"),
],
)
def test_order_total_rejects_invalid_values(subtotal, region, coupon, message):
with pytest.raises(ValueError, match=message):
order_total(subtotal, region, coupon)
Each row names one reason for rejection. Avoid a single row with several invalid arguments: the function rejects the first invalid field it encounters, so that row cannot establish behavior for later fields. A 0.00 subtotal was already accepted in Step 2, while -0.01 is now rejected; together they pin down the zero boundary. NaN is included because Decimal("NaN") parses successfully, yet should never become an order total.
Keep boundary cases close to the business contract. If the product later imposes a maximum subtotal, test the maximum accepted amount and the first rejected amount as a pair. Do not scatter random large numbers through the table and call them boundary tests. Input validation also has an order: this implementation parses the subtotal, checks finiteness and sign, validates region, then validates coupon. The table therefore gives each error a valid value for every other field, which isolates exactly the validation branch under test.
python -m pytest -q test_pricing.py::test_order_total_rejects_invalid_values
python -m pytest -q test_pricing.py
Verification: the focused command reports 5 passed; the full file reports 9 passed. To inspect a single rejection, append its bracketed ID to the node path or use -k negative-cent. Prefer -k when a shell treats brackets as a glob. The Pytest versus unittest guide can help you explain why one compact test function can still yield five independently reported cases.
Step 4: Add Row-Specific Marks and Stable Selection IDs
A mark can select a fast subset without duplicating test functions. Put the following pyproject.toml in the project root. Registering smoke makes its purpose discoverable and prevents an unknown-mark warning in a strict suite.
# pyproject.toml
[tool.pytest.ini_options]
markers = [
"smoke: small pricing cases for fast feedback",
]
Now replace only the us-shipping row in the Step 2 decorator with this equivalent row. The marks argument belongs to pytest.param; the function-level test remains unchanged.
pytest.param(
"10.00", "US", None, Decimal("15.00"),
marks=pytest.mark.smoke,
id="us-shipping",
),
Run the marked subset, then select the same case by its full node ID. The node ID is useful when reproducing exactly one CI failure; -m smoke is useful for an intentional suite subset. Marker selection and keyword selection are different: -m evaluates marks, while -k matches names and IDs.
python -m pytest -q -m smoke
python -m pytest -q 'test_pricing.py::test_order_total_examples[us-shipping]'
python -m pytest --markers
Verification: each of the first two commands reports 1 passed; --markers includes the registered smoke description. The other eight cases are deselected by -m smoke, not skipped. A skip is an outcome attached to a collected item, whereas deselection excludes it from that particular run. Avoid marking a flaky row as xfail merely to make CI green: use an expected-failure mark only for a tracked, specific defect and remove it once fixed. Official marking guidance explains custom marker registration and selection.
Step 5: Use Pytest Parametrize Examples for Independent Dimensions
Some inputs form a genuine cross-product. Add this test to test_pricing.py to compare regional shipping for two subtotals and both coupon states. The expected difference between Canada and the US is always 3.00 because the regional fees are 8.00 and 5.00. This checks an invariant without restating every final total from the first table.
@pytest.mark.parametrize("subtotal", ["0.00", "25.00"], ids=["zero", "twenty-five"])
@pytest.mark.parametrize("coupon", [None, "SAVE10"], ids=["full-price", "discounted"])
def test_ca_shipping_is_three_more_than_us(subtotal, coupon):
ca_total = order_total(subtotal, "CA", coupon)
us_total = order_total(subtotal, "US", coupon)
assert ca_total - us_total == Decimal("3.00")
The two decorators produce four cases, not two. Stacking is appropriate because each subtotal matters under either coupon state and the same invariant applies to all four combinations. For a sparse matrix, use explicit pytest.param rows instead: a cross-product of ten browsers, twenty accounts, and five regions would create 1,000 cases whether or not each combination has value. Also avoid passing the same argument name through two stacked decorators; each layer must parameterize a different dimension.
python -m pytest --collect-only -q test_pricing.py::test_ca_shipping_is_three_more_than_us
python -m pytest -q test_pricing.py
Verification: collection prints four bracketed IDs combining coupon and subtotal labels, and the complete file reports 13 passed. Do not depend on the display order of those label parts in downstream tooling. Match by the explicit test name or run the complete function when the data matrix changes. These cross-product ideas also apply to data-driven Postman tests, though Postman uses a different runner and data lifecycle.
Step 6: Defer Resource Creation With an Indirect Fixture
Plain values are best for most rows. When each case needs a temporary file, a fixture can create it during test setup. Put the following in conftest.py. Under indirect=True, pytest sends the parameter to subtotal_file as request.param; the test receives the fixture's returned Path, not the raw string. tmp_path gives each generated case its own isolated directory.
# conftest.py
import pytest
@pytest.fixture
def subtotal_file(request, tmp_path):
path = tmp_path / "subtotal.txt"
path.write_text(request.param, encoding="utf-8")
return path
Append this test to test_pricing.py. The test reads exactly the file its case received and still calls the same order_total function defined in Step 1. The fixture illustrates lifecycle control without claiming that a local text file is expensive. For real suites, the deferred resource might be a database record or service client, with teardown in a yield fixture.
@pytest.mark.parametrize(
"subtotal_file",
[
pytest.param("2.00", id="small-file"),
pytest.param("200.00", id="large-file"),
],
indirect=True,
)
def test_total_from_subtotal_file(subtotal_file):
subtotal = subtotal_file.read_text(encoding="utf-8")
assert order_total(subtotal, "LOCAL") == Decimal(subtotal)
python -m pytest -q test_pricing.py::test_total_from_subtotal_file
python -m pytest -q test_pricing.py
Verification: the focused run reports 2 passed, then the complete file reports 15 passed. If you remove indirect=True, subtotal_file becomes the literal string row, so .read_text() fails. That is a useful demonstration of the distinction between a parameter value and fixture setup. Do not use indirect parametrization merely to rename a constant; use it when the setup phase has real ownership or cleanup. See Playwright Python fixtures with pytest for a browser-oriented fixture lifecycle.
Step 7: Load a Reviewable JSON Case Catalog
Large but stable data tables can live outside Python. Create pricing_cases.json beside test_pricing.py. Keep expected totals in the file because an oracle calculated with order_total would only echo the implementation. Every record has a unique id, and all values remain strings so parsing is explicit.
[
{"id": "local-cent", "subtotal": "0.01", "region": "LOCAL", "coupon": null, "expected": "0.01"},
{"id": "ca-no-coupon", "subtotal": "12.00", "region": "CA", "coupon": null, "expected": "20.00"},
{"id": "us-ten-percent", "subtotal": "50.00", "region": "US", "coupon": "SAVE10", "expected": "50.00"}
]
Append the following code to test_pricing.py. The module imports go at the top with the existing imports, before the test functions. The Path(__file__) lookup makes the file location independent of the shell's current directory, provided pytest can import the test module. The list is read during collection, which is why the file should be small, committed, and deterministic.
# Add these imports at the top of test_pricing.py:
import json
from pathlib import Path
# Add this definition and test after the imports and existing tests:
CASES = json.loads(Path(__file__).with_name("pricing_cases.json").read_text(encoding="utf-8"))
@pytest.mark.parametrize(
"case",
[pytest.param(case, id=case["id"]) for case in CASES],
)
def test_order_total_from_catalog(case):
actual = order_total(case["subtotal"], case["region"], case["coupon"])
assert actual == Decimal(case["expected"])
python -m pytest --collect-only -q test_pricing.py::test_order_total_from_catalog
python -m pytest -q
Verification: collection shows local-cent, ca-no-coupon, and us-ten-percent; the whole example suite reports 18 passed. If JSON syntax is broken, pytest fails during collection before any test runs. That early failure is useful for a committed case catalog, but inappropriate for a changing remote feed: fetch and validate remote data in a controlled setup phase instead. Review the catalog for duplicate IDs, stale expected values, and sensitive customer data before committing it.
A data file does not automatically improve maintenance. Three cases are easy to read as Python rows; externalize them when the catalog is edited by a different group, reused across several checks, or large enough that a separate review makes sense. Add a lightweight schema check before expanding it: require the keys id, subtotal, region, coupon, and expected, require unique IDs, and reject records with unexpected fields. That validation belongs near the loader because a misspelled key should fail collection with a clear message instead of producing a late KeyError inside an assertion. Keep credentials and personal data out of the catalog, including IDs that might be printed in CI. If you later move the same approach into an HTTP project, the Python API automation framework guide covers the surrounding client and configuration choices.
Troubleshooting
Problem: ModuleNotFoundError: No module named 'pricing' -> Run from the project root and confirm pricing.py sits beside test_pricing.py. Use python -m pytest from the active virtual environment so the interpreter that installed pytest runs the suite.
Problem: fixture 'expected' not found -> Check spelling and tuple width in the decorator. Every named argument must receive a value from each row, except fixtures deliberately supplied by pytest. A missing name is not inferred from a dictionary key or a nearby variable.
Problem: Collection fails before tests run -> Read the first collection traceback. A malformed pricing_cases.json, an import error, or a decorator row with the wrong number of values can all stop discovery. Validate the JSON and rerun --collect-only -q before debugging assertions.
Problem: A selected node ID is not found -> Copy the ID from --collect-only -q and quote bracketed node IDs in your shell. IDs can change when generated from values, so explicitly name cases you intend to rerun or cite in a bug report.
Problem: PytestUnknownMarkWarning mentions smoke -> Place the pyproject.toml from Step 4 in the root pytest detects, and keep the marker registration under [tool.pytest.ini_options]. Use python -m pytest --markers to inspect the active configuration.
Problem: The JSON-backed test reports a surprising expected total -> Calculate discount and shipping separately on paper. For 50.00 in the US with SAVE10, merchandise becomes 45.00, then shipping makes 50.00. If the product rule changes, update the implementation and independent expectation together after confirming the requirement.
Interview Questions and Answers
The questions below are mirrored in interviewQnA for readers practicing a concise spoken explanation. First run the examples, then describe what collection creates, which rows isolate a defect, and why the expected values are trustworthy. A good answer connects the API to a testing decision rather than merely naming a decorator.
For example, if asked why the suite separates exact totals from the shipping invariant, explain that exact rows prove specific outputs while the invariant covers a relationship across combinations. If asked why the file fixture is indirect, explain that pytest should create a separate path for each case during setup. The distinction between collection-time data and execution-time resources is particularly useful when diagnosing slow or order-dependent suites.
Common Mistakes
- Use explicit row IDs when the raw values are hard to recognize. Auto-generated IDs can be useful for tiny tables, but a complex object often becomes an unhelpful argument label.
- Keep one behavior per invalid-input row. Two bad fields in the same row only prove whichever validation runs first.
- Resist producing a huge Cartesian product to look thorough. Count the generated cases and ask which combinations exercise distinct risk.
- Do not compute
expectedwith the function under test. A duplicated defect can make every row pass. - Keep fixture setup independent per case. A reused mutable list, file, or account can make case order affect results.
- Treat
skipandxfailas explicit product knowledge, with a reason and follow-up. They are not substitutes for debugging a broken assertion. - Keep real credentials and personal records out of parameter tables and IDs, since IDs appear in logs and reports.
Where To Go Next
Move one real business rule into the same pattern: define inputs, output, rejection behavior, and a small set of boundary rows. Run --collect-only -q during review to inspect the generated suite before trusting the pass count. For scenario-level collaboration, compare this approach with pytest-bdd in Python. To rehearse design choices under interview pressure, use the top pytest interview questions. The techniques also transfer to UI checks, but keep browser state isolated and avoid multiplying cases that all follow the same path.
The authoritative references for APIs used here are the Pytest parametrization guide, the parametrization examples, and the Pytest command-line usage guide. They cover IDs, pytest.param, indirect, selection, and other options when your suite grows beyond this example.
Conclusion
These Pytest parametrize examples turn one pricing rule into 18 independently reported cases without copying 18 test functions. The first table documents exact totals, the invalid rows protect rejection behavior, stacked decorators check a cross-product invariant, and the indirect fixture plus JSON catalog show two distinct ways to manage data. Start with the smallest table that explains the contract, inspect its collected IDs, and expand only where a new row can catch a different defect.
Interview Questions and Answers
What does @pytest.mark.parametrize do during collection?
It creates a separate test item for each parameter set before test execution. That gives every data row its own outcome and node ID. I inspect `--collect-only` when reviewing a suite because a single test function can expand into far more cases than its source suggests.
When would you use pytest.param instead of a plain tuple?
I use `pytest.param` when a row needs an explicit ID or a row-specific mark. For example, a shipping edge case can be named `discount-before-ca-shipping` so a failed CI line identifies the rule. Plain tuples are sufficient for a short table whose generated IDs are already clear.
How do you avoid weak expected values in a data-driven test?
I derive expected values from the requirement or an independent example, not from the function under test. For monetary rules I write down the order of discount and shipping, then encode the final amount as a literal expected value. Otherwise the same defect can appear in both calculation and oracle.
What is the purpose of indirect parametrization?
It passes a case value to a fixture as `request.param`, allowing the fixture to create a resource during setup. The test then receives the fixture's result. I use it for files, accounts, or other owned state, not as a wrapper around a simple constant.
How do stacked parametrize decorators affect case count?
They form a Cartesian product across distinct parameter names. A two-value coupon dimension and a two-value subtotal dimension yield four items. I count the resulting cases and prefer explicit rows if the product generates combinations without distinct risk.
How do you debug a pytest failure in one parameter row?
I first read the bracketed case ID and assertion diff, then rerun that exact node ID. I check the row data and the expected business rule before changing implementation code. If the failure happens during collection, I inspect imports, row widths, and external data syntax instead.
Why can JSON-backed parametrization fail before any test runs?
A module-level JSON load executes while pytest imports the test module for collection. Malformed JSON or a missing file therefore prevents test items from being created. I keep local catalogs versioned and validated, and avoid fetching changing remote data at import time.
Frequently Asked Questions
How do I parametrize multiple arguments in pytest?
Pass a tuple of argument names and a sequence of rows to `@pytest.mark.parametrize`. Each row must supply one value per name in the same order. Pytest collects a separate test item for every row.
How do I give each pytest parameter set a readable name?
Wrap a row in `pytest.param(..., id="case-name")` or supply an `ids` sequence to the decorator. Inspect the results with `python -m pytest --collect-only -q`. Unique IDs make reruns and failure reports easier to interpret.
Can I parametrize a fixture in pytest?
Yes. Use `indirect=True` for a parameter whose name matches a fixture. The fixture reads the raw value through `request.param` and returns the resource the test uses. This is helpful when setup should occur during test execution.
Do stacked parametrize decorators create every combination?
Yes. Two dimensions with two values each produce four collected cases. Choose explicit rows instead when most combinations are irrelevant or too costly to run.
Can pytest read parameter values from JSON?
Yes. Load a local JSON file and build `pytest.param` rows from its records. When that load happens at module import, invalid JSON stops collection, so keep the file small, versioned, and deterministic.
What is the difference between a marked row and a skipped row?
A custom mark such as `smoke` labels a row for selection with `-m`. A skip records that a collected case was not executed. Selecting only smoke rows deselects the other rows; it does not mark them skipped.
How can I rerun one parametrized case?
Use the full node ID shown by `--collect-only -q`, including the bracketed case ID, and quote it in the shell. You can also use `-k` with a distinctive ID substring when the exact node path is inconvenient.