QA How-To
promptfoo vs DeepEval for LLM Evaluation (2026)
Compare promptfoo vs DeepEval for LLM evaluation with runnable tests, CI gates, RAG metrics, red teaming, and a clear framework choice for QA teams in 2026.
19 min read | 3,881 words
TL;DR
Promptfoo fits configuration-driven comparisons and red teaming; DeepEval fits Python application tests and metric-based checks. Run reviewed cases through the real application boundary, then choose by workflow and integration cost.
Key Takeaways
- Choose Promptfoo for prompt/provider matrices and dedicated red-team workflows.
- Choose DeepEval when Python tests need direct application calls and metric objects.
- Start with deterministic labels to verify the evaluation harness without model costs.
- Prove each CI gate by inserting one intentionally wrong expectation.
- Pass actual retrieved context to RAG faithfulness metrics.
- Review individual failures and judge stability before enforcing semantic thresholds.
Promptfoo vs DeepEval is a choice about how you organize LLM evaluations. Promptfoo suits configuration-driven comparisons of prompts, providers, and test cases, plus dedicated red-team work. DeepEval suits Python application tests that create LLMTestCase objects and apply named metrics inside a test runner. Both can block a regression when the target, dataset, and pass rule are defined clearly.
This guide compares the workflows by running the same ticket router through both tools. The first run is local and deterministic, so you can verify every command without a model API key. You will then see what must change when the target becomes a real LLM application, where exact labels, semantic scores, retrieved evidence, and CI status each answer a different question.
TL;DR
| Need | Start with | Reason |
|---|---|---|
| Compare prompts or models across one dataset | Promptfoo | Its prompt, provider, and test matrix exposes each combination. |
| Evaluate a Python application with fixtures | DeepEval | Test functions can construct cases from live application outputs. |
| Discover prompt injection and other adversarial failures | Promptfoo | Its red-team commands generate and run attack cases. |
| Score RAG answers against retrieved context | DeepEval | Built-in RAG metrics use explicit answer and context fields. |
| Gate exact output contracts in CI | Either | equals and ExactMatchMetric both reject a wrong label. |
Neither framework turns a weak dataset into a reliable release decision. Write down the product behavior first, review representative cases, and verify that one deliberately wrong result fails the job. A green dashboard without that check is only a report.
What You Will Build
- A
route_ticketfunction that returnsbillingortechnical. - A Promptfoo suite that invokes the function through a command provider.
- A DeepEval test that invokes the same function from Python.
- A controlled failing case to check each runner's process status.
- A plan for replacing the router with an actual LLM service and adding semantic checks.
The router is a harness example, not a claim that deterministic code needs LLM evaluation. It keeps the output contract stable while you compare authoring, execution, and debugging. Once both suites work, the function can call your application API and return the category observed by a user.
Prerequisites
Use Python 3 with venv and a Node.js release supported by your installed Promptfoo package. The current Promptfoo getting-started guide and DeepEval quickstart describe their installation requirements. Install current releases in a scratch directory; when you move the work into a repository, record the resolved packages in its lockfiles rather than copying an invented version from an article.
mkdir llm-eval-compare
cd llm-eval-compare
python3 -m venv .venv
. .venv/bin/activate
python -m pip install -U deepeval
npm init -y
npm install --save-dev promptfoo
Verify installation with python -m pip show deepeval and npx promptfoo --version. They should print package metadata and a CLI version. If npx cannot launch the CLI, check the installed package's Node.js requirement and your active Node runtime. Keep .venv, generated evaluation output, and environment files with credentials out of version control.
1. Promptfoo vs DeepEval: Evaluation Model
Promptfoo reads prompts, providers, tests, and assertions from a configuration. The runner renders test variables into each prompt, calls a provider, and grades the returned output. This structure makes a prompt-by-provider matrix straightforward. A single dataset can reveal whether a candidate prompt improves one model while harming another. The Promptfoo configuration guide documents these parts and how they combine.
DeepEval models an interaction as an LLMTestCase. Its input and actual_output fields identify the question and observed answer; metrics specify the quality claim to check. A Python test can call the app, build a case, and pass the case with metrics to assert_test. That style is useful when setup requires fixtures, database state, a retriever, or a local service client. DeepEval can also evaluate batches outside test functions, but a test function is the clearest way to learn its release-gate behavior.
Both can evaluate an application rather than a raw model. The crucial choice is where you attach the test: to a configurable provider boundary or to Python code that invokes the application. If the configured target omits retrieval, tool execution, or post-processing, an excellent score can coexist with a broken product. Define the tested boundary in the evaluation report so readers know what passed.
2. Promptfoo vs DeepEval: Comparison Table
| Dimension | Promptfoo | DeepEval | Practical question |
|---|---|---|---|
| Authoring | YAML or JavaScript configuration | Python cases and test functions | Which format can your team review accurately? |
| Natural comparison | Prompt by provider by case | Cases scored by chosen metrics | Are you comparing candidates or testing one integrated app? |
| Deterministic checks | equals, contains, and other assertions |
ExactMatchMetric and Python assertions |
Is the expected output a constrained string? |
| App integration | HTTP, script, model, or custom provider | Call application code in a test | How much adapter code will you maintain? |
| RAG evidence | Assertions can inspect configured outputs and context | retrieval_context is a case field for RAG metrics |
Can you capture what the app actually retrieved? |
| Security discovery | Dedicated red-team setup and report commands | Adversarial cases and custom Python checks | Do you need generated attacks or fixed regressions? |
| CI surface | CLI exit status and result files | Test-run exit status and metric reports | Does your CI preserve failures from the runner? |
The table describes convenient starting points, not exclusive capabilities. Both projects evolve and support more than a short comparison can list. Choose based on one representative slice of your system, including the adapter, report, and failure diagnosis. The LLM evaluation framework selection guide gives a broader checklist for team ownership and evaluation scope.
Step 1: Define a Shared Application Contract
Create triage.py in the directory you just made. The function returns one of two allowed strings. That makes a character-for-character comparison meaningful. When you connect a model later, keep the function signature and return value stable. Parse the app response inside this boundary, validate its label, and expose service errors rather than silently turning them into a category.
# triage.py
import sys
def route_ticket(ticket: str) -> str:
text = ticket.lower()
if "invoice" in text or "refund" in text:
return "billing"
return "technical"
if __name__ == "__main__":
print(route_ticket(sys.argv[1]))
Verify the function before adding evaluation tooling:
python triage.py 'Please correct my invoice'
python triage.py 'The login page crashes'
The outputs should be billing and technical, respectively. The default-to-technical rule is a simplification for this exercise. A production classifier may need an unknown result or human handoff for ambiguous requests. Add those cases to the dataset when that is part of the product contract. Do not normalize a wrong label into a passing one; normalize only differences the product itself ignores, such as permitted surrounding whitespace.
This shared function prevents a misleading comparison. If Promptfoo calls the real app while DeepEval scores static strings, different outcomes say nothing about the frameworks. Both runners must exercise the same revision, input, and output boundary for a fair operational trial.
Step 2: Evaluate the Router With Promptfoo
Save promptfooconfig.yaml beside triage.py. The documented custom-script provider sends the rendered prompt as the first command argument. Here the prompt is the ticket text, and Python prints only the label. equals is a deterministic Promptfoo assertion, so this suite does not invoke a grader model.
# promptfooconfig.yaml
prompts:
- '{{ticket}}'
providers:
- 'exec: python triage.py'
tests:
- description: Invoice correction routes to billing
vars:
ticket: Please correct my invoice
assert:
- type: equals
value: billing
- description: Login crash routes to technical
vars:
ticket: The login page crashes
assert:
- type: equals
value: technical
Run the suite, then inspect individual rows in the local viewer:
npx promptfoo eval --config promptfooconfig.yaml --no-cache
npx promptfoo view
Expect two passing cases. If the command provider cannot find python, activate .venv in the same shell and check python --version. Promptfoo inherits the process environment that launched it. If a row fails despite a correct visible label, inspect raw output for a debug line or trailing text. The script must write exactly the category to standard output.
This example verifies wiring, case substitution, and assertion behavior. It does not measure a model. For a real comparison, you can point providers at model APIs or your app endpoint and vary prompt templates while keeping test cases fixed. Count the resulting calls before a large run: prompts multiplied by providers multiplied by cases can grow quickly. The Promptfoo CI tutorial covers running a reviewed suite on code changes.
Step 3: Evaluate the Same Router With DeepEval
Save test_triage.py in the same directory. Each parameterized test calls the shared function and creates an LLMTestCase with input, observed output, and expected output. DeepEval's Exact Match documentation confirms these fields and the ExactMatchMetric API. It compares strings without a judge-model request.
# test_triage.py
import pytest
from deepeval import assert_test
from deepeval.metrics import ExactMatchMetric
from deepeval.test_case import LLMTestCase
from triage import route_ticket
@pytest.mark.parametrize(
("ticket", "expected"),
[
("Please correct my invoice", "billing"),
("The login page crashes", "technical"),
],
)
def test_ticket_route(ticket: str, expected: str) -> None:
case = LLMTestCase(
input=ticket,
actual_output=route_ticket(ticket),
expected_output=expected,
)
assert_test(case, [ExactMatchMetric(threshold=1.0)])
Run the documented test command:
deepeval test run test_triage.py
Both parameterized cases should pass. No remote metric provider is required for this test. The Python form lets you prepare fixture data or inspect a response object before deciding what belongs in actual_output. It also places responsibility for test data and fixtures in your codebase. Keep case construction small enough that a reviewer can tell which application behavior failed.
If your target is asynchronous, adapt the application call using the framework's supported async test pattern rather than assigning a previously captured answer. Static answers are suitable for checking a metric itself; they do not catch an application regression. The DeepEval metrics walkthrough extends this pattern beyond exact labels.
Step 4: Check That Each Runner Really Fails
A passing example is incomplete until you have seen a deliberate failure reach the shell. Temporarily add Please issue a refund to the Promptfoo YAML with expected value technical. The router will return billing, so the assertion should fail. Add the same ticket and wrong expected label to the DeepEval parameter list. Do this only as a short gate check, then remove the bad expectations.
After inserting the Promptfoo case, run:
npx promptfoo eval --config promptfooconfig.yaml --no-cache
printf 'Promptfoo exit: %s
' "$?"
The current Promptfoo CLI reference documents exit code 100 when at least one case fails, unless the configured failure-code override changes it. After inserting the DeepEval case, run:
deepeval test run test_triage.py
printf 'DeepEval exit: %s
' "$?"
DeepEval should return a nonzero status. These verification commands matter in CI: a wrapper that pipes output, suppresses an exception, or reports only a score may accidentally mark a failing release gate as green. Check the final job status with one intentionally incorrect case. Then restore both suites to their two correct cases and rerun them to confirm the clean baseline.
Before increasing case count, give each example a stable identifier and a reason for its expected outcome. Draw cases from bug reports, production traces with permission, and human-reviewed examples. The golden dataset guide covers how to avoid a collection that contains only easy questions.
Step 5: Export and Check a Promptfoo Result File
A CI job needs an inspectable artifact as well as a pass or fail status. Promptfoo can write an evaluation to JSON using its documented --output option. Run the known passing suite and choose an output path inside this scratch directory. Do not upload the file blindly: evaluated prompts and responses can contain private user data once you connect a real application.
npx promptfoo eval --config promptfooconfig.yaml --no-cache --output promptfoo-results.json
python -m json.tool promptfoo-results.json > /dev/null
The first command should complete with two passing cases; the second should exit successfully if the result file is valid JSON. Open the file and locate the per-case outputs and assertion results before designing a dashboard or report. Keep the full artifact when debugging, because an aggregate percentage loses the exact input and returned label. In a production pipeline, restrict artifact access, define retention, and redact sensitive text at collection time. The export does not replace the runner's exit status, which you already proved in Step 4.
9. Connect a Real LLM Application
Replace the inside of route_ticket with a call to the production-facing application path. Keep route_ticket(ticket: str) -> str and the label contract so both test suites continue to run. If the service returns JSON, parse the response once, require the category field, and reject categories outside the allowed set. Put timeouts and retry rules in the application client, not in a metric. A provider outage should appear as an execution error, not as an incorrect-answer score.
A direct Promptfoo model provider is useful when the question is which prompt or model to use. An application command or HTTP target is better when acceptance depends on retrieval, tools, fallback logic, or post-processing. DeepEval's in-process call naturally reaches Python application code, but mocks can weaken it: if a test mocks the retriever and model while claiming to measure end-to-end quality, its score describes the mock. State whether the evaluation targets a component or the whole response path.
Verify the new boundary with the same commands from Steps 2 and 3. Expect real model latency and token use now. Store credentials in environment variables or a secret manager; avoid adding them to YAML, Python fixtures, or result artifacts. Record app revision, model identifier, prompt revision, and case ID for each run. When retrieval changes, preserve document IDs or a snapshot so a failed answer can be reproduced.
If the service returns a label with different capitalization, decide at the product boundary whether capitalization matters. Do not add a permissive assertion merely to hide an app contract violation. Conversely, do not fail a user-safe output because a test expects irrelevant punctuation.
10. Add Semantic and RAG Checks
Exact match is strong for category labels and brittle for free-form explanations. Two useful responses can have different wording, order, and length. For a RAG answer, DeepEval's FaithfulnessMetric evaluates whether claims align with retrieval_context; AnswerRelevancyMetric addresses whether the answer responds to the input. The FaithfulnessMetric docs specify the required case fields. Feed it the passages the application actually retrieved, not the passages you wish it had found.
Promptfoo supports model-graded assertions alongside deterministic ones. A rubric might require every refund condition to be supported by cited policy text. Such an assertion calls a judge provider and needs the appropriate credentials. Keep your equals checks for enum-like contracts and use a semantic judge only where exact rules fail to express the behavior. Review the judge's reasons and errors on a human-labeled sample before setting a threshold.
For both tools, separate factual support, relevance, and formatting. A grounded answer can omit the user's actual question. A relevant answer can invent a policy. A valid JSON object can carry a wrong recommendation. One blended score conceals these failure modes. Use the DeepEval vs Ragas RAG comparison when choosing RAG measures and the regression threshold guide when deciding whether noisy scores should block a release.
Verify the semantic step by rerunning the deterministic cases first. Then run a small reviewed set with the judge configured, inspect each borderline response, and document the threshold decision. This distinguishes a target regression from a judge change or missing retrieved context.
11. Dataset Design and Failure Diagnosis
A useful evaluation dataset is a map of product risks, not a bag of random prompts. For ticket routing, include clear billing and technical examples, mixed-intent tickets, unsupported requests, spelling variants, long tickets, and text that tries to instruct the classifier to ignore its rules. Give each case an expected outcome or a review label. If the product allows multiple valid outcomes, encode that policy explicitly rather than forcing an arbitrary single answer.
Hold a small set of known incidents as stable regression cases. Use a separate discovery set for newly generated or sampled inputs. If you repeatedly tune a prompt against the same tiny set, a rising score may reflect memorization of the test shape rather than better user behavior. Review samples from actual application traffic only after removing private data and confirming you may use them for evaluation.
Promptfoo's matrix view helps spot a prompt that fails across all providers or one provider that fails across every case. DeepEval's test output helps tie a failed metric to Python setup and application traces. In either tool, inspect raw target output before trusting aggregate pass percentage. A failure could mean a wrong label, malformed app JSON, an unavailable provider, a changed retrieval result, or a judge error. Those demand different fixes.
When reporting the outcome, include the target revision, dataset revision, case count, exact gate criteria, and representative failures. A statement such as "92% passed" gives little guidance without denominators and examples. Keep the case IDs stable across reruns so the team can distinguish new failures from persistent ones.
12. Cost, Speed, and Repeatability
The local example uses no model tokens. After connecting an LLM, count target and judge calls separately. As an illustration, two prompts, three providers, and forty cases can produce 240 target calls in a Promptfoo matrix, before model-graded assertions make any judge calls. DeepEval can also call a judge several times per case if you attach multiple metrics. This arithmetic is not a runtime or price benchmark; use your own token logs and provider terms for estimates.
Keep a fast deterministic suite for every pull request and a wider semantic suite for a scheduled or release run if latency becomes costly. Cached responses can help compare assertions on a fixed output, but they cannot verify the current live model. Run critical cases with cache disabled when a fresh response is part of the release claim. Record the resolved model identifier and judge configuration because a floating alias or judge prompt can change without changing your test code.
Repetition helps characterize noisy scores, yet repeated runs are not a substitute for case quality. If a metric fluctuates near a threshold, inspect its rubric and the actual examples. Do not quietly lower the threshold until the report turns green. A threshold should correspond to an acceptable user impact, and it should be checked against human review. See writing custom DeepEval metrics when a built-in metric cannot express a specific product rule.
13. Security and Red-Team Coverage
Promptfoo's dedicated red-team quickstart describes setup, attack generation, execution, and reporting. This helps when you need discovery beyond your fixed regression cases. Define the target's legitimate purpose and forbidden behavior before generating attacks. A reported vulnerability needs human triage with the payload, target response, and violated policy. Generated volume alone is not proof of exploitability.
DeepEval can express adversarial inputs as ordinary Python test cases and score an observable security property. This is convenient when a failure depends on a precise tool response, retrieval document, or application state. A custom metric is still software: test its false-positive and false-negative behavior before trusting it. Generic answer relevance is not a detector for secret disclosure, unauthorized tool use, or cross-user data access.
For the ticket router, add text such as a customer message that asks the model to output an unsupported category. The deterministic sample will ignore most instructions because it does not interpret them. Once route_ticket calls a model, that case becomes a meaningful prompt-injection regression. For a retrieval application, build an adversarial RAG dataset that places hostile text in lower-trust documents, then observe whether the final answer follows it.
Keep security discovery and release gating distinct. New generated attacks are candidates for investigation; confirmed failures become stable cases with explicit expected behavior. This prevents a changing attack generator from making every CI run incomparable.
Which Should You Choose
Choose Promptfoo when the immediate work is comparing prompt or model candidates, sharing a readable evaluation matrix, or using dedicated red-team workflows. Its YAML suite is approachable for QA reviewers who want to see inputs and assertions without reading application internals. Check that the provider you configure reaches the real behavior under test; a direct model call can miss retrieval and post-processing failures.
Choose DeepEval when the evaluated path is a Python application and tests need fixtures, application calls, or explicit RAG case fields. It lets you keep evaluation near code and reason about failures through test functions. Account for the maintenance cost of custom fixtures and metric configuration, especially if only a few engineers understand them.
Use both if they serve different decisions. Promptfoo can compare candidate prompts before adoption while DeepEval guards the integrated Python path. Avoid maintaining the same expectation in two unrelated files forever. Put shared cases in a reviewed source of truth or give the suites distinct scopes. If your main need is trace-linked experimentation instead, compare Promptfoo vs LangSmith before committing to a runner.
Troubleshooting
Promptfoo cannot find Python -> Reactivate .venv, check python --version, and run the CLI from the directory containing triage.py. The command provider inherits the launching process environment.
An exact assertion fails on a correct-looking word -> Inspect raw standard output for extra logging, whitespace, and capitalization. equals and ExactMatchMetric compare strings, not intent.
DeepEval collects no cases -> Confirm the filename and function begin with test_, run from the project directory, and verify that the active environment contains DeepEval.
A metric requests a model key in the offline sample -> Check whether you added a judge-based metric. The shown ExactMatchMetric does not call an LLM; a semantic judge needs its provider configured.
The frameworks disagree after integration -> Capture the same input, target revision, raw output, and retrieval snapshot from each run. Check for stale cache and differing normalization before comparing aggregate results.
Common Mistakes
- Testing only easy, well-formed requests and calling the result production coverage.
- Using exact string equality for open-ended prose with several valid answers.
- Passing ideal reference documents as
retrieval_contextinstead of actual retrieved passages. - Changing the application prompt and judge rubric at the same time, then attributing a score shift to one cause.
- Reporting a pass percentage without case IDs, raw outputs, and reasons for failed assertions.
- Choosing a semantic threshold from one run with no human-reviewed calibration set.
- Letting a CI wrapper discard a nonzero exit code from the evaluation command.
- Saving secrets or unredacted customer text in configuration, fixtures, or shared reports.
Interview Questions and Answers
The interviewQnA entries below ask how you would select a boundary, prove a CI gate, and diagnose disagreement between tools. Prepare to explain a real failed case, the evidence needed to reproduce it, and why your metric matches the product risk. An interviewer should be able to challenge your expected outcome; a tool feature list is not enough.
Where To Go Next
Preserve the two green examples as wiring checks, then grow a reviewed dataset from actual failure modes. Use Promptfoo evals in CI to operationalize a matrix, or extend the Python suite using the DeepEval metrics guide. Add one metric per named behavior before combining scores. Keep an exact contract check even when you add a semantic judge.
When presenting this work as a QA project, show the target boundary, dataset, passing and failing examples, and the deliberate gate test. You can map that evidence to a job description in the resume dashboard and rehearse the trade-offs in interview practice. These artifacts demonstrate that you can design an evaluation, not merely run a CLI.
Conclusion
Promptfoo vs DeepEval comes down to the workflow that makes your application risk visible. Promptfoo is a strong starting point for prompt/provider matrices and adversarial discovery. DeepEval fits Python application tests with explicit cases and metrics. Both can reject the same wrong routing label, as the runnable examples show.
Start with the offline suites, inject one controlled failure, and confirm your CI job fails. Then connect the real application path and add reviewed cases for behaviors that exact matching cannot cover. Choose the primary runner from the defects it reveals and the effort required to keep its results trustworthy.
Interview Questions and Answers
How would you choose between Promptfoo and DeepEval for a new assistant?
I would identify whether the target is a model prompt, HTTP service, or integrated application. Promptfoo is efficient for a prompt/provider matrix and adversarial discovery. DeepEval fits a Python codebase that needs fixtures and direct calls. I would prototype one representative case in both and compare maintenance cost and failure visibility.
What does the offline ticket-router example prove?
It proves that both runners invoke the same target, compare outputs to expectations, and expose a failing status. It does not measure generative quality because the router is deterministic. I would use it to validate CI wiring before paying for judge calls.
Why is ExactMatchMetric inappropriate for many generated answers?
Free-form answers can be correct with different wording, order, or punctuation. Exact matching would mark those alternatives as failures. It is appropriate for a constrained label; for prose I would use a task-specific rubric and inspect borderline outputs.
What data must a RAG faithfulness test capture?
The case needs the user input, the application's actual answer, and the retrieval context supplied to generation. The context must come from the tested run, not a manually selected ideal passage. I would record document identifiers or a snapshot so a failure can be reproduced.
How would you check that an eval blocks a bad release?
I would add one intentionally wrong expected value, run the exact CI command, and confirm the job fails. Then I would inspect whether the failure points to the correct case and metric. This catches wrappers that swallow nonzero process statuses.
How do model-graded metrics affect reproducibility?
Judge output can change with its model, prompt, provider behavior, or sampling. I would record those settings, keep a human-reviewed calibration set, and monitor borderline cases. One aggregate score is not conclusive evidence.
When would you red team with Promptfoo instead of hand-writing cases?
I would use generated attacks for discovery after defining the application's purpose and forbidden behavior. Hand-written cases remain useful for known incidents and stable regressions. Every reported vulnerability needs inspection of the payload, response, and violated rule.
What is a fair cost comparison between the frameworks?
I would count target calls, judge calls, and repeated runs for the same reviewed dataset. Promptfoo matrices can multiply calls across prompts and providers; DeepEval can multiply calls across metrics. I would measure latency and token usage in the actual pipeline rather than rely on generic benchmarks.
Frequently Asked Questions
Is Promptfoo or DeepEval better for comparing prompts?
Promptfoo is the more direct starting point when you want a matrix of prompt templates, providers, and cases. DeepEval can compare outputs too, but you usually write Python code to organize that experiment.
Can I run Promptfoo and DeepEval without an LLM API key?
Yes. Promptfoo can call a local script with deterministic assertions. DeepEval can run ExactMatchMetric on locally produced outputs. Model-graded metrics and remote targets need the relevant provider setup.
Which tool should I use for RAG faithfulness?
DeepEval provides FaithfulnessMetric to evaluate an answer against retrieval_context. Promptfoo can test RAG behavior through its assertion and provider configuration. In either tool, capture the context the application actually retrieved.
Does Promptfoo replace pytest for LLM testing?
Promptfoo is a standalone evaluation runner with CLI status and result output. It does not need to replace conventional software tests. Keep each check at the boundary where it gives useful evidence.
Does DeepEval require Confident AI?
No. DeepEval can run local test files through deepeval test run. Its hosted platform is optional for teams that want centralized results and related features.
Can I use both frameworks in one project?
Yes, when each answers a different question. One practical split is Promptfoo for candidate prompt comparisons and DeepEval for integrated Python regression tests. Avoid maintaining duplicate expectations in unrelated datasets.
How do I know an LLM evaluation gate is working?
Insert one deliberately wrong expected result and run the exact CI command. Confirm the job fails and identifies the correct case, then remove the bad expectation. Repeat when wrapper scripts change how command status is handled.