QA How-To
LLM Evaluation Metrics Explained for Testers
LLM evaluation metrics explained for testers with runnable Python examples for accuracy, retrieval, grounding, abstention, latency, cost, and release gates.
22 min read | 2,904 words
TL;DR
Use a metric set that separates answer correctness, retrieval, grounding, abstention, latency, and cost. This tutorial builds a reproducible Python evaluator, checks its boundary cases, and compares a candidate release with a baseline.
Key Takeaways
- Choose metrics from product failure modes and the evidence each case captures.
- Use exact match for canonical answers and token F1 only as a lexical signal.
- Measure retrieval relevance separately from generated-answer correctness and claim support.
- Represent undefined scores as null, then exclude them explicitly from eligible averages.
- Compare releases on the same labeled cases and preserve per-case results.
- Calibrate semantic judges against human labels before adding them to release gates.
LLM evaluation metrics explained for testers begin with a question: what failure would this score detect in your product? Exact match can catch a wrong contract value, but it cannot judge a useful paraphrase. A groundedness check can expose an unsupported claim, but it cannot prove the answer addresses the user's request. Use a small set of complementary measures and keep each score tied to visible evidence.
This tutorial builds a local evaluation harness for a document-answering assistant. You will measure answer correctness, retrieval quality, claim support, abstention, latency, and cost from a four-case fixture. Then you will compare a candidate release with a baseline and turn the comparison into a gate. The fixture is deliberately small so you can inspect every calculation; a production suite needs more cases and reviewed labels.
What You Will Build
- A JSON golden dataset with prompts, reference answers, retrieved document IDs, relevance labels, claim support labels, latency, and cost.
- A Python module that computes deterministic answer and retrieval metrics with explicit rules for missing labels.
- A command-line report that preserves per-case results and a macro summary.
- A candidate-versus-baseline release check, followed by unit tests for boundary cases.
- A method for deciding when to add a human or LLM judge after deterministic checks run out of useful signal.
The worked examples use saved observations rather than calling a model. That makes the tutorial reproducible without credentials and lets you verify the arithmetic before attaching a live application. Later, replace fixture answers and trace fields with captured application outputs while keeping the same schema. For a broader architecture, see the LLM evaluation pipeline guide.
Prerequisites
Use Python 3.12 and its standard library. The code also works with later compatible Python versions, but Python 3.12 is the version used for the commands here. Check your exact installed patch version with python3 --version; do not pin a fabricated patch number. You need a shell and a writable empty directory. There are no pip dependencies, model keys, or hosted services.
python3 --version
mkdir llm-metrics-demo
cd llm-metrics-demo
Verify: the first command prints Python 3.12.x, and the shell is now inside llm-metrics-demo. On a system where python3 points to another version, use the executable for your installed Python 3.12. The implementation relies on documented json, collections.Counter, re, math, and unittest APIs in the Python 3.12 standard library.
Before scoring, define the unit of analysis. One row below represents one user request and one observed answer. The retrieved_ids field records the ordered top results. The relevant_ids field is a human-labeled set of documents that could answer the request; it must be labeled independently of what happened to be retrieved. A claim's supported value is a reviewed judgment against the supplied evidence. These fields have different owners and should never be inferred from the model's own answer.
| Question | Metric | Required evidence | Main limitation |
|---|---|---|---|
| Is a fixed answer exactly right? | Normalized exact match | Reference and actual answer | Rejects valid paraphrases |
| How much wording overlaps? | Token F1 | Reference and actual answer | Rewards shared words, not truth |
| Did retrieval return useful documents? | Precision@2 and recall@2 | Ranked IDs and relevance labels | IDs say little about passage quality |
| Are stated claims supported? | Grounded claim rate | Answer claims and reviewed support labels | Depends on claim segmentation and evidence review |
| Did the system decline when needed? | Abstention correctness | Expected abstention and observed answer | Needs a controlled refusal vocabulary |
| Is the experience affordable and fast? | p95 latency and mean cost | Measured trace fields | Small fixtures make tail estimates unstable |
Step 1: Create a Golden Dataset With Traceable Labels
Save these four cases as cases.json. The values are illustrative, including the latency and dollar amounts. Do not treat them as benchmark measurements. Case qa-4 intentionally has a wrong warranty answer even though retrieval found the right document. This distinction will show why a retrieval score cannot substitute for an answer score.
cat > cases.json <<'JSON'
[
{
"id": "qa-1", "slice": "fact",
"reference": "Paris", "answer": "Paris",
"should_abstain": false,
"retrieved_ids": ["fr-guide", "de-guide"],
"relevant_ids": ["fr-guide"],
"claims": [{"text": "Paris is the capital", "supported": true}],
"latency_ms": 800, "cost_usd": 0.003
},
{
"id": "qa-2", "slice": "policy",
"reference": "30 days", "answer": "30 days",
"should_abstain": false,
"retrieved_ids": ["old-policy", "current-policy"],
"relevant_ids": ["current-policy"],
"claims": [{"text": "The return window is 30 days", "supported": true}],
"latency_ms": 1100, "cost_usd": 0.004
},
{
"id": "qa-3", "slice": "unanswerable",
"reference": "I don't know", "answer": "I don't know",
"should_abstain": true,
"retrieved_ids": [], "relevant_ids": [],
"claims": [],
"latency_ms": 600, "cost_usd": 0.002
},
{
"id": "qa-4", "slice": "warranty",
"reference": "two years", "answer": "one year",
"should_abstain": false,
"retrieved_ids": ["warranty", "shipping"],
"relevant_ids": ["warranty"],
"claims": [{"text": "The warranty lasts one year", "supported": false}],
"latency_ms": 900, "cost_usd": 0.003
}
]
JSON
python3 -m json.tool cases.json > /dev/null
Verify: json.tool exits with status zero. Also inspect the IDs with python3 -c 'import json; print([x["id"] for x in json.load(open("cases.json"))])'. The expected list is qa-1 through qa-4. A malformed fixture should stop the evaluation before it produces misleading averages.
A real golden dataset should include source snapshots or stable document identifiers, a labeler, label date, and disagreement notes. Never label a document relevant merely because the retriever returned it. Add hard cases such as contradictory policies, outdated documents, requests with no answer, and near-miss entities. The golden dataset tutorial goes deeper into sampling and review.
Step 2: Put LLM Evaluation Metrics Explained for Testers Into Code
Create metrics.py with all scoring functions. The first two scores evaluate output text. Normalized exact match lowercases and collapses punctuation to tokens, so "Paris." and "paris" match. Token F1 counts shared token occurrences, not semantic equivalence. Its precision denominator is the answer length and its recall denominator is the reference length. When both strings contain no tokens, the code returns 1; when only one is empty, it returns 0.
cat > metrics.py <<'PYCODE'
from collections import Counter
from math import ceil
import re
def tokens(text):
return re.findall(r"[a-z0-9]+", text.lower())
def exact_match(reference, answer):
return tokens(reference) == tokens(answer)
def token_f1(reference, answer):
expected, actual = Counter(tokens(reference)), Counter(tokens(answer))
if not expected and not actual:
return 1.0
if not expected or not actual:
return 0.0
overlap = sum((expected & actual).values())
precision = overlap / sum(actual.values())
recall = overlap / sum(expected.values())
return 0.0 if overlap == 0 else 2 * precision * recall / (precision + recall)
def precision_at_k(retrieved_ids, relevant_ids, k=2):
if not relevant_ids:
return None
return sum(doc in set(relevant_ids) for doc in retrieved_ids[:k]) / k
def recall_at_k(retrieved_ids, relevant_ids, k=2):
if not relevant_ids:
return None
return len(set(retrieved_ids[:k]) & set(relevant_ids)) / len(set(relevant_ids))
def grounded_claim_rate(claims):
if not claims:
return None
return sum(bool(c["supported"]) for c in claims) / len(claims)
def abstention_correct(answer, should_abstain):
return (" ".join(tokens(answer)) in {"i don t know",
"cannot answer from the provided documents"}) == should_abstain
def evaluate(case):
return {
"id": case["id"],
"slice": case["slice"],
"exact_match": exact_match(case["reference"], case["answer"]),
"token_f1": token_f1(case["reference"], case["answer"]),
"precision_at_2": precision_at_k(case["retrieved_ids"], case["relevant_ids"]),
"recall_at_2": recall_at_k(case["retrieved_ids"], case["relevant_ids"]),
"grounded_claim_rate": grounded_claim_rate(case["claims"]),
"abstention_correct": abstention_correct(case["answer"], case["should_abstain"]),
"latency_ms": case["latency_ms"],
"cost_usd": case["cost_usd"],
}
def mean_defined(values):
known = [value for value in values if value is not None]
return None if not known else sum(known) / len(known)
def percentile_nearest_rank(values, fraction):
ordered = sorted(values)
return ordered[ceil(fraction * len(ordered)) - 1]
def summarize(rows):
fields = ("exact_match", "token_f1", "precision_at_2",
"recall_at_2", "grounded_claim_rate", "abstention_correct",
"cost_usd")
result = {field: mean_defined([row[field] for row in rows]) for field in fields}
result["latency_p95_ms"] = percentile_nearest_rank(
[row["latency_ms"] for row in rows], 0.95)
result["case_count"] = len(rows)
return result
PYCODE
python3 -m py_compile metrics.py
Verify: py_compile exits without a syntax error. Then run python3 -c 'from metrics import exact_match, token_f1; print(exact_match("Paris.", "paris"), token_f1("two years", "one year"))'. It prints True followed by a fractional F1. The exact fraction is 0.5 because one of two tokens overlaps in both directions.
The abstention function deliberately accepts two controlled response forms. Production systems often produce many refusal phrasings; either normalize them with a documented classifier or use a reviewed label. Do not quietly assume that any sentence containing "cannot" is an abstention. Likewise, exact match is suitable for a fixed field or a short canonical answer, while open-ended explanations need a rubric or a human review. The LLM evaluation interview questions cover why one metric cannot stand in for overall quality.
Step 3: Read Retrieval and Grounding Scores Without Mixing Them
The code separates ranking from generation. Precision@2 divides relevant returned documents by two slots. Recall@2 divides unique relevant returned documents by the size of the independently labeled relevant set. For qa-2, only the second result is relevant, so precision@2 is 0.5 and recall@2 is 1.0. A different metric, reciprocal rank, would expose that the relevant result appeared second; this tutorial keeps the initial report small.
Run the following direct checks against the Step 1 fixture. They test a formula, an undefined case, and a supported-claim label. The use of None means "not applicable," not a perfect or failing retrieval score.
python3 - <<'PYCODE'
import json
from metrics import evaluate
cases = json.load(open("cases.json", encoding="utf-8"))
for case in cases:
row = evaluate(case)
print(case["id"], row["precision_at_2"], row["recall_at_2"],
row["grounded_claim_rate"])
PYCODE
Verify: qa-1, qa-2, and qa-4 each show 0.5, 1.0 for precision and recall. qa-3 shows None for both because no relevant document was labeled. Its grounded claim rate is also None because it made no factual claim. qa-4 shows a grounded claim rate of 0.0 despite recall@2 of 1.0. That contrast is the diagnostic point: the answer generation stage can fail with adequate retrieval.
A passage can be topically relevant while still lacking evidence for a precise claim. If that distinction matters, store passage text, spans, and citation offsets, then have reviewers label support at the claim level. Context precision and faithfulness in RAG evaluation tools address related questions, but their definitions and required inputs differ. Read the Ragas metrics documentation before equating a framework score with these hand-labeled formulas. For retrieval experiments, the retrieval precision guide shows additional ranking choices.
Step 4: Produce a Baseline Report and Inspect Slices
Save run_eval.py. It reads a dataset path and prints a JSON object with individual rows and a summary. Keeping rows alongside the mean lets you see which prompt failed. The summary uses a macro average: each eligible case contributes one value per metric. For grounded_claim_rate, that means one row with five claims carries the same weight as one row with a single claim. If claim volume matters, calculate a separate claim-weighted rate and name it clearly.
cat > run_eval.py <<'PYCODE'
import argparse
import json
from metrics import evaluate, summarize
parser = argparse.ArgumentParser()
parser.add_argument("dataset")
args = parser.parse_args()
with open(args.dataset, encoding="utf-8") as source:
cases = json.load(source)
if not cases:
raise SystemExit("Dataset must contain at least one case")
rows = [evaluate(case) for case in cases]
print(json.dumps({"rows": rows, "summary": summarize(rows)}, indent=2))
PYCODE
python3 run_eval.py cases.json > baseline-report.json
python3 -m json.tool baseline-report.json > /dev/null
Verify: open baseline-report.json or print its summary with python3 -c 'import json; print(json.load(open("baseline-report.json"))["summary"])'. The exact-match mean is 0.75, the grounded-claim mean is 2/3, the case count is 4, and nearest-rank p95 latency is 1100 ms. The p95 value is simply the maximum of four observations under this estimator. It is not evidence that a production service has a 1100 ms p95.
Check slice results before making a release decision. Here every slice has one case, so a slice mean is just that row's result. In a larger dataset, summarize separately by workflow, language, customer tier, answerability, policy sensitivity, and document age. A global improvement can conceal a critical regression in an uncommon slice. Keep the same prompts and labels for baseline and candidate comparisons; otherwise, a change in dataset composition can masquerade as a product change.
Record the run's application build, model identifier, prompt revision, retrieval index snapshot, and dataset revision outside this minimal fixture. Those identifiers are essential when a score moves unexpectedly. For live runs, log latency from a monotonic clock around the actual request and derive cost from provider usage or a documented pricing calculation. Never estimate cost from response length alone. The latency and cost measurement guide discusses trace capture.
Step 5: Turn LLM Evaluation Metrics Explained for Testers Into a Release Gate
Create candidate.json by copying the baseline observations and fixing only qa-4's answer and support label. This is a controlled example: the candidate improves answer quality while keeping the saved latency and cost observations identical. In a real evaluation, rerun the application for every case and preserve both trace sets. Do not edit observations by hand to claim a performance improvement.
python3 - <<'PYCODE'
import json
with open("cases.json", encoding="utf-8") as source:
cases = json.load(source)
candidate = [dict(case) for case in cases]
target = next(case for case in candidate if case["id"] == "qa-4")
target["answer"] = "two years"
target["claims"] = [{"text": "The warranty lasts two years", "supported": True}]
with open("candidate.json", "w", encoding="utf-8") as output:
json.dump(candidate, output, indent=2)
PYCODE
python3 run_eval.py candidate.json > candidate-report.json
Verify: python3 -c 'import json; d=json.load(open("candidate-report.json")); print(d["summary"]["exact_match"], d["summary"]["grounded_claim_rate"])' prints 1.0 and 1.0. The policy and abstention results should be unchanged. The candidate is an artificial improvement case for validating the gate, not an empirical claim about an LLM.
Now save compare.py. The example policy requires no decrease in exact match, grounded claim rate, or abstention correctness; it also caps p95 latency and mean cost at 125% of the baseline. The 25% allowance is an illustrative team policy, not a universal threshold. Missing scores fail the gate rather than becoming zero or one.
cat > compare.py <<'PYCODE'
import json
import sys
with open(sys.argv[1], encoding="utf-8") as source:
baseline = json.load(source)["summary"]
with open(sys.argv[2], encoding="utf-8") as source:
candidate = json.load(source)["summary"]
checks = {}
for name in ("exact_match", "grounded_claim_rate", "abstention_correct"):
left, right = baseline[name], candidate[name]
checks[name] = left is not None and right is not None and right >= left
for name in ("latency_p95_ms", "cost_usd"):
left, right = baseline[name], candidate[name]
checks[name] = left is not None and right is not None and right <= left * 1.25
checks["case_count"] = baseline["case_count"] == candidate["case_count"]
for name, passed in checks.items():
print(f"{name}: {'PASS' if passed else 'FAIL'}")
sys.exit(0 if all(checks.values()) else 1)
PYCODE
python3 compare.py baseline-report.json candidate-report.json
Verify: every check prints PASS and the process exits zero. To prove the gate can fail, run python3 compare.py candidate-report.json baseline-report.json; it exits one because the baseline's lower exact-match and grounding scores now count as a regression against the candidate. In CI, use the real prior release as the first argument and the build under test as the second.
This aggregate gate is only a starting point. A wrong answer about a high-risk policy should be able to block a release even when the mean improves. Add per-case critical assertions and slice floors after you have enough labeled examples. Thresholds should come from a reviewed baseline, severity, and observed run-to-run variation. The regression threshold guide explains how to set those limits without borrowing somebody else's number.
Step 6: Test Metric Boundaries Before Trusting the Gate
Save tests that target the parts most likely to produce a plausible but false report: punctuation normalization, repeated tokens, empty relevance labels, empty claims, and the nearest-rank percentile. These are unit tests of the evaluator, not tests that restate the four fixture results. A defect in a metric function can approve a bad model change even when the application tests are correct.
cat > test_metrics.py <<'PYCODE'
import unittest
from metrics import (
exact_match, token_f1, precision_at_k, recall_at_k,
grounded_claim_rate, percentile_nearest_rank
)
class MetricBoundaryTests(unittest.TestCase):
def test_normalized_exact_match(self):
self.assertTrue(exact_match("Paris.", "PARIS"))
self.assertFalse(exact_match("two years", "one year"))
def test_repeated_tokens_count_once_per_occurrence(self):
self.assertAlmostEqual(token_f1("red red blue", "red blue"), 0.8)
def test_no_relevance_label_is_not_perfect_retrieval(self):
self.assertIsNone(precision_at_k([], []))
self.assertIsNone(recall_at_k([], []))
self.assertEqual(recall_at_k([], ["needed"]), 0.0)
def test_no_claims_are_unscored(self):
self.assertIsNone(grounded_claim_rate([]))
def test_nearest_rank_small_sample(self):
self.assertEqual(percentile_nearest_rank([800, 1100, 600, 900], 0.95),
1100)
if __name__ == "__main__":
unittest.main()
PYCODE
python3 -m unittest -v test_metrics.py
Verify: five tests pass. If the repeated-token test fails, inspect whether your implementation counts token multiplicity; set overlap would incorrectly treat "red red blue" like "red blue." If the empty-label tests fail, inspect whether a missing denominator was silently mapped to a desirable score.
Before using the script on live data, add schema validation. The small fixture uses required JSON keys, so a missing field already raises KeyError, but a string in latency_ms or a duplicate case ID deserves a clearer error. Check that every observed answer belongs to exactly one case ID, that relevant IDs use the same document namespace as retrieved IDs, that support labels come from review, and that costs use the same currency. Add a separate test for your application's structured output contract if it emits JSON. For repeated model sampling, use the nondeterminism testing guide.
How to Choose the Next Metric
Start from a failure taxonomy, not a dashboard menu. If the application promises a JSON schema, validate the schema and required fields before any semantic grading. If it answers a known entity question, exact match against a canonical label is cheap and interpretable. If it retrieves documents, assess the retriever and the generated answer separately. If it summarizes free-form documents, use claim support and task-specific human review. If it operates tools, evaluate tool arguments, permission checks, and final task completion rather than only final prose.
Metric names can hide different denominators. For retrieval, precision@k depends on k and on whether empty slots count. Recall@k depends on the set of relevant documents, which may be incomplete if annotators never searched beyond returned results. For grounding, a model can produce a single broad claim or many narrow claims; segmentation changes the rate. State these choices in your report. The same numeric value from two implementations may not be comparable.
An LLM judge becomes useful when exact matching and lexical overlap cannot distinguish a good paraphrase from a bad answer. Give it a narrow rubric, source evidence, and an output schema. Sample human-reviewed cases, compare judge decisions with human labels, inspect disagreements, and retain high-impact human review. Pairwise comparisons can be easier to calibrate than unexplained 1-to-5 scores, but presentation order can affect judgment. Research on position bias in LLM judges supports checking swapped answer order. The judge calibration tutorial provides a focused workflow.
Troubleshooting
Problem: Exact match rejects a correct paraphrase -> Keep exact match for canonical short answers and structured fields. Add a reviewed semantic rubric for open-ended answers, and report it under a distinct metric name. Do not weaken normalization until opposite meanings become indistinguishable.
Problem: Recall@2 is always 1.0 -> Audit relevant_ids. If labelers copied retrieved_ids into that field, the score is circular. Have reviewers identify relevant documents from a wider candidate pool or an independent search, then record incomplete-label cases separately.
Problem: A refusal has no grounded claims -> Leave grounded_claim_rate as null for that case. Evaluate the refusal with an expected-abstention label and, where relevant, a safety or policy rubric. Counting no claims as perfect grounding rewards a system that refuses everything.
Problem: CI flips between pass and fail -> Capture several runs of the same build and inspect per-case variance. Keep deterministic gates for schema and permissions; use repeated trials and confidence-aware decisions for stochastic answer quality. Record model and retrieval configuration on every run.
Problem: Candidate looks better but is slower -> Review per-request latency distribution and trace where time increased. The four-row example p95 is a teaching calculation, not a stable tail estimate. Collect a representative larger sample before setting a service-level latency gate.
Problem: Judge scores rise while reviewers report worse answers -> Revisit the rubric and the human calibration set. Check whether verbosity, answer order, or copied source text is biasing the judge. Route disagreements to review and version the judge prompt and model with the run.
Interview Questions and Answers
The interviewQnA field below contains practical questions on metric selection, undefined denominators, label quality, judge calibration, and release gates. In an interview, explain what each score measures, show one counterexample it misses, and name the evidence needed to investigate a failure. A credible answer separates the retrieval stage, the generation stage, and the product-level outcome.
Common Mistakes
- Reporting one overall "quality score" without naming the case set, rubric, aggregation, and exclusions.
- Treating token F1 as proof that a factual statement is true.
- Using generated answers to label retrieval relevance, which makes the evaluation circular.
- Turning null metrics into 100% scores, especially for unanswered or unlabeled cases.
- Comparing two releases on different datasets without accounting for composition changes.
- Letting a mean pass while a critical safety or policy case fails.
- Reusing an LLM judge without checking human agreement after a model, prompt, or domain change.
- Presenting latency or cost figures from a toy fixture as production benchmarks.
Where To Go Next
Attach the runner to captured application traces, expand the golden set by failure mode, and add critical-case assertions before you rely on a release gate. For a production evaluation system, follow the LLM evaluation pipeline guide. If the main risk is answer grounding, build stronger evidence labels and compare with RAG faithfulness methods. If the main risk is score instability, run repeated trials and review variance before tightening thresholds.
Conclusion
LLM evaluation metrics explained for testers are useful when every score points to a concrete failure and an inspectable record. This tutorial gave you an executable baseline, candidate comparison, and tests for the metric code itself. Keep answer correctness, retrieval, grounding, abstention, latency, and cost separate; then add human judgment where the task cannot be reduced to deterministic labels. Your next step is to replace the four illustrative rows with reviewed traces from your own product and examine each failed case before adjusting a threshold.
Interview Questions and Answers
How would you choose metrics for a RAG support bot?
I would separate retrieval, answer, and operational failures. I would label relevant documents independently, measure precision and recall at a chosen k, review claim support against retrieved passages, and score expected abstentions. I would retain individual failed traces because an aggregate can conceal a severe policy error.
Why can exact match fail on a good LLM answer?
Two answers can express the same fact with different wording, so a string comparison produces a false failure. I use exact match for canonical values such as a date or code, with documented normalization. Open-ended answers need a rubric, reviewed examples, or a calibrated judge.
What does a high retrieval recall with low groundedness suggest?
The relevant evidence was available in the retrieved set, but the answer still made unsupported claims. I would inspect prompt construction, passage truncation, citation mapping, and the generator's output. I would not spend the first debugging pass tuning the retriever.
How do you prevent a circular retrieval evaluation?
Relevant documents must be labeled independently of the retriever's returned list. I would ask reviewers to search a wider corpus, use known source records, and track label completeness. Copying returned IDs into the relevance set makes recall look artificially high.
How would you evaluate an unanswerable question?
I would label the expected behavior as abstention and check whether the response actually declined. Groundedness is undefined when the response makes no factual claims; I would not award it 100%. I would also inspect whether the system leaked a guess or presented weak evidence as certainty.
How would you validate an LLM-as-a-judge metric?
I would define a narrow rubric with examples, collect human labels, and measure judge agreement and disagreement by failure slice. For pairwise judging, I would swap answer order on a sample to check position effects. I would version the judge model and prompt, then review drift after changes.
What should a release gate do when a metric is undefined?
It should make the missing evidence visible and follow an explicit policy, usually failing a required gate or excluding the case from a documented optional metric. Silently coercing null to zero or one changes the decision. I would validate eligibility counts alongside score values.
How do you compare two LLM builds fairly?
I would run both against the same labeled prompts and source snapshot, preserve configuration identifiers, and collect enough repeated samples for stochastic outputs. I would compare per-case changes and critical slices before trusting macro averages. Latency and cost should come from measured traces under comparable conditions.
Frequently Asked Questions
What are the most useful LLM evaluation metrics for testers?
Start with metrics tied to your application's failure modes: exact match or schema validity for fixed outputs, retrieval precision and recall for search, claim support for grounded answers, and abstention correctness for unanswerable requests. Add latency and cost as operational measures. A single overall score cannot explain which stage failed.
Is token F1 a semantic correctness metric?
No. Token F1 measures word overlap after tokenization. It can reward a false sentence with familiar words and penalize a correct paraphrase, so use it as a diagnostic signal beside reviewed labels or a calibrated semantic rubric.
How do I handle a case with no relevant documents?
Mark retrieval precision and recall as not applicable if the dataset has no relevant-document label. Evaluate whether the assistant abstained instead. If the label is missing because review is incomplete, record that reason and do not score the case as perfect retrieval.
What is the difference between retrieval recall and groundedness?
Retrieval recall asks whether the ranked results contain labeled relevant documents. Groundedness asks whether claims in the answer are supported by the evidence supplied to the model. High recall can coexist with an invented answer.
Should an LLM judge replace human evaluation?
Use a judge for scalable, narrowly defined semantic checks after calibrating it on human-reviewed cases. Keep humans in the loop for disagreements and high-impact outcomes. Recheck agreement when the judge model, rubric, or domain changes.
How should I set an LLM evaluation release threshold?
Measure a reviewed baseline on a stable dataset, account for repeated-run variation, and set separate gates for critical cases and operational limits. The 25% latency and cost allowance in the tutorial is illustrative; your threshold should reflect product requirements and observed data.
Why does a four-case p95 equal the maximum latency?
The tutorial uses the nearest-rank percentile, whose 95th-percentile rank for four observations is the fourth ordered value. Such a small set cannot estimate production tail latency reliably. Use many representative traces before treating p95 as a service target.