QA How-To
Ragas Tutorial: Evaluate a RAG Pipeline
Ragas tutorial to evaluate a RAG pipeline with runnable Python, trace validation, retrieval and answer metrics, CSV reports, and calibrated release gates.
24 min read | 3,003 words
TL;DR
Build a reviewed JSONL trace set, validate it, and score each case with Ragas collections metrics for retrieval precision, retrieval recall, faithfulness, and answer relevancy. Inspect individual failures, then calibrate CI thresholds against human labels before treating scores as a release gate.
Key Takeaways
- Capture the question, ranked retrieved chunks, final response, and reviewed reference from the same RAG request.
- Validate trace shape before paying for judge calls.
- Use collections-based ContextPrecision, ContextRecall, Faithfulness, and AnswerRelevancy metrics for distinct failure modes.
- Keep per-case scores and source artifacts instead of relying on one average.
- Calibrate release cutoffs against human labels and keep provider failures separate from quality failures.
- Record evaluator, model, dataset, prompt, and index versions with every comparison.
A Ragas Tutorial to Evaluate a RAG Pipeline should show what each score measures, which trace fields it needs, and how to turn a failed score into a specific investigation. This tutorial builds a repeatable evaluation of saved retrieval augmented generation traces. You will score retrieval ranking, retrieval coverage, answer grounding, and answer relevance using the current Ragas collections API.
The sample domain is a small returns policy. Its records are deliberately simple so you can inspect every retrieved chunk and expected answer. Once the script runs, replace those records with traces from your own pipeline. A saved trace still represents an end-to-end RAG execution: it contains the user's question, the contexts actually returned by retrieval, and the final response. Scoring that fixed boundary makes prompt and retriever changes comparable.
What You Will Build
- A JSONL set of RAG traces with stable case IDs, references, responses, and ranked contexts.
- A Python validator that rejects incomplete or incorrectly shaped cases before calling a judge.
- A scorer using
ContextPrecision,ContextRecall,Faithfulness, andAnswerRelevancyfromragas.metrics.collections. - A CSV report with per-case scores and a small release gate that identifies the failed dimension.
- A workflow for replacing example traces with real application output and reviewing score changes.
This is a small evaluation harness, not a benchmark claim. No example threshold here is a universal RAG quality standard. Ragas describes these metrics in its official metric documentation. For a broader test strategy, see the RAG application evaluation guide.
Prerequisites
Use Python 3.12, a working python3.12 executable, an OpenAI API key with access to the judge and embedding models selected below, and an isolated virtual environment. The commands install the current compatible Ragas and OpenAI packages available to your environment. Print their actual installed versions and record those values with each run; do not copy an invented package pin from an article. If your organization pins dependencies, use the exact versions validated in its lockfile. The example model identifiers are names accepted by the official Ragas examples, but provider availability and entitlements can change, so substitute models available in your account if needed.
Judge calls and embeddings incur provider usage. The fixture has three cases, but each metric can make several model calls. Run the one-case smoke evaluation before the full set. Keep the API key in an environment variable or secret store and redact personal data from production traces. A retrieval chunk can contain the same sensitive data as the original document.
The required fields are precise: user_input is the user's question, retrieved_contexts is the ranked list delivered to generation, response is what the application produced, and reference is a human-reviewed acceptable answer. A reference is needed for the retrieval metrics in this exercise. It is not a copy of the generated answer. If you cannot supply a trustworthy reference, omit those reference-based metrics instead of filling the field with guessed text.
Step 1: Install and inspect the evaluator
Create a clean working directory for the tutorial and install dependencies. The python3.12 command makes the interpreter requirement explicit. On a system where the executable has a different path, use the path to your installed Python 3.12 binary.
mkdir ragas-rag-eval
cd ragas-rag-eval
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install ragas openai
export OPENAI_API_KEY='replace-with-your-evaluation-key'
Verify the installed interpreter and distribution metadata before writing test data:
python --version
python -m pip show ragas openai
python -c "from ragas.metrics.collections import ContextPrecision, ContextRecall, Faithfulness, AnswerRelevancy; print('collections imports OK')"
The last command should print collections imports OK. If it fails on an import, check python -m pip show ragas inside the active environment. Ragas documents the collections imports for new projects; older examples using ragas.metrics and single_turn_ascore belong to its legacy interface. Match your installed package to the Ragas migration documentation before changing the tutorial code. Record the package versions from pip show beside later results, because a metric implementation change can alter scores even when the application has not changed.
Step 2: Create a trace dataset with references
Save the following as make_cases.py. Every object is one observed or deliberately constructed RAG execution. The second case places an irrelevant shipping chunk before a useful returns chunk, making rank quality visible. The third case contains an unsupported claim about an automatic refund; it is useful for observing faithfulness without relying on a fabricated numeric score.
import json
from pathlib import Path
cases = [
{
"id": "return-window",
"user_input": "How long do I have to return an unused item?",
"reference": "Unused items can be returned within 30 days of delivery.",
"retrieved_contexts": [
"Returns policy: unused items may be returned within 30 days of delivery.",
"Shipping policy: standard delivery takes three to five business days."
],
"response": "You can return an unused item within 30 days of delivery."
},
{
"id": "start-return",
"user_input": "Where do I start a return for an eligible order?",
"reference": "Open the order details page and select Start a return.",
"retrieved_contexts": [
"Shipping policy: tracking appears after the parcel is dispatched.",
"Returns policy: open the order details page and select Start a return."
],
"response": "Open the order details page and select Start a return."
},
{
"id": "refund-timing",
"user_input": "When is a refund issued after a return?",
"reference": "The refund is issued after the returned item passes inspection.",
"retrieved_contexts": [
"Refund policy: a refund is issued after the returned item passes inspection.",
"Returns policy: keep the parcel receipt until processing is complete."
],
"response": "Your refund is issued immediately when you post the parcel."
}
]
with Path("cases.jsonl").open("w", encoding="utf-8") as output:
for case in cases:
output.write(json.dumps(case, ensure_ascii=False) + "\n")
print(f"Wrote {len(cases)} cases to cases.jsonl")
Run and verify this step without an API call:
python make_cases.py
python -c "import json; rows=[json.loads(line) for line in open('cases.jsonl', encoding='utf-8')]; assert len(rows)==3; assert len({r['id'] for r in rows})==3; print([r['id'] for r in rows])"
Expect the three IDs in the order shown. Keep order in retrieved_contexts: context precision depends on rank. Do not concatenate chunks into one string or sort them alphabetically. In a real capture, also retain document IDs, index version, query rewrite, and retrieval scores in separate trace metadata. The scorer needs the text list, while those extra fields explain why a particular chunk appeared.
Step 3: Validate the measurement boundary
Create evaluate_rag.py with the complete script below. Its validation path runs without a provider key, so you can catch malformed records before paying for evaluation. The scorer is deliberately sequential: three cases are enough to learn the output shape, and sequential calls avoid accidental bursts against a new provider quota. It records every dimension separately instead of collapsing the case to a single average.
import argparse
import asyncio
import csv
import json
import math
import os
from pathlib import Path
from openai import AsyncOpenAI
from ragas.embeddings.base import embedding_factory
from ragas.llms import llm_factory
from ragas.metrics.collections import (
AnswerRelevancy,
ContextPrecision,
ContextRecall,
Faithfulness,
)
REQUIRED = ("id", "user_input", "reference", "retrieved_contexts", "response")
METRICS = ("context_precision", "context_recall", "faithfulness", "answer_relevancy")
def load_cases(path: Path) -> list[dict]:
cases = []
seen = set()
with path.open(encoding="utf-8") as source:
for number, line in enumerate(source, start=1):
if not line.strip():
continue
case = json.loads(line)
if not isinstance(case, dict):
raise ValueError(f"line {number}: expected a JSON object")
missing = [field for field in REQUIRED if field not in case]
if missing:
raise ValueError(f"line {number}: missing {missing}")
if any(not isinstance(case[field], str) or not case[field].strip()
for field in ("id", "user_input", "reference", "response")):
raise ValueError(f"line {number}: text fields must be nonempty strings")
contexts = case["retrieved_contexts"]
if not isinstance(contexts, list) or not contexts or any(
not isinstance(chunk, str) or not chunk.strip() for chunk in contexts
):
raise ValueError(f"line {number}: retrieved_contexts needs nonempty strings")
if case["id"] in seen:
raise ValueError(f"line {number}: duplicate ID {case['id']}")
seen.add(case["id"])
cases.append(case)
if not cases:
raise ValueError("the dataset is empty")
return cases
async def score_cases(cases: list[dict]) -> list[dict]:
if not os.getenv("OPENAI_API_KEY"):
raise RuntimeError("Set OPENAI_API_KEY before scoring")
client = AsyncOpenAI()
llm = llm_factory("gpt-4o-mini", client=client)
embeddings = embedding_factory(
"openai", model="text-embedding-3-small", client=client
)
precision = ContextPrecision(llm=llm)
recall = ContextRecall(llm=llm)
faithfulness = Faithfulness(llm=llm)
relevancy = AnswerRelevancy(llm=llm, embeddings=embeddings)
rows = []
for case in cases:
common = {
"user_input": case["user_input"],
"retrieved_contexts": case["retrieved_contexts"],
}
results = {
"context_precision": await precision.ascore(
**common, reference=case["reference"]
),
"context_recall": await recall.ascore(
**common, reference=case["reference"]
),
"faithfulness": await faithfulness.ascore(
**common, response=case["response"]
),
"answer_relevancy": await relevancy.ascore(
user_input=case["user_input"], response=case["response"]
),
}
row = {"id": case["id"]}
for name, result in results.items():
value = float(result.value)
if not math.isfinite(value):
raise ValueError(f"{case['id']}: {name} returned a nonfinite score")
row[name] = value
rows.append(row)
print(case["id"], {name: round(row[name], 3) for name in METRICS})
return rows
def save_csv(rows: list[dict], path: Path) -> None:
with path.open("w", newline="", encoding="utf-8") as output:
writer = csv.DictWriter(output, fieldnames=("id", *METRICS))
writer.writeheader()
writer.writerows(rows)
print(f"Saved {len(rows)} scored cases to {path}")
def main() -> None:
parser = argparse.ArgumentParser()
parser.add_argument("--dataset", type=Path, default=Path("cases.jsonl"))
parser.add_argument("--output", type=Path, default=Path("scores.csv"))
parser.add_argument("--limit", type=int)
parser.add_argument("--validate-only", action="store_true")
args = parser.parse_args()
cases = load_cases(args.dataset)
print(f"Validated {len(cases)} cases")
if args.validate_only:
return
if args.limit is not None:
if args.limit < 1:
parser.error("--limit must be positive")
cases = cases[:args.limit]
rows = asyncio.run(score_cases(cases))
save_csv(rows, args.output)
if __name__ == "__main__":
main()
Verify parsing and field validation now:
python evaluate_rag.py --validate-only
python -m py_compile evaluate_rag.py
Expect Validated 3 cases and no compile error. This check tests the trace contract, not answer quality. A malformed retrieved_contexts value is a data failure; letting it reach a model metric can produce misleading output or a harder-to-debug exception. The reference requirement is intentional for this particular four-metric suite. If your pipeline has no gold answers yet, collect and review them with domain experts before adding reference-based retrieval scores. The golden dataset guide helps structure that review.
Step 4: Ragas Tutorial Evaluate RAG Pipeline Metrics
Run the first trace only. The script uses Ragas's documented collections scorers and calls ascore with named arguments matching each metric. The faithfulness documentation defines grounding against retrieved contexts. The context precision documentation evaluates whether useful chunks are ranked early. The context recall documentation checks whether the retrieved text supports information in the reference. Answer relevancy compares the question with what the response appears to answer; it also needs embeddings.
python evaluate_rag.py --limit 1 --output smoke.csv
python -c "import csv; rows=list(csv.DictReader(open('smoke.csv', newline='', encoding='utf-8'))); assert len(rows)==1 and rows[0]['id']=='return-window'; print(rows[0])"
Expect one printed score dictionary and one CSV row. The exact numbers can vary with the judge, embeddings, metric version, and provider behavior. A successful smoke run proves that authentication, API compatibility, dataset shape, metric calls, and CSV writing work together. It does not prove the model judged the case correctly. Open the row and compare it with the trace: the returns chunk supports the answer and appears before an unrelated shipping chunk. If a value is surprising, inspect those inputs before changing a threshold.
Here is the field map that prevents a common evaluation mistake:
| Dimension | Inputs from the saved trace | What a low value suggests | What it cannot establish alone |
|---|---|---|---|
| Context precision | Question, reference, ranked contexts | Relevant chunks appear too late or irrelevant chunks dominate | Whether the final response is truthful |
| Context recall | Question, reference, retrieved contexts | Evidence for the reference was not retrieved | Whether the response used the evidence |
| Faithfulness | Question, response, retrieved contexts | Response claims lack support in supplied chunks | Whether those chunks contain all required facts |
| Answer relevancy | Question, response, judge embeddings | Response misses the user's intent | Whether an on-topic claim is factually true |
These metrics are complementary. A grounded answer to the wrong question can score well on faithfulness and poorly on relevancy. A retriever can find every needed fact while ranking a distracting document first. A fluent, relevant response can still claim something unsupported. Use each score to choose a component investigation, not as four votes on one undefined concept of quality.
Step 5: Score the full set and inspect each failure
Run all saved executions and check the CSV schema. The third response contradicts its only refund source, so inspect its faithfulness result first. Do not promise that a model judge will return a particular decimal or classify every sample perfectly. Treat the intentionally bad example as a reason to compare judgment with a human label.
python evaluate_rag.py --output scores.csv
python -c "import csv; rows=list(csv.DictReader(open('scores.csv', newline='', encoding='utf-8'))); assert len(rows)==3; assert rows[2]['id']=='refund-timing'; print('Rows:', len(rows), 'Columns:', list(rows[0]))"
The score file is useful only with its inputs. Store the original cases.jsonl, the application revision that produced responses, prompt identifier, retrieval index snapshot, model identifiers, installed package versions, and run date beside it in your normal experiment storage. This tutorial keeps the report small, but those identities are mandatory for a real comparison. If a team changes chunking and the generator prompt in the same run, a lower faithfulness score cannot tell them which edit caused it.
Read case start-return in light of rank. The shipping chunk is first and the return instruction second. Context precision should call attention to poor ordering, while recall can remain healthy because the needed fact is still present. Read refund-timing differently: the source says inspection precedes refund, but the response says posting the parcel triggers an immediate refund. That is a generation defect or a failure in grounding controls, provided the saved context truly matches what the generator saw. For a deeper retrieval diagnosis, use the Ragas context precision walkthrough and the retrieval precision guide.
When a result looks wrong, make a small manual annotation: case ID, metric, human judgment, supporting chunk, and judge reason if exposed by the result object. Examine whether the reference is complete, whether the chunks were truncated at capture time, and whether the user question needs conversation history. A single-turn question stripped of its preceding turn is a bad sample for a single-turn metric, even if the application itself behaved correctly.
Step 6: Read the Ragas Tutorial Evaluate RAG Pipeline Scorecard
Aggregate carefully. A mean over three heterogeneous cases can improve while a critical refund error gets worse. The command below prints the lowest case for each metric, which is more useful for a first triage pass than a lone overall score. It operates on the CSV produced in Step 5 and makes no additional judge calls.
python - <<'PYCODE'
import csv
rows = list(csv.DictReader(open("scores.csv", newline="", encoding="utf-8")))
metrics = ("context_precision", "context_recall", "faithfulness", "answer_relevancy")
for metric in metrics:
worst = min(rows, key=lambda row: float(row[metric]))
mean = sum(float(row[metric]) for row in rows) / len(rows)
print(f"{metric}: mean={mean:.3f}, lowest={worst['id']} ({float(worst[metric]):.3f})")
PYCODE
Verify that four lines print and every lowest ID exists in cases.jsonl. If refund-timing is not lowest for faithfulness, the correct response is to inspect the judge's reasoning and the exact captured texts, not to edit the fixture until the metric behaves as expected. Model-based metrics can make mistakes. The article on measuring answer faithfulness shows why claim-level review matters.
Segment reports by failure risk and query type once you have more than a few records. Track product questions, policy questions, multi-hop questions, ambiguous questions, empty retrieval, and out-of-scope requests separately. Averages can conceal a weak slice, and a balanced dataset matters more than a large collection of near-duplicate easy questions. Keep an untouched holdout for release decisions if prompts are repeatedly tuned against the visible evaluation set.
For comparison runs, score the same case IDs under baseline and candidate configurations. Compare paired changes per metric and manually review large negative moves. Do not compare a candidate's aggregate with a baseline built from a different dataset or a different judge model. If the judge changes, rescore both candidates with the new judge before attributing movement to your RAG pipeline.
Step 7: Add a calibrated release gate
The next script demonstrates mechanics for a case-level gate. Its example thresholds are deliberately illustrative. Replace them with values chosen from human-reviewed passes and failures for your own product. A quality gate should show the exact case and metric that failed. It should also distinguish a provider error from a low quality score; evaluate_rag.py exits before writing a complete CSV if scoring raises an exception.
Save this as gate.py:
import csv
import sys
from pathlib import Path
limits = {
"context_precision": 0.50,
"context_recall": 0.50,
"faithfulness": 0.50,
"answer_relevancy": 0.50,
}
path = Path(sys.argv[1] if len(sys.argv) > 1 else "scores.csv")
with path.open(newline="", encoding="utf-8") as source:
rows = list(csv.DictReader(source))
if not rows:
raise SystemExit("No scored cases found")
failures = []
for row in rows:
for metric, minimum in limits.items():
score = float(row[metric])
if score < minimum:
failures.append(f"{row['id']}: {metric}={score:.3f} < {minimum:.2f}")
for failure in failures:
print(failure)
print(f"Checked {len(rows)} cases; {len(failures)} metric failures")
raise SystemExit(1 if failures else 0)
Verify the gate as an executable script:
python -m py_compile gate.py
python gate.py smoke.csv
The command prints a count and may return exit code 1, depending on the actual judge scores. That is valid behavior for a calibrated gate demonstration. Run python gate.py scores.csv after the full evaluation to list failures across all three cases. For CI, require an expected case count and store the CSV only after scoring finishes successfully. An empty, partial, or stale report must not count as a pass.
Calibrate with domain reviewers before enforcing the threshold. Label a representative sample as acceptable or unacceptable, compare labels with each metric, inspect false passes and false failures, and choose a cutoff according to the cost of each error. For a refund policy, a false pass on an unsupported promise may be more serious than a false fail on mildly awkward wording. Keep deterministic checks for exact policy fields, citations, schema, and permissions alongside Ragas. The LLM judge calibration guide and evals in CI guide give more detail on review and release workflow.
Troubleshooting
Problem: ModuleNotFoundError for a collections metric -> Activate .venv, run python -m pip show ragas, and compare the installed package with the official collections documentation. A notebook kernel or system Python may be importing a different installation. Do not mix legacy single_turn_ascore examples into this script.
Problem: authentication or model access error -> Confirm OPENAI_API_KEY is set in the process running the script and that the account permits the chosen judge and embedding models. Replace both model identifiers with approved alternatives supported by your installed Ragas version if your provider configuration differs. Treat this as an execution failure, not a quality zero.
Problem: missing reference or empty contexts -> Repair the captured trace or define a separate metric set that does not need a reference. Context recall cannot establish missing evidence without an answer or source to compare against. Empty retrieval is a meaningful application event, but record and evaluate it through an explicit empty-retrieval scenario rather than slipping it into this fixture's nonempty-context contract.
Problem: scores change between identical runs -> Save the judge model, embeddings model, Ragas version, trace, and prompt configuration, then repeat a small fixed subset. Inspect cases near the threshold and measure label flips. A score change without an input change can come from nondeterministic judgment or provider behavior.
Problem: apparently high faithfulness for a bad answer -> Check whether the misleading statement is actually supported by an outdated or incorrect retrieved chunk. Faithfulness tests grounding relative to supplied evidence, not truth in the outside world. Review the document source and its version, then compare with a trusted reference.
Problem: context precision is low while the answer is correct -> Inspect ranking and chunk redundancy. The answer may succeed because the useful chunk arrived late, while irrelevant material still consumed the context budget. Improve retrieval order and test the same questions again rather than treating the correct response as proof that retrieval needs no work.
Interview Questions and Answers
Q: Which four dimensions does this tutorial score?
Context precision examines rank usefulness, context recall examines coverage of reference facts, faithfulness examines support for response claims, and answer relevancy examines whether the response addresses the question. Each points to a different inspection path.
Q: Why is a reference answer included?
Context precision and context recall in this example use a reviewed reference. It supplies a target for judging retrieval utility and coverage; copying the application's own response into that field would hide mistakes.
Q: Can a faithful RAG answer still be wrong?
Yes. A response can faithfully repeat an outdated retrieved policy. Faithfulness establishes support in the delivered contexts, so source freshness and correctness require separate controls.
Q: What should be recorded with a score?
Keep the case ID, question, ranked chunks, response, reference, metric version, judge and embedding models, application revision, and index snapshot. Without those, a regression is difficult to reproduce or attribute.
Q: Why not use one average as the release gate?
A high score on common easy questions can mask a critical failure on a policy question. Gate important cases or slices explicitly and review per-case regressions before looking at an overall trend.
Q: How do you decide whether a low score is a retrieval or generation defect?
Compare context precision and recall with faithfulness and relevancy, then inspect the actual chunks and claims. Missing reference facts suggest retrieval; unsupported claims despite good evidence suggest generation. More than one component can fail on the same trace.
Common Mistakes
- Scoring a rewritten or manually improved response instead of the output the application actually showed.
- Replacing ranked retrieved chunks with an unordered document collection before calculating precision.
- Treating
referenceandresponseas interchangeable fields. - Declaring a release blocked from an illustrative cutoff without human calibration.
- Letting a provider timeout create an empty CSV that CI interprets as a pass.
- Comparing scores across different judges, embedding models, or metric implementations without a paired rescore.
- Ignoring privacy when production documents are included in evaluation traces.
Conclusion
A useful Ragas evaluation starts with a trustworthy trace and an explicit question for each metric. This tutorial validated saved RAG executions, scored retrieval and generation separately, and preserved case-level evidence so a low value can lead to a concrete investigation. Run it first on a reviewed sample from your own pipeline, compare scores with human judgments, and make only calibrated thresholds part of a release decision.
Where To Go Next
Replace the three example records with captured executions from your RAG pipeline. Start with a small reviewed set covering normal, ambiguous, unsupported, and policy-sensitive questions. Keep the question, ranked contexts, and user-visible response from the same request; attach a human-reviewed reference. Run the smoke command, score the full set, examine per-case changes, and only then calibrate a gate.
Expand retrieval diagnostics with RAG retrieval recall at k, test source attribution with RAG citation correctness examples, and broaden the data with an adversarial RAG evaluation dataset. The RAG application evaluation guide connects those checks into a larger test plan. The immediate next step is concrete: export ten representative traces from your own application, review their references, and run the same four-metric scorecard against them.
Interview Questions and Answers
How would you construct a Ragas evaluation dataset for a production RAG app?
I would sample real user tasks across risk and query types, redact sensitive content, and retain the question, ranked chunks, response, and source metadata from the same execution. Domain reviewers would write or approve references. I would reserve a holdout so prompt tuning does not optimize only visible cases.
What does context precision measure in Ragas?
It evaluates whether chunks judged useful for answering a question appear high in the ranked retrieved list. I would use it to diagnose noisy top results and compare retriever changes on paired queries. It does not establish that the generated answer is correct.
Why can context recall be high when the answer is bad?
The retriever may have supplied all evidence needed for the reference answer while the generator ignored or contradicted it. I would inspect the retrieved chunks and then compare faithfulness and the actual response. Recall is a retrieval signal, not a generation outcome.
How is faithfulness different from factual correctness?
Faithfulness asks whether response claims are supported by the retrieved context. Factual correctness compares the response with a trusted reference or ground truth. A stale policy document can support an answer that is faithful to retrieval but wrong for the current policy.
What would you do when answer relevancy passes but faithfulness fails?
The response may directly address the user's question while inventing unsupported details. I would enumerate its claims, locate support in the retrieved chunks, and check whether the missing support is a retrieval gap or a generation grounding failure. The two metrics should not be averaged into a pass.
How would you set a CI threshold for a Ragas score?
I would label representative acceptable and unacceptable cases, run a fixed evaluator configuration, and study false passes and false failures by risk slice. The cutoff would reflect the cost of those errors. I would also require complete case counts and treat provider errors as indeterminate runs.
Why must RAG evaluation keep model and index versions?
Judge and embedding model changes can alter the measurement instrument, while index changes alter retrieved evidence. Without those identifiers, a score movement cannot be attributed to the application change. I would rescore baseline and candidate under the same evaluator when comparing them.
Frequently Asked Questions
How do I evaluate a RAG pipeline with Ragas?
Capture the question, ranked retrieved contexts, generated response, and a reviewed reference for each case. Run Ragas metrics on those fields and inspect per-case results before setting any release threshold.
Does Ragas require reference answers?
Not every metric does. The context precision and context recall metrics used in this tutorial use a reference, while faithfulness and answer relevancy can score without one. Choose metrics that match evidence you can trust.
What is the difference between context precision and context recall?
Context precision tests whether useful retrieved chunks are ranked ahead of less useful ones. Context recall tests whether the retrieved material covers information in the reference answer. A case can have strong recall and weak ranking.
What does a low faithfulness score mean?
It suggests claims in the response lack support in the retrieved contexts supplied to the metric. Inspect the exact chunks and claims before blaming the generator, because a missing or truncated capture can distort the result.
Can I run this Ragas example without an API key?
The dataset creation, validation, and CSV inspection steps do not need a provider key. The four model-based scoring calls need access to the configured judge and embedding services.
Are the threshold values in the example production ready?
No. They demonstrate gate mechanics only. Label representative cases with domain reviewers, measure false passes and false failures, and choose cutoffs for your product's risk.
Why should I preserve retrieved chunk order?
Context precision evaluates ranking, so moving a useful chunk from first to last changes the question being tested. Save the list exactly as the generator received it.