Adds the deepeval eval-loop skill via `npx skills add` with skills-lock.json for reproducible reinstalls. Symlinked to Claude Code via .claude/skills/. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016z2ZFYHQCex8yZAMVMTZzZ
9.3 KiB
Metrics
Use 3-5 metrics for the first eval suite when the user is unsure. More metrics make iteration slower and harder to interpret. Reuse existing project metrics and thresholds before adding new ones.
Keep metric instances in a separate metrics.py module (or the project's
existing metrics module). Eval test files should import metric lists rather than
constructing several ad hoc metrics inline.
Name component/span metric lists after the exact component they evaluate. Avoid
generic names like COMPONENT_METRICS because one suite can evaluate several
components with different metric requirements.
Required Rule
Single-turn LLMTestCase evals must use single-turn metrics.
Multi-turn ConversationalTestCase evals must use multi-turn conversational
metrics. Do not use AnswerRelevancyMetric, FaithfulnessMetric, or other
single-turn LLMTestCase metrics on multi-turn end-to-end evals.
Metric Types
Choose metrics by what the user wants to measure, not only by app type.
| Type | Use when | Examples |
|---|---|---|
| Custom criteria | The success criteria is product- or domain-specific | GEval, DAGMetric, ConversationalGEval, ConversationalDAGMetric |
| RAG retriever | You need to evaluate retrieved context quality | ContextualRelevancyMetric, ContextualPrecisionMetric, ContextualRecallMetric |
| RAG generator | You need to evaluate the final answer against context | AnswerRelevancyMetric, FaithfulnessMetric |
| Agentic flow | You need to evaluate task completion, plans, steps, tools, or arguments | TaskCompletionMetric, ToolCorrectnessMetric, ArgumentCorrectnessMetric, PlanAdherenceMetric, PlanQualityMetric, StepEfficiencyMetric |
| Multi-turn chatbot | You need to evaluate an entire conversation | ConversationCompletenessMetric, RoleAdherenceMetric, TurnRelevancyMetric, ConversationalGEval |
| Safety and compliance | You need to detect risky or policy-violating outputs | BiasMetric, ToxicityMetric, PIILeakageMetric, MisuseMetric, RoleViolationMetric, NonAdviceMetric |
| Format / structure | You need output to match a schema or instruction set | JsonCorrectnessMetric, PromptAlignmentMetric |
| Other task-specific quality | The app is summarization, hallucination-sensitive, image-based, or otherwise specialized | SummarizationMetric, HallucinationMetric, multimodal metrics |
Aim to include at least one custom metric when the user's definition of success
is not fully captured by a predefined metric. In practice, custom metrics should
usually be GEval for single-turn evals or ConversationalGEval for multi-turn
evals.
Default If User Is Unsure
If the user says "I don't know" or gives no metric preference:
- Use 3-5 metrics.
- Put metrics on the end-to-end eval first.
- Do not add safety metrics by default unless the app is safety/compliance sensitive or the user asks for them.
- Use about half custom metrics and half system-specific metrics.
- Add component/span metrics only after E2E/traces show component failures, or if the user explicitly wants component-level scoring.
Good system-specific defaults:
- Single-turn tracing E2E: strongly prefer
TaskCompletionMetricandStepEfficiencyMetricas the baseline pair, especially for agents and multi-step AI apps. - Agent:
TaskCompletionMetricplus tool/argument correctness only whentools_calleddata exists. - RAG:
FaithfulnessMetric,AnswerRelevancyMetric, andContextualRelevancyMetricare strong candidates. - Multi-turn chatbot: use conversational metrics only, plus a
ConversationalGEvalcustom criterion when product-specific behavior matters.
For custom metrics, assume GEval for single-turn or ConversationalGEval for
multi-turn. There is a very high chance this is the right custom metric type.
Do not start with DAG unless the user already has a DAG metric or specifically
needs decision-tree scoring.
Use GEval when scoring is subjective or there is no predefined metric for the
thing the user cares about. Correctness is a common example: there is no generic
"correctness metric" because correctness depends on the task. Define a GEval
named Correctness and write criteria that explain what correct means for this
app.
Use DAGMetric only when the metric is decision-based: the score should follow
explicit branches, checks, or deterministic rubric paths. DAG is useful when the
metric is more like a decision tree than a subjective judge. Do not start with
DAG for ordinary subjective scoring.
When choosing GEval.evaluation_params, include only fields the test case will
actually have. Be especially careful with reference-space params like
expected_output, context, retrieval_context, or expected_tools; if the
dataset or app does not provide them, the metric will fail at runtime. Prefer
input and actual_output unless the eval plan explicitly creates the
reference fields.
If existing project metrics are present, use them first. If there are too many, tell the user: "You already have a lot of metrics here, which may make evals slow or hard to interpret. I recommend narrowing the first run to the highest signal metrics."
Reference-Based Metrics
Some metrics require reference fields. Use them sparingly unless the plan includes those expected values, because missing fields will cause metric errors.
Reference-based fields include:
expected_outputexpected_outcomeexpected_toolscontextretrieval_context
Examples:
ContextualPrecisionMetricandContextualRecallMetricneedexpected_output.ToolCorrectnessMetricneedsexpected_tools.- Multi-turn outcome metrics may depend on
expected_outcome. - RAG grounding metrics need
retrieval_context.
If the dataset does not include the required fields, choose metrics that match available fields or update the dataset generation/loading plan first.
Common Single-Turn Metrics
| Metric | What it checks | Required test case fields |
|---|---|---|
AnswerRelevancyMetric |
Output answers the input | input, actual_output |
FaithfulnessMetric |
Output is grounded in retrieved context | input, actual_output, retrieval_context |
ContextualRelevancyMetric |
Retrieved context is relevant to input | input, retrieval_context |
ContextualPrecisionMetric |
Relevant context is ranked highly | input, retrieval_context, expected_output |
ContextualRecallMetric |
Retrieved context covers expected answer | input, retrieval_context, expected_output |
TaskCompletionMetric |
Agent/app completed the task | input, actual_output |
StepEfficiencyMetric |
Agent/app completed the task efficiently without unnecessary steps | trace steps/tool activity |
ToolCorrectnessMetric |
Called tools match expected tools | input, tools_called, expected_tools |
ArgumentCorrectnessMetric |
Tool arguments are correct | input, tools_called |
JsonCorrectnessMetric |
Output matches expected schema | input, actual_output; constructor needs expected_schema |
PromptAlignmentMetric |
Output follows prompt instructions | input, actual_output; constructor needs prompt_instructions |
GEval |
Custom single-turn criteria | constructor needs name, criteria or evaluation_steps, and evaluation_params |
Common Multi-Turn Metrics
| Metric | What it checks | Required test case fields |
|---|---|---|
ConversationCompletenessMetric |
Conversation achieved the expected outcome | turns with role, content |
RoleAdherenceMetric |
Assistant stayed in role across turns | turns with role, content |
TurnRelevancyMetric |
Assistant turns are relevant | turns with role, content |
TurnFaithfulnessMetric |
Turns are faithful to retrieval context | turns with role, content, retrieval_context |
TurnContextualRelevancyMetric |
Turn retrieval context is relevant | turns with role, content, retrieval context |
GoalAccuracyMetric |
Conversation achieved the user's goal | turns with role, content |
TopicAdherenceMetric |
Conversation stayed on allowed topics | turns with role, content; constructor needs relevant_topics |
ConversationalGEval |
Custom multi-turn criteria | constructor needs name and criteria or evaluation_steps |
Choosing Metrics
Ask what the user cares about in product terms first. Then map that to metrics.
Ask:
- What failure would be unacceptable in production?
- Is success about final answer quality, retrieved context, tool use, safety, conversation completion, or output format?
- Do we need a custom criterion because the product definition of "good" is domain-specific?
- Which fields does the dataset/test case actually contain?
Mappings:
- "Does it answer correctly?" ->
AnswerRelevancyMetricor task-specificGEval - "Is it grounded in docs?" ->
FaithfulnessMetricplus contextual metrics - "Did the agent finish the task?" ->
TaskCompletionMetric - "Did the agent take efficient steps?" ->
StepEfficiencyMetric - "Did it use the right tool?" ->
ToolCorrectnessMetric - "Did the chatbot complete the conversation?" ->
ConversationCompletenessMetric - "Did it stay in character?" ->
RoleAdherenceMetric
If unsure for single-turn tracing, start with TaskCompletionMetric and
StepEfficiencyMetric, then add 1-3 more E2E metrics only when the app's
success criteria need them. Add component/span metrics only after the first run
reveals where the app is failing.