feat(skills): install deepeval skill from confident-ai/deepeval
Adds the deepeval eval-loop skill via `npx skills add` with skills-lock.json for reproducible reinstalls. Symlinked to Claude Code via .claude/skills/. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016z2ZFYHQCex8yZAMVMTZzZ
This commit is contained in:
175
.agents/skills/deepeval/references/metrics.md
Normal file
175
.agents/skills/deepeval/references/metrics.md
Normal file
@@ -0,0 +1,175 @@
|
||||
# Metrics
|
||||
|
||||
Use 3-5 metrics for the first eval suite when the user is unsure. More metrics
|
||||
make iteration slower and harder to interpret. Reuse existing project metrics
|
||||
and thresholds before adding new ones.
|
||||
|
||||
Keep metric instances in a separate `metrics.py` module (or the project's
|
||||
existing metrics module). Eval test files should import metric lists rather than
|
||||
constructing several ad hoc metrics inline.
|
||||
|
||||
Name component/span metric lists after the exact component they evaluate. Avoid
|
||||
generic names like `COMPONENT_METRICS` because one suite can evaluate several
|
||||
components with different metric requirements.
|
||||
|
||||
## Required Rule
|
||||
|
||||
Single-turn `LLMTestCase` evals must use single-turn metrics.
|
||||
|
||||
Multi-turn `ConversationalTestCase` evals must use multi-turn conversational
|
||||
metrics. Do not use `AnswerRelevancyMetric`, `FaithfulnessMetric`, or other
|
||||
single-turn `LLMTestCase` metrics on multi-turn end-to-end evals.
|
||||
|
||||
## Metric Types
|
||||
|
||||
Choose metrics by what the user wants to measure, not only by app type.
|
||||
|
||||
| Type | Use when | Examples |
|
||||
| --- | --- | --- |
|
||||
| Custom criteria | The success criteria is product- or domain-specific | `GEval`, `DAGMetric`, `ConversationalGEval`, `ConversationalDAGMetric` |
|
||||
| RAG retriever | You need to evaluate retrieved context quality | `ContextualRelevancyMetric`, `ContextualPrecisionMetric`, `ContextualRecallMetric` |
|
||||
| RAG generator | You need to evaluate the final answer against context | `AnswerRelevancyMetric`, `FaithfulnessMetric` |
|
||||
| Agentic flow | You need to evaluate task completion, plans, steps, tools, or arguments | `TaskCompletionMetric`, `ToolCorrectnessMetric`, `ArgumentCorrectnessMetric`, `PlanAdherenceMetric`, `PlanQualityMetric`, `StepEfficiencyMetric` |
|
||||
| Multi-turn chatbot | You need to evaluate an entire conversation | `ConversationCompletenessMetric`, `RoleAdherenceMetric`, `TurnRelevancyMetric`, `ConversationalGEval` |
|
||||
| Safety and compliance | You need to detect risky or policy-violating outputs | `BiasMetric`, `ToxicityMetric`, `PIILeakageMetric`, `MisuseMetric`, `RoleViolationMetric`, `NonAdviceMetric` |
|
||||
| Format / structure | You need output to match a schema or instruction set | `JsonCorrectnessMetric`, `PromptAlignmentMetric` |
|
||||
| Other task-specific quality | The app is summarization, hallucination-sensitive, image-based, or otherwise specialized | `SummarizationMetric`, `HallucinationMetric`, multimodal metrics |
|
||||
|
||||
Aim to include at least one custom metric when the user's definition of success
|
||||
is not fully captured by a predefined metric. In practice, custom metrics should
|
||||
usually be `GEval` for single-turn evals or `ConversationalGEval` for multi-turn
|
||||
evals.
|
||||
|
||||
## Default If User Is Unsure
|
||||
|
||||
If the user says "I don't know" or gives no metric preference:
|
||||
|
||||
- Use 3-5 metrics.
|
||||
- Put metrics on the end-to-end eval first.
|
||||
- Do not add safety metrics by default unless the app is safety/compliance
|
||||
sensitive or the user asks for them.
|
||||
- Use about half custom metrics and half system-specific metrics.
|
||||
- Add component/span metrics only after E2E/traces show component failures, or
|
||||
if the user explicitly wants component-level scoring.
|
||||
|
||||
Good system-specific defaults:
|
||||
|
||||
- Single-turn tracing E2E: strongly prefer `TaskCompletionMetric` and
|
||||
`StepEfficiencyMetric` as the baseline pair, especially for agents and
|
||||
multi-step AI apps.
|
||||
- Agent: `TaskCompletionMetric` plus tool/argument correctness only when
|
||||
`tools_called` data exists.
|
||||
- RAG: `FaithfulnessMetric`, `AnswerRelevancyMetric`, and
|
||||
`ContextualRelevancyMetric` are strong candidates.
|
||||
- Multi-turn chatbot: use conversational metrics only, plus a
|
||||
`ConversationalGEval` custom criterion when product-specific behavior matters.
|
||||
|
||||
For custom metrics, assume `GEval` for single-turn or `ConversationalGEval` for
|
||||
multi-turn. There is a very high chance this is the right custom metric type.
|
||||
Do not start with DAG unless the user already has a DAG metric or specifically
|
||||
needs decision-tree scoring.
|
||||
|
||||
Use `GEval` when scoring is subjective or there is no predefined metric for the
|
||||
thing the user cares about. Correctness is a common example: there is no generic
|
||||
"correctness metric" because correctness depends on the task. Define a `GEval`
|
||||
named `Correctness` and write criteria that explain what correct means for this
|
||||
app.
|
||||
|
||||
Use `DAGMetric` only when the metric is decision-based: the score should follow
|
||||
explicit branches, checks, or deterministic rubric paths. DAG is useful when the
|
||||
metric is more like a decision tree than a subjective judge. Do not start with
|
||||
DAG for ordinary subjective scoring.
|
||||
|
||||
When choosing `GEval.evaluation_params`, include only fields the test case will
|
||||
actually have. Be especially careful with reference-space params like
|
||||
`expected_output`, `context`, `retrieval_context`, or `expected_tools`; if the
|
||||
dataset or app does not provide them, the metric will fail at runtime. Prefer
|
||||
`input` and `actual_output` unless the eval plan explicitly creates the
|
||||
reference fields.
|
||||
|
||||
If existing project metrics are present, use them first. If there are too many,
|
||||
tell the user: "You already have a lot of metrics here, which may make evals
|
||||
slow or hard to interpret. I recommend narrowing the first run to the highest
|
||||
signal metrics."
|
||||
|
||||
## Reference-Based Metrics
|
||||
|
||||
Some metrics require reference fields. Use them sparingly unless the plan
|
||||
includes those expected values, because missing fields will cause metric errors.
|
||||
|
||||
Reference-based fields include:
|
||||
|
||||
- `expected_output`
|
||||
- `expected_outcome`
|
||||
- `expected_tools`
|
||||
- `context`
|
||||
- `retrieval_context`
|
||||
|
||||
Examples:
|
||||
|
||||
- `ContextualPrecisionMetric` and `ContextualRecallMetric` need
|
||||
`expected_output`.
|
||||
- `ToolCorrectnessMetric` needs `expected_tools`.
|
||||
- Multi-turn outcome metrics may depend on `expected_outcome`.
|
||||
- RAG grounding metrics need `retrieval_context`.
|
||||
|
||||
If the dataset does not include the required fields, choose metrics that match
|
||||
available fields or update the dataset generation/loading plan first.
|
||||
|
||||
## Common Single-Turn Metrics
|
||||
|
||||
| Metric | What it checks | Required test case fields |
|
||||
| --- | --- | --- |
|
||||
| `AnswerRelevancyMetric` | Output answers the input | `input`, `actual_output` |
|
||||
| `FaithfulnessMetric` | Output is grounded in retrieved context | `input`, `actual_output`, `retrieval_context` |
|
||||
| `ContextualRelevancyMetric` | Retrieved context is relevant to input | `input`, `retrieval_context` |
|
||||
| `ContextualPrecisionMetric` | Relevant context is ranked highly | `input`, `retrieval_context`, `expected_output` |
|
||||
| `ContextualRecallMetric` | Retrieved context covers expected answer | `input`, `retrieval_context`, `expected_output` |
|
||||
| `TaskCompletionMetric` | Agent/app completed the task | `input`, `actual_output` |
|
||||
| `StepEfficiencyMetric` | Agent/app completed the task efficiently without unnecessary steps | trace steps/tool activity |
|
||||
| `ToolCorrectnessMetric` | Called tools match expected tools | `input`, `tools_called`, `expected_tools` |
|
||||
| `ArgumentCorrectnessMetric` | Tool arguments are correct | `input`, `tools_called` |
|
||||
| `JsonCorrectnessMetric` | Output matches expected schema | `input`, `actual_output`; constructor needs `expected_schema` |
|
||||
| `PromptAlignmentMetric` | Output follows prompt instructions | `input`, `actual_output`; constructor needs `prompt_instructions` |
|
||||
| `GEval` | Custom single-turn criteria | constructor needs `name`, `criteria` or `evaluation_steps`, and `evaluation_params` |
|
||||
|
||||
## Common Multi-Turn Metrics
|
||||
|
||||
| Metric | What it checks | Required test case fields |
|
||||
| --- | --- | --- |
|
||||
| `ConversationCompletenessMetric` | Conversation achieved the expected outcome | `turns` with `role`, `content` |
|
||||
| `RoleAdherenceMetric` | Assistant stayed in role across turns | `turns` with `role`, `content` |
|
||||
| `TurnRelevancyMetric` | Assistant turns are relevant | `turns` with `role`, `content` |
|
||||
| `TurnFaithfulnessMetric` | Turns are faithful to retrieval context | `turns` with `role`, `content`, `retrieval_context` |
|
||||
| `TurnContextualRelevancyMetric` | Turn retrieval context is relevant | `turns` with `role`, `content`, retrieval context |
|
||||
| `GoalAccuracyMetric` | Conversation achieved the user's goal | `turns` with `role`, `content` |
|
||||
| `TopicAdherenceMetric` | Conversation stayed on allowed topics | `turns` with `role`, `content`; constructor needs `relevant_topics` |
|
||||
| `ConversationalGEval` | Custom multi-turn criteria | constructor needs `name` and `criteria` or `evaluation_steps` |
|
||||
|
||||
## Choosing Metrics
|
||||
|
||||
Ask what the user cares about in product terms first. Then map that to metrics.
|
||||
|
||||
Ask:
|
||||
|
||||
- What failure would be unacceptable in production?
|
||||
- Is success about final answer quality, retrieved context, tool use, safety,
|
||||
conversation completion, or output format?
|
||||
- Do we need a custom criterion because the product definition of "good" is
|
||||
domain-specific?
|
||||
- Which fields does the dataset/test case actually contain?
|
||||
|
||||
Mappings:
|
||||
|
||||
- "Does it answer correctly?" -> `AnswerRelevancyMetric` or task-specific `GEval`
|
||||
- "Is it grounded in docs?" -> `FaithfulnessMetric` plus contextual metrics
|
||||
- "Did the agent finish the task?" -> `TaskCompletionMetric`
|
||||
- "Did the agent take efficient steps?" -> `StepEfficiencyMetric`
|
||||
- "Did it use the right tool?" -> `ToolCorrectnessMetric`
|
||||
- "Did the chatbot complete the conversation?" -> `ConversationCompletenessMetric`
|
||||
- "Did it stay in character?" -> `RoleAdherenceMetric`
|
||||
|
||||
If unsure for single-turn tracing, start with `TaskCompletionMetric` and
|
||||
`StepEfficiencyMetric`, then add 1-3 more E2E metrics only when the app's
|
||||
success criteria need them. Add component/span metrics only after the first run
|
||||
reveals where the app is failing.
|
||||
Reference in New Issue
Block a user