feat(skills): install deepeval skill from confident-ai/deepeval

Adds the deepeval eval-loop skill via `npx skills add` with skills-lock.json
for reproducible reinstalls. Symlinked to Claude Code via .claude/skills/.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016z2ZFYHQCex8yZAMVMTZzZ
This commit is contained in:
2026-06-24 17:06:58 +00:00
parent 9369d97934
commit 7afa1f3f02
18 changed files with 1624 additions and 0 deletions

View File

@@ -0,0 +1,84 @@
# Datasets
Use documented `EvaluationDataset` APIs directly. Do not invent wrapper helpers
for dataset loading in templates.
Dataset source order is strict:
1. Ask whether the user already has a dataset.
2. If they do, load it with the documented `EvaluationDataset` API.
3. If they do not, read `synthetic-data.md` and generate one with
`deepeval generate`.
Do not hand-create or make up goldens. For a useful first generated eval
dataset, target about 30-50 goldens. If the user insists on manual goldens,
warn that generated goldens are usually less biased and more reproducible, then
recommend augmenting any manual seed set with `deepeval generate --method
goldens`.
If the user has a dataset, check its size before accepting it as sufficient.
Fewer than 10 goldens is very likely too small. If the dataset is small or the
user is unhappy with it, read `synthetic-data.md` and consider augmenting from
existing goldens with `deepeval generate`.
## Local JSON
```python
from deepeval.dataset import EvaluationDataset
dataset = EvaluationDataset()
dataset.add_goldens_from_json_file(file_path="tests/evals/.dataset.json")
```
## Local JSONL
```python
dataset = EvaluationDataset()
dataset.add_goldens_from_jsonl_file(file_path="tests/evals/.dataset.jsonl")
```
## Local CSV
```python
dataset = EvaluationDataset()
dataset.add_goldens_from_csv_file(file_path="tests/evals/.dataset.csv")
```
If the CSV uses custom column names, set the documented column arguments when
adapting the template.
## Confident AI
```python
dataset = EvaluationDataset()
dataset.pull(alias="My Evals Dataset")
```
Use this when the user says the dataset is on Confident AI and credentials or
MCP/API access are available.
## Pytest Convention
Load the dataset directly in the test file immediately before parametrization.
Do not hide dataset loading in `conftest.py` or custom fixture wrappers:
```python
dataset = EvaluationDataset()
dataset.add_goldens_from_json_file(file_path="tests/evals/.dataset.json")
@pytest.mark.parametrize("golden", dataset.goldens)
def test_llm_app(golden):
...
```
For end-to-end test cases that are built before assertion, add them back to the
dataset with `dataset.add_test_case(...)`, then parametrize over
`dataset.test_cases` if that better matches the app.
Datasets are either single-turn or multi-turn once loaded. Do not mix `Golden`
and `ConversationalGolden` items in one dataset.
For chatbot / multi-turn agent evals, the loaded dataset contains
`ConversationalGolden`s. After loading, pass `dataset.goldens` to
`ConversationSimulator.simulate(...)` to create `ConversationalTestCase`s for
pytest.