Adds the deepeval eval-loop skill via `npx skills add` with skills-lock.json for reproducible reinstalls. Symlinked to Claude Code via .claude/skills/. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016z2ZFYHQCex8yZAMVMTZzZ
2.2 KiB
Artifact Contracts
Create eval artifacts that users can inspect, edit, commit, and rerun without an agent.
Preferred Layout
tests/
evals/
test_<app>.py
metrics.py
.dataset.json
Use an existing eval directory if the project already has one.
First look for an existing test folder. If one exists, put the eval suite there.
If none exists, create tests/evals/.
Prefer one eval test file for the first setup. Component/span metrics belong in the same single-turn tracing file. Add more files only for a clearly distinct use case.
Dataset Files
Preferred generated dataset path:
tests/evals/.dataset.json
Use .dataset.json, not goldens.json. The mental model is: a dataset contains
goldens.
Supported input formats:
.json.jsonl.csv
The dataset should contain the fields needed by the chosen template and metrics. For RAG, include context or enough information to reconstruct context from the app. For multi-turn evals, use conversational goldens.
Pytest Files
Eval tests should:
- load the dataset from
tests/evals/.dataset.jsonby default - call the real app entry point
- prefer native DeepEval integrations and traced
Goldenassertions - build
LLMTestCases only in explicit no-tracing evals - import a small, explicit metric list from
metrics.py - add span-level metrics only for useful component diagnostics
- use existing metrics and thresholds when found
- avoid network calls unrelated to the app or evaluation model
- be run with
deepeval test run, not the rawpytestcommand
Placeholder Contract
Templates intentionally contain placeholders:
- dataset file paths in
add_goldens_from_*_file(...) - AI app module/function names such as
import_module("ai_app").run_traced_ai_app - metric lists in
metrics.py - integration callback/instrumentation setup when applicable
Replace every placeholder before running evals. If a placeholder remains, stop and adapt the template instead of running a broken suite.
Result Artifacts
Do not create hidden result caches unless DeepEval already does so. The durable artifacts are the test files, dataset files, tracing integration, and optional Confident AI hosted reports.