Files
holocron/.agents/skills/deepeval/references/artifact-contracts.md
Defame1297 7afa1f3f02 feat(skills): install deepeval skill from confident-ai/deepeval
Adds the deepeval eval-loop skill via `npx skills add` with skills-lock.json
for reproducible reinstalls. Symlinked to Claude Code via .claude/skills/.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016z2ZFYHQCex8yZAMVMTZzZ
2026-06-24 17:06:58 +00:00

2.2 KiB

Artifact Contracts

Create eval artifacts that users can inspect, edit, commit, and rerun without an agent.

Preferred Layout

tests/
  evals/
    test_<app>.py
    metrics.py
    .dataset.json

Use an existing eval directory if the project already has one.

First look for an existing test folder. If one exists, put the eval suite there. If none exists, create tests/evals/.

Prefer one eval test file for the first setup. Component/span metrics belong in the same single-turn tracing file. Add more files only for a clearly distinct use case.

Dataset Files

Preferred generated dataset path:

tests/evals/.dataset.json

Use .dataset.json, not goldens.json. The mental model is: a dataset contains goldens.

Supported input formats:

  • .json
  • .jsonl
  • .csv

The dataset should contain the fields needed by the chosen template and metrics. For RAG, include context or enough information to reconstruct context from the app. For multi-turn evals, use conversational goldens.

Pytest Files

Eval tests should:

  • load the dataset from tests/evals/.dataset.json by default
  • call the real app entry point
  • prefer native DeepEval integrations and traced Golden assertions
  • build LLMTestCases only in explicit no-tracing evals
  • import a small, explicit metric list from metrics.py
  • add span-level metrics only for useful component diagnostics
  • use existing metrics and thresholds when found
  • avoid network calls unrelated to the app or evaluation model
  • be run with deepeval test run, not the raw pytest command

Placeholder Contract

Templates intentionally contain placeholders:

  • dataset file paths in add_goldens_from_*_file(...)
  • AI app module/function names such as import_module("ai_app").run_traced_ai_app
  • metric lists in metrics.py
  • integration callback/instrumentation setup when applicable

Replace every placeholder before running evals. If a placeholder remains, stop and adapt the template instead of running a broken suite.

Result Artifacts

Do not create hidden result caches unless DeepEval already does so. The durable artifacts are the test files, dataset files, tracing integration, and optional Confident AI hosted reports.