Files
holocron/.agents/skills/deepeval/references/traced-evals.md
Defame1297 7afa1f3f02 feat(skills): install deepeval skill from confident-ai/deepeval
Adds the deepeval eval-loop skill via `npx skills add` with skills-lock.json
for reproducible reinstalls. Symlinked to Claude Code via .claude/skills/.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016z2ZFYHQCex8yZAMVMTZzZ
2026-06-24 17:06:58 +00:00

2.7 KiB

Traced Evals

Tracing is the default single-turn eval path when the app can produce traces through a DeepEval integration or manual instrumentation. The trace is the end-to-end execution and spans are the components; component-level metrics are attached to specific spans inside the same single-turn tracing eval, not split into a separate test shape.

This reference covers the eval-coupled side of tracing: attaching metrics to spans and the pytest/script shapes for traced evals. To instrument the app — add @observe, wire framework integrations, set span types, tags, and metadata — use the deepeval-tracing skill.

Component / Span Metrics

When metrics belong to a specific component, keep them in the single-turn tracing eval and attach them to the exact span they evaluate.

If a supported integration creates the span, stage metrics for the next span of that type:

from deepeval.tracing import next_retriever_span

from metrics import RETRIEVER_SPAN_METRICS


with next_retriever_span(metrics=RETRIEVER_SPAN_METRICS):
    run_ai_app_with_integration_tracing(golden.input)

If manual instrumentation or the integration supports observed component spans, attach metrics directly to @observe:

from deepeval.tracing import observe

from metrics import GENERATOR_LLM_SPAN_METRICS


@observe(type="llm", metrics=GENERATOR_LLM_SPAN_METRICS)
def call_model(messages):
    ...

Name span metric lists after the component, such as RETRIEVER_SPAN_METRICS, GENERATOR_LLM_SPAN_METRICS, or ORDER_LOOKUP_TOOL_SPAN_METRICS. Do not create one global component metric list for the app. Use next_agent_span, next_llm_span, next_tool_span, or next_retriever_span to match the span type the integration creates.

Pytest vs Script Shapes

For CI/CD, prefer the pytest shape shown in each integration doc — pass the Golden directly through the traced app and assert:

@pytest.mark.parametrize("golden", dataset.goldens)
def test_agent(golden: Golden):
    run_ai_app_with_integration_tracing(golden.input)
    assert_test(golden=golden, metrics=TRACE_METRICS)

For scripts or iteration loops, use evals_iterator and pass the Golden through the traced app:

for golden in dataset.evals_iterator(metrics=TRACE_METRICS):
    run_ai_app_with_integration_tracing(golden.input)

Do not convert a traced single-turn eval into a hand-built LLMTestCase unless the user explicitly chooses no tracing.

Confident AI

If the user chooses Confident AI results, confirm either deepeval login has been run or CONFIDENT_API_KEY is exported. Prefer CONFIDENT_API_KEY for CI and other non-interactive runs. After evals, use deepeval view to open the latest hosted report when appropriate.