Adds the deepeval eval-loop skill via `npx skills add` with skills-lock.json for reproducible reinstalls. Symlinked to Claude Code via .claude/skills/. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016z2ZFYHQCex8yZAMVMTZzZ
1.6 KiB
Choose Use Case
Classify the target app before choosing templates, datasets, or metrics. Infer from code first; ask only when the code is ambiguous.
Top-Level Use Case
Choose exactly one top-level use case:
- Chatbot or multi-turn agent
- Agent
- RAG
- Plain LLM
Precedence rule:
chatbot / multi-turn agent > agent > RAG > plain LLM
If the app is both RAG and agentic, classify it as an agent.
If the app is both chatbot and agentic, classify it as chatbot / multi-turn agent.
If the app is a chatbot backed by RAG, classify it as chatbot / multi-turn agent.
Signals
| Use case | Signals in code | Test shape |
|---|---|---|
| Chatbot / multi-turn agent | message history, chat endpoint, user session, turns, assistant role, multi-turn state | Multi-turn E2E |
| Agent | tools, function calling, MCP tools, actions, planner, graph, LangGraph, CrewAI, PydanticAI | Single-turn E2E by default |
| RAG | retriever, vector store, documents, chunks, context, citations, no higher-precedence chatbot or agent behavior | Single-turn E2E by default |
| Plain LLM | one prompt in, one answer out, no tools or retrieval | Single-turn E2E |
Use cases guide metrics and required trace fields. Templates are separated by test shape: single-turn tracing, single-turn no-tracing, and multi-turn E2E. Optional component/span metrics stay inside the single-turn tracing shape.
Dataset Default
For chatbot or multi-turn agent use cases, generated datasets should be multi-turn by default. Use single-turn QA pairs only if the user explicitly says they want QA pairs for testing for now.