Back to the AI glossary Workflows
Evaluation/Evals
Systematic measurement and evaluation of AI quality.
Explanation
Evals are tests and metrics used to measure the quality of AI outputs. They help to identify whether a system works reliably and where weaknesses lie.
How it works
You define test cases with expected results and let the AI process them. Automatic and manual ratings show how correct and helpful the answers are.
Example
Define 100 typical customer requests as a test set. The AI agent answers them and the results are checked for correctness, tonality and completeness.
Why it matters
Without Evals you fly blind. Systematic evaluation is the basis for continuous improvement and trust in AI systems.