AI Glossary · Definition
What is Evaluation?
Evaluation is the structured testing of an AI system's outputs against agreed examples and criteria to measure quality before and after launch.
By DAIDU EditorialUpdated
Evaluation, explained
Instead of judging an AI system on a few demos, teams build a set of realistic test cases with expected results, including difficult and unusual ones, and score outputs for accuracy, tone, safety and format. Re-running the same evaluation after any change to prompts, models or data shows whether quality improved or slipped. Evaluation turns 'it seems fine' into evidence.
Example
Before launch, a clinic tests its FAQ assistant on 100 real patient questions and checks that every clinical question is escalated.
