Evaluation
The systematic measurement of whether a system produces acceptable outputs against defined criteria, on a held-out or adversarial set, before and after deployment.
Why it matters
Evaluation is what makes a governance claim falsifiable. Without it, 'the agent performs well' is an assertion. The hard part for agents is that the unit under test is a trajectory rather than an output: an agent can produce individually reasonable steps and an unreasonable outcome, and single-turn evaluation will not see it.
What it is not
These are routinely confused with Evaluation. The distinctions are not pedantic — each one has consequences for how a system is governed.
Validation in model risk terms is an independent assessment of a model's fitness for purpose. Evaluation is measurement, and is one input to validation.
Relationships
Typed edges into the rest of the ontology. These are what make the canon traversable rather than merely readable.
| Verb | Target | Meaning |
|---|---|---|
relatedTo | Reproducibility | An association too weak or too general for a stronger verb. |
addresses | marque:standards | The subject speaks to the object as a question or concern. |
Record
| Canonical identifier | QIS-TERM-00042 |
| Status | Canonical industry term |
| Adoption | Widely used |
| Domain · Layer | Stack categories · Intelligence |
| Origin | Machine learning practice; extended to LLM and agent systems. |
| Semantic aliases | None recorded. |
| First published | 2026-08-02 |
| Last reviewed | 2026-08-02 · 180-day cycle |
Cite the identifier, not the URL. Identifiers are stable; URLs may change. This entry is free to read, quote and index under the dual license. Corrections to [email protected] are published.