QIS-TERM-00042 QISTRUST.COM THE GOVERNANCE LAYER REV 2026-08-02 · BUILD 14.4
The QIS Canon · Stack categories

Evaluation

The systematic measurement of whether a system produces acceptable outputs against defined criteria, on a held-out or adversarial set, before and after deployment.

QIS-TERM-00042·Canonical industry term·Widely used

Why it matters

Evaluation is what makes a governance claim falsifiable. Without it, 'the agent performs well' is an assertion. The hard part for agents is that the unit under test is a trajectory rather than an output: an agent can produce individually reasonable steps and an unreasonable outcome, and single-turn evaluation will not see it.

What it is not

These are routinely confused with Evaluation. The distinctions are not pedantic — each one has consequences for how a system is governed.

≠ Validation
Commonly conflated

Validation in model risk terms is an independent assessment of a model's fitness for purpose. Evaluation is measurement, and is one input to validation.

Relationships

Typed edges into the rest of the ontology. These are what make the canon traversable rather than merely readable.

QIS-TERM-00042 — outbound relationships
VerbTargetMeaning
relatedToReproducibilityAn association too weak or too general for a stronger verb.
addressesmarque:standardsThe subject speaks to the object as a question or concern.

Record

Canonical identifierQIS-TERM-00042
StatusCanonical industry term
AdoptionWidely used
Domain · LayerStack categories · Intelligence
OriginMachine learning practice; extended to LLM and agent systems.
Semantic aliasesNone recorded.
First published2026-08-02
Last reviewed2026-08-02 · 180-day cycle

Cite the identifier, not the URL. Identifiers are stable; URLs may change. This entry is free to read, quote and index under the dual license. Corrections to [email protected] are published.