Skip to content
AI13 min readupdated 11/08/2026

Measuring AI quality: test set, release limits and monitoring in operation

By Kevin Kröger, Geschäftsführer, Software und Plattformbetrieb

Räumliche Visualisierung eines neuronalen Netzes
Header image: Unsplash
THE SHORT ANSWER

AI quality is not assessed by individual convincing answers. Reliable proof requires a versioned test set from real specialist cases, clear evaluation criteria, separate limits for critical errors, documented model and prompt statuses as well as ongoing monitoring in operation. Only when quality, costs and the consequences of errors are visible together can an AI system be taken into account.

Why isn't a good demo enough?

A demo usually shows well-known, well-prepared cases. In everyday life, the system encounters incomplete entries, contradictory documents, new terms and rare exceptions. Therefore, the evaluation begins with a list of real tasks and the question of what errors can arise in each case. An incorrect wording proposal has a different effect than a fabricated release, an overlooked contractual clause or an inadmissible data transfer. The NIST AI RMF explicitly connects AI assessment with context, measurement, and ongoing management of risk.

How is a representative test set constructed?

Collects normal processes, difficult edge cases, outdated information, intentionally misleading questions, and cases where the system should safely reject. Each case is given an expected property rather than just a pattern set. Evaluates, for example, technical accuracy, completeness, reference to sources, permissible use of data and helpful uncertainty. The test set is versioned, checked by those responsible and must not be quietly incorporated into development. Otherwise the team will optimize for known tests and overestimate the subsequent quality.

What quality limits does the company need?

An average value masks rare but serious errors. Therefore, define at least one overall limit and separate zero or low limits for critical events. Determines when a response may be used automatically, when a person must review it, and when the system should abort. This decision is part of the process design and not just the model. Also documents who releases a new model version and what evidence must be provided for this.

What needs to be monitored after go-live?

Observes response quality on samples, rejection rate, necessary corrections, response time, costs per process and security-relevant events. Captures model, system instruction, knowledge level and relevant configuration per issue without unnecessarily logging confidential content. User feedback is a signal, but not proof of truth. Changes to sources, roles, models or tools trigger targeted regression tests. In this way, monitoring becomes a technical control loop instead of a pure infrastructure dashboard.

How is deterioration treated?

Quality waste requires the same seriousness as technical malfunctions. Defines warning thresholds, responsible parties, a way back to the last checked status and a way to stop automatic actions. Separates root cause analysis into model, data, retrieval, prompt, authorization and process. A quick model change can postpone the problem if the root cause is an outdated knowledge source or an unclear work order. After correction, the affected sample is included in the test set.

Next steps

From the answer to implementation

Sources and basis

The central statements in this article were reviewed against the following primary sources.

Frequently asked questions

How big does an AI test set have to be?
There is no general minimum number. What is crucial is coverage of important tasks, critical error consequences, user groups and real exceptions as well as sufficient repeatability.
Can an automatic key figure replace the specialist examination?
No. Automatic evaluations help with repetitions, but must themselves be checked against professional judgments. Critical cases require comprehensible human evaluation.
When does it have to be tested again?
For changes to the model, prompt, knowledge base, roles, tools or process and regularly based on current production cases.
Continue reading

More specialist articles about AI

AI

What would this look like in your organisation?

We apply the specialist assessment to your situation and clarify a concrete next step.

Request a meeting