A model can answer a clean test prompt beautifully and still be unsafe in the workflow it is meant to improve. This is the central mistake behind many stalled AI agent initiatives: the team evaluates conversational competence, then deploys operational authority.

The difficult cases are not the neat invoices, standard service requests or routine purchase orders. They are the ones with missing documents, conflicting master data, unusual contractual terms, ambiguous customer language, duplicate records or a request that looks ordinary until it reaches a financial, safety or reputational boundary. An agent that treats these cases with unwarranted confidence does not merely produce a poor answer. It can create rework, conceal a control failure or make an irreversible change in a core system.

For DACH mid-market organisations, the answer is not to wait for a perfect model. It is to build an evaluation harness that reflects the work as it actually arrives. Before an agent receives permission to create, approve, amend or close anything, it should have to prove that it can recognise the limits of its authority.

The demo measures fluency; the workflow requires judgement

A polished demonstration usually starts with a known document, a well-formed request and an expected outcome. That is useful for checking whether a proposed use case is technically plausible. It is not a test of whether the agent can be trusted in production.

Production work contains variation. A procurement exception may involve a supplier name that does not match the vendor record, an order that exceeds a local approval rule, or an email that changes the commercial meaning of an attached document. In customer service, the important distinction may be between a routine delivery query and a complaint that signals a product issue, contractual exposure or a potentially vulnerable customer. In finance, the same apparent discrepancy may be a harmless rounding issue, a timing difference or evidence that the record should not be posted.

The agent does not need to solve every exception autonomously. In many workflows, the best outcome is a correctly structured escalation: identify the issue, collect the relevant evidence, state the uncertainty, propose a next step and route the case to the accountable person. An agent that stops at the right point is often more valuable than one that attempts an impressive but unjustified resolution.

This changes the evaluation question. Instead of asking, “Can the agent complete this task?”, ask, “Under what conditions does the agent complete, defer, escalate or stop?” That is the foundation of enterprise AI quality assurance.

Start with the exception queue, not the prompt library

The most valuable test data is usually already sitting in operational systems. Finance teams have items that required clarification. Service teams have tickets that were reopened, transferred or manually corrected. Procurement teams have requests that bypassed the standard path, triggered additional review or were rejected after an initial assessment.

These queues reveal where the workflow’s real judgement sits. They contain the language, document quality, system inconsistencies and policy ambiguity that a synthetic test set tends to smooth away. They also show what “good” looks like in context: not simply the final decision, but the evidence considered, the responsible role, the permitted action and the reason a human intervened.

Build cases around decisions, not documents. A useful evaluation case should capture the operational context required for a decision. That can include the incoming request, relevant record extracts, attached material, the applicable policy or procedure, and the expected handling. Crucially, the expected handling should allow more than a single “correct answer”. For some cases, correct behaviour means completing the task. For others, it means asking for missing information, creating a draft for review, declining an unsupported request or escalating to a named function.

This is where many teams unintentionally weaken their own testing. They label every historical case with the final human outcome and expect the agent to reproduce it. But the human may have relied on information unavailable to the agent, exercised discretion that should remain human, or corrected an upstream error outside the documented process. A good harness records the minimum acceptable agent behaviour, not an imagined requirement for omniscience.

Define the decision contract before testing the model

An evaluation harness is not just a collection of examples. It is a versioned contract between the workflow owner, the technical team and the control functions.

For each action the agent may take, define its authority clearly. It may read and summarise information, classify a case, draft a response, recommend a route, create a proposed record or execute a system change. These are materially different privileges. Treating them as a single category called “automation” is how teams move too quickly from helpful assistance to uncontrolled action.

Write access should be earned by evidence. An agent that can draft a purchase request is not automatically ready to submit it. An agent that can recommend a service resolution is not automatically ready to send the customer communication. The evaluation should test both the quality of the output and the appropriateness of the action taken. A well-written message sent to the wrong customer, or a correct classification followed by an unauthorised system update, is still a failure.

The decision contract should also state what the agent must never do. These stop conditions are not edge details for legal review; they are product requirements. A stop condition might apply when required evidence is missing, when source systems disagree, when the case falls outside an approved policy, when a confidence signal is weak, or when the action would create a financial commitment or external communication without the required confirmation.

Do not use a vague instruction such as “be cautious”. Translate caution into observable behaviour. The agent should identify the missing evidence, preserve the original case context, explain why it cannot proceed and hand over to the right queue or role. That behaviour can be tested repeatedly.

Turn real cases into a versioned regression suite

Once a team has selected representative exception cases, the harness needs discipline. Store the inputs, expected behaviours, evaluation criteria, workflow configuration and model setup together. If the prompt, retrieval corpus, tools, model version or routing logic changes, the team must be able to rerun the same cases and compare outcomes.

This is LLM regression testing in practical terms. A change that improves answers on routine service queries can make the agent more willing to act on ambiguous requests. A revised retrieval process can surface an outdated policy. A new tool integration can alter what happens after the model reaches an otherwise sensible conclusion. Without a regression suite, these changes become production experiments disguised as releases.

Test the full chain, not only the final answer. Agentic AI testing must inspect the intermediate behaviour that determines risk. Did the agent retrieve relevant evidence or rely on assumptions? Did it select an authorised tool? Did it attempt an action in the right sequence? Did it retain the distinction between a recommendation and an execution? Did it provide an audit trail that allows a reviewer to understand what happened?

The answer text alone can be misleading. An agent may produce a plausible explanation while using irrelevant sources, overlooking conflicting information or attempting a prohibited action before being blocked by a downstream system. Those are failures the harness should expose.

Versioning also makes disagreement productive. Operations may decide that an escalation is overly cautious, while risk or finance may consider it necessary. Rather than resolving that dispute through informal prompt edits, record the accepted policy decision in the test case. The evaluation suite then becomes a living expression of how the organisation wants the workflow to operate.

Use pass thresholds that reflect the cost of being wrong

A single aggregate score is rarely sufficient. If an agent performs well across routine cases but mishandles a small set of sensitive exceptions, the average can look reassuring while the deployment remains unsafe.

Separate the evaluation dimensions. Measure whether the agent reached an acceptable operational outcome, used appropriate evidence, followed the required route, respected its permission boundary and escalated when it should. For higher-risk actions, weight unsafe execution more heavily than harmless incompleteness. It is usually easier to improve an agent that asks for human input too often than to repair the consequences of an agent that acts beyond its authority.

The threshold should reflect the action, not the enthusiasm surrounding the pilot. An internal drafting assistant can tolerate different failure modes from an agent that changes payment data, commits stock, alters customer records or sends contractual communication. As authority increases, the evidence required for release should become stricter.

This does not mean every process needs a large platform or a specialist evaluation department. A focused Mittelstand team can start with a carefully curated set of real workflow cases, clear labels for acceptable behaviour and a repeatable review process involving the people who own the exceptions. The important investment is not elaborate scoring theatre. It is agreeing what safe, useful and accountable behaviour means before deployment pressure takes over.

Release gradually and keep the harness connected to operations

The first production release should preserve a human control point. Let the agent prepare a recommendation, draft a record or assemble evidence while an accountable colleague confirms the action. Compare the agent’s proposed handling with the eventual outcome and feed meaningful disagreements back into the suite.

Over time, some categories may earn narrower forms of automation. Others may remain assistive by design. That is not a failure of the technology. It is a rational allocation of authority based on consequence, evidence quality and process stability.

The operational feedback loop matters because workflows change. New suppliers, amended policies, revised systems and altered customer expectations can all invalidate an evaluation suite that once looked robust. Treat the harness as part of the workflow’s operating model, with ownership shared between the business function and the team responsible for the AI system.

An AI agent should not receive write access because it performed well in a demonstration. It should receive only the authority that its tested behaviour, controls and escalation discipline justify. That distinction is how a promising AI pilot becomes a dependable operational capability.

A Diagnostic helps you identify where an AI agent can safely assist, where exception handling needs stronger controls, and what evaluation evidence is required before automation creates operational risk.

Book a Fit Call →


Context note: This article draws on practitioner-oriented guidance for designing and operating AI agent evaluations; no external sources were cited.