Buying an AI system is not the same as buying conventional software. A conventional application may process important data, but its core behaviour is usually specified, testable and relatively stable between releases. A model, agent or copilot introduces a more complicated operating reality. Its behaviour depends on the model version, system instructions, connected data, retrieval logic, tools, user inputs and the controls surrounding it. If one of those components changes, the system’s risk profile can change with it.
That is why a standard vendor questionnaire is not enough. It may collect useful declarations, but it rarely creates an operating control that the organisation can use at go-live, during an incident or when a supplier relationship ends. The better procurement mechanism is an AI supplier evidence pack: a structured, reusable set of artefacts that shows what is being deployed, how it handles data, which decisions remain with people, and what evidence supports the claim that it is ready for the intended use.
For the Geschäftsführung, this is not administrative overhead. It is a way to turn a vague assurance such as “the supplier is compliant” into a decision that can be understood, challenged and owned. For IT, security and data teams, it replaces repeated email exchanges with a baseline that can be reviewed when the use case expands or the underlying technology changes.
Why contractual promises do not provide operational control
Many procurement processes still put AI governance in the contract. The supplier confirms that it follows applicable requirements, signs data-processing terms, provides security language and perhaps supplies a generic product security document. Those documents matter, but they do not answer the questions that arise once an AI system starts affecting work.
Consider a copilot used by a customer service team. The buyer needs to know not simply that the supplier uses “secure AI”, but which model is called, where prompts and outputs travel, whether customer records are retrieved at runtime, how long logs are retained, and whether agents can take actions in connected systems. The same questions become more consequential when an agent can draft contract amendments, update a CRM record, create a purchase request or trigger a workflow.
A contract is usually written to allocate liability between parties. An evidence pack is written to enable safe operation. It helps the organisation establish what it has approved, what assumptions were made and which controls must continue to work after launch.
The central procurement question is therefore not whether the supplier can complete a compliance questionnaire. It is whether the supplier can provide evidence that the specific deployment is known, bounded and governable.
Start with a deployment identity, not a product name
“AI assistant” is not a useful description for governance. Neither is a product name alone. A supplier evidence pack should establish a deployment identity that captures the actual system being introduced into the organisation.
This means identifying the supplier, any material subcontractors, the model provider or providers, the relevant model family and version where that information is available, and the hosting arrangement. It should also state whether the application uses a general-purpose model directly, routes requests across several models, uses retrieval-augmented generation, or fine-tunes a model for a defined task.
This detail matters because buyers often assume they are assessing one system when they are in fact assessing a chain of systems. A vendor may provide the workflow layer, another provider may host the model, a separate service may handle document extraction, and a cloud platform may store operational logs. Each relationship can shape data handling, availability, security responsibilities and the practical ability to investigate an incident.
The evidence pack should be specific to the deployment, rather than recycled marketing material. If the supplier cannot identify the components that process data or influence outputs, the buyer does not yet have enough clarity to approve production use.
Make the data-processing map usable by operators
Data maps are often produced as a compliance exercise and then filed away. For AI systems, that is a missed opportunity. The map should be understandable to the people who run the service, not only to legal specialists.
It should show what enters the system, including user prompts, uploaded documents, retrieved internal knowledge, conversation history and metadata. It should show where each category is processed and stored, whether it is used for service improvement or model training, and how deletion requests or retention rules are handled. It should also explain what leaves the system: generated text, tool calls, audit records, notifications and exports to other business applications.
The decisive distinction is between content and control data. A supplier may promise not to train on customer content while still retaining prompts, outputs or metadata for operational purposes. That can be reasonable, but it must be visible. The organisation needs to know what it can switch off, what is essential for troubleshooting, and what is retained beyond the user session.
For a Mittelstand organisation, this map does not need to become an enterprise architecture thesis. It needs to answer practical questions quickly when a business owner asks whether HR records, engineering documents, customer information or commercially sensitive material may be used in a particular workflow.
Classify the role before debating the paperwork
AI governance discussions often start with a long list of documents. Start instead with the organisation’s role and the intended use. A business using an off-the-shelf copilot internally has a different responsibility profile from a business adapting an AI system for customers, embedding it into a product, or developing an agent that makes consequential workflow decisions.
The evidence pack should record the supplier’s view of its role, the buyer’s intended role and the reasoning behind that assessment. It should also document the purpose, user group, deployment context and foreseeable misuse cases. This is particularly important where the system touches employment, finance, safety, customer eligibility or other sensitive decisions.
The point is not to force every modest internal assistant through the same process as a high-impact system. It is to avoid the opposite mistake: treating a system as low-risk because it has been labelled “productivity”. A drafting assistant that merely proposes a first version of a meeting summary is materially different from an agent that acts on that summary by updating commercial records or sending customer communications.
A procurement team should insist that the supplier describes these boundaries in plain language. If the boundary cannot be explained, it cannot be monitored.
Ask for evaluation evidence that matches the real task
Generic model benchmarks are a weak basis for a buying decision. They may show broad capability, but they say little about whether the system works reliably with the organisation’s documents, terminology, workflows and tolerance for error.
A useful evidence pack includes an evaluation approach for the actual use case. For a knowledge assistant, that means testing whether answers are grounded in approved sources, whether citations or links lead users to the right material, and what happens when the source base does not contain an answer. For a document-processing workflow, it means examining extraction quality, exception handling and the route for uncertain cases. For an agent, it means testing tool permissions, action boundaries and failure behaviour.
Evaluation must include refusal and escalation behaviour. Buyers should not only ask when the system performs well. They should ask what it does when information is missing, instructions conflict, an external tool fails or a user requests an action beyond the approved scope.
The results do not need to imply impossible certainty. AI systems remain probabilistic. What matters is whether the supplier can show a disciplined test process, disclose known limitations and identify the control that prevents an ordinary error from becoming a business incident.
Treat logging and human override as product requirements
A system that cannot be investigated cannot be governed. Before go-live, the buyer should understand what operational records exist and whether they allow the organisation to reconstruct a material event. This normally includes relevant user activity, model or workflow configuration, retrieved sources where applicable, tool calls, approvals, errors and changes to permissions or prompts.
Logging must be designed with data minimisation in mind. Capturing every input indefinitely is not a sign of maturity. The goal is a proportionate audit trail that supports incident investigation, quality improvement and accountability without creating an unnecessary secondary data store.
Human oversight needs similar precision. “A human remains in the loop” is insufficient if no one knows who that human is, when they intervene or whether they have authority to stop the process. The evidence pack should show the escalation path, the ability to disable the system or a high-risk tool, and the fallback process when the AI service is unavailable or unreliable.
For agentic systems, this is where procurement becomes architecture. A credible design usually separates recommendation from execution, constrains tool access, applies approval gates to material actions and makes it possible to revoke permissions quickly. An agent with broad access and no practical kill switch is not an innovative procurement choice; it is an unmanaged operational dependency.
Include security evidence and an exit route
Security evidence should be connected to the actual service, not limited to a supplier’s corporate posture. Buyers need clarity on identity and access management, tenant separation, encryption arrangements, vulnerability handling, administrative access and incident communication. Where the system connects to internal tools, the integration path deserves at least as much scrutiny as the model itself. A well-protected model interface does not compensate for an overly privileged connector.
The final part of the evidence pack is often neglected: exit. AI suppliers can become embedded quickly because prompts, evaluations, retrieval structures, workflow logic and user habits accumulate around their platform. Buyers should establish what can be exported, how business data will be returned or deleted, which configurations can be documented, and what happens to logs and access credentials at termination.
Exit planning is not a prediction that the supplier will fail. It is a discipline that prevents convenience from becoming lock-in. It also improves the buying decision today, because a supplier that can explain portability and deletion clearly usually understands its own operating model.
Make the evidence pack a living go-live control
The pack should not disappear after contract signature. It should become the baseline for production approval and subsequent change management. A new model version, additional agent capability, new data source, revised retention setting or expanded user group may require the evidence to be updated. That is not bureaucracy for its own sake; it recognises that an AI deployment is a changing system.
Ownership should be explicit. Procurement can collect the evidence, but IT, security, data protection, legal and the business owner each need to validate the parts they are qualified to assess. The Geschäftsführung should not be expected to review technical logs, but it should be able to see whether the decision has a defined owner, known limitations, a credible override route and a manageable supplier dependency.
A supplier evidence pack turns AI vendor due diligence into something more useful than a gate. It creates a shared operating record. For organisations moving from isolated experimentation to production workflows, that record is often the difference between a controlled capability and a black box with a purchase order.
A Diagnostic helps your organisation define the evidence, ownership and go-live controls needed for an AI deployment — before supplier assumptions become operational risk.
Context note: This article provides practitioner guidance and does not cite external sources.
