The certificate problem

When the EU AI Act's training requirements landed on compliance calendars, most organisations reached for the same playbook they used for GDPR awareness or information security basics: a learning management system, a set of modules, a quiz at the end, and a certificate that confirms someone clicked through the material. The mechanics are familiar, the reporting is tidy, and the artefact—a timestamped completion record—looks like evidence. It is not.

The Act's requirement that organisations "carry out regular AI training with their staff to ensure they have high levels of AI literacy" is not satisfied by documentation of exposure. It is satisfied by evidence that people can act differently when AI outputs appear in their workflow. Recent research suggests many European CIOs report progress on training delivery but focus on completion metrics rather than behavioural evidence. The gap between those two concepts is where operational risk lives, and it is the gap that auditors, incident reviewers, and works councils will probe when something goes wrong.

The practical question is not whether your staff completed a course. It is whether a customer service representative, faced with an AI-generated response that contradicts their experience, knows how to escalate it rather than send it. Whether a procurement analyst, presented with a supplier risk score from an automated system, understands which inputs are probabilistic and which are deterministic. Whether a claims handler, offered a recommended decision path, can articulate why they accepted or rejected it. If those behaviours have not changed, the training was a communications exercise, not a control.

What makes AI literacy an operating control

An operating control is a mechanism embedded in the workflow that reduces the likelihood or impact of a specific risk. It is not a policy statement or a training record; it is a decision point, a checklist, a logged review, or an escalation path that changes what happens next. For AI literacy to function as a control, it must alter the behaviour of people who interact with AI outputs in ways that are observable, repeatable, and auditable.

Role-based training anchored to specific workflows. Generic AI awareness training—what a large language model is, why bias matters, how probabilistic systems work—provides context but does not change behaviour. Effective training starts with the question: in this role, when does AI produce an output that this person will act on, and what decision does that output inform? A customer service agent needs to recognise when an AI-drafted reply misses regulatory language or contradicts a known exception. A financial controller needs to know which reconciliation steps can be automated and which require human judgement because the underlying data quality is uneven. A hiring manager needs to understand that an AI-ranked candidate list is not a merit order but a pattern-matching exercise that may encode historical bias. The training must be specific enough that the learner can point to a moment in their day when they would apply it.

Workflow-specific decision checklists. If AI literacy is to function as a control, the knowledge must be encoded into the workflow at the point of use. This does not mean adding a pop-up reminder; it means designing a decision checkpoint where the user must answer a small set of questions before proceeding. For an AI-generated contract clause: does this clause reflect the current legal standard, or is it based on outdated precedent? For an AI-recommended inventory reorder: does this recommendation account for the supply chain disruption flagged last week, or is it extrapolating from historical patterns? For an AI-assisted diagnostic suggestion: does this suggestion align with the patient's presenting symptoms, or is it anchored to the most common condition in the training data? The checklist does not need to be exhaustive; it needs to force the user to pause and apply judgement rather than accept the output by default.

Logged human review and escalation paths. Behaviour change becomes auditable when the system records not just what was decided but how the decision was made. If a user accepts an AI output, the log should capture whether they reviewed it, whether they modified it, and whether they escalated it for a second opinion. If they reject it, the log should capture why. This is not surveillance; it is the creation of an evidence trail that demonstrates the organisation's human oversight mechanisms are functioning. When an incident occurs—a mis-sold product, a discriminatory hiring decision, a compliance breach—the ability to show that staff were trained, that decision points existed, and that people used them is the difference between a control failure and a process failure. The former is a design problem; the latter is an execution problem. Regulators and works councils understand the distinction.

Evidence that people can challenge AI outputs. The most important behavioural shift is the willingness to question an AI recommendation when it does not align with professional judgement or situational context. This willingness is cultural, but it must be operationalised. It requires that escalation paths are clear, that raising a concern does not slow the workflow to the point of dysfunction, and that the organisation treats challenges as valuable signals rather than friction. If staff consistently accept AI outputs without modification, one of two things is true: either the system is performing exceptionally well, or the users have learned that questioning it is not worth the effort. The second scenario is far more common, and it is the one that creates liability.

How to build the evidence base

The shift from training-as-certificate to training-as-control requires a different kind of evidence. Completion rates and quiz scores are inputs; they tell you who was exposed to the material. What you need is evidence of changed behaviour: decision logs, escalation rates, modification patterns, and feedback loops that show people are engaging with AI outputs critically rather than passively.

Start with a workflow map that identifies AI touchpoints. For each role that interacts with AI-generated outputs, document the specific moments where a decision is made based on that output. This is not a comprehensive process map; it is a focused inventory of the points where human judgement is supposed to intervene. If you cannot identify those points, you do not have a control framework; you have automation with a human in the loop by accident rather than design.

Design role-based training that teaches recognition, not theory. The training should present realistic examples of AI outputs that require judgement: an output that is plausible but wrong, an output that is correct but incomplete, an output that reflects a pattern that no longer applies. The learner should practice recognising these cases and articulating why they would escalate, modify, or reject the output. The goal is not to make everyone an AI expert; it is to make them competent in the specific judgements their role requires.

Embed decision checkpoints in the workflow and log the responses. This is where the control becomes operational. The system should require the user to confirm that they have reviewed the AI output, that they understand its limitations, and that they have applied the relevant decision criteria. The log should capture their responses, any modifications they made, and any escalations they triggered. This is not a compliance theatre exercise; it is the creation of a dataset that shows whether the training changed behaviour. If escalation rates are near zero, either the AI system is flawless or the users are not engaging. The latter is far more likely.

Monitor patterns and close the loop. The evidence base is not static. If you see that certain roles consistently accept AI outputs without modification, that is a signal that either the training was ineffective or the workflow design does not encourage critical engagement. If you see that escalations cluster around specific types of outputs, that is a signal that the AI system may have a systematic weakness or that the training needs to address a gap. The organisation that treats this data as a feedback loop—adjusting training, refining decision criteria, and improving the AI system—is the one that can demonstrate continuous improvement rather than one-time compliance.

Why this matters beyond the audit

The practical value of treating AI literacy as an operating control extends beyond regulatory compliance. It changes the economics of AI deployment. When staff are trained to recognise when an AI output is trustworthy and when it requires scrutiny, the organisation can deploy AI more aggressively in low-risk contexts and more cautiously in high-risk ones. The alternative—treating all AI outputs as equally uncertain and requiring blanket human review—eliminates the efficiency gains that justified the investment in the first place.

It also changes the risk profile. Based on the Act's structure and enforcement patterns from similar regulations, organisations that deploy AI widely but cannot demonstrate that their staff understand its limitations are likely to face greater scrutiny than those that avoid AI entirely. The incident that triggers regulatory attention is rarely the first failure; it is the failure that reveals a pattern of passive acceptance. The ability to show that decision points existed, that people used them, and that the organisation responded to the signals they generated is the difference between an isolated incident and evidence of systemic neglect.

Finally, it changes the relationship between the organisation and its works council or employee representatives. The question of whether AI systems are being used to monitor, evaluate, or replace staff is a live issue in many DACH mid-market firms. The ability to demonstrate that AI literacy training is not a box-ticking exercise but a genuine effort to equip staff to work with these systems on their terms—challenging outputs, escalating concerns, and retaining decision authority—builds trust in a way that policy documents cannot.


A Diagnostic maps your current AI touchpoints, identifies where human oversight is nominal rather than operational, and designs role-based training that turns literacy into a control—before an incident forces you to prove it worked.

Request a Diagnostic →


This article draws on findings from recent research into European CIO perspectives on EU AI Act compliance readiness, which noted that whilst many organisations report progress on training delivery, the focus remains on completion metrics rather than behavioural evidence.