Three decisions determine whether AI delivers anything in a company: which areas get automated, at what point in time, and to what quality. All three require that the people deciding understand how these systems work.
This piece works through German corporate law. The mechanism it describes is not specific to Germany.
While researching this text, an AI system gave me a legal citation: article, journal, volume, author. The author died in 2023; the article does not exist. Everything about it was formatted correctly, none of it was true. Anyone who copies that citation into a memo will not notice, until somebody looks it up.
That is the competence at issue. It does not consist in programming. It consists in seeing what an output is worth, and deriving from that what a system is good for.
Which areas. The question sounds like one of prioritisation, but is in truth about suitability. Linguistic tasks with a tolerance for variance succeed: drafts, summaries, pre-sorting. Tasks with a strict correctness requirement, where verifying the output costs more than doing the work by hand, fail. Anyone who does not grasp this distinction selects by visibility rather than by suitability, and ends up where a demonstration is easy rather than where the leverage lies.
What that mistake costs is shown by the industry’s most quoted report, though not in the way it is usually cited. Between January and June 2025, Project NANDA at the MIT Media Lab examined more than 300 publicly documented initiatives, held interviews in 52 organisations, and surveyed 153 senior leaders. For bespoke, task-specific tools, 20 per cent of the initiatives examined reached a pilot and 5 per cent reached production. From pilot to regular operation, then, one in four made it. Generic chatbots cleared the same hurdle in around 83 per cent of cases. The same technology yielded three times the success rate, decided entirely by the choice of tool type.
What made that report famous was a different figure: “95 per cent of all AI pilots fail”. That number comes from the same funnel, measures a reported effect six months after the pilot, and is not an audited return calculation. The report is not peer reviewed, it comes from an initiative that recommends its own infrastructure in its conclusion, and it is no longer freely available at its original address. Anyone quoting it without knowing this fails to carry out the very check at issue here.
A word on the reach of these figures: they show what the choice of tool type does. That a lack of competence at the top causes the failures is not something they show, and the report does not claim it either. It names workflow, learning capability, and organisation as obstacles. The point here is a different one. Whoever decides on tool type, timing, and quality threshold decides the exact variables along which these rates differ.
At what point in time. This decision commits capital and restructuring effort. Moving too early means building a process on capabilities that do not yet exist, only to rebuild it twelve months later. Moving too late means the competition already has that rebuild behind it. The estimate of what will be possible in a year cannot be passed downwards, because it is tied to investments decided at the top.
To what quality. The hardest of the three. A language model works with probabilities; its output is a distribution whose edges you have to know. The question is therefore not whether the system makes mistakes, it is which error rate a given process tolerates and what an error costs there. In pre-sorting job applications an error means something different from an error in billing. That weighing is a commercial decision, and it requires somebody who understands how the errors arise and whether they fall randomly or hit particular cases systematically.
The strongest objection to all of this runs: a board member does not do the bookkeeping either. True, and it proves the opposite. A board member does not book, but must be able to read a balance sheet. That is exactly what is asked here. Nobody expects a manager to train a model. What is expected is that they can read a system’s output like a balance sheet, recognising what is in it, what is missing, and where estimates were made. Whoever cannot do that waves through whatever is put in front of them, and the strategy is then set by whoever wrote the paper. Often that is a vendor.
Deloitte surveyed the state of that literacy in early 2025 among 695 board and C-suite members in 56 countries. Two thirds say their own body has limited or no AI knowledge and experience; in Germany 59 per cent, though from a sample of only 49 respondents. Formal AI education for the body is offered by 48 per cent of companies.
The counter-position is well put. One of the leading US corporate law firms wrote in May 2026 that directors need not necessarily develop individual expertise or approve every AI tool, and may rely on management and qualified experts. The conditions stand in the same paragraph: clear visibility of the core tools in use, of the workflows that could be materially affected, and of the reporting and escalation processes. Anyone wanting to decide which tools count as critical and which workflows are materially affected needs their own judgement about the systems to do so. The condition presupposes what the main clause relieves them of. And the passage speaks to supervisory boards under US law. It relieves oversight; about executive management it says nothing.
Legally, none of this works as a lever, and it should therefore not be used as one. Article 4 of the AI Act required a sufficient level of AI literacy from February 2025, and since 27 July 2026 only measures that foster its development. A particular level for individuals expressly no longer has to be ensured, and the addressee is the organisation as provider or deployer.
What remains is a question of competence in the formal sense. The decision on whether to use AI is, according to the relevant commentary, a management decision that the executive management has to take itself as part of its responsibility for the principles of business policy. What can be delegated, on that reading, is the execution. The decision stays at the top.
The business judgement rule is often invoked here: whoever relies on expert advice and documents it is protected. That is correct, and it carries a condition that is rarely read along with it. Protected is whoever acts on an adequate basis of information. Whether a basis is adequate has to be judged by the person taking the decision. Whoever cannot do that also does not know whether they are protected. They hope so.
What to do is smaller than it sounds. For reading a balance sheet there is formal training, a standard, and an auditor who answers for it. For reading an AI output there is none of that, which is why the route runs through your own observation. First: work yourself, regularly, with the systems being decided about. Use a task of your own whose result you can judge, rather than a prepared demonstration. This is not an end in itself, it is the only available route to judgement. Second: look for errors deliberately instead of being shown successes. A system’s worth shows at its edges. Third: have every vendor answer how you would recognise that their tool does not work in your house. Whoever has no answer to that is selling.
This is not operational work and replaces neither the specialist department nor external advice. It is a sample taken to calibrate your own judgement, and it has to be repeated after every model change, because otherwise the judgement goes stale. Whoever does not take it decides on automation, timing, and quality thresholds on somebody else’s judgement. That is the difference between a strategy and a collection of pilot projects that nobody stops, because nobody can say how their failure would be recognised.