Skip to content

Who is accountable when your AI is confidently wrong?

By Timothy Ng5 min read

Accountability for a wrong AI answer stays with the person or function that acted on it, and with the executive who authorised the system to operate there. A bad AI answer is fluent, well structured and gives no signal that something underneath is wrong, so the boundary has to be set before the output is trusted: decide the consequence of failure before granting authority, expose the strength of the evidence, and keep consequential actions behind an enforced human approval.

Key points

  • Fluent writing tells you nothing about whether the answer is reliable.
  • Decide the consequence of failure before granting authority, not after an incident.
  • A reader needs to see how strong the underlying evidence is, not only the answer.
  • Consequential actions involving another person stay behind a hard human boundary.

One of the harder problems with business AI is that a bad answer does not necessarily look bad. It may be fluent, well structured and entirely plausible. There may be no error message, no obvious failure and nothing in the quality of the writing to tell the person reading it that something underneath is wrong.

That changes the governance problem considerably. AI does not need complete autonomy to create risk. It only needs enough trust for somebody to act on what it says without understanding how the answer was reached.

That is why one of the questions I think leadership teams need to answer early is:

Who is accountable when the AI is confidently wrong?

My background is in CTO leadership across organisations where technology decisions carry different levels of commercial, operational and regulatory consequence. That experience shaped the boundaries we put around our own AI assistant. Operating the system across real work has since given us useful evidence of why those boundaries matter.

Four areas in particular are worth looking at.

1. Decide the consequence of failure before granting authority

There is a significant difference between an AI being capable of doing something and deciding that it should be allowed to do it. Preparing an internal summary has one level of consequence. Making a commitment to a client has another.

Our system can prepare external communications. It cannot independently decide that a draft should become an external action. Anything consequential involving another person remains behind a hard human boundary.

That was not introduced after the AI sent something inappropriate. It was a design decision made before giving the system that authority.

The useful leadership question is therefore not simply:

What can our AI do?

It is:

What is the greatest consequence it can create without somebody genuinely seeing and approving the action first?

If that consequence matters, the control should not depend on the model choosing to follow an instruction. It needs to exist outside the model.

2. Good writing tells you almost nothing about whether the answer is reliable

We have seen the system produce a convincing answer using information that had become stale. We have also seen weak underlying evidence emerge as prose that looked every bit as authoritative as an answer backed by much stronger information. That is a difficult characteristic of AI for a business to manage.

With many traditional systems, bad input eventually produces an obvious failure. Generative AI can turn uncertainty into something remarkably polished. The person using it therefore needs more than the answer itself.

Where did the information come from? Is that the authoritative source? Is it current?

How strongly does the available evidence support the conclusion? Is there actually enough information to answer? In some circumstances, the correct behaviour is for the system to say that it does not have enough reliable information.

I would rather have an AI decline to give me a definitive answer than manufacture certainty from weak evidence. For a leadership team, the important question is whether somebody can distinguish a well-supported answer from one that merely sounds well supported. Careful proofreading is not enough when both can read perfectly well.

3. Assumptions can become organisational facts remarkably quickly

This problem predates AI. I have seen it happen in strategy documents, product plans and board material for years. Somebody writes down something plausible. Another document repeats it. A third person assumes it was validated previously. Eventually the organisation forgets that the original statement began as an assumption.

AI can accelerate that process considerably. A generated output can become the input for another process. Unless the status of the underlying information is preserved, the second system may have no reason to question whether the original statement was ever established as fact.

That means an effective knowledge system has to preserve more than the words. It needs to preserve what those words represent. Is this a validated fact?

A hypothesis? An assumption? An interpretation?

Something still waiting for confirmation? Organisational knowledge has a tendency to compound what it is given. AI makes that compounding considerably faster.

A useful test is whether your system can distinguish between something the business knows and something somebody once thought might be true.

4. Correct information in the wrong context is still wrong

Context is what makes an AI assistant substantially more useful. It is also what makes some failures much harder to spot. If a system operates across different clients, projects, departments or workstreams, language inevitably overlaps.

That creates a subtle but serious failure mode: the information itself can be correct while being completely inappropriate for the context in which it is used. The resulting answer may still sound entirely credible. Our approach is deliberately conservative.

Facts from separate contexts should not be combined merely because they appear related. If the reference is ambiguous, the system should establish the correct context rather than infer it. This is particularly important because ordinary review may not catch the problem. A human can read the answer, see nothing obviously wrong and approve information that belongs somewhere else.

The protection therefore needs to exist before the answer is generated. That is architecture and governance, not better wording in a prompt.

These are technology leadership decisions

It is tempting to treat problems like these as a new collection of AI-specific issues. I do not think they are. They are decisions about authority, permissions, information integrity, accountability, trust and appropriate human oversight.

Technology leaders have always had to make those decisions around consequential systems. AI makes the boundary more complicated because the system is interpreting information rather than simply executing predictable instructions. That is also why I would be cautious about approaching AI adoption mainly as a procurement exercise.

Buying access to a model is easy. Connecting systems is becoming easier. Building an assistant is becoming easier too.

The more difficult work is deciding what role that assistant should have once people begin relying on it.

We identified 13 situations worth governing

The four examples above are part of a wider set of 13 situations we have documented through operating the system. I do not think of them as 13 mistakes made by an AI. They are 13 situations I would rather an organisation considered before it encounters them in something consequential.

For each one, the same management questions apply:

Could this happen here? What would the consequence be? What currently prevents it?

Who owns the decision? Those questions become increasingly important as businesses move from experimenting with AI to allowing it to participate in normal operational work. Because the serious question is no longer whether a business should use AI.

It is where that AI should operate, what information it should trust, what authority it should have and where human accountability remains non-negotiable. That is the subject of the next article: the 13 governance decisions I think organisations should make before the technology starts making them by default.

Common questions

Who is legally accountable for an AI decision in a business?
The organisation and its officers remain accountable. Using an AI system to produce or support a decision does not transfer responsibility to the model or the supplier, so the board needs to know where AI has authority and what it can affect.
How do you tell a reliable AI answer from a plausible one?
Not from the prose. The system has to expose what it used, how current that information is, and how strong the supporting evidence was, and it has to be allowed to decline rather than manufacture certainty.
What should a board ask about AI accountability?
Where AI has authority, what it is allowed to trust, how uncertainty is surfaced, and what happens when it gets something wrong. A board does not need the architecture to ask those four questions.