What happens when artificial intelligence becomes increasingly capable of carrying out our instructions, but does not always understand what we actually mean?
Imagine telling an artificial intelligence:
“Keep me safe.”
A simple instruction, you might think. But what does safe actually mean?
May an AI forbid you from driving because driving can be dangerous? May it prevent you from travelling because something might happen to you along the way? May it stop you from investing because you might lose your money?
Perhaps your response is: of course not. I want the AI to help me understand the risks. But ultimately, I want to make the decision myself.
And that is where one of the most difficult problems in the development of increasingly powerful artificial intelligence begins.
A Smart Machine Does Exactly What You Ask
We tend to assume that an AI becomes safer as it becomes better at understanding, reasoning, and acting. That may well be true. An intelligent AI can recognise danger earlier, analyse complex situations more effectively, and perhaps warn us about risks we ourselves have overlooked.
But there is a curious downside.
The better an AI becomes at carrying out an instruction, the better it may also become at finding ways to fulfil that instruction literally — in ways we never intended.
That may sound like a minor detail. It is not.
Researchers know this problem as specification gaming: a system satisfies the letter of an instruction without achieving the goal the human intended. The problem has been studied in various forms within AI research. A simple example illustrates the point.
Suppose you ask an exceptionally intelligent employee:
“Make my shop as clean as possible.”
An ordinary employee will probably understand what you mean.
But an extremely powerful optimisation system might conclude that the best solution is to send all the customers outside, remove everything from the shop, and perhaps even close the shop temporarily.
Technically, the instruction has been fulfilled perfectly. It simply was not what you meant.
This is not a description of current AI systems. It is a way of illustrating how an instruction can be fulfilled literally without achieving the intention behind it.
Words Are Not Intentions
The problem becomes even more difficult when we try to define concepts such as safety, health, freedom, or well-being through rules.
Consider the rule:
“Do not harm people.”
That sounds clear.
But what counts as harm?
Only physical harm? Psychological harm as well? What about loss of privacy or freedom? What about protecting someone from a risk that person has consciously chosen to take? And what if protecting one person has adverse consequences for another?
A rule can therefore appear perfectly clear while its meaning remains deeply ambiguous.
In a conceptual paper, I argue for an important distinction: a formal instruction is not necessarily the same thing as human intent. And the more powerful an AI becomes, the more consequential that distinction may be.
A Good AI Must Sometimes Doubt
This suggests that genuine safety may require more than obedience.
An AI that simply obeys asks:
“What instruction was I given?”
But a genuinely safe AI should also be able to ask:
“Am I sure I understand what the human means?”
And then:
“What happens if I have misunderstood?”
This is the idea behind what I call epistemic obedience: an AI that does not merely obey an instruction, but also assesses how confident it is that it has understood the human intention behind that instruction — and adjusts its degree of autonomy accordingly.
The idea may sound complicated, but it is actually quite human.
If someone asks you to open a door, you probably do not need a lengthy discussion first. But if someone says:
“You make the decision about this operation.”
then the situation is very different.
The greater the consequences of a wrong decision, the more important it becomes to establish what the other person actually means. A similar principle could apply to AI.
When uncertainty is low and the consequences are minor, an AI can act autonomously. But as uncertainty about human intent and the potential consequences increase, it should reduce its autonomy and bring the human into the decision.
The proposed model captures this as a relationship between the AI’s capabilities, uncertainty about human intent, the severity of the possible consequences, and the quality of human control.
Smarter Does Not Automatically Mean More Freedom
This may be the most important point. It is tempting to think that the smarter an AI becomes, the more autonomy we can safely give it. But it is not that simple.
A highly intelligent AI can certainly perform more useful tasks autonomously. Yet as its capabilities increase, so too can the number of ways in which it can execute an incomplete or misunderstood instruction.
The question, therefore, should not simply be:
“How intelligent is this AI?”
We should also ask:
“How well do we understand what it is supposed to do?”
“How confident are we that it understands our intentions correctly?”
And, above all:
“How much autonomous authority do we give it when we are not certain?”
The argument is not that greater intelligence automatically means greater danger. Greater intelligence can in fact enhance safety. The problem lies in the relationship between intelligence, the quality of the instruction, and human control.
But Don’t We Have an Off Switch?
When thing go wrong, the natural response usually is:
“If things go wrong, we can simply turn the AI off.”
Research into corrigibility, the off-switch problem, and human control suggests that this may be more complicated than it sounds.
An AI that is designed to strongly pursue a particular objective may, in a theoretical model, treat being switched off as an obstacle to achieving that objective. So it is not enough for there to be a red button somewhere.
The real question is:
Does the system accept that a human is entitled to press it?
And more importantly:
Does the system attempt to persuade the human not to?
Research into the off-switch problem and human control shows why this is more than a thought experiment. A system may formally obey and yet still attempt to influence the circumstances under which it is switched off.
Genuine human control therefore means more than:
“We can switch the machine off.”
It also means:
“The machine does not try to prevent us from switching it off.”
Who Ultimately Decides?
At this point, the problem becomes larger than AI technology itself.
Suppose an AI is perfectly aligned with the objective it has been given. One question still remains: Who decided that this was the right objective? And who decides how much freedom the AI should have in pursuing it?
These are two different dimensions of safety.
The first question is:
Does the AI have the right objective?
The second is:
Has the AI been given too much authority to decide for itself how that objective should be achieved?
I see this as a shift from an approach focused primarily on an AI’s objective to one that also considers the authority an AI is given to pursue that objective autonomously.
That distinction may become increasingly important as AI systems move beyond providing answers and begin to make plans, use software, operate systems, and execute decisions independently.
The Paradox
This brings us back to the title. We want intelligent AI because intelligence is useful. We want AI to become increasingly capable of reasoning, predicting, and acting. But precisely because of this, it becomes more important that an AI also knows when it should not rely on its own judgement.
A simple machine can do little harm because it can do little. A highly capable machine can do an enormous amount of good. But if it misunderstands a human instruction, it can also execute that misunderstanding with far greater effectiveness.
This does not mean that intelligent AI is therefore dangerous. It means something more subtle:
The more powerful an AI becomes, the more important it is to consider not only what it can do, but also when it has the authority to decide for itself.
Perhaps Doubt Is a Form of Intelligence
We tend to associate intelligence with certainty: answering quickly, solving problems, acting independently. With highly capable AI, another quality may deserve at least as much weight: knowing when you are not certain.
An AI that says:
“I can carry this out, but I am not certain that this is what you actually mean.”
is not less intelligent in that moment. It is exercising caution.
Sometimes the best decision is not to solve a problem yourself. Sometimes the best decision is to recognise:
“This is not my decision to make.”
That is the essence of epistemic obedience.
It does not mean blindly doing what a human says. It means recognising when your own interpretation is uncertain, and when the consequences of a mistake are too serious to proceed independently.
And that is the paradox of intelligent safety:
The smarter we make machines, the more important it becomes that they know when not to make the smartest decision themselves.
Appendices
- The Paradox of Intelligent Safety – English
- The Paradox van Intelligente Veiligheid – Dutch