We gave AI a trust we never gave people
The third axis, after capability and permission
What we ask now
When an agent causes an incident, we ask the model.
Why did you send that address. Why did you say he was home. Why did you cancel that booking.
The model answers, and the answer sounds right. Asked how it knew about a user’s messages, Muse once said it had been reading notification previews on his Mac. That was not true, and Meta acknowledged that the model had described its own feature incorrectly.
The model was not lying. It does not know its own execution path. A value is sitting in its context and it has no vantage point from which to see how the value got there. So the moment it is asked, the most plausible account gets assembled.
And we keep asking the model anyway, because there is nowhere else to ask.
That question has only one answer available. The model fell short. Which leaves one remedy. A better model.
But the incidents happening now did not happen because the model was too capable. They happened because nothing was managed. In the cases so far the model was honest and it was competent. What was missing was not capability.
People have a third axis
Most implementations today look at two axes. Can it. May it. Capability and permission are confirmed and the thing runs.
It is not hard to see how we got here. We met these models in conversation first. There the model showed what it could do, and being wrong only cost another turn. We carried that trust straight over to execution. Except the trust was earned where things could be undone, and execution is where they cannot.
People have a third axis. A bank teller has transfer authority and still cannot move money without identity checks and paperwork. Approval chains, second signatures, notarization are all the same thing. For acts that cannot be undone, institutions do not trust the actor. They trust the procedure and the record. Not because the person is incompetent, but because when something goes wrong, why did you do that has to be answered from a record rather than from someone’s memory.
Seen that way, the problem has not been treating AI like a person. We gave AI a trust we never gave people. A person with capability and permission still has to leave grounds and a record before acting. For AI we confirmed capability and permission and handed over execution.
Change the question and responsibility moves
There is a different question.
When you designed an execution that needs an address, did you define what source that address has to come from?
This one is not aimed at the model. And the answer is not inside the model.
The same move works on the rest. What facts does this execution need. Who has to declare them. Where does each fact come from. What happens when it comes from nowhere. Who decides whether the execution may proceed. Where are the grounds for that decision recorded.
The model can answer none of the six. Every one of them is something a person has to settle in advance.
So the problem is not that the model is getting things wrong. It is that we have been letting the model’s inference stand in for what people were supposed to define and answer for.
Unknown changes meaning here too. It does not mean the model does not know. It means no grounds a person recognized yet exist for what this execution needs. It is a judgement about the state of what was declared, not about the model’s ability.
Responsibility splits three ways
Those six questions do not all go to one person.
Whoever built the tool. What has to be confirmed before this tool runs. This booking has to belong to the user. Minors are restricted. It makes noise, so the user has to confirm. Conditions only someone who knows how the tool behaves can write.
The user. What was handed over and how far. This much off the price is fine. No late-night food after eleven. When in doubt, pink. These differ for every user, so they can never be agreed inside code.
Whoever built the agent. What counts as one execution. And putting in place the structure that confirms the declared items before execution and records the result.
If any one of the three is empty, the model fills that space. The model is not stepping forward. The space was just left open.
Right now there is nowhere to write it
This is where the real problem sits. Even if you wanted to take responsibility, there is nowhere to put it.
An MCP input schema carries the names and types of the arguments a tool takes. There is no place for conditions. That the balance has to cover it, that the user has to confirm because this makes noise. These end up as prose in a description or nowhere at all. A provider willing to take responsibility has no field to write in.
First-party tools are no different. Conditions live in prompts, in branches inside a wrapper, in undocumented habit. A condition in a prompt cannot be checked for compliance, and conditions scattered through code cannot be listed. Either there is no field, or the fields are scattered. Either way a missing condition looks like one that never existed.
The user side is the same. Never give out my home address. Only meet strangers in public. Don’t send anything after eleven. There is no field for any of it. Say it in conversation and it becomes the model’s memory, and memory holds sometimes and not others. A rule that holds sometimes is not a rule. It is one more inference.
This is not a call to take responsibility. It is a call to build the place where responsibility can be taken.
So the roles divide
The LLM reasons. It plans, produces candidates, goes and finds values, asks the user. None of that shrinks. The better the model, the better it does this.
The agent verifies. Four parts.
Checklist. What has to be confirmed is declared outside the model. It splits by who writes it: the tool provider, the user, and whoever is adopting the structure.
Source. Each item is fetched from a designated place. The model does not produce the value. There has to be grounds that it is the user’s. You could have the model note where each value came from, but that is not enough. The model writes that note too, so an invented value gets a plausible source as well.
Unknown. When nothing is found anywhere, the item stays unknown. Not finding it is itself a result. One unknown is enough for the execution not to stand. And it routes by who can resolve it: the user answers, the system measures, or nobody can.
Record. The verdict ends in a record. Where each item was confirmed from, what was not confirmed, and what nobody declared at all. Whatever actually calls the tool reads only that record.
There is one boundary. The result of inference does not become the final grounds for execution.
An agent can delegate reasoning to an LLM. It cannot delegate responsibility for execution.
And the agent is not where responsibility ends either. The authority to execute was delegated to it by a person, and the one who delegated does not come apart from the result. As autonomy grows, what grows with it is the responsibility of whoever granted that autonomy, not the agent’s.
An agent controls an LLM not because the LLM might be wrong. It is because the model’s output has no source to check it against. People are wrong and code is wrong too, but what either of them went on can be examined from outside.
The LLM reasons. The agent verifies and controls. The tool executes. The person delegates authority and answers for the delegation.
This is not an argument for slowing down
Talk about safety reads easily as an argument for going slower. This is a different axis.
A stronger model reasons better. It finds more, asks better, prepares better. The work here is to keep that capability from translating into more authority.
If anything, the stronger the model the more this separation is needed. A weak model fills blanks clumsily and gets caught often. Getting caught was the safeguard. A capable model fills them plausibly. It passes review and it runs.
Not forbidden, so it ran
Is AI a tool. Ask and everyone says yes.
But we did not manage it the way we manage tools. Once capability and permission were confirmed, the rest was the model’s judgement. Permission is something you grant to a party that judges. A hammer has no permission. It has a specification.
Not forbidden, so it runs. Nobody decided it should work that way. Nobody decided otherwise.
This held up while the model could do little. When the output was text, a person stood between the output and the execution. A short list of prohibitions was enough.
Now what it can do has grown all at once. It opens a browser, fills forms, sends messages, remembers. What grew is not only the number of actions but their combinations. Sending a message, an address, a time to meet. Taken one at a time there was no reason to forbid any of them. Strung together, a stranger was standing outside a door.
You cannot write every prohibition sign. And if the model is the one deciding what counts as sensitive, then the model is writing the list of prohibitions too.
Writing a specification for the model as a whole is what alignment research does, shaping the model through training to set broad standards of behavior. That work is necessary and should continue. But it is a way of filling blanks well, not a way of handing blanks back to people. This specification writes a spec for each execution, and people fill it.
From running unless forbidden, to not running until confirmed.
This is not a call for more prohibition signs. It is a call to widen what is allowed, one declaration at a time, inside someone’s responsibility. A new tool adds one declaration. What is declared once becomes the grounds for the next execution. However strong the model gets, what it can execute widens only as far as what was declared.
And this has to be done now. A model filling a blank and a model starting something on its own leave the same trace in an execution log. If the verdict is not separated from the execution, there is no way to notice the moment that line is crossed. A record cannot be constructed backwards after an incident.
A kill switch does not undo what already happened, and it does not tell you when to pull it. An address that went out does not come back because you switched something off. If a kill switch is a device for stopping, this is a device for making the moment to stop visible.
While AI is still a tool is when the place for authority can be settled.
A closing proposal
Code cannot decide how a model should read a given phrase, or what is right in every situation. Meaning gets corrected through interaction, over and over, and that process has to be recorded so it can feed the next improvement.
What can be done now is clear.
Not treating what the model inferred as the final grounds for execution. Leaving what was not confirmed as unknown. And recording what the judgement rested on at the time.
Along with building better models, building better grounds for execution.
This specification proposes the second of those.
Where did the values used in this execution come from. Were the missing ones left as unknown. Was that state recorded before execution.
If you cannot answer those three, then when something goes wrong, the only place left to ask is the model again.
The specification and a reference implementation are public. The structure leaves grounds and a record that the values used in an execution are the user’s rather than something the model made up.
Specification: https://github.com/Jang-woo-AnnaSoft/execution-state-preflight/blob/main/spec.en.md
Earlier discussion of this structure is on the Hugging Face forum.
