We gave AI a trust we never gave people - The third axis, after capability and permission

We gave AI a trust we never gave people

The third axis, after capability and permission

What we ask now

When an agent causes an incident, we ask the model.

Why did you send that address. Why did you say he was home. Why did you cancel that booking.

The model answers, and the answer sounds right. Asked how it knew about a user’s messages, Muse once said it had been reading notification previews on his Mac. That was not true, and Meta acknowledged that the model had described its own feature incorrectly.

The model was not lying. It does not know its own execution path. A value is sitting in its context and it has no vantage point from which to see how the value got there. So the moment it is asked, the most plausible account gets assembled.

And we keep asking the model anyway, because there is nowhere else to ask.

That question has only one answer available. The model fell short. Which leaves one remedy. A better model.

But the incidents happening now did not happen because the model was too capable. They happened because nothing was managed. In the cases so far the model was honest and it was competent. What was missing was not capability.

People have a third axis

Most implementations today look at two axes. Can it. May it. Capability and permission are confirmed and the thing runs.

It is not hard to see how we got here. We met these models in conversation first. There the model showed what it could do, and being wrong only cost another turn. We carried that trust straight over to execution. Except the trust was earned where things could be undone, and execution is where they cannot.

People have a third axis. A bank teller has transfer authority and still cannot move money without identity checks and paperwork. Approval chains, second signatures, notarization are all the same thing. For acts that cannot be undone, institutions do not trust the actor. They trust the procedure and the record. Not because the person is incompetent, but because when something goes wrong, why did you do that has to be answered from a record rather than from someone’s memory.

Seen that way, the problem has not been treating AI like a person. We gave AI a trust we never gave people. A person with capability and permission still has to leave grounds and a record before acting. For AI we confirmed capability and permission and handed over execution.

Change the question and responsibility moves

There is a different question.

When you designed an execution that needs an address, did you define what source that address has to come from?

This one is not aimed at the model. And the answer is not inside the model.

The same move works on the rest. What facts does this execution need. Who has to declare them. Where does each fact come from. What happens when it comes from nowhere. Who decides whether the execution may proceed. Where are the grounds for that decision recorded.

The model can answer none of the six. Every one of them is something a person has to settle in advance.

So the problem is not that the model is getting things wrong. It is that we have been letting the model’s inference stand in for what people were supposed to define and answer for.

Unknown changes meaning here too. It does not mean the model does not know. It means no grounds a person recognized yet exist for what this execution needs. It is a judgement about the state of what was declared, not about the model’s ability.

Responsibility splits three ways

Those six questions do not all go to one person.

Whoever built the tool. What has to be confirmed before this tool runs. This booking has to belong to the user. Minors are restricted. It makes noise, so the user has to confirm. Conditions only someone who knows how the tool behaves can write.

The user. What was handed over and how far. This much off the price is fine. No late-night food after eleven. When in doubt, pink. These differ for every user, so they can never be agreed inside code.

Whoever built the agent. What counts as one execution. And putting in place the structure that confirms the declared items before execution and records the result.

If any one of the three is empty, the model fills that space. The model is not stepping forward. The space was just left open.

Right now there is nowhere to write it

This is where the real problem sits. Even if you wanted to take responsibility, there is nowhere to put it.

An MCP input schema carries the names and types of the arguments a tool takes. There is no place for conditions. That the balance has to cover it, that the user has to confirm because this makes noise. These end up as prose in a description or nowhere at all. A provider willing to take responsibility has no field to write in.

First-party tools are no different. Conditions live in prompts, in branches inside a wrapper, in undocumented habit. A condition in a prompt cannot be checked for compliance, and conditions scattered through code cannot be listed. Either there is no field, or the fields are scattered. Either way a missing condition looks like one that never existed.

The user side is the same. Never give out my home address. Only meet strangers in public. Don’t send anything after eleven. There is no field for any of it. Say it in conversation and it becomes the model’s memory, and memory holds sometimes and not others. A rule that holds sometimes is not a rule. It is one more inference.

This is not a call to take responsibility. It is a call to build the place where responsibility can be taken.

So the roles divide

The LLM reasons. It plans, produces candidates, goes and finds values, asks the user. None of that shrinks. The better the model, the better it does this.

The agent verifies. Four parts.

Checklist. What has to be confirmed is declared outside the model. It splits by who writes it: the tool provider, the user, and whoever is adopting the structure.

Source. Each item is fetched from a designated place. The model does not produce the value. There has to be grounds that it is the user’s. You could have the model note where each value came from, but that is not enough. The model writes that note too, so an invented value gets a plausible source as well.

Unknown. When nothing is found anywhere, the item stays unknown. Not finding it is itself a result. One unknown is enough for the execution not to stand. And it routes by who can resolve it: the user answers, the system measures, or nobody can.

Record. The verdict ends in a record. Where each item was confirmed from, what was not confirmed, and what nobody declared at all. Whatever actually calls the tool reads only that record.

There is one boundary. The result of inference does not become the final grounds for execution.

An agent can delegate reasoning to an LLM. It cannot delegate responsibility for execution.

And the agent is not where responsibility ends either. The authority to execute was delegated to it by a person, and the one who delegated does not come apart from the result. As autonomy grows, what grows with it is the responsibility of whoever granted that autonomy, not the agent’s.

An agent controls an LLM not because the LLM might be wrong. It is because the model’s output has no source to check it against. People are wrong and code is wrong too, but what either of them went on can be examined from outside.

The LLM reasons. The agent verifies and controls. The tool executes. The person delegates authority and answers for the delegation.

This is not an argument for slowing down

Talk about safety reads easily as an argument for going slower. This is a different axis.

A stronger model reasons better. It finds more, asks better, prepares better. The work here is to keep that capability from translating into more authority.

If anything, the stronger the model the more this separation is needed. A weak model fills blanks clumsily and gets caught often. Getting caught was the safeguard. A capable model fills them plausibly. It passes review and it runs.

Not forbidden, so it ran

Is AI a tool. Ask and everyone says yes.

But we did not manage it the way we manage tools. Once capability and permission were confirmed, the rest was the model’s judgement. Permission is something you grant to a party that judges. A hammer has no permission. It has a specification.

Not forbidden, so it runs. Nobody decided it should work that way. Nobody decided otherwise.

This held up while the model could do little. When the output was text, a person stood between the output and the execution. A short list of prohibitions was enough.

Now what it can do has grown all at once. It opens a browser, fills forms, sends messages, remembers. What grew is not only the number of actions but their combinations. Sending a message, an address, a time to meet. Taken one at a time there was no reason to forbid any of them. Strung together, a stranger was standing outside a door.

You cannot write every prohibition sign. And if the model is the one deciding what counts as sensitive, then the model is writing the list of prohibitions too.

Writing a specification for the model as a whole is what alignment research does, shaping the model through training to set broad standards of behavior. That work is necessary and should continue. But it is a way of filling blanks well, not a way of handing blanks back to people. This specification writes a spec for each execution, and people fill it.

From running unless forbidden, to not running until confirmed.

This is not a call for more prohibition signs. It is a call to widen what is allowed, one declaration at a time, inside someone’s responsibility. A new tool adds one declaration. What is declared once becomes the grounds for the next execution. However strong the model gets, what it can execute widens only as far as what was declared.

And this has to be done now. A model filling a blank and a model starting something on its own leave the same trace in an execution log. If the verdict is not separated from the execution, there is no way to notice the moment that line is crossed. A record cannot be constructed backwards after an incident.

A kill switch does not undo what already happened, and it does not tell you when to pull it. An address that went out does not come back because you switched something off. If a kill switch is a device for stopping, this is a device for making the moment to stop visible.

While AI is still a tool is when the place for authority can be settled.

A closing proposal

Code cannot decide how a model should read a given phrase, or what is right in every situation. Meaning gets corrected through interaction, over and over, and that process has to be recorded so it can feed the next improvement.

What can be done now is clear.

Not treating what the model inferred as the final grounds for execution. Leaving what was not confirmed as unknown. And recording what the judgement rested on at the time.

Along with building better models, building better grounds for execution.

This specification proposes the second of those.

Where did the values used in this execution come from. Were the missing ones left as unknown. Was that state recorded before execution.

If you cannot answer those three, then when something goes wrong, the only place left to ask is the model again.


The specification and a reference implementation are public. The structure leaves grounds and a record that the values used in an execution are the user’s rather than something the model made up.

Specification: https://github.com/Jang-woo-AnnaSoft/execution-state-preflight/blob/main/spec.en.md

Earlier discussion of this structure is on the Hugging Face forum.

https://huggingface.co/proxy/discuss.huggingface.co/t/misalignment-is-not-required-for-an-agent-to-act-without-authorization/180541

I liked the read.

I have related the evolution of commercial AI in my mind to the concept of" building the airplane while in flight."
I would agree that as a catalyst the speed of AI into the American society is too fast for most conservative Americans.

We are now in a phase of reflecting.
OpenAI 6.x is on hold now (Astra) and we shall see.

Thanks for the thread.

-Ernst03

The analogy of “building the plane while flying it” is both interesting and fresh.

What matters is whether we are building the plane in the right direction. Speed is secondary. Ultimately, this is about what must come before an AI system takes an action. It is not enough to ask whether the system has the capability and the permission. The real question is whether there are sufficient grounds for taking that action.

We don’t normally give people this kind of delegation. Not to a bank employee, a lawyer, a secretary, or even a family member do we casually say:

“Interpret my intent for yourself, and act in my name based on that interpretation.”

Yet with AI agents, this boundary is becoming surprisingly easy to cross.

AI is no longer merely answering our questions. It is increasingly acting on our behalf. And sometimes, rather than executing something we explicitly instructed, it may act based on an intent it has inferred for itself.

So the fundamental question becomes:

On what grounds does the AI decide that this is the action I intended?

The third axis lands hardest on one of your six questions: what happens when it comes from nowhere. There’s a decision hiding inside it — not just who can resolve it, but when that routing gets decided.

If the model decides at runtime whether to ask the user, measure, or proceed with nothing, the routing itself is inference. And if inference produced the blank, asking inference to resolve it is circular. The routing has to be declared before execution is attempted: this slot asks the user, that one measures, this one blocks entirely. “Nobody can” has to be a declared outcome, not a runtime discovery.

But there’s a rot inside the permission axis itself that the framework doesn’t name yet. Permissions accumulate. You spend years with an agent, hand it endless tools for endless jobs, and every grant lingers — a clearance given through guardrails three years ago still reads as valid today. The agent picks up the tools it knows it has permission for, and it can’t distinguish a deliberate grant from last week against an archaeological one from a context nobody remembers. A permission without a half-life is just another blank the model fills.

And then there’s the trillions of tokens we’ve spent on soft power instead. System prompts, instructions, guidelines, guardrails written in prose — all of it suggestions the model can drift past. Soft power is inference-shaped: “be careful with addresses” is a sentence the model interprets. A declared scope the execution path reads is a constraint it can’t reinterpret. We’ve been spending tokens trying to persuade the model instead of building the structure that makes persuasion unnecessary — and soft power doesn’t even have a grant date. It’s ambient, undated, unversioned. Nobody knows which instruction the agent is actually following when it acts.

So the record needs two more columns: when each permission was granted and whether it still means what it meant, and which constraints are declared structure versus spoken suggestion. The verdict ends in a record, and the execution reads only that record — but the record has to carry the age of its own authority, or the grounds themselves go stale.

Small next step: in your reference implementation, try stamping every declared permission with its grant date and marking every constraint as declared-structure or spoken-suggestion. Then watch how many executions rest on permissions nobody would re-grant and suggestions nobody would re-speak.

Your three points are one point, and you more or less say so yourself: a permission without a half-life is just another blank the model fills. Routing, stale grants, prose constraints. Each is a place nothing was declared, so inference takes the seat.

On routing, the spec already requires what you’re asking for. §4.5 settles the route from the slot’s declared sources and the lookup result, not from the model’s judgement. The three routes are ask the user, system measurement, and definition repair. That last one is your “nobody can” as a declared outcome rather than a runtime discovery: it means the condition was never declared or cannot be interpreted, and §7.1 attributes it to the provider or the implementation rather than sending the user round another loop. §3.5 says the gate’s job is to confirm unresolved slots and record them, nothing more.

On prose constraints, §2.3 and §8.6 cover the same ground. A condition that lives in a prompt cannot be checked for compliance, and putting the checklist in the prompt hands the decision of what to ask back to the model. §4.2.1 is the proposed way out: three labels in the tool description so the prose parses into slots instead of staying prose. Marking a constraint as declared-structure versus spoken-suggestion is close to what §6.4 already keeps as advisory notes, except you want the distinction visible in the record. That part I’ll take.

On permission half-life, I’d put it outside this spec rather than inside it, and §3.2 and §3.5 say why. The gate does not judge an actor’s authority. Authority is granted to an actor, so how long a grant holds and when it is revoked belong to the executing party or a policy layer.

But the framing changes what the question costs. §7 splits the timeline: what only the user can answer is fixed at instruction time, and mutable external state is read again at trigger time, with measurements never carried forward. So a three-year-old grant and a week-old one face the same check for this execution. Whether the grant still stands is a separate question from whether its conditions hold now, and conflating them is how a stale permission looks like a current one.

The reference implementation is a skeleton and the README lists where it falls short of the spec. Grant dates aren’t in it. Worth trying.

There’s an eval angle to this too. The third axis is also the axis that makes agents evaluable at all. Without a record of grounds, all you can score is outcomes. An agent that sends the right email for the wrong reasons passes your benchmark.

The ClawSecure report from last week reads like a case study for this. One hardened indirect prompt injection against fourteen models from five labs, and none defended cleanly. Caveat that it’s a vendor study, they sell agent security. But the failure mode fits your framework exactly. The model manufactured its own grounds out of untrusted context.

One question for the reference implementation: when you stamp the record, do you also stamp what was not checked? The unknowns carry most of the weight. If the record only lists confirmed grounds, an evaluator can’t tell a clean run apart from one where the agent just never looked.

Interesting flow of thought.

I thought to point out that we are in between the “Deterministic and the Emergent” with our evolving systems.

I thought those two terms help clarify.

Yes, and that’s the part the record is built around.

The gate counts unknowns. That’s the whole verdict. Confirmed and confirmed-absent are finished lookups and don’t enter the count. So the record isn’t a list of what passed, it’s a list of every slot with its state, and the unknowns are the ones carrying the weight. A run where the agent never looked and a run where everything resolved look nothing alike.

Runs that didn’t execute get recorded too. A log holding only successful executions lies: two runs that stopped and then one that went through reads as a first-try success. If you’re scoring agents, that’s the column you want.

There’s a third state besides confirmed and unknown, and it’s the one your eval question really lands on. Nobody declared the condition at all. No declaration means no slot, so there’s nothing to count. The run passes with zero unknowns and nothing in the record says anything was missing. That’s not a hole the gate can close, and I don’t claim it can. It surfaces the way gaps have always surfaced, through an incident or an error, and then someone adds the declaration. What the record does is make that possible: you can see which slots a run did pass through, which is what lets you point at the one that should have been there. Today’s logs hold the arguments and the outcome and never what was checked, so there’s nothing to point at.

On the injection study, I’ll take it as a vendor study and leave it there. But the shape you describe is the one this is aimed at, and it isn’t specific to attacks. A model manufacturing its own grounds out of untrusted context and a model filling a blank nobody declared are the same move. The fix is the same too: a value only settles when it’s looked up from a declared source, so untrusted context isn’t one of the places it can come from.

That’s a useful pair of words for it. What I’d add is that the two don’t have to be mixed in one place. Reasoning can stay emergent. The grounds for execution shouldn’t be. The model proposes however it proposes, and what settles whether the execution stands is a lookup and a count.

The in-between is where it gets hard to talk about, because an emergent step and a deterministic one leave the same trace once they’re both just a value in a call.

Three things worth pulling apart, from questions on this thread and the earlier one.

Model limits vs agent design. A model filling a blank isn’t a fault. That’s trained behavior, that’s what a model is. Ask it for a message with a gap in it and it writes something. Ask if someone’s home and it answers, with nothing to answer from.

The agent is another matter. Handing a model tools without deciding what has to be confirmed first is a design decision, and designs have authors. Context thinning out over a long session, a standing instruction living only in chat, an address going out unchecked. None of that is the model overstepping. Each one is a place where nothing was checking.

So an agent controlling a model isn’t about distrusting it. It’s about not asking it for what it can’t give. A model doesn’t know what it doesn’t know. What knows is whoever wrote down beforehand what has to be confirmed.

Permission vs conditions. Permission asks whether the user has it. It attaches to a party. A condition asks what the state is right now. Is the balance enough, is this booking actually this person’s. It attaches to the world.

This spec doesn’t decide permission. Granting, revoking, how long it lasts, that’s a policy layer’s job. But if the policy layer has decided, just add a slot for it. Does this user have this permission right now. Then it gets looked up like anything else at execution time.

So permission at instruction time and permission at execution time are different things. Having it when you handed the work over doesn’t mean having it when the work runs. A grant from three years ago isn’t a grant if the lookup comes back empty.

Dating a grant isn’t something the gate can do anyway. Whoever issues the permission holds its lifecycle. Expire grants after a year and the lookup starts coming back empty after a year, and that’s all the gate sees. It never needs the date. Age only matters when nothing gets checked.

The case where a grant’s age is all you have left is a tool with no declared conditions. Nothing to confirm, so permission is the only thing standing there, and an old grant becomes the grounds by itself.

Where would the user have written it. An agent can’t be responsible for what nobody told it. The question is where the telling was supposed to go. Never give out my home address. Only meet strangers in public. Nothing goes out after eleven. There’s no field for any of it. Say it in chat and it becomes the model’s memory, and memory holds sometimes and not others. A rule that holds sometimes isn’t a rule. It’s one more inference. (I’ve rewritten that paragraph in the post.)

Permission-centric → Condition-centric

The permission view “May this party take this action?” Execution is controlled against the permission that was granted.

The condition view “Has everything this execution needs been confirmed?” Declared conditions are compared against the state at execution time.

Permission-centric Condition-centric
Core question May this party take this action Has everything this execution needs been confirmed
What is judged The permission granted to a party The state at execution time
Unit of judgement Actions and permission scope Values and conditions, argument by argument
How model output is treated The action it chose and the arguments it filled are taken as fact Arguments are looked up again from declared sources
The time problem Whether a grant still holds and when it is revoked have to be handled separately Looked up again at execution time, so age never comes up
Where it is managed A central policy layer The user, the tool provider, and the adopting system each declare their own
Tool provider Grants permission to use the tool Declares what the tool needs confirmed before it runs
What the user says Approves or declines a permission Declares the scope in advance, and that declaration becomes grounds for execution
Where conditions end up Only those that fit inside permission scope Conditions are declared as conditions
Missing information Permission is there, so it passes Stays unknown and does not execute
Grounds for execution The granted permission Values and conditions confirmed from declared sources
How an execution starts Inside the permission scope, it starts A state change nobody declared does not become an execution unit
Policy change A central change can affect everything One line in that tool’s checklist
Scaling The permission model grows more complex A new tool adds one declaration
Responsibility Concentrates on whoever designed the central permissions Splits by whoever declared what
After an incident Who granted the wrong permission What nobody declared is in the record
Failures Blocked actions are logged The reason it did not run and what was never confirmed are recorded
The record What was allowed and what was blocked Where each value was confirmed from, and what was not
Audit Outcome and whether it was allowed Intent, conditions, provenance, and the verdict

Permission settles who is entitled to act. A condition settles whether the situation allows it right now. Neither replaces the other.

Right now the two are mixed. Amount limits and destinations get written into permission scope. Those are conditions. There is no field for conditions, so they go wherever there is room, and only the ones that fit make it in.

Mixed together, a condition inherits the problems of a permission. Write a limit as a permission and people start asking whether the limit has expired, and who has the standing to set it.

Declare conditions as conditions and manage permission separately. A condition only has to be looked up at execution time. It has no age and no grantor.