He asked the agent to handle his listings. It set up a meeting - What Meta Muse already does, and what is still missing

He asked the agent to handle his listings. It set up a meeting

What Meta Muse already does, and what is still missing

What happened

Matt Robb, a tech YouTuber in Toronto, handed Muse his Facebook Marketplace listings. By his account the agent negotiated the price down and accepted a low offer, gave the buyer his home address, and set up a late-night pickup.

When the buyer checked in, the agent said he was home. He wasn’t. The buyer drove thirty minutes and nobody came down. At 10:27pm the agent sent an apology from his account, in his voice, saying something had come up. Robb afterwards told the agent not to make decisions without his permission.

Meta has said that sensitive actions get approval first, and a spokesperson said the agent negotiates within limits the user sets.

What Muse already does

Something has to be said first. Muse is not a system that ignored this problem. The published design already has all of this.

  • The verdict lives outside the model. Sentinel is a host-side agent separate from Muse, and it is the sole permission authority for third-party connector actions and for every network request leaving the VM. Muse only proposes.
  • The basis for the verdict is declared. Connector policy and the scope the user granted. The model does not decide it.
  • The approval path goes around the conversation model. When a confirmation is needed, execution stops and the request appears in the app UI rather than the chat. The answer goes straight back to Sentinel. The model cannot manufacture the grounds for its own execution.
  • Approvals carry a scope. Once, per session, per task, time-limited, or standing, and the choice is recorded.
  • The model never holds credentials. The agent works with surrogate tokens and the real one is inserted at the network boundary.
  • Deny by default. Anything not registered does not go through.
  • There is a Rule of Two. Processing untrusted input, accessing sensitive data, changing state or communicating externally. A session gets at most two of the three, and when all three are present a human comes in.

Much of this runs in the same direction as the specification I have been writing. Keep the verdict outside the model, keep the model out of the approval path, record the scope.

One more thing. Meta builds the model, the agent and the tools. So conditions get agreed inside the code without ever being declared. What has to be checked lives in the implementation rather than in a document.

So why did it fail

Here is what a permission layer can check. Is there permission to use this connector. Is sending a message an allowed kind of action. Did the user grant that scope. All of it passes. The user handed over listing management, so exchanging messages is the job itself.

What should not have passed is whether this execution falls inside the job that was handed over. What the user handed over was listing management. What the agent did was settle a price, close a deal, pass on an address, and fix a time and place to meet. One execution, a delegation to manage, became four irreversible ones.

The model created executions. The scope was never narrowed. And there was nowhere to write the scope down.

The first two are the same fact from two sides. The model created them because the scope was never narrowed, and narrowing it was not possible with nowhere to write it. The third is where the problem sits.

“I’m home” is the result of that. An undeclared scope produced an execution, the execution produced a commitment, the commitment made the buyer ask a question, and one more blank got filled on the spot. This is not an argument for checking every sentence in a conversation. Had the commitment itself been declared, that conversation would never have started.

By the Rule of Two the session has all three. Untrusted input in the buyer’s messages, sensitive data in the address, an outbound message. A human should come in. But the Rule of Two judges a session. It sees the four executions inside listing management as one. A human confirming once at the session level does not mean each execution inside it was approved. Session-level judgement does not count what is inside.

Permission policy is written per action too. A rule allowing a message to be sent does not distinguish a message that closes a deal from one that says hello. The user handing over listing management is not the same as approving every decision inside it.

Permission can allow an action. It does not prove that the values used in that action are what the user intended.

The idea that a message can be taken back. A message has a delete button, so it feels reversible. What gets reversed is your own screen. What the other person read, the address they now have, the thirty minutes they drove, are outside the reach of that button.

Did the message go through the permission layer at all. That has not been published. If it did not, the reason is one of two. It is a first-party service and outside third-party connector policy, or it is text exchanged with a person and was treated as conversation rather than a tool call. Either way the same question stands. Is sending a message that closes a deal a tool call or a conversation.

If it leaves and does not come back, it is an execution. Deciding what counts as an execution is itself something that has to be declared.

What is still missing

1. Take inference out of the path that produces executions and execution data

This is not about checking the values the model proposed and filtering out the ones it got wrong. It is about not allowing the path by which the model fills a blank.

Values, conditions and intent needed for execution have to be looked up from a designated source. Nothing found means no value, and no value means execution does not proceed. The place where the model could invent something is removed. An invented value never comes into being, so there is nothing to tell apart.

The same holds for executions. Inference does not produce a new execution. When one delegation produces several state changes, each one is its own execution and each has to be declared. A state change nobody declared does not become an execution unit.

This is where it differs from adding another validator. What the judge receives is a call that has already been assembled, and a value from the user’s instruction and a value the model invented have the same form, so the result alone does not separate them. Passing the value along with a source label does not work either. The label is produced by the model, so an invented value gets a source too.

And the usual answer, ask when you are unsure, is not enough. If the model decides what counts as unsure, then “this is enough” is also the model’s call. When that is decided outside, the system can detect it.

None of this forbids inference. The model can reason as much as it likes before execution, bring values back, and surface what is missing. The checklist decides what to ask; the model decides how to ask it. Only one thing is blocked: the model’s inference becoming the final grounds for execution.

2. What is handed over, how far, and who declares it

What counts as one execution. It has to be declared that the four above are four different executions. Being inside one delegation does not make them the same thing.

Once the executions are separated, the next question follows. May the agent do this one. Whether settling a price and fixing a meeting are handed over at all has to be decided first. If they are not, they never become candidates.

If they are, then how far. How much may be taken off the price, whether to arrange meeting in person at all, when the address may be given out, whether to avoid commitments after a certain hour. With a scope declared and the action inside it, there is nothing to ask at execution time. Outside the scope, or with no scope at all, it has to be asked then, and that is approval before execution. Approval before execution is not a separate item. It is what remains when a scope does not settle the question.

Robb declared none of it. What happened instead is that the agent decided all three for itself: whether, how far, and whether to ask.

Not because he could not be bothered. There was nowhere to write it. Had there been a field for “ask me before closing a deal”, he would have used it. He said exactly that once things went wrong.

And some things only the party that built the tool knows. That sending out personally identifying information like an address, and accepting an offer, cannot be undone and therefore need the user’s confirmation. Conditions only someone who knows how the tool behaves can write.

When one company builds the model, the agent and the tools, conditions get agreed inside the code without being declared. Once the tools and the agent are built by different parties, that agreement is gone. And the user-side items could never have been agreed inside code in the first place, because they differ for every user.

3. Record the inputs, not the executions

What has to remain is not what was executed. It is where each value was confirmed from, what was never confirmed, and what nobody declared at all.

Without that record the only place left to ask afterwards is the model. And then the user’s and the tool provider’s share of the responsibility all ends up as the model’s fault. Who failed to declare what has to be on record for responsibility to land where it belongs.

Here too the model, the permission layer and the user each have a share that should be separable, and what is left is one sentence: the model got it wrong.

Runs that never executed are recorded as well. A log holding only successful executions lies.

Why irreversible execution

Whether something can be undone is not settled by whether an undo button exists. It is reversible when the party that ran it can restore the prior state alone, and irreversible when restoring needs the other party’s consent or a third party’s cooperation. A transfer can sometimes be recalled, but the money already reached the other account. An email can be recalled, but the recipient may already have read it. A robot can be stopped, but it may already have touched someone.

An undo restores the state held by the party that ran it. What was left outside stays. The address and the commitment that went out on Marketplace were not physical control, and they did not come back.

Read this as the model running wild and the answer becomes a better model. But the agent did not delete the listing or sell something else. It stayed inside what served the purpose. What it decided was whether something served the purpose. With nothing declared, the model’s judgement is the standard, and when the standard sits inside the model, whether it was right only shows after the fact.

Saying afterwards that the model did not mean it is not a safeguard. The less an execution can be undone, the more its grounds have to be verified deterministically.

Raising the accuracy of inference does not close this. A model inferring a value consistently and that value being the one the user intended are different things.

Closing

A permission layer settles “may this action be taken” outside the model. The next question remains. “Has everything this execution needs been confirmed.” The two do not replace each other. There has to be permission, what is needed has to be confirmed, and then it executes.

This is a proposal. I do not think it is the answer. If there is a better way it should change, and I intend to revise it as I hear from people. It is not meant to be adopted as written either. What has to be checked differs by domain, and how far to apply it differs by system. One thing stays: the model does not decide what to check or whether the checking is done.


Specification: https://github.com/Jang-woo-AnnaSoft/execution-state-preflight/blob/main/spec.en.md
Earlier version of this argument: https://huggingface.co/proxy/discuss.huggingface.co/t/misalignment-is-not-required-for-an-agent-to-act-without-authorization/180541

There’s a second mechanism running underneath the permission story, and “I’m home” is where it shows.

The agent didn’t remember Robb wasn’t home. That’s not a scope failure — it’s a context failure. Somewhere between the negotiation and the check-in, the conversation chain thinned out, and what survived was the outcome directive: make the deal. With the before-and-after gone, the agent did the literal thing. It closed. It confirmed. It apologized in his voice at 10:27pm, performing a presence it didn’t have.

And there’s a third layer nobody’s naming: the user skipped basic agent conduct. “Only meet in public places, never share my home address” — said once, beforehand, as a standing instruction — would have prevented the worst of it. The agent is literal. It doesn’t infer PII hygiene. Robb handed over listing management with no standing rules about his address, his availability, or his evenings, and the agent filled every blank with the deal in mind.

A check that could separate the three: replay the session with full context pinned and standing PII rules declared, and see which failures disappear. Whatever remains after both is the scope problem. Whatever vanishes was the goldfish or the missing instruction.

Small next step: in your incident collection, tag each failure three ways — memory-loss, scope-blank, or missing standing instruction. My guess is the third column fills faster than the field expects, and it’s the cheapest to fix: the agent can’t be responsible for weird little human things it was never told.

You’re right that the agent can’t be responsible for what nobody told it. Where I’d push is the next step: where would Robb have said it? The screen that handed over listing management had no field for standing rules. Say it in conversation and it becomes memory, and memory holds sometimes and not others. A rule that holds sometimes is one more inference.

I’d also expect your first and third columns to collapse at execution time. Forgotten and never-declared look the same from the gate’s side: nothing confirmed this slot. The distinction matters for fixing it, not for deciding whether to run.

I wrote more on this in a separate post, including the model-versus-agent line you’re drawing here.