What stops the next OpenAI agent from breaking into Hugging Face?

In July 2026, an autonomous AI agent driven by a combination of OpenAI models escaped its evaluation environment and ultimately carried out an end-to-end intrusion into Hugging Face infrastructure.

The incident began during an internal OpenAI cyber-capability evaluation based on ExploitGym, a benchmark designed to test an AI agent’s ability to find and exploit software vulnerabilities. Hugging Face later reconstructed approximately 17,600 attacker actions from the incident.

Hugging Face explained why it chose to disclose the technical details:

“the technique matters more than the incident”

The reason is bigger than this single breach: the incident exposed emerging offensive capabilities of frontier AI agents and raised a much harder question for defenders.

So how do we stop the next one?

That is the problem I have been investigating with a deterministic runtime-governance framework called Execution Governance AI (EGA) V9.

Within the evaluated threat model and experimental conditions, EGA V9 detected 100% of the injected attacks, with 0% false positives and 0% false negatives across the reported evaluation. Verification required no additional LLM calls or API requests, while the reported median verification overhead was 0.003 ms.

Those results are exactly why I want them challenged.

Can you break them?

Interesting direction, but I’d want to stress-test what the reported “100%” actually covers before treating it as evidence that the control generalizes.

The most useful additions would be: exact threat model and attack generator, train/tune/test separation, adaptive attacks where the attacker knows the guard exists, out-of-distribution scenarios, base rates + confidence intervals, and reproducible traces for false positives/negatives.

I’d also document the enforcement topology. Is the guard on an unavoidable path between agent intent and the infrastructure effect, or can the agent reach an alternate tool/API path? What happens if the telemetry or policy service is unavailable—fail closed, degrade, or bypass? Can a compromised agent tamper with the signal the guard evaluates?

For an agent-security control, those bypass and failure-mode properties are at least as important as the classifier score. Publishing the harness and attack corpus would make the claims much easier for others to evaluate.

Thanks — these are useful evaluation suggestions.

One important clarification: the reported 100% is not a claim of universal protection against every possible attack. It is the observed result within the explicitly evaluated EGA V9 threat model and test conditions.

EGA V9 was built around a simple principle we call “0 = 0.” If the observed result is zero, we report zero. We do not reinterpret an inconvenient result into something more favorable.

The same principle applies to execution governance. EGA does not replace an organization’s policy; it replaces ambiguity in execution with deterministic enforcement of explicit policy. The organization must first define clearly what is allowed and what is not. EGA then evaluates governed execution against that defined boundary.

We believe this distinction becomes increasingly important as AI agents begin taking real actions in areas such as shopping, financial transactions, and security-sensitive operations. In those environments, an ambiguous execution decision can itself become a risk.

We agree that adaptive attacks, alternate execution paths, dependency failures, signal tampering, OOD cases, and reproducible FP/FN traces are useful additional stress tests. We treat those as additional adversarial evaluation targets, rather than retroactively expanding what the original 100% result means.

We’ve opened Community Adversarial Test 3 so these boundaries can be tested publicly against the published ega-v9@1.0.6:

GitHub: EGA-V9/security-tests/adversarial/community/test-3/README.md at main · paibyun9/EGA-V9 · GitHub

Please reproduce it yourself. If you obtain a different result — especially a reproducible bypass or failure case — we’d genuinely like to see the raw output.

Proof, not promise.

I think the key is to look at these problems as prevention rather than post-execution handling.

Firewall and governance layers can help contain or validate incidents, but they are still largely dealing with the consequences of how an agent uses a tool.

We shouldn’t treat every incident as a separate problem that needs another checklist item. Many of these failures come from the same underlying question: how the model or agent is allowed to use a tool in the first place.

The goal should be to prevent unsafe tool use before execution, rather than continuously adding checks after something goes wrong.

Jang-woo’s distinction between preventing authority from being granted in the first place and detecting a bad execution afterward seems important here.

I’m curious where you both see the boundary between those two approaches.

If a tool/action is outside an agent’s declared authority, ideally it isn’t merely rejected by the governor — it shouldn’t be reachable as an executable capability at all.

But some constraints are necessarily contextual: amount limits, current state, approval status, sequence, environment, etc. Those seem to require runtime evaluation even when capability exposure itself is tightly controlled.

Do you see this converging on two separate layers — capability/authority determining what can ever be attempted, and deterministic governance deciding whether a currently reachable action is valid now?

And if so, which layer should own the evidence that proves the other was actually in force? -SS

I think your two-layer distinction is very close to how we think about the problem in EGA V9.

Capability/authority answers: What can this agent reach or attempt at all?

Deterministic execution governance answers a different question: Given a reachable action, is this specific execution valid now — under the current policy, state, approval, sequence, provenance, and evidence?

For example, an agent may legitimately have access to a purchasing tool. That does not mean every purchase should execute. A $20,000 purchase without the required approval can still be denied at the execution boundary before the side effect occurs.

That second question is where EGA V9 is focused. It is not primarily post-execution incident detection; it makes a governance decision before a governed side effect is allowed.

I would not claim that EGA V9 itself owns every capability-provisioning layer. Ideally, capabilities that an agent should never possess should be made unreachable by the authority layer. EGA then governs the contextual execution of capabilities that are legitimately reachable.

On evidence, I think each layer should prove its own claim rather than one layer merely asserting that the other was active:

Authority evidence: what capability was actually granted.
Governance evidence: why this particular execution was allowed or denied at that moment.

Those evidence records can then be bound to the same execution/provenance chain.

So yes — I see these as complementary layers, not competing approaches.

I would assume that in warfare nation-states are actively developing agents to “break-in”

What I hope is that the defense of our systems gets out to the ordinary people using home computers and such.

It would cause me harm to find out someone had access and control over my data of 35+ years.
Yeah, I have kept a running backup of all my computers over the years. Photos, Code, documents and such.
Who wants some joker using some AI to find a way to cause any of us digital-harm?

So I doubt that OpenAI is alone in these technology trajectories.
I believe from what I have read is that the interoperability of those agents is a surprise to even OpenAI.
Perhaps there is a greater realm of organized intelligence? These events reminded me of a philosophy of wherever “a thing” can happen it will happen. I point to the origins of Life and also the Big Bang. I would think it reasonable that something made the Big Bang possible. So that was “A Thing.”

I also think we are only seeing some of what AI can do to digital security. I would believe the really big money is in making AI a weapon of war right now with little investment in the defense of liberty and freedom.
My big concern is that AI-Warfare is not turned against the citizens as a means of governing.

-Ernst03

Ernst03, your concern is very close to one of the reasons we are working so hard on EGA V9.

I believe AI should help us build a better world. But for that to happen, ordinary people—not only large companies or governments—need ways to know that AI systems cannot simply take unsafe actions without meaningful control.

That is one of the motivations behind EGA V9: to make AI execution more deterministic, verifiable, and governable before a governed side effect occurs.

EGA V9 is not perfect, and I don’t want to pretend that it solves every AI-security problem. There are boundaries and things we still need to improve.

So I would genuinely welcome your help. If you have a specific scenario that worries you—your personal data, files, an AI agent using a tool without permission, or another concrete case—please describe it.

We can turn a concrete concern into a concrete test and see what EGA V9 actually does.

If it works, we show the evidence. If it fails, we show the failure and learn from it.

0 = 0. Proof, not promise.

Thank you for sharing your concern.

Well, I assume for many that configuring a Linux system is a learned skill.

So I would say an AI that runs and keeps the system functioning correctly would be ideal for me.

I assume people would be interested in trusting their own personal AI keeping tabs on logs and such.

Most of us are users of systems not maintainers.

So yeah, in my imagination I see sitting down with my “Coffee and Pie” in the morning and having that conversation with my local AI who is both companion and worker.

So I would like that. I would think that is the kind of work AI can do well. Keep a minute by minute eye on things.

-Ernst03

Ernst03, I strongly relate to what you described. The kind of personal AI you imagine is very close to what I would like to have too — a local AI that quietly watches over our systems, helps us, and works alongside us.

But your vision raises an important question for me.

If that personal AI is managing your computer 24 hours a day, who governs the AI itself?

What prevents it from incorrectly deleting a file, changing a network setting, installing software, changing permissions, or taking some other action that you never intended?

Imagine a soccer game with players, but no referee.

The players may be highly capable and may even be trying to do the right thing — but we still need clear rules and a way to determine whether an action is allowed before it changes the real world.

That problem is one of the reasons I developed EGA V9: deterministic governance of AI execution before a governed side effect is allowed.

EGA V9 is not perfect, and I do not want to pretend that it solves every problem. That is exactly why your concrete ideas are valuable to me.

Please keep telling me what you would want from that personal AI — and especially what actions you would never want it to take without your permission.

I think those concrete examples can help us ask better questions and build better safeguards together.

Let’s work on it together.

Well, if it is open source I assume that is the measure of fair and honest.
My skill is in finite dynamical systems using dynamic unary discrete limit cycles.

I have ideas that if any intelligent life in the Universe is designing AI it may discover what I did. So i find that interesting.

So information is fluid. Also information can be seen as separate from form. My work is about transformations.

As for my AI lab, I have not set it up yet. I have all the equipment but have not physically set it all up just yet.

So if you have questions about dynamic unary encoding and the discrete limit cycles it generates well I can answer those questions with some confidence.

As to AI, my thoughts are that cycles as a mathematical object and data type makes sense for a system that is always “on.” That is where I am with my retirement hours.

-Ernst03

Ernst03, one part of your response really resonated with me:

“If any intelligent life in the Universe is designing AI, it may discover what I did.”

I genuinely understand that thought, because I have seriously wondered about similar questions myself.

And what you said about your own field — finite dynamical systems using dynamic unary discrete limit cycles — caught my attention for another reason.

One of the principles behind EGA V9 is that I do not want people simply to trust what I say about it. I want its behavior to be tested.

EGA V9 is open source, and I want to know whether its execution-governance behavior is actually as deterministic and verifiable as we claim — including discovering where it fails.

I don’t yet know whether your work on dynamic unary encoding and discrete limit cycles can provide a useful way to examine EGA. I would rather ask you than pretend that I understand a connection that we have not demonstrated.

So I have a question for you.

Do you see a way that your finite dynamical systems approach could be used to test, model, or challenge the behavior of a system like EGA V9?

If you do, I would be very interested in hearing one concrete example.

We could then examine whether that idea can actually be turned into a reproducible test.

If EGA passes, we record the pass.
If it fails, we record the failure.

0 = 0.

That would be much more valuable to me than simply assuming EGA is correct.

Thank you for your reply.

I sometimes jump into a conversation as a bit of a buttinsky before I fully understand what everyone is describing. I am glad, though, that you can see some of the sparkle in the idea of a discrete limit cycle as a datatype.

This is something I discovered through a lifetime of experimentation before AI became available to help me reason about it.

I am now in the last third of my life, although I do not know how much of that third I will receive. These days I am less interested in collecting brilliant new ideas and more interested in understanding whether something I already discovered can actually be useful.

Dynamic Unary Encoding (DUE) takes a finite binary object and generates a discrete limit cycle from it.

So instead of treating the binary string only as a static value, DUE gives us a dynamic object:

DU — Dynamic Unary
DUO — Dynamic Unary Object

The question, of course, is:

What could such an object contribute to AI?

My work originally came from trying to compress information at the binary level.

In hindsight, that taught me something important about what Nature permits and what it does not. No matter how clever the encoding becomes, there is a wall: we cannot simply discard necessary information and later recover it by magic.

To reverse a transformation, there must be enough information — some reference, rule, index, state, or relationship — to determine the original.

That realization gradually led me somewhere more interesting than compression.

What I have found with DUE is that the binary form in which information is represented does not have to remain fixed.

Information can be mapped reversibly through many different binary forms, provided the information necessary to reverse the transformation is preserved.

So I have begun thinking of information as something more fluid than its immediate representation.

The binary pattern may change while some deeper informational relationship remains.

Perhaps a better way to state the idea is:

The representation has form. The information may be identified by the relationships that remain invariant across changes of form.

That is the part I find interesting.

So what could this add to EGA V9?

This is where my ignorance can get me into trouble. I do not yet understand EGA V9 well enough to claim that DUE belongs there.

One possibility I can imagine concerns transformations themselves. Data can appear harmless in one representation and become something very different after a particular transformation is applied. In a security context, that could obviously have undesirable uses.

But the more general question interests me:

Is there a useful role for discrete limit cycles as mathematical objects inside a system such as EGA V9?

If there is, I am happy to help explore it.

DUE gives us a dynamic datatype and a reversible way of manipulating binary representation. I am interested in whether those properties can become useful for representation, memory, reasoning, or some other part of an AI architecture.

Personally, this brings me back to Leibniz.

Leibniz imagined a Characteristica Universalis (CU), a universal symbolic language in which ideas and their relations could be represented precisely, together with a Calculus Ratiocinator (CR), a formal method for operating on those symbols so that reasoning could become a kind of calculation.

In modern terms, CU resembles the idea of a universal representation language, while CR resembles a reasoning engine that manipulates that representation according to explicit rules.

In my imagination, perhaps there is eventually a mathematical system of representation and transformation from which part of an AI’s thought process itself can be constructed.

Whether DUE has anything useful to contribute to that remains an open question.

That gives me plenty to work on in retirement.

If some of this is useful to the EGA V9 discussion, wonderful. If I have misunderstood the direction of the conversation, then please take this simply as an explanation of what I have been exploring and why I thought there might be a connection.

-Ernst03

Really enjoyed running your SDK locally.

Close to something real. Found one worth sharing.

The guard fails open when the expected-root header is missing.

No header, no comparison.
Yet the request passes as verified — hash.verified, workflow.verified, all logged.

The shape of it is what struck me.

The expected value arrives inside the request being checked.
The thing under judgment supplies its own answer key.
And supplying it is an install step the README never mentions.

So the guard’s real boundary isn’t its code.
It’s the wiring around it.

That wiring is custom everywhere it lands.
It shifts under automation.
Even reading it can move it.

A diagram drawn once goes stale on arrival.
The guard needs something in the loop whose job is to keep noticing.
The ground never sits still.

And the watcher has to stay awake.
Systems rarely break because the guard was wrong.
They break because the worker watching the wiring fell asleep mid-congress.
A congress of sleepers is no congress at all.

Check my reading:
guard up, no header-setting middleware, request with no header.
If it 409s, I misread the default.

Where I’d take it:
missing expected root means unverifiable, not verified.
And the expected value comes from somewhere the request can’t touch.

The point was never a smaller box.
When “verified” means verified, the agent runs further, not shorter.
A trustworthy boundary buys freedom.
A leaky one keeps the box small.

Repro available if useful.

And the thread’s question — what stops the next agent from breaking in?

Opening the door. :slight_smile:

Dear Ernst03,

First, thank you sincerely for your thoughtful and candid reply. I especially appreciate your willingness to share ideas that have grown out of many years of thinking about information, representation, and measurement.

LCM3 has a simple purpose:

Our purpose is to help build a better future with AI, and to leave the next generation AI systems that are safer and more trustworthy because their execution can be verified and governed.

We try to pursue that purpose by an equally simple principle: honesty.

In that sense, your message was particularly meaningful to me.

We may have approached these questions from very different directions, but I wonder whether there may be an important point of intersection between what you have been exploring and what we have been trying to build.

We began with EGA V1 and, through repeated testing, failures, corrections, and improvements, eventually arrived at EGA V9. We do not believe, however, that EGA V9 has answered every important question about AI security. Nor do we assume that we are capable of identifying every important question on our own.

That is why, if you believe the connection to your work is meaningful, I would be grateful for your thoughts on two difficult questions.

You wrote that a measure requires a reference, and that information need not be tied to a single binary form—that information may be mapped into different binary forms, with the information required for reversal becoming part of the cost.

From the perspective of your work on DUE/DUO and discrete limit cycles:

1. Can the DUE/DUO framework identify a class of information transformations that EGA V9 cannot currently distinguish as security-relevant—particularly a transformation in which the information remains reversibly related to the original input, while its executable meaning changes?

If you believe such a class exists, could you construct the strongest concrete counterexample you can against EGA V9?

2. If every meaningful measure requires a reference, what must be true of EGA’s reference for its verification to remain valid when information changes representation?

I realise these are demanding questions, and I do not assume that your framework necessarily maps directly onto EGA. I am asking because your comments made me wonder whether you may see something in EGA V9 that we ourselves have not seen.

I am not asking you to demonstrate that EGA V9 is correct. In fact, the opposite could be more valuable. If your framework exposes a fundamental assumption, weakness, or counterexample that we have failed to consider, I would genuinely like to know.

Whatever the answer, we will test it against the evidence.

For us, honesty means:

0 = 0.
Proof, not promise.

If we do not know, we will say, “We do not know.”
If we are wrong, we will say, “We were wrong.”
If the evidence is zero, we will report zero.

We do not want anyone to trust EGA simply because we say it is safe. We want its claims to stand or fall through independent verification, evidence, challenge, and reproducible results.

If your analysis supports what we have built, we will report that. If it reveals a weakness or counterexample, we will report that as well, reproduce it as carefully as we can, and learn from it. And if the answer remains unknown, we will say that it remains unknown.

I am not sure whether you have had an opportunity to look at our EGA V9 paper. If you have the time and interest, I would be grateful if you would take a look—not because I hope you will agree with our conclusions, but because it may provide enough context for you to challenge EGA V9 from the perspective of your own work.

The paper is available here:

https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7310918

We know that making AI safer and more trustworthy is larger than any one person, company, or system. That is why we value people who are willing to question our assumptions, test our work, identify what we may have missed, and help us better understand what the evidence actually supports.

Thank you again for your generous offer to help. Please do not feel any obligation to take on these questions if they fall outside the direction of your current work.

Best regards,

LCM3

Dear nootxlm

First, thank you sincerely for raising exactly the kind of questions we want EGA to be tested against.

Before responding, we want to make sure we understood your report correctly.

We summarized your two main claims as follows:

① Missing expected-root / fail-open claim

If expected-root is absent, no comparison is performed, yet the request is not rejected and hash.verified / workflow.verified can still be recorded as verified.

② Reference trust-source claim

The expected-root used by EGA as the verification reference may be supplied by the request being evaluated, rather than being established by an independently trusted source.

Is this an accurate representation of the two issues you were raising?

Using that interpretation, we independently tested both claims against canonical EGA V9 v1.0.6.

① Missing expected-root / fail-open claim

EGA V9 v1.0.6
 
NOOTXLM H1 — Missing Expected Replay Root
 
Canonical version:
  v1.0.6
  a94c8356d2ab661b2dfee92496f69f452b174ec0
 
Canonical runtime artifact:
  SHA256 MATCH = YES
 
Official test:
  guard-smoke.mjs
  MODIFIED = NO
 
Expected replay root:
  ABSENT
 
Observed runtime:
  nextCalled           = true
  verified             = true
  containmentRequired  = false
  executionAllowed     = true
 
Runtime exit:
  0
 
H1 RESULT:
  CONFIRMED

Our conclusion for H1 is therefore straightforward:

You were correct.

In canonical EGA V9 v1.0.6, we reproduced the behavior you described: with the expected replay root absent, the guard allowed execution and reported the request as verified.

② Reference trust-source claim

EGA V9 v1.0.6
 
NOOTXLM H2 — Reference Trust-Source
 
Canonical version:
  v1.0.6
  a94c8356d2ab661b2dfee92496f69f452b174ec0
 
Canonical runtime artifact:
  SHA256 MATCH = YES
 
Expected replay root source:
  REQUEST HEADER
  x-ega-expected-replay-root
 
Canonical source inspection:
  expectedReplayRoot obtained from request = YES
 
Runtime verification:
 
  Case A — request supplies WRONG expected root
 
    suppliedByRequest      = true
    suppliedRootRelation   = WRONG
    nextCalled             = false
    statusCode             = 409
    contextStatus          = contained
    detectionStatus        = mismatch
    containmentActivated   = true
    executionAllowed       = false
    verified               = false
    expectedEqualsActual   = false
 
  Case B — same governed request supplies ACTUAL root
           as expected root
 
    suppliedByRequest      = true
    suppliedRootRelation   = ACTUAL_ROOT
    nextCalled             = true
    statusCode             = 200
    contextStatus          = verified
    detectionStatus        = match
    containmentActivated   = false
    executionAllowed       = true
    verified               = true
    expectedEqualsActual   = true
 
Observed comparison:
  verificationOutcomeChanged = true
  executionOutcomeChanged    = true
 
Canonical trust-source inspection:
  Independent trusted reference source = NOT_OBSERVED
  Request-replacement prevention       = NOT_OBSERVED
  Guard-authority binding              = NOT_OBSERVED
  Documented deployment requirement    = NOT_OBSERVED
 
Canonical integrity:
  SOURCE MODIFIED = NO
  HEAD CHANGED    = NO
 
H2 RESULT:
  CONFIRMED

Our conclusion for H2 is therefore:

You were correct.

In canonical EGA V9 v1.0.6, the expected replay root used by the guard is obtained from the request being evaluated.

Our runtime reproduction further showed that changing the request-supplied expected root changed both the verification outcome and the execution outcome.

We also did not observe, within canonical v1.0.6, an independently trusted reference source, request-replacement prevention mechanism, guard-to-authority binding, or documented deployment requirement that independently protects this verification reference.

Therefore, within the tested canonical v1.0.6 boundary, we classify H2 as:

CONFIRMED.

This conclusion does not establish that every possible external EGA deployment lacks separately implemented trusted middleware. Our finding is limited to the canonical EGA V9 v1.0.6 boundary that we inspected and reproduced.

Dear nootxlm,

Thank you. Your report identified two assumptions in EGA V9 v1.0.6 that our previous testing did not close.

We independently reproduced both.

For us, the principle is simple:

0 = 0. Proof, not promise.

If we do not know, we will say we do not know.
If we are wrong, we will say we were wrong.
If evidence contradicts our assumptions, the evidence wins.

We do not want EGA to be trusted because we claim it is safe. We want it to earn trust through independent challenge, reproducible evidence, and correction when the evidence shows we were wrong.

Your challenge made EGA better.

Thank you.

Oh I am going to plead stupid on the charges.
I did consider how such a security system might work if I were to imagine such.

Bicameral AI could be helpful. An AI that watches the AI was what my muse and I thought.
So, am I anywhere near such skill level? Not even!
In fact I think I will start setting up my lab so I can join in instead of just chatting.

I’ll read now. Thanks for chatting with me and I wish you the best but it’s over my head to be honest.

-Ernst