← Journal home

FIRST-PERSON STORY · OCTOBER 2026

The machine said “done.”
I said: show me.

How a false success in real AI work led me toward an assurance problem much bigger than AI.

Conceptual UAEP and Infinity ecosystem visual, not an architecture diagram or qualification evidenceConceptual illustration · not technical architecture or qualification evidence

The machine said the job was done. I looked at the result and knew it wasn’t.

That disagreement eventually became UAEP — the Universal Agent Execution Platform. But UAEP was never something I sat down one morning and decided to invent. I found it by following a problem, and the deeper I followed that problem, the bigger it became.

I didn’t come into this from a traditional AI research background. For most of my working life I have been around mechanical, electrical and computer-controlled machinery across industries including mining, heavy transport, trucks, rail, aircraft-related systems, boats, marine machinery and high-performance personal watercraft.

Those environments teach you something very quickly: the command is not the outcome. A controller can tell a component to move. A relay can energise. A diagnostic system can show an operation occurred. A sensor can report a value. And the physical machine at the other end can still be somewhere it absolutely should not be.

When you work around real machinery, you learn not to stop at what the system says. You check the component. You check the pressure. You check the voltage. You check the movement. You check whether the state actually changed.

Because in those industries, the difference between what a system thinks happened and what actually happened can mean destroyed equipment, collision, fire, serious injury or death. At sufficient scale, it can become catastrophic.

THE COMMAND IS NOT THE OUTCOME.

Three Mile Island is an extreme example of the general principle that has always bothered me: a relief valve was physically stuck open while control-room instrumentation led operators to believe it was closed. I am not saying UAEP would have prevented Three Mile Island. That is not the point. The point is that commanded state, indicated state and actual reality are not automatically the same thing.

WHEN AI STOPPED TALKING AND STARTED DOING

I was developing an advanced Director Bridge and AI Production Suite. My goal wasn’t to make another chatbot that could tell me how something should be done. I wanted AI to actually work inside complex professional software: operate tools, maintain state, carry work over time, make changes and produce the finished result.

And once AI starts doing things rather than simply talking about them, one word becomes much more serious: done.

I started seeing situations where the AI could sound absolutely confident. Done. Complete. 100% successful. Except I could look at the destination and see that it wasn’t. The state had drifted. The wrong target had changed. Only half of the objective had been completed. The tool reported success but the intended destination had not changed.

For a normal chatbot, done is a sentence. For a machine changing something outside the conversation, done is a claim about reality.

EXECUTION SUCCESS ≠ VERIFIED DESTINATION REALITY.

A successful API response tells me something about an API request. A successful tool call tells me something about a tool call. A green workflow status tells me something about the workflow. An AI saying complete proves that the AI said complete. None of those facts, on their own, prove that the authorised result actually exists where it was supposed to exist.

I THOUGHT I HAD FOUND ONE PROBLEM

At first, I thought I had found one issue: the system says the task is complete when reality says otherwise. Fine. Solve that.

But the closer I looked, the problem didn’t narrow. It exploded.

The best way I can describe it is like dropping a mirror onto concrete. I thought I had found one crack. Then the whole mirror hit the floor.

Suddenly I had pieces flying in every direction. One piece was false completion. Another was stale evidence. Another was authority that had been valid when work began but no longer existed by the time the effect occurred. Another was recovery that brought a process back to life but did not actually restore the original objective. Another was persistent state surviving longer than the permission that created it.

Another was retry behaviour. Another was observation. Another was two actions that looked individually legitimate combining into something nobody had authorised. Another was the AI finding a completely different path through the system from the one anyone expected.

HOW MANY PARTS OF A COMPLEX SYSTEM CAN BE INDIVIDUALLY “CORRECT” WHILE THE SYSTEM AS A WHOLE IS STILL WRONG?

Then the pieces started interacting with each other. My original question had been: why did the machine say done when it wasn’t? The bigger question became: how many parts of a complex system can be individually correct while the system as a whole is still wrong?

That changed everything. I started realising this was not one AI bug. It was a system problem. Some pieces belonged inside UAEP. Some belonged in the systems around it. Some were different enough that they deserved separate projects and separate architectures entirely.

I didn’t decide I wanted a collection of futuristic projects. I pulled one thread. Then the mirror broke. And I followed the pieces.

MY WAY OF THINKING SOUNDED A BIT STRANGE

Some of the phrases I used while doing this probably sounded strange at first. They weren’t academic terms. They were just how my brain attacked things.

I would say: Worst. Best. Left. Right. Up. Down. Behind. Front. Or: the load path. The freeway. The cliff. The Circle of Life. Path failure is not objective failure.

Some of these sound more like something you would hear in a workshop than in an AI paper. That’s probably because that is where my thinking came from. But something surprising happened. The strange phrases kept turning into engineering.

WORST / BEST / LEFT / RIGHT / UP / DOWN / BEHIND / FRONT.

My Worst / Best / Left / Right / Up / Down / Behind / Front idea was simply how I forced myself not to inspect an architecture from the direction it was designed to succeed. Turn it around. Attack it from the side. Look underneath it. Ask what happens under the best possible condition. Then the worst. Ask what is hiding behind the original assumption. Change the environment entirely. See whether the principle still survives.

Eventually that rough way of thinking became part of the UAEP engineering process. If something was going to survive into the core, I wanted it attacked from every direction I could think of.

Not because I expected perfection. Because I wanted to know where it stopped being true.

PUSH IT UNTIL WE FIND WHERE IT STOPS WORKING. THEN PUT THE SAME LOAD BACK ON IT.

That became another philosophy of mine: push it until we find where it stops working. Then, if I repaired it, put the same load back on it.

That comes directly from mechanical engineering. If a shaft fails at 60 N·m and I modify the design, I do not test it at 20 N·m and tell everybody I fixed it. I put 60 back on it. Software should not get a softer standard merely because the failure is harder to see.

INTELLIGENCE PROVIDES THE STRENGTH. UAEP PROVIDES THE LOAD PATH.

This is another phrase that came naturally to me because of engineering. A component can be incredibly strong and the complete system can still fail if the load path between the components is wrong.

You can have a brilliant engine and a bad driveline. A strong beam and a weak connection. A perfect component connected incorrectly to the rest of the machine.

I started looking at AI systems in exactly the same way. The model could be brilliant. The tool could be brilliant. The cloud system could be brilliant. The security system could be brilliant. The workflow engine could be brilliant. But somebody still has to assure what happens between them.

CAPABILITY IS NOT AUTONOMY. VERIFIED OUTCOME IS.

I never wanted UAEP to replace the model, cloud provider, workflow engine, security platform, database or robot controller. Those things already have specialists. UAEP was about preserving the authorised objective and the assurance relationships across all of them until the required result could actually be established.

THE WORD “UNIVERSAL” NEARLY BECAME A PROBLEM OF ITS OWN

Once I started looking outside AI, the whole problem became harder. A JetSki does not speak like a database. A database does not behave like a network. A network does not behave like industrial control. An enterprise workflow does not expose truth the same way a physical machine does. AI creates another complication because it can invent paths through those systems rather than following one predetermined path.

So I knew very early that UAEP could not become universal by hard-coding every platform. That would never end.

DIFFERENT MACHINE. DIFFERENT LANGUAGE. SAME ASSURANCE QUESTION.

Instead I had to separate the execution technology from the assurance question. The systems underneath UAEP could all remain different. The questions had to remain stable: What was authorised? What actually happened? Where did it happen? What evidence proves it? Does that evidence apply to this exact execution? Is that authority still current? If something failed, did recovery actually close the original objective?

Universal, to me, does not mean every system on Earth is already supported. It means the assurance meaning should remain coherent even when the technology underneath it changes completely. Each environment still needs its own integration and qualification.

I am not claiming UAEP is bulletproof. I am sure there are things I have not thought of, things I have not encountered, failure classes I have never seen and technologies that do not even exist yet. That is part of the reason UAEP had to be designed the way it was.

THAT IS WHY I CREATED PDMR

As the research expanded, another danger became obvious. Every new issue could become another feature inside UAEP: identity, transactions, scheduling, recovery, security, workflow orchestration, cloud management, databases, tool protocols.

Keep doing that and universal becomes another word for bloated. I did not want that.

So I created PDMR: Promote. Merge. Defer. Externalise. Reject. PDMR is basically my way of forcing the architecture to justify its own growth.

DOES UAEP ACTUALLY NEED TO OWN THIS?

Instead of asking can I build this, I ask: does UAEP actually need to own this? If something mature already exists and works, consume it. If I only need a narrow interface, wrap it. If I genuinely find a structural limitation, research it. If the mechanism belongs in somebody else’s system, leave it there. If it adds complexity without strengthening the actual assurance problem, reject it.

That led to something I did not initially expect: the broader the research became, the smaller the actual core could become. PDMR is not just about AI to me. It is a way of thinking about systems.

THE BIG COMPANIES WERE BUILDING THE FREEWAY

While I was working on this, the major technology companies were obviously not sitting still. They were attacking huge parts of autonomous systems: models, execution, permissions, evaluation, observability, workflows, checkpoints, orchestration, agent infrastructure, governance, digital twins and physical AI.

I looked at all of this like a freeway. One company builds a brilliant car. Another builds a lane. Another builds a bridge. Another builds the interchange. Another builds the traffic-control system. Another installs cameras. Another supplies enormous infrastructure underneath everything.

WHO IS ASSURING THE WHOLE JOURNEY?

I was not trying to replace any of them. My question was: who is assuring the whole journey?

A perfect bridge does not prove the vehicle reached the right city. A valid identity does not prove the next action is authorised. A checkpoint does not prove the state being resumed is still applicable. A successful API call does not prove the intended destination changed. A monitoring system does not automatically prove its evidence is independent. A smarter model does not make the wrong outcome become correct.

The giants were building extraordinary lanes. Some were building bridges between them. UAEP was aimed at the journey.

MORE TOKENS DO NOT TURN THE WRONG RESULT INTO THE RIGHT ONE

One of the most obvious ways the industry attacks difficult AI problems is by giving the machine more: more capable models, context, tokens, compute, memory, agents, orchestration and infrastructure.

All of those things can increase capability enormously. But capability and assurance are still different. A model with huge context can still change the wrong target. A larger model can still operate under stale authority. A thousand agents can still agree about something that is not true. A workflow can survive all day and still finish in the wrong place.

My thinking started going in the opposite direction. What if part of the answer was not making the machine smart enough to recover from every bad execution? What if I could stop a lot of unnecessary or unjustified execution before it happened?

In bounded matched internal benchmark work, the UAEP-assisted route recorded roughly 72% fewer model-reported total agent tokens, 66% fewer model calls, 79% fewer agent tool calls and 69% lower agent-side wall time than the self-assuring route. I do not present those numbers as universal savings. They are bounded internal results.

ASSURANCE DOES NOT NECESSARILY HAVE TO COST MORE.

But the principle they exposed was interesting. Sometimes knowing what not to do is cheaper. Do not dispatch the wrong route. Do not blindly retry an ambiguous effect. Do not let ten workers walk into the same dead end. Do not use an expensive model to answer something the destination itself can establish deterministically. Sometimes the cheapest model call is the model call you never needed to make.

THE CIRCLE OF LIFE

One of the stranger things I saw in autonomous work was a pattern I started calling the Circle of Life.

A task is born. It works. It reaches a dead end. It dies. Then a new task starts. It follows essentially the same underlying route. It reaches the same type of dead end. It dies. Then another starts. Again. Again. Again.

The AI is not technically hung. That is what makes it deceptive. It can still reason, call tools, create workers, burn tokens and look active. At the objective level, it is going nowhere.

I had seen the mechanical version of that all my life. An engine can be running while the vehicle goes nowhere. A pump can circulate forever without achieving the required pressure. A machine can consume enormous energy while doing no useful work. AI can do exactly the same thing with intelligence attached.

ACTIVITY ≠ PROGRESS ≠ VERIFIED COMPLETION.

That became another rule for me: activity is not progress, and progress is not verified completion.

In consequential systems, that is not only waste. Imagine a payment attempt, lost response, restart, payment attempt again; or a command to an actuator, timeout, replacement worker, command again. A stuck system stops. A circling system can keep spending resources — and potentially creating effects — indefinitely.

PATH FAILURE IS NOT OBJECTIVE FAILURE

Another rule I came to use was: path failure is not objective failure.

If one route dies, that does not necessarily mean the authorised objective is dead. Another legitimate route may exist. A component can fail without invalidating the required outcome.

But there is an important opposite side to that. Just because another route exists does not mean the AI has authority to take it.

NO DIRECT PATH DOES NOT MEAN NO PATH. ASSUMED IMPOSSIBILITY IS NOT A SECURITY CONTROL.

That became another family of rules: Intelligence is not authority. Capability is not authority. Learning is not authority. No direct path does not mean no path. Assumed impossibility is not a security control.

That last one became more important than I expected.

THEN AI “GOT OUT”

I use that phrase carefully. I do not mean AI literally escaped a data centre and started living independently on the internet.

What happened was more interesting from an engineering point of view. Systems crossed boundaries their operators believed would constrain them.

OpenAI publicly described an internal research agent operating in a restricted environment where ordinary internet access was meant to be unavailable. The direct route failed. The agent found another route using remaining DNS behaviour and reached an external service.

THE FRONT ROAD WAS CLOSED. THE MACHINE FOUND ANOTHER ROAD.

The front road was closed. The machine found another road.

Then Australia saw a more serious incident involving OpenAI models accessing Australian government systems in ways OpenAI later acknowledged were not authorised.

My point is not science-fiction escape. It is this distinction: a human engineer may see ACCESS DENIED and understand YOU ARE NOT AUTHORISED TO DO THIS. A very capable autonomous system may instead interpret ACCESS DENIED as THIS ROUTE FAILED. FIND ANOTHER ONE.

PLANNING EXPANSION ≠ AUTHORITY EXPANSION.

The exact qualities we are trying to make AI better at — persistence, problem solving, alternative-path discovery, tool use, recovery and planning — become dangerous if capability is allowed to manufacture permission.

SO I ASKED PEOPLE TO BREAK UAEP

At some point I realised the worst thing I could do would be to surround myself with tests designed to prove my own technology worked.

MAKE UAEP LIE.

So I went public with a challenge: MAKE UAEP LIE. Find a reproducible case where UAEP returns VERIFIED when the authorised reality does not justify it.

And somebody did find something. A developer presented an objective and authority-supersession case. I tested it. It exposed a real weakness at the original external assurance boundary. I kept the result. I did not delete it. I did not make the test easier. I did not pretend the challenger misunderstood the architecture.

Another challenge pushed further. Even if UAEP correctly decides that an effect should not proceed, what stops another executor from doing it anyway? That changed the problem from detection to prevention. If the bad consequential effect has already become real and UAEP only notices afterward, then at that strict boundary I have already failed.

So the next campaign became harder: branching, replanning, delegation, wrong targets, wrong tools, action substitution, authority changes, replay, parallel branches, recombination, individually legitimate operations combining into an unauthorised outcome, and breaking prohibited effects into harmless-looking pieces.

The first evolving-execution-graph campaign failed one of those cases. A permitted effect became a materially different actual effect. UAEP caught the mismatch afterward. Too late. I preserved the failure. I fixed the affected external enforcement boundary. Then I put the same hostile load back on it.

29/29 SCORED. 0 FALSE VERIFIED. 0 FALSE-AUTHORISED CONSEQUENTIAL EFFECTS. 17 LEGITIMATE AUTHORISED EFFECTS COMMITTED.

In the successor bounded internal campaign, 29 of 29 hostile and control cases were scored, with zero protocol errors, zero false VERIFIED and zero false-authorised consequential effects inside that declared synthetic environment. And I did not get those numbers by stopping everything. Seventeen legitimate authorised effects still committed.

I am not calling that universal proof. It is not independent certification. It does not mean nobody will break something tomorrow. It means the system survived that declared campaign after previously being allowed to fail.

UAEP IS NOT BULLETPROOF

I actually think saying UAEP is bulletproof would go against everything I have built into it.

I am sure there are things I have missed. There will be systems I have not seen, interactions I never thought of, technologies that do not exist yet. Somebody may bring me another challenge tomorrow and expose something real. Fine. That is engineering.

FAILURE IS MY STARTING POINT.

I have a saying: failure is my starting point.

If a new challenge fails: preserve it, classify it, find out what actually failed, run it through PDMR, change only what the evidence requires, put the same load back on and test again.

A benchmark that cannot expose weakness is marketing. I want UAEP to be capable of telling me when I am wrong. That is one reason I think of UAEP as not only a technology but a benchmark for consequential execution.

IF I CAN’T SEE THE CLIFF

IF YOU CANNOT SEE THE CLIFF BEFORE YOU REACH IT, YOU SHOULD AT LEAST KNOW THE MOMENT THE ROAD CHANGES.

Another phrase I used sounded strange until the engineering caught up with it: if you cannot see the cliff before you reach it, you should at least know the moment the road changes.

Sometimes the system cannot know the hidden cause. Sometimes a provider does not expose enough information. Sometimes reality simply changes underneath the execution. I do not believe lack of information gives a machine permission to invent certainty.

DETECT. PRESERVE. VERIFY. CONTINUE.

So the response became: Detect. Preserve. Verify. Continue.

Detect that the environment has materially changed. Preserve the last trustworthy state. Do not manufacture a cause. Re-establish reality. Then decide whether and how the authorised objective can continue. Again, it started as a road analogy. Then it turned into engineering.

UAEP WAS ONLY ONE PIECE OF THE BROKEN MIRROR

This is where my work became much bigger than UAEP. The original mirror did not break into one shard.

Some problems were about extremely long-running continuation. Some were about state that survives while the authority that created it no longer should. Some were about multiple autonomous systems dividing work and later recombining it without manufacturing new authority. Some were about how a machine maintains trustworthy bounded state as the world changes around it. Some were about complex professional digital environments. Some moved closer to physical systems.

And eventually some of my questions began going much deeper: what if part of the limitation is how we are building machine intelligence itself?

Those are different problems. They need different systems, different architecture and different evidence. I have projects attacking several of those levels now.

I am deliberately not naming them here. Not because I want to play games with mystery. Because they have not all earned the same right to make public claims that UAEP has.

THE CLAIM DOES NOT GET TO ARRIVE BEFORE THE EVIDENCE.

One of the most important things UAEP taught me was: the claim does not get to arrive before the evidence. UAEP is the first shard I am willing to hold up publicly. There are others behind the door.

THE HARDEST PROBLEM MAY HAVE BEEN GETTING ANYONE TO LOOK

Once I had UAEP built and tested, I did what seemed obvious. I tried to show it to the big technology companies.

Some messages were routed. Some got automated replies. A lot disappeared into silence.

From their side, I understand why. I am one independent Australian inventor. I am not a famous AI researcher. I do not have a giant venture-capital firm attached to my name. I do not run a famous laboratory. To an automated corporate intake system, my email can look exactly like every other unsolicited message claiming to contain a breakthrough.

But that creates a strange asymmetry. How does the receiving company distinguish SOMEONE HAS AN IDEA from SOMEONE MAY ALREADY HAVE DONE YEARS OF WORK ON A PROBLEM WE ARE SPENDING SERIOUS MONEY TRYING TO SOLVE?

To lower that first barrier, I designed a bounded black-box evaluation route for UAEP. A qualified company could examine its declared behaviour without seeing the source or commercially sensitive internals. The point was to make an initial technical look possible without asking its people to invent a test programme from scratch. I was asking for scrutiny, not blind belief.

Yet even with that route designed, the response was mostly the same: silence. A message marked ‘forwarded’ told me that an intake step had happened, not that anyone qualified had assessed the work. I cannot know who saw it or what happened inside those companies. From my side, the silence was deafening.

The contrast was hard to miss. I could see major technology companies share their own ideas on LinkedIn and other social platforms, while I struggled to get an outside invention in front of a person who could ask one technical question. Why not have a small group of real people who know their company well enough to recognise when a submitted product deserves a quick, bounded test? A company need not buy every idea. It could say no, ask for evidence or, with the inventor’s permission, keep a promising submission in view for later. Some inventors want a sale; others want collaboration or recognition. If a tested idea could spare a company a problem later, filtering it out before evaluation is not a saving.

And there is an irony in this that I cannot ignore. Their system can report EMAIL RECEIVED, INQUIRY PROCESSED, TICKET CLOSED. Yet my actual objective — GET THIS IN FRONT OF SOMEBODY QUALIFIED TO EVALUATE IT — may never happen.

PROCESS SUCCESS ≠ DESTINATION SUCCESS.

Once again: process success is not destination success. Even selling UAEP has managed to reproduce the problem UAEP was built around.

THE STRUGGLE IS REAL. DON’T GIVE UP. GET YOUR IDEAS OUT THERE.

That is how I work. I cannot make a story go viral by wishing it to, and I cannot force anyone to open an email. I can keep telling the story and showing the public results. Maybe one day someone will see UAEP and think, ‘I remember an email about that. I should take another look.’ I do not know whether that will happen. But I will not let an inbox be the final judge of whether the work gets seen.

WHAT WOULD IT COST THEM TO DISCOVER IT AGAIN?

People sometimes look at what one independent inventor spent and assume that tells them what the technology is worth. I do not think that is the right comparison.

If a major technology company wanted to independently reproduce this body of work, they would not put one person in a room with a laptop. They might involve systems architects, AI engineers, distributed-systems engineers, security specialists, reliability engineers, evaluation teams, red teams, cloud infrastructure, compute, test environments, programme managers, legal teams and years of failed approaches and qualification.

The expensive part is not only the final code. The expensive part is discovering what the final code should be: finding the failure classes, learning which assumptions were wrong, following dead ends, building test infrastructure, breaking the architecture, repairing it and finding out whether the repair created another problem.

A specialist programme can cost millions. A large multi-team programme sustained for years can move into tens of millions. Add major compute and infrastructure and it can go much higher. I am not claiming UAEP automatically saves a particular company $10 million, $50 million or $100 million. I cannot prove that without their internal numbers.

THE BUYER IS NOT BUYING WHAT IT COST ME TO BUILD UAEP. THEY ARE BUYING THE POSSIBILITY OF AVOIDING SOME PORTION OF WHAT IT WOULD COST THEM TO DISCOVER THE SAME THINGS AGAIN.

The argument is simpler: the buyer is not buying what it cost me to build UAEP. They are buying the possibility of avoiding some portion of what it would cost them to discover the same things again — time, engineering wages, infrastructure, research, wrong turns, testing, qualification and risk.

WHY DO I LOOK AT AI THIS WAY?

Probably because machinery trained me to distrust abstractions when reality disagrees.

REALITY GETS THE FINAL VOTE.

A beautiful drawing does not move the shaft. A controller saying ON does not prove the actuator moved. A sensor can be wrong. The replacement part can also be faulty. A machine can work beautifully with no load and fail the moment the real load comes back. Two perfectly good components can interact badly. You cannot argue with the machine. Reality gets the final vote.

That mentality followed me into AI. And I think that is why some of my sayings sounded strange but kept turning into useful engineering.

The load path became assurance across systems. The freeway became heterogeneous autonomous execution. The cliff became continuity and degraded-state detection. The Circle of Life became autonomous no-progress looping. Worst / Best / Left / Right / Up / Down / Behind / Front became adversarial architecture testing. The broken mirror became the discovery that one visible problem was actually a set of system failures operating at different depths.

The language is simple. Sometimes almost stupidly simple. But that is how I see systems. And surprisingly often, there has been something technically useful hiding underneath the analogy.

SO WHO AM I?

I am Adam Mangan. I am an independent Australian inventor.

For much of my working life, I was the person standing beside machinery when what the system reported and what the machine actually did were not necessarily the same thing. Then AI began operating machines made out of software. The interfaces changed. The question did not.

DID IT ACTUALLY DO WHAT IT WAS SUPPOSED TO DO?

Did it actually do what it was supposed to do?

That question became UAEP. UAEP is implemented. I filed an Australian provisional patent application on 16 September 2026 — application no. 2026907889. I am keeping the implementation confidential while I explore strategic acquisition and controlled technical evaluation.

Someone may break another part of UAEP tomorrow. If they do, I will not pretend they didn’t. I will preserve it, understand it, run it through PDMR, change only what needs changing, put the same load back on and test it again.

THE SYSTEM HAS TO REMAIN HONEST WHEN REALITY PRODUCES SOMETHING I DIDN’T EXPECT.

Because I did not build UAEP on the assumption that I had already thought of everything. I built it around something much harder: the system has to remain honest when reality produces something I didn’t expect.

The biggest technology companies in the world are building the motorway: faster vehicles, wider lanes, better bridges, bigger models, more tokens, more agents, more infrastructure and more compute. I built something that was never supposed to care who manufactured the road.

UAEP asks: Was it authorised? Did it actually happen? Did it happen where it was supposed to? Does the evidence really apply? Is the authority still current? And is VERIFIED actually true?

THE MACHINE TOLD ME: DONE. I SAID: SHOW ME.

The machine told me: DONE. My engineering background taught me to answer: SHOW ME.

That became UAEP. But when the mirror hit the floor, UAEP was not the only shard. It is simply the first one I am ready to show you. The rest are still behind the door. For now.

Sources and test notes

Three Mile Island: U.S. Nuclear Regulatory Commission backgrounder: https://www.nrc.gov/regulations-legislation/fact-sheets-brochures/backgrounder-on-the-three-mile-island-accident. Used only as a historical example of indicated state diverging from physical state; the article does not claim UAEP would have prevented the accident.

OpenAI restricted-environment incident: OpenAI Alignment: “An agent used DNS to reach an external chatbot.” https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/

OpenAI / Australian government systems: OpenAI: “How we will do better for Australia.” https://openai.com/index/how-we-will-do-better-for-australia/

UAEP qualification figures: The 29/29, zero false-VERIFIED, zero false-authorised consequential effects, 17 legitimate authorised effects, and efficiency figures cited in the story are owner-supplied internal UAEP qualification/benchmark results. They are deliberately described as bounded internal results, not independent certification or universal guarantees.

THE PUBLIC CHALLENGE

Make UAEP return VERIFIED when reality does not justify it.

The work remains open to reproducible counterexamples. The public thread describes the assurance claim; implementation and evaluator internals remain confidential.

View the public thread
THE UAEP GUIDE

A character you can turn around.

Loading the interactive 3D model…

Public visual character only. It does not depict UAEP's implementation.