Autonomy turns on what a system knew when it acted
Diogo Almeida, who built TypeSafe's Jev, said recently on the Latent Space podcast that we have "this supercharged engine of automation" that "just does not have the right plugs" to connect to economically valuable work. I think he's right. And I'd name the plug: the ability to check what the system did after it acted alone.
Nobody turns on "always allow" because the system is perfect. They turn it on because approving every action by hand has become unbearable. The setting is only sane if someone can later ask what the system knew when it acted and get a true answer. That means the facts the model is handed at the moment it acts have to be recoverable later, exactly as they were. Building that is the job I mean by a feature engine.
The facts are the decision

Every automated decision rests on a few facts about the world, such as how many new people an account has paid this week, or how far a payment sits from the customer's usual. The model reasons over those facts and the workflow carries the case forward, but the facts are what make the decision right or wrong.
They also get the least attention. They are read from whatever the database holds at that moment or computed overnight in a batch, and they are rarely kept once the decision is made. That was workable while a person made each decision. The person looked at the facts, and if anyone asked later, the person remembered. When a system acts alone thousands of times a day, nobody remembers.
Mark Zuckerberg, describing Meta's Muse agent on the Sources podcast, explained that logins and payments need a person's approval each time, and that you can tell the agent to "always allow this kind of thing." That setting is the reason to automate in the first place. It also means most actions are never looked at.
That leap of faith already has a parable. Last week, The Verge reported that a Meta Muse user set "Allow Always" for Facebook Marketplace; the agent then shared his pickup address with a buyer and agreed to a below-ask price, and he only found out after the buyer had left. Nothing about the permission was ambiguous. The system did what it was allowed to do, and nobody was watching while it did it.
The cases that do get reviewed are the ones that went wrong, the ones someone complained about, and, in a careful shop, a handful picked at random. Each of those reviews happens days or months later and asks what the system knew when it acted. By then the data has moved on. The account has new transactions and the customer's baseline has shifted. Recomputing the facts today and judging the old decision by them gives an answer shaped by hindsight. A decision that was sound on the evidence of the day can look careless, and a careless one can look fine.
Almeida made a related point about language models that applies to any automated system: "It's really easy to see when an error happens. It's very hard to see when a subtle thing that looks correct happens." Decisions that look correct are the ones nobody pulls for review, so when one finally is pulled, it has to be seen as it was, or the review learns nothing.
Two kinds of state

Part of the confusion is the word "state." Many agent frameworks organize work as state machines, and the good workflow engines keep a careful history of each run: which step it reached and what it passed along. That is the machine's state. It is small, and it matters most while the work is running.
The other kind of state is the condition of the world: what an account or a device looked like at a given moment, derived from everything recorded up to then. That is what a review has to be checked against, and producing it is a different job from running the workflow. A history of steps can hold the value that tipped a decision. It cannot say whether that value was right at the time, because it doesn't hold how the value came to be.
Event sourcing keeps the raw events. What a review needs is the derived facts as they stood at a given moment, recomputable later to the same answer. Keeping the events alone doesn't supply that; the derivation is the missing piece.
Why the moment matters

The common answer is a snapshot: facts computed on a schedule, stored, and read by whatever needs them. Snapshots are cheap, and for reporting they work well. For a system that acts alone they fall short twice. At decision time they are stale, so the system acts on yesterday's picture. At review time they have been overwritten or recomputed, so the reviewer judges with hindsight.
What a system acting alone needs is facts computed at the moment it acts, from what was known at that moment, in a way someone else can repeat later and get the same answer. That is the job I mean by a feature engine. It also has to be honest about gaps. A history too short to judge, or a source that had gone quiet, should show up as missing rather than as a zero. If every field always has a value, some of those values are guesses, and a review can't tell which.
This costs more than a nightly batch, which is why the batch is the default. The comparison that matters, though, is with the alternative. If the facts behind a decision can't be recovered, the only review you can rely on is looking at everything, and in practice nobody does.
None of the underlying ideas are new. Banks and trading firms have kept replayable records for decades, and machine learning teams learned long ago to keep hindsight out of their training data. What has changed is how many systems are now allowed to act alone.
Where the value moves
Models get better and cheaper every few months, and workflows are well served by mature tools. What decides whether a company can let a system act on its own is quieter: whether it can say, for any decision months later, what the system knew when it acted. I expect the companies that can answer that to hand more of their work to automated systems, and sooner, than the companies still asking a person to approve each time.
That commoditization is already underway. The vLLM Semantic Router team just released Decision 2.0, a family of decision models in six sizes from 0.6B to 27B, explicitly framed as "Towards Open Foundation Decision Models." The value accrues to what the commodity plugs into, and that's the state and learning loop machinery AI systems need.
If you're handing work to a system that acts alone, three questions are worth asking:
Could you rebuild a decision from last quarter and get the same facts it acted on?
Can the system show what it did not know when it acted?
Are the facts used in review the same facts used at decision time because the system was built that way, or because a nightly snapshot usually works?
When the answers are yes, checking a small fraction of what a system did tells you something true about the rest, because each case you pick up still carries its evidence. Sampling becomes real oversight, and "always allow" becomes something you can keep checking rather than a leap of faith.