PPrompt an agent, watch it work, and the demo is genuinely impressive. Put it in front of a customer's money and the question changes shape entirely. Not can it do this, but what exactly will it do, who approved that, and what happens the day it is wrong.
The industry's answer so far has been to watch more closely. Log every step, score every output, add a guardrail model to check the first model, and escalate when confidence drops. All of it is inspection after the fact — a way of noticing that something went wrong slightly faster than a customer would have.
Inspection cannot produce a guarantee. If the decision about what an agent may do is made in the same breath as the doing, there is no earlier moment to point at and no artefact to review. Nothing can be signed, because there was never a fixed thing to sign.
So move the decision. Let the model do its thinking once, up front, and commit the result to something concrete enough to be read, checked and approved. Then let ordinary software carry it out. The creativity stays; the uncertainty leaves the runtime.