The demo was never the hard part.
A demonstration is built to succeed on examples somebody chose. A system has to fail safely on the ones nobody chose, in front of people whose work depends on it, on a Tuesday, when the person who understands it is on leave.
Everything that makes that difference is unglamorous and gets cut first: the written boundary, the approval gates, the audit trail, the evaluation set, the behaviour when the thing cannot proceed, the handover to people who did not build it. Cut them and you have a demo that reached production, which is a different and worse thing than a system.
We build apps and the agents inside them for organisations where a wrong automated action is expensive: money moves, a patient record changes, a container sits. In that setting the question is never whether a model can do it. It is whether anyone can tell what it did, and undo it.