A useful agent has a job: “create the refund if policy matches, else escalate.” It has tools that are typed and logged. It has a stop: max steps, missing data, or a class of action that is irreversible.
We measure completed jobs, handed-off jobs, and refused jobs. A high “conversation length” is not a success metric. An audit trail that a supervisor can read on Monday is.
If you cannot name the system of record the agent writes to, you do not have an automation project. You have a demo.

