An AI agent should be measured by the business outcomes people accept, not by how many messages it writes, tools it calls, or steps it completes. A useful scorecard connects each case to an accountable owner, a clear acceptance decision, the effort required to correct the result, and any customer or safety consequence. That structure exposes an agent that looks busy while creating rework.
Build the metric tree from accepted outcomes
Begin with the event that creates the work and the business state that closes it. For a support agent, the unit might be a case that remains resolved after review. For a sales operations agent, it might be a record accepted into the next stage without cleanup. Define the unit before choosing a dashboard. Otherwise, the easiest activity to count will quietly become the goal.
Choose a named business owner who can accept or reject each outcome. Segment results by case type, risk and difficulty so that easy volume cannot hide failures on important cases. The three checks for invented agent metrics are useful here because every reported gain should be traceable to a source record, a calculation and an independent read-back.
A practical metric tree has one primary measure and several protections. The primary measure can be cost per accepted business outcome. The protections show whether that result was achieved by shifting work or risk elsewhere. Compare accepted outcomes with reopened cases, human corrections, exceptions, cycle time, operating cost and harmful results. Do not combine them into one flattering score.
- Outcome: the target business state was reached and accepted by the owner.
- Correction: a person had to change the agent output before acceptance.
- Exception: the case left the automated path because a rule, input or tool failed.
- Consequence: the result affected a customer, obligation, payment or protected record.
Count correction and exception debt
Correction work is part of the agent cost, even when it happens in another team. Record who corrected the case, what changed and how long the correction took. Review exceptions separately. A high exception rate may be appropriate for rare or risky cases, but it should not be disguised as successful automation. The goal is a stable operating boundary, not maximum autonomy.
Use representative cases rather than a smooth demonstration. Preserve the original input, the agent proposal, the tool response, the reviewer decision and the final accepted state. If the agent closes many tickets that customers reopen, the reopen belongs in the same measurement window. A local or lower-cost model can change the economics, but the deployment cost discussion only matters after outcome quality is measured on the real workflow.
Review the metric tree by consequence. Low-impact classification errors can tolerate a different threshold from an incorrect payment, contract change or customer promise. Set a stop condition for each band. If quality falls, corrections rise or a safety boundary is crossed, the agent should lose authority until the cause is understood.
Make the decision from evidence, not activity
At each review, ask three questions. Are accepted outcomes increasing without hidden correction work? Are exceptions moving toward the cases the team intended to reserve for people? Are cost and cycle time improving without weaker quality or safety? A dashboard that cannot answer those questions is reporting activity, not performance.
Completion evidence also matters. The verified completion example shows why an agent statement is not enough. The measurement record should include the target state and an independent check of the system that owns it. This prevents confident narratives from entering the accepted-outcome count.
Frequently Asked Questions
Lead with the cost per business outcome accepted by the process owner, rather than message volume, tool calls or completed steps.
Add correction effort to operating cost, track exceptions separately by case type and consequence, and never count either as an accepted outcome.
High message, tool-call or step counts without accepted business outcomes show that the dashboard measures activity instead of value.
The tree should be led by cost per accepted business outcome, with corrections, exceptions, quality and safety as explicit protections.
The final deliverable is a metric tree with definitions, owners, data sources, review cadence and rejection conditions. It should be small enough for an operator to use on the next batch of cases. Teams that need help defining the evaluation, authority boundary and production checks can review AI4SALE AI agent development. The service link is a next step, not evidence that the current agent performs well.
