AI Agent ROI: Measure Completed Work, Not Model Calls
Build an AI agent business case with cost per completed case, safe completion, cycle time, rework, escalation, and verified value.

AI agent ROI should not begin with the price of a million tokens or the number of agent runs. A run may produce no outcome, create rework, or merely shift cost to another team. A better unit is one safely completed case: a customer request resolved, an application verified, an invoice moved to a valid state, or a report delivered on time.
The business case must compare baseline and pilot across the same type of work, quality bar, and time window. Total cost includes models, integrations, infrastructure, monitoring, review, rework, and incidents. Benefit is recognized only when the business system verifies the outcome.
Unit
Move from runs to completed cases
A polished output is not a completed job.
Define case_id, starting state, completion criteria, and the authoritative system. In support, completion is not the sent answer; it is a resolved ticket that does not reopen during a defined observation window. In accounts payable, completion is not OCR output; it is a valid, nonduplicate invoice that passes the correct approval policy.
Keep the denominator transparent: all eligible cases, not only those the agent accepted. Separate completed_by_agent, completed_with_human, handed_off, failed, cancelled, and policy_blocked. Excluding difficult work creates an attractive but misleading completion rate.
OpenAI recommends establishing an evaluation baseline, meeting the accuracy target, and then optimizing cost and latency. See the agent guide. The sequence matters: do not substitute a cheaper model before measuring the outcome impact.




