Define success before the first call
Before running a pilot, write down what the team counts as “done”: required fields complete, no invented data, internal rules respected. Without that criterion you will compare draft speed with already verified work — and overestimate the gain.
Choose a single task type with enough volume for 20–50 real or representative cases. Expanding to several flows at once makes it hard to tell what worked and what it cost.
Measure the current way of working first
On a small sample, time the average until an accepted result: reading, searching, drafting, correcting, sending. Include typical interruptions. That baseline is the reference, not a figure from a public demo.
Also note hidden costs: email clarifications, duplicated data, exceptions that always reach the same person. A pilot that only moves work to another stage does not improve ROI.
Three metrics worth tracking together
Acceptance rate: how many results pass review without major corrections. Review time: minutes a person spends validating or correcting. Service cost: the bill for the pilot period, including retries.
A cheap call with low acceptance can cost more than a stronger call that passes first time. That is why ROI is calculated per accepted outcome, not per number of calls. Details on choosing effort level and models by role appear in our guide on cost discipline in AI projects.
A simple calculation, without universal promises
Illustrative: if you have 80 cases a month, the baseline is 10 minutes per case, and after the pilot 4 minutes of review remain at a service cost equivalent to 1 minute of work per case, the gross “saving” is 5 minutes × 80. Subtract pilot preparation and maintenance. If the figure is negative or fragile, the process is not worth automating yet — or first needs a clearer form, without a model.
Do not use a marketing growth percentage as a target. Compare similar periods and case types. A successful pilot can also mean deciding not to scale: you learned quickly that review costs as much as the original work.
What you decide after the pilot
Scale only if acceptance is stable, review stays sustainable and cost per accepted outcome is clearly below the value of time saved. Document the policy: what runs on rules, what on an economical model, what on a more capable one, and who approves exceptions.
Bring the baseline, the case sample and the “done” criterion to a project discussion. We can propose a bounded pilot with the same metrics so you know whether the investment deserves the next step.

