A model can produce a test case in seconds. Whether that test case is useful is a different question. Someone still has to check the assertion, inspect the data assumptions and decide whether it protects a meaningful product behaviour.
Our starting point is the accepted outcome: a piece of work that passes an agreed quality bar and can be used for its intended purpose. Measuring that unit makes comparisons between human work, deterministic automation and AI assistance more useful.
Define acceptance before the pilot
Choose a bounded workflow. For a test design pilot, an accepted outcome might be a reviewed test specification with traceable requirements, realistic data, meaningful assertions and no duplicate coverage. A generated paragraph is not automatically an outcome.
Write the acceptance rubric before evaluating the sample. Include reasons for rejection and what requires correction. Keep the same rubric for the baseline and the proposed workflow. Otherwise a lower quality bar can make an experiment look cheaper without improving the work.
Count the full workflow
Start with direct execution costs: model usage, infrastructure and any tools used specifically for the workflow. Convert input and output usage into money using the actual model, price and billing arrangement for that run. Record that price basis because it can change.
Then add human preparation, review, correction, retries and final acceptance. Failed attempts still consumed resources, so their costs remain in the numerator even when their outputs never become accepted work.
Cost per accepted outcome = total in-scope workflow cost / number of accepted outcomes.
Record setup and integration separately, then explain how you allocate them over an expected volume. If there are no accepted outcomes, report the experiment as unsuccessful; do not hide the failure behind a zero or undefined unit cost.
An illustrative calculation
Suppose an experiment produces 100 candidate outputs. Model usage costs €80, review costs €240 and correction costs €180. Eighty outputs meet the acceptance criteria.
The operating cost is €500, and the cost per accepted outcome is €6.25. Dividing only model cost by all candidates would suggest €0.80. Those numbers describe different things.
These are hypothetical values, not client results or a promised return. A real comparison also needs any applicable setup, licence and infrastructure allocation. The human-only baseline needs an equivalent accounting boundary.
Separate money from capacity
A shorter task can free capacity without reducing payroll or external spend. Both can matter, but they support different business decisions. Report cash cost changes separately from hours released, and show where the released capacity can realistically go.
Use sensitivity ranges when review effort or future volume is uncertain. A workflow that works for a large, repetitive workload may not justify its setup cost for an occasional task. Acceptance rate and error consequences also matter alongside the average cost.
Make a decision you can revisit
End the pilot with a decision: expand, revise or stop. Preserve the sample, acceptance rubric, cost assumptions and known limitations so the result can be challenged later. Re-evaluate when the workflow, model or quality bar changes.
Our QA & AI Value Engineering service turns this approach into a baseline, bounded experiment and a decision grounded in your own workflow.
