Skip to content
ABTO Guide

How-to guides

Switch to a cheaper model

Switch to a cheaper model only after confirming that outcomes hold.

The pass mark for a cheaper model is lower than the incumbent’s. It does not have to come out ahead; cutting cost while outcomes hold is enough to pass. So before moving the review summary feature to a model at half the price, two numbers decide it: does it really get that much cheaper, and do outcomes stay where they were.

In Testing, put the variant currently in production and the cheaper candidate side by side, then run them on real customer inputs via “Load from history”. Comparing the cost numbers in the results shows right away whether “half the price” holds up on your traffic too.

Swap in a handful of different inputs and eyeball the response quality as well. If quality visibly collapses, stop here. A few test calls’ worth of cost weeded out a bad candidate before it reached any user.

Save the candidate that passed as a variant and give it 10% in Routing. For a few days, watch the new variant’s error rate and p50 latency on Compare, and drop the ratio back to zero if either spikes.

Once Compare’s low-sample warning clears, read in this order.

CheckBar
Error rateIs it at the same level as the incumbent
Success metricIs it the same as before. If slightly lower, let revenue per AI dollar settle it
Cost per callDo the savings from the first step show up on real traffic too

Once it passes, raise the ratio in Routing and drop the incumbent to zero. The before and after shows in the cost trend on Overview’s 30-day range, and if you have a “revenue per AI dollar” metric set up, you can report to the team in one sentence like “success rate held, revenue per AI dollar up 40%”.

If that metric does not exist yet, set it up first with Prove the impact.