How to compare AI models, effort settings and instructions fairly

6 minute read · Updated 4 October 2026

Should a task use a different model, a lower effort setting or a new version of its instructions? The honest way to decide is to run both setups on the same work, count how often each one met your success criterion, and note what each cost in tokens, money and time, retries included. This guide shows how to do that, and how Vaze keeps the result next to the workflow it belongs to.

1. Fix the task, the workload and the success criterion

Before comparing anything, write down three things and keep them the same for both setups:

  • Task: the job, for example summarising support tickets.
  • Workload: the exact set of inputs both setups will run on, such as the same 40 archived tickets.
  • Success criterion: how a reviewer decides an output is good enough, for example “names the customer’s issue and the next step”.

If the workload or criterion changes between runs, the comparison stops being fair.

2. Run both setups and count attempts, including retries

Run the baseline (what you use today) and the candidate (what you want to try) on the same workload. Count every attempt, including retries after a failure or timeout, and count how many attempts met the criterion. Use the same number of attempts for both, so the pass counts are comparable.

3. Total the tokens, cost and latency, retries included

Add up total input tokens, total output tokens, total API cost and total latency across all attempts, retries included. If outputs also go through an automated check or a second review pass, include that overhead in the totals too. Retries and review are part of the real cost of a setup, so leaving them out flatters whichever setup failed more often.

Keep this to usage-based API cost. Do not fold subscription or seat charges into the comparison; they are billed differently and belong in your spend records.

4. Keep measured values and estimates apart

Mark each setup as an observation you measured, or as an estimate you worked out from list prices or a smaller sample. Both are useful, but they should never be mixed in one conclusion. An estimate is a reason to run a proper test, not a result.

A worked example (fictional)

A fictional company wants to know whether its support-summary workflow can use a lower effort setting with a revised instruction. Both setups ran on the same 40 archived tickets, with 4 retries each, so 44 attempts each. All figures below are made up for illustration.

Fictional worked example: two setups for the same support-summary task, both observed and entered by the team
FieldBaselineCandidate
Model and dated versionModel A · 2026-09Model A · 2026-09
Effort settingStandardLow
Instruction / prompt revision labelv3v4
EvidenceObserved · entered by youObserved · entered by you
Attempts (including retries)4444
Attempts meeting the criterion3738
Total input tokens52,80061,600
Total output tokens9,2407,920
Total API costGBP 0.62GBP 0.48
Total latency (ms)396,000264,000

The candidate met the criterion slightly more often, used more input tokens because its instructions are longer, and finished faster at a lower recorded cost for this workload. A one-off result like this describes this workload on this date. It does not prove equal quality on other work or a lasting saving, so the team keeps both setups recorded and retests when anything changes.

Recording a comparison in Vaze

Run the tests outside Vaze, with tools and data you are authorised to use. In Vaze, enter labels and aggregate counts only. Never paste prompts, outputs, ticket contents or private test sets into it.

  1. Open More, then Environment, and select the asset the comparison belongs to, such as the support-summary instruction set. If it is not recorded yet, choose Add asset.
  2. Under Optimisation comparisons, choose Record a comparison.
  3. Fill in Task label, Workload / test-set label, Success criterion and Observation date (UTC).
  4. For both Baseline and Candidate, enter the Model and dated version, Effort setting, Instruction / prompt revision label and Evidence (Observed · entered by you or Estimated scenario), then Attempts (including retries) and Attempts meeting the criterion.
  5. Optionally add Total input tokens, Total output tokens, Total API cost and Total latency (ms), then choose Save comparison.

Recording a comparison needs write access to your company’s dashboard, and the asset must not be retired. Entering costs also needs Spend access: cost fields and the currency are only shown to people with it, and others see Withheld · Spend access required. Vaze asks for the same number of attempts on both sides. Saved observations cannot be edited, so to correct or repeat a test you record a new comparison.

When something changes, retest

Each comparison remembers the recorded setup it was saved against: the asset’s record in Environment and the records of the assets linked directly to it. If any of those records is edited or retired, the comparison shows Setup changed · retest instead of Recorded setup unchanged. That is the prompt to run the workload again and record a fresh comparison.

Vaze compares its own records only. It does not see changes inside the underlying files, prompts or systems, so update the record when you change one of those, and retest.

Common questions

Does Vaze run the comparison for me?

No. You run both setups and enter the attempts, passes and totals you observed or estimated. Vaze records them, keeps them with the right asset and tells you when the setup has changed.

Why count retries?

Retries cost tokens, money and time. Leaving them out makes a setup that fails more often look cheaper and faster than it really is.

Can I compare an estimate with a measurement?

You can record either, and each side is labelled. Treat an estimate as a reason to run a proper test, not as a result to compare against a measurement.

See your company’s AI in one place.

Sign in with Google or Microsoft, name your organisation and add the tools you already use. Free, with no card.

Get started free
Help & support