Buyer guide

Define the Metrics Before the Pilot, Not After It

ByClunic Research Team10 min read

Free tool

AI Scribe ROI Calculator

Put your own numbers in and see the range, not one flattering figure.

Need it signed off?

Thirty free minutes with an analyst on the vendor, the workflow and the rule you are unsure about.

Book an evaluation call

Why does the baseline have to come first?

Because a baseline measured after the change is not a baseline, and a pilot without one cannot produce a decision. It produces an anecdote that people argue about.

This is the most common failure we see, and it is not a sophistication problem. Organisations that run rigorous clinical audits will start an AI pilot on a Monday with no record of what the preceding month looked like, because the vendor was ready and the enthusiasm was there. Six weeks later the question is whether documentation time fell, and the only available answer is that it feels better.

Two weeks of baseline is usually enough at practice scale, four at hospital scale. It does not need to be sophisticated. Self-reported minutes per clinic session, collected consistently, beats an elegant measure collected once. The point is that the number exists, was recorded before anything changed, and has a date attached to it.

The second reason to do this first is political. A baseline agreed in advance removes the argument about whether the result counts. When the number moves, nobody relitigates the measurement, and when it does not move, the decision to stop is already made. This is the Chart phase of the CARE method, and it is the phase people skip because it produces no visible progress.

What is the difference between adoption and outcome metrics?

Adoption tells you whether the tool was used. Outcome tells you whether using it changed anything worth paying for. You need both, and you must not mistake the first for the second.

LayerExample metricsWhat it provesFailure mode
AdoptionPercent of eligible encounters run through the agent, weekly active clinicians, sessions per userPeople are actually using itHigh adoption with no outcome change means the work moved rather than reduced
QualityEdit rate, percent of outputs accepted without substantive change, error escalationsThe output is good enough to trustLow edit rate can mean high quality or low scrutiny. Sample manually to tell them apart
OutcomeDocumentation minutes per session, days in accounts receivable, first-pass approval rate, calls answered within 30 secondsThe thing you are buying actually movedMoves too slowly to see in a six week pilot unless the baseline is solid
ExperienceClinician net satisfaction, patient complaint volume, staff intent to continueIt will survive contact with the rest of the organisationEarly adopter enthusiasm generalises badly

The rule of thumb: one outcome metric, two adoption metrics, one quality metric, one experience question. Five numbers. A pilot dashboard with fifteen metrics is a pilot nobody will read the dashboard of, and the discipline of choosing one outcome metric forces the conversation about what you are actually buying.

Choose the outcome metric by workflow. For ambient documentation it is documentation minutes per session or pyjama time. For prior authorization it is first-pass approval rate or touch time per authorization. For a phone agent it is abandonment rate or calls answered within a target. For denial management it is overturn rate on appealed denials.

How long before a pilot should show something?

It depends on the workflow, and the honest planning number is not the vendor's number.

WorkflowFirst adoption signalFirst credible outcome signalWhy the lag
Ambient documentationWeek 1Weeks 4 to 6Clinicians need a fortnight to stop fighting the template
Phone and scheduling agentsWeek 1Weeks 3 to 4Call volume is high, so the sample builds quickly
Prior authorizationWeeks 2 to 3Weeks 8 to 12Payer decisions have their own clock, and the outcome is downstream
Denial management and codingWeeks 2 to 4One full billing cycle, often 90 daysThe money arrives after the claim, not after the change

The practical consequence is that pilot length should be set by the workflow rather than by the calendar quarter. A six week prior authorization pilot cannot show an outcome, so either run it for a quarter or accept in advance that you are testing feasibility and adoption only, and say so.

Be explicit about that distinction when you design it. A feasibility pilot asks whether this can work here. An outcome pilot asks whether it is worth paying for. Both are legitimate. Running the first and reporting it as the second is how organisations end up scaling something that never had a business case.

When should you kill a pilot?

When the condition you wrote down in advance is met. The value of a stopping rule is entirely in its being written before anyone is invested.

A workable rule has three parts: a date, a threshold, and a named decision maker. For example: by the end of week six, if median documentation minutes per session have not fallen by at least fifteen percent against baseline, the medical director stops the pilot and we do not proceed to contract. That is specific enough to be uncomfortable, which is the point.

Three signals that should trigger an early stop regardless of the numbers. Silent abandonment, where a participant quietly stops using the tool and nobody notices until you check the logs, which usually means editing burden exceeded time saved. A quality incident where an output reached a patient or a chart without adequate review, which is a process failure that a better model will not fix. And integration drift, where the workaround that made the pilot possible turns out not to be sustainable at volume.

The hardest case is the pilot that is going fine but not well. Adoption is decent, everyone is mildly positive, the outcome metric has moved four percent. That is the pilot that gets extended twice and eventually bought out of fatigue. The stopping rule exists specifically for that case, and honouring it is the difference between a portfolio of deployments and a portfolio of subscriptions.

Who should own the measurement?

Not the vendor, and not the person who championed the purchase.

Vendors will offer to supply the metrics, and their dashboards are usually good. The problem is not accuracy, it is selection: a vendor dashboard reports what the vendor instruments, which is adoption and usage, because that is what the product can see. It cannot see your baseline, your loaded cost per hour, or the work that moved somewhere else. Take the vendor's data as one input and own the outcome metric yourself.

The champion should not own it either, for the ordinary reason. Give measurement to someone who is neutral about the result: a practice manager, a quality lead, an analyst. They do not need to be senior. They need to be uninvolved in whether the answer is yes.

One more role worth naming: someone who talks to the participants weekly. Numbers tell you what happened and conversations tell you why, and in a pilot of five clinicians the conversation is often the more reliable instrument. Ask what they stopped doing, not just what they think of the tool. The answer to what they stopped doing is where the time actually went.

What do you measure for safety rather than value?

Three things, and they are separate from the value metrics because they have a different threshold: value metrics need to move, safety metrics need to stay at zero.

Sample manually. Pull a fixed number of outputs each week, twenty is usually enough, and have a clinician read them against the source. Edit rate from the product tells you what people changed. Manual sampling tells you what they should have changed and did not, which is the only way to distinguish a genuinely good output from an unscrutinised one.

Track near misses explicitly, with a channel that takes ten seconds to use. An output that was wrong and was caught is the most valuable data in a pilot, and it will not be reported if reporting it feels like raising an incident. Ask for them in the weekly conversation rather than waiting for a form.

Record every case where the agent acted outside its intended scope, however trivial. A scheduling agent that booked the wrong slot type, a phone agent that failed to transfer when it should have. These are the incidents that scale badly: harmless at five clinicians per week and material at five hundred. Where the workflow touches a patient directly, this is also the evidence base for the disclosure and human review obligations described in the HIPAA and AI page, and the reason governance is designed alongside the agent rather than after it in the governance engagement.

How do you count the savings honestly?

Decide in advance what reclaimed time converts into, and only count it once.

Time saved is not money until something happens to it. Three legitimate conversions: additional visits, reduced overtime or agency cost, or retained clinicians who would otherwise have cut sessions. A fourth, going home earlier with no financial effect, is a genuine and often sufficient reason to buy, but it is not a cost saving and calling it one will not survive a finance review.

Pick one conversion, state the assumption, and expose it. The AI scribe ROI calculator and the prior authorization cost calculator both work this way deliberately: the assumptions are visible and changeable, so a sceptical CFO can move them rather than reject the model. A business case with hidden assumptions loses the argument even when it is right.

Two counting errors worth naming. Double counting, where the same reclaimed hour appears as both added visit revenue and reduced overtime. And attributing to the agent a change that came from something else that happened in the same quarter, which is why the baseline period should be recent and the pilot should not coincide with a staffing change if you can avoid it.

What does a complete pilot definition look like?

One page, written before the vendor has access to anything, containing eight lines.

  1. The workflow, named precisely, and the sites and clinicians in scope.
  2. The baseline: what is measured, by whom, over which two or four weeks.
  3. One outcome metric, with the threshold that would justify the spend.
  4. Two adoption metrics and one quality metric.
  5. The pilot length, set by the workflow rather than the quarter.
  6. The stopping rule: date, threshold, named decision maker.
  7. What reclaimed time converts into, and the assumption behind it.
  8. Who owns measurement, and who talks to participants weekly.

That page is the Chart and Evaluate ends of the CARE method, and it is what makes the loop work: chart what the work really is, architect the agent and the governance together, run it in a bounded pilot, and evaluate against the baseline before anyone decides to scale. The evaluation is what tells you whether the design was right, which is why it cannot be defined afterwards. The full method sits on the services overview.

If you are still choosing between vendors, the compliance screen and this metric definition should run in parallel rather than in sequence, and the questions for the first are in the fifteen vendor questions. If you have not yet decided which workflow to pilot, that is the prior question and the wrong one to answer by pilot. A three week AI readiness audit produces a ranked shortlist with the stopping rule written down in advance, and it is deliberately capable of concluding that you should not buy anything yet.

Sources

Primary material behind the claims above. Read the source before acting on any summary of it.

Questions we get asked

What metrics should I use for an AI pilot in healthcare?

Five: one outcome metric tied to why you are buying, two adoption metrics, one quality metric such as edit rate, and one experience question. More than that and nobody reads the dashboard. The outcome metric should be chosen by workflow, and it must have a baseline measured before the pilot starts.

How long should an AI agent pilot run?

Set the length by workflow rather than by calendar. Ambient documentation shows an outcome signal in four to six weeks. Prior authorization needs eight to twelve because payer decisions have their own clock. Denial management and coding need a full billing cycle, often ninety days.

What is a stopping rule and why write it in advance?

A stopping rule is a date, a threshold and a named decision maker, agreed before the pilot begins. Written in advance it makes the decision automatic. Written afterwards it becomes a negotiation, and the pilot that is going fine but not well gets extended twice and bought out of fatigue.

Can we use the vendor's dashboard as our metrics?

Use it, but do not rely on it alone. A vendor dashboard reports what the product can instrument, which is adoption and usage. It cannot see your baseline, your loaded cost per hour, or work that moved elsewhere rather than disappearing. Own the outcome metric yourself.

How do we count time savings in a business case?

Decide in advance what reclaimed time converts into: additional visits, reduced overtime or agency cost, or clinician retention. Count it once, and state the assumption openly so a finance reviewer can change it rather than reject the model. Going home earlier is a good reason to buy but is not a cost saving.

What if adoption is high but nothing else moved?

That usually means work moved rather than reduced. Ask participants what they stopped doing, and look for the task that grew: more editing, more review, more chasing. It is a real result and it is the signal to renegotiate scope or stop, not to extend the pilot in the hope that the outcome catches up.