Analysis

The Pilot Worked. Why Did Nothing Reach Production?

ByClunic Research Team12 min read

Free tool

AI Readiness Assessment

Twelve factual questions on data access, governance and change capacity.

Need it signed off?

Thirty free minutes with an analyst on the vendor, the workflow and the rule you are unsure about.

Book an evaluation call

How common is the pilot to production gap?

Widely enough that it has become the defining pattern of enterprise AI, though the most quoted number deserves care. MIT's Project NANDA published a working report in 2025, The GenAI Divide: State of AI in Business, reporting that around 95 percent of generative AI pilots produced no measurable profit and loss impact, based on roughly 150 leadership interviews, a survey of 350 employees and an analysis of 300 public deployments. It is not peer reviewed and the sampling is not representative, so treat the figure as directional rather than precise. What the report describes is nevertheless recognisable: a small minority of deployments capturing most of the value, and a majority stuck at the pilot stage.

Healthcare has its own version, and the causes are more specific than general enterprise inertia. Clinical workflows are already loaded, integration is genuinely hard, the governance path is often undefined, and the people who benefit from a change are frequently not the people who pay for it. The result is an organisation with a portfolio of successful pilots and no production systems, which is a worse position than having done nothing, because the credibility is spent.

The four causes below account for most of what we see. None of them is about model quality. All of them are decided in the weeks before a pilot starts, which is the useful part, because it means they are preventable rather than merely diagnosable.

What is integration debt, and why does it kill pilots?

A pilot is usually allowed to skip the hard integration. Users log into a separate application, data is exported nightly, results are pasted back by a coordinator, and the vendor's implementation team quietly does by hand what would need automating at scale. The pilot succeeds. The work that was skipped is the work that determines whether production is possible, and nobody sized it.

The specific debts recur:

  • Write back. Reading from the EHR is comparatively easy. Writing structured data back into it, in a way that satisfies clinical governance and audit, is the expensive half, and it is the half pilots defer.
  • Identity and single sign on. A pilot with fifteen users can manage separate credentials. Two thousand users cannot, and provisioning at that scale is its own project.
  • Context passing. Launching the tool with the right patient already in context is what makes it usable in a seven minute encounter. Without it, adoption dies regardless of output quality.
  • Interface queues and error handling. Pilots ignore the failure path. Production is mostly failure path.
  • Vendor change control. Your EHR vendor's upgrade cycle will break something. Who tests, who fixes, and inside what commitment.

Ask one question before any pilot begins: what will be done by hand during this pilot that would have to be automated in production, and who has estimated that work? If nobody can answer, the pilot is testing the model rather than the deployment, and those are different questions. The realistic scope varies substantially by platform, which is why we keep separate integration notes for Epic, Oracle Health, athenahealth and eClinicalWorks.

Why does a missing baseline end a programme?

Because without one, the year end review becomes an argument about impressions, and impressions lose to budget pressure every time.

The sequence is predictable. A pilot runs for ten weeks. Clinicians report liking it. The finance team asks what changed, and the only available answers are survey sentiment and a vendor dashboard measuring activity inside the vendor's own product. Nobody captured total EHR time, after hours time, note turnaround, denial rates or whatever the relevant operational measure was, in the eight weeks before go live, using a definition that was written down. The programme cannot prove anything, so it is renewed on faith once and cancelled the second time.

The fix is unglamorous and cheap. Before the pilot: name three metrics, define them precisely including any timeout or attribution rule, pull at least eight weeks of history, and segment by the unit that will vary, usually specialty or department. Keep a concurrent comparison group of similar clinicians who are not enabled, which a wave rollout gives you at no cost. Write down in advance what result would count as success and what result would cause you to stop.

That last clause is the one that gets omitted, and it is the one that makes the rest work. A pilot with no stated failure criterion cannot fail; it can only be extended. Extended pilots are how organisations spend two years and reach no decision.

The published evidence base is thinner than most business cases assume, which raises the value of your own numbers rather than lowering it. We set out what the documentation literature does and does not support in the pajama time evidence base, and the modelling assumptions we would defend in the AI scribe ROI calculator.

What happens when there is no governance route?

The pilot reaches its end, produces a good result, and then discovers that nobody knows who approves moving it into production. Security has not completed a review because it was told this was a pilot. Legal has not read the agreement because the pilot ran on a trial term. Privacy has not seen a risk analysis. The clinical governance committee meets monthly and its agenda is full.

What follows is a four month approval crawl during which the pilot is switched off, the enthusiasts move on, and the eventual approval arrives for a deployment nobody is still positioned to run. The most common outcome is not rejection. It is expiry.

The preventable version of this is to run the governance path in parallel with the pilot rather than after it. Register the system before the pilot starts, assign its risk tier, complete the security and privacy reviews against the production configuration, and get the production decision criteria agreed in advance so the committee is ratifying a pre agreed threshold rather than opening a fresh debate. A committee that has published turnaround targets makes this straightforward, which is one of several reasons to build one properly, as set out in our AI governance committee playbook.

There is a second, subtler governance failure worth naming. Organisations frequently approve a pilot under one model of human oversight and then scale it under a different one, usually because the review step that made the pilot safe does not survive contact with volume. If the production configuration removes a human from the loop, it is a different system and needs a different review. That distinction is the subject of agents, copilots and automation, and it is where the HIPAA analysis and any device question also have to be revisited.

Who actually pays, and who actually benefits?

The most durable failure cause, because it is structural rather than procedural.

Consider a documentation tool. The budget sits with IT or with a system level innovation fund. The benefit accrues to individual clinicians as time returned in the evening. There is no mechanism by which that benefit becomes a line item, so at renewal the cost is visible and the benefit is not. The tool is cut, and everyone involved reports that it worked.

The same asymmetry appears elsewhere. A prior authorisation agent reduces work in a clinical department while the savings, if any, land in revenue cycle. A scheduling agent reduces no shows, which improves a metric owned by operations while the cost sits with the call centre budget. An AI phone agent reduces abandoned calls, which nobody was being measured on.

Three practical responses, in ascending order of effort:

  1. Name the beneficiary budget before the pilot. Not the sponsor, the budget that improves. If you cannot name one, the deployment is a wellbeing investment and should be argued as one, honestly, rather than as a financial case that will fail its own audit.
  2. Agree the conversion mechanism in advance. If saved clinician time is meant to become visits, the schedule change that converts it has to be agreed with the department before go live. Time saved without a conversion mechanism does not become revenue, and assuming it will is the most common modelling error we see.
  3. Put the renewal decision with the beneficiary. A tool whose renewal is decided by the department that experiences the benefit survives at a very different rate from one decided by the department that carries the cost.

How do the four failure modes compare?

The early warning signs are visible well before the pilot ends, if anyone is looking for them.

Failure modeWarning sign during the pilotWhen it usually surfacesThe preventive step
Integration debtManual steps performed by the vendor or a coordinator, separate login, nightly exportsMonth 4 to 8, at scale upList every manual step and size its automation before the pilot begins
Missing baselineSuccess discussed in terms of sentiment and vendor dashboardsFirst renewal reviewEight weeks of defined pre deployment metrics plus a stated failure criterion
Governance vacuumNobody can name who approves production, security review deferred as out of scopeThe week the pilot endsRegister and tier before the pilot, run reviews against the production configuration in parallel
Incentive misalignmentThe sponsor cannot name which budget improvesMonth 12 to 18, at renewalName the beneficiary budget and the conversion mechanism, and place the renewal decision with the beneficiary

Notice the timing. Integration debt and governance vacuums surface within months and are usually recoverable at some cost. Baseline and incentive failures surface a year later and are usually not, because by then the decision is being made by people who were not in the room and have only the numbers in front of them.

How does the CARE method avoid this?

CARE is how we run engagements: chart, architect, run, evaluate. It is not a novel framework so much as an insistence that the four steps happen in that order and that none of them is skipped because the technology looks straightforward.

PhaseWhat happensWhich failure it preventsArtifact
ChartMap the workflow as performed, measure the current state, register what AI is already running, and name the beneficiary budgetMissing baseline, incentive misalignmentBaseline metrics with definitions, workflow map, system register
ArchitectDesign the production configuration, size the integration work including write back and identity, set the risk tier and the oversight model, agree the decision criteriaIntegration debt, governance vacuumIntegration scope, risk tier and control set, written success and stop criteria
RunWave rollout with champions, a live feedback loop and a concurrent comparison groupAdoption collapseSegmented usage and quality reporting
EvaluateMeasure against the baseline on the pre agreed criteria and make an explicit expand, hold or stop decisionThe extended pilot that never concludesDecision memo with the numbers, readable without the consultant present

The load bearing element is that evaluate feeds back into run rather than closing the engagement. An engagement that cannot be measured is not finished, and a deployment whose measurement stops at go live will decay quietly until somebody notices at renewal.

The other deliberate feature is that the stop criterion is written during architect, before anyone is invested. Deciding in advance what result would cause you to stop is the cheapest governance control available, and it is the one most consistently omitted, because writing it feels like anticipating failure. It is the opposite: it is what allows a programme to stop one thing and keep its credibility for the next.

What should you do differently on the next one?

Six changes, none of which require a larger budget.

  1. Stop running pilots that test the model. The models work. Run pilots that test the deployment: the integration, the oversight model, the exception path and the workflows you expect to fail.
  2. Capture the baseline first. Eight weeks, defined metrics, segmented, with a comparison group. If this is not possible, that finding is itself worth having before you spend anything.
  3. Write the stop criterion. One paragraph, agreed by the sponsor, before go live.
  4. Run governance in parallel. Register, tier and review against the production configuration during the pilot, not after it.
  5. Name the beneficiary budget. And if there is not one, argue the case you actually have rather than the one that scores better.
  6. Pilot the worst case. Include the difficult specialty, the noisy environment and the non English encounter. A pilot that only covers favourable conditions tells you what you already assumed.

If you would rather not learn these in sequence, that is what our engagements are for. An AI readiness audit produces the chart phase, the deployment roadmap produces the architect phase, and AI governance and compliance makes sure the route from pilot to production exists before anyone needs it. The full picture of how we work sits on our services page, and the underlying decision about whether to buy any of this at all is covered in build versus buy for healthcare AI agents.

Sources

Primary material behind the claims above. Read the source before acting on any summary of it.

Questions we get asked

Is it true that 95 percent of AI pilots fail?

That figure comes from MIT Project NANDA's 2025 report The GenAI Divide, which found around 95 percent of generative AI pilots showed no measurable profit and loss impact. It is a working report based on roughly 150 interviews, a 350 person survey and 300 public deployments, not a peer reviewed study with a representative sample, so treat it as directional. The pattern it describes is consistent with what we see in healthcare, but the precise number should not be quoted as established fact.

How long should a healthcare AI pilot run?

Eight to twelve weeks of live use, after at least eight weeks of baseline capture. Shorter than eight weeks and you are measuring novelty; longer than twelve and the pilot has usually become a permanent state that avoids a decision. Fix the end date and the decision criteria in writing at the start, and treat any extension as requiring the same approval as a new pilot.

What should we measure during a pilot?

Three operational metrics defined before go live, segmented by the unit that will vary, plus adoption and quality. For documentation that usually means total EHR time per scheduled hour, time outside scheduled hours and note turnaround, alongside weekly active clinicians and a sampled note quality review. Avoid metrics that only exist inside the vendor's dashboard, because they measure use of the product rather than the outcome you are buying.

Should we pilot with enthusiasts or with a representative group?

Both, deliberately separated. Enthusiasts tell you the ceiling and surface product defects fast. A representative group tells you what a full rollout will look like, which is the number that matters for the business case. Reporting the enthusiast result as the expected outcome is the single most common way a pilot produces a forecast that does not survive contact with the second wave.

Who should own a healthcare AI pilot?

An operational owner in the department where the work actually happens, with a technical owner alongside them and an executive sponsor who can release budget. Pilots owned solely by IT or by an innovation function tend to optimise for launching rather than for finishing, because that is what those functions are measured on. The owner should be someone whose own numbers move if the deployment works.

What is the cheapest way to reduce the risk of a failed deployment?

Write the stop criterion before you start. It costs a paragraph and it forces the sponsor, the operational owner and the finance partner to agree in advance what the deployment is for and what result would end it. Almost every other control on this list is more expensive, and none of them helps an organisation that cannot bring itself to conclude a pilot.