Nobody Is Using It: The Clinician AI Adoption Playbook for After Go Live
ByClunic Research Team11 min read
Free tool
Clinician Turnover Cost Calculator
This is the number a documentation-burden or burnout business case has to beat: what nurse and physician turnover already costs your organization every year, before anyone proposes a fix for it.
Need it signed off?
Thirty free minutes with an analyst on the vendor, the workflow and the rule you are unsure about.
Book an evaluation callWhy does adoption stall after go live?
The pattern repeats across health systems with unhelpful regularity. Selection takes four months and involves everyone. Go live is a launch event with training slots and a support line. Six weeks later, the licence report shows that a third of the enabled clinicians have used the tool in the past week, and the finance team asks what happened to the business case.
Nothing unusual happened. The project treated go live as the finish line, when for a clinician it is the start of a personal cost benefit calculation that runs every session. A clinician who tries an ambient documentation tool on a complex visit, gets a note that takes eight minutes to fix, and has twenty two patients that day, has learned something and will act on it. No amount of launch communication reaches that clinician afterwards.
Adoption work is therefore mostly failure detection: finding the specific workflows where the tool produces more work than it saves, and either fixing them or excluding them honestly. Systems that do this well hold usage; systems that respond with more training do not, because the problem was never that people did not know how to press the button.
This is the operational half of a broader problem we cover in why AI pilots fail in healthcare. That piece is about the gap between pilot and production; this one is about the gap between production and use.
How do you measure real usage rather than licensed seats?
Almost every AI adoption report that looks good is measuring the wrong thing. Seats enabled is a procurement number. Logins is an attendance number. Neither tells you whether the tool is doing the job.
| Metric | What it actually tells you | How to measure it | Warning sign |
|---|---|---|---|
| Seats enabled | What you are paying for. Nothing else. | Contract and licence report | Used as the headline adoption number |
| Weekly active clinicians | Whether the tool survived first contact | Distinct clinicians with at least one completed use in a rolling 7 days, over seats | Below 50 percent at week 8 |
| Encounter penetration | Whether it is the default or the exception | Encounters using the tool over eligible encounters, per clinician | High user count with low penetration means selective use on easy visits |
| Sustained users at 90 days | Whether the value was real or novelty | Clinicians active in week 1 who are still active in week 13 | Below 60 percent retention |
| Edit burden | Whether the output is usable as drafted | Time from draft available to signature, and character level edit distance if the vendor exposes it | Rising over time rather than falling |
| Abandonment rate | Where the tool breaks | Sessions started and not completed, segmented by visit type | Clustering in one specialty or visit type |
The segmentation matters more than the totals. An aggregate of 62 percent weekly active hides the fact that primary care is at 85 percent and orthopaedics is at 12 percent, and those are two different problems with two different fixes. Report by department, by specialty and by visit type from the first week, and insist the vendor provides the data at that grain before you sign. If they cannot, that is a selection finding, and one worth raising during vendor selection rather than after.
Capture the baseline before go live. Systems that skip this cannot tell whether after hours documentation time fell, and end up arguing about survey sentiment instead. The measurement plan belongs in the deployment roadmap, not in a retrospective.
Which champion model works?
Champions are the most reliable adoption lever available and the most commonly implemented badly. The failure is appointing enthusiasts, announcing them, and giving them nothing.
Three models, in descending order of how well they hold up:
- Embedded peer champion with protected time. One clinician per department, four hours a month of released clinical time, an explicit remit to sit with colleagues during real clinics rather than to run training sessions. This is the model that works, and the protected time is what makes it work. A champion doing this on goodwill lasts about five weeks.
- Super user network without protected time. Cheaper, and adequate for a tool with low workflow complexity. Expect the network to decay within a quarter and plan a refresh rather than pretending it will not.
- Central specialist team. A small central group that rotates through departments. Useful for the first six weeks and for difficult specialties, and weak afterwards, because the credibility of a champion comes largely from being a colleague who carries the same panel.
Choose champions for credibility rather than enthusiasm. The most useful champion in a department is often a respected sceptic who has been persuaded by their own data, because their endorsement carries information that an early adopter's does not. Ask department leads who their colleagues actually listen to, and accept the answer even when it is inconvenient.
Give champions a real feedback channel upward. In the AMA's 2024 survey of physician sentiment on AI, a designated feedback channel was the most cited requirement for advancing adoption, named by 88 percent of physicians surveyed, ahead of data privacy assurances and EHR integration. Champions who collect complaints that visibly go nowhere stop collecting them.
How do you close the note quality feedback loop?
The single highest leverage mechanism, and the one most often run on a quarterly cadence when it needs to run on a weekly one.
The loop has four steps and each has a failure mode:
- Capture. A one click way to flag a bad output from inside the workflow, with an optional comment. If flagging requires leaving the EHR or filling a form, you will receive perhaps a tenth of the signal. Route the flag to a named human, not a shared mailbox.
- Triage. Someone reads every flag within two working days and classifies it: configuration, training gap, product defect, or unsuitable workflow. This role is roughly a day a week during the first two months at a mid sized system.
- Act. Configuration and training issues get fixed locally. Product defects go to the vendor with examples and a tracked reference. Unsuitable workflows get excluded, publicly, which is the step organisations resist and the one that buys the most credibility.
- Report back. Tell the person who flagged it what happened, by name, within a week. This is the step that determines whether the loop keeps producing input.
Set up a small standing review of note quality itself, separate from the flags. Five notes per specialty per week, read by a clinician against a short rubric: is it accurate, is it complete, does it contain anything that was not said, is it the length a reviewer would want. This catches the drift that individual clinicians accept because each note was only slightly off. Systems that only track flags never see it.
Where the tool is drafting patient facing content, such as replies in the clinical inbox, add a stricter review for the first eight weeks and keep the sample size honest. A drafting tool with a human reviewer sits in a different risk position from an autonomous one, a distinction we set out in agents, copilots and automation.
Why do mandates fail?
Because a mandate converts a product problem into a compliance problem and then declares victory when the compliance number moves.
The specific mechanism is worth understanding. Mandating use of a documentation tool produces exactly the behaviour you asked for: sessions started. Clinicians who find the output unusable start the session, discard the draft and write the note themselves, and your weekly active number reaches 95 percent while the actual documentation burden rises. You have now lost the measurement that would have told you something was wrong, and you have taught a department that the programme is not interested in their experience.
Mandates also destroy the champion model. A champion in a mandated rollout is enforcement, not peer support, and colleagues stop telling them the truth about what is broken.
There is a narrow case where a requirement is appropriate: where the tool is demonstrably better on the metrics the department cares about, adoption has plateaued above 70 percent, and the remaining gap is habit rather than fit. Even then, frame it as a workflow standard with a documented exception route rather than a mandate, and keep the abandonment metric visible so you can see the discard behaviour if it starts.
What works instead of a mandate is embarrassingly ordinary: fix the specific things clinicians complained about, tell them you fixed them, show each department its own numbers next to the peer distribution, and let the sceptics watch the early adopters go home earlier. Peer data moves clinicians. Policy memos do not.
What should the first eight weeks look like?
A schedule that assumes adoption is a project with a plan rather than an outcome that emerges.
- Weeks minus 4 to 0. Baseline capture: after hours EHR time, note turnaround, inbox time, and a short burnout or workload measure. Champions named and released time confirmed in writing. Feedback route built and tested.
- Weeks 1 to 2. Go live in two or three departments, not everywhere. Champions present in clinic. Daily flag triage. Expect the first product defects here and log them formally.
- Weeks 3 to 4. First segmented usage report to department leads, showing their own numbers. First round of configuration changes shipped and communicated. Publicly exclude any workflow the tool clearly cannot handle.
- Weeks 5 to 6. Expand to the next wave, carrying the fixes forward. Start the weekly note quality sample. Report the loop's turnaround time alongside usage.
- Weeks 7 to 8. First honest read on sustained use and edit burden. Decide, with evidence, whether to expand, hold or narrow. A decision to narrow scope at week eight is a good outcome, not a failure.
Resist the temptation to launch system wide on day one because the licence is already paid for. Wave rollouts cost a few weeks and save the credibility that a bad first impression destroys permanently in a clinical workforce.
Why does adoption differ so much by specialty?
Because the underlying work differs, and a tool tuned on ambulatory primary care encounters behaves differently in a procedural clinic. Expect and plan for the variance rather than treating a low adopting department as resistant.
The predictable drivers are visit length, structure and vocabulary. Short high volume visits with repetitive structure, such as much of orthopaedics and dermatology, often see less benefit from ambient documentation because the clinician's own template is already fast, while the ambient draft needs editing. Long, discursive visits in primary care, behavioural health and complex chronic care usually see more. Procedural specialties frequently need a different tool or no tool at all.
Environment matters too. Noisy rooms, multiple speakers, family members interpreting and non English consultations all degrade performance, and the departments with the most of these will report the worst experience regardless of enthusiasm. Ask about this during selection and again during the pilot; it is the most consistently under tested condition, and it correlates with the patient populations you least want to serve worse.
Where a whole specialty is a poor fit, a different product may be the answer rather than a different rollout, which is the question our AI medical scribe comparison is built to answer. The right response to a low adopting specialty is a specific investigation, not an escalation. If the answer is that the tool does not fit, say so, release the seats and put them where they earn their cost. That decision is cheap at week eight and expensive at month eighteen.
How do you sustain adoption past ninety days?
Three habits, all boring, all frequently dropped once the launch team disbands.
Keep the report running. Segmented usage and edit burden, monthly, to department leads, indefinitely. Adoption decays quietly, and a system that stops measuring finds out at renewal.
Re run the loop on every vendor release. Model and product updates change behaviour, sometimes materially, and the clinician who notices first will assume it is their imagination unless there is somewhere to report it. Put vendor release notes on a named person's desk.
Onboard new joiners properly. A year in, a meaningful share of your clinicians will have arrived after go live and received none of the launch support. Fold the tool into standard onboarding with a champion contact, or watch adoption erode by attrition.
Finally, revisit the business case with the numbers you now have rather than the ones you projected, rerunning the ROI model on measured inputs instead of projected ones. If after hours time fell by a smaller amount than the vendor's material suggested, say so plainly and decide on the real figure. The published evidence on documentation burden is more mixed than most sales decks imply, as we set out in the pajama time evidence base, and an honest internal number is worth more than a favourable external one.
Adoption work is not glamorous and it is usually the part nobody scoped. If you would rather have it planned before go live than improvised after it, that sequencing is what our deployment roadmap and staff training engagements exist to produce, and it starts from the same baseline capture described above.
Sources
Primary material behind the claims above. Read the source before acting on any summary of it.
- OtherAMA Augmented Intelligence Research, physician sentiment on AI adoption requirements (opens in a new tab)
- OtherAdoption of AI in Healthcare Delivery Systems: Early Applications and Impacts, Peterson Health Technology Institute, March 2025 (opens in a new tab)
- OtherTethered to the EHR: primary care physician workload assessment using EHR event log data, Annals of Family Medicine, 2017 (opens in a new tab)
- NISTAI Risk Management Framework (AI RMF 1.0), NIST (opens in a new tab)
Questions we get asked
What is a realistic adoption rate for an AI scribe at ninety days?
Among clinicians who opted in and whose visit types suit the tool, 70 to 85 percent weekly active is a good result. Across a whole enabled population including specialties the tool fits poorly, expect considerably lower and treat the aggregate as close to meaningless. Publish the segmented numbers instead, because that is where the decisions are.
Should we make AI documentation tools mandatory?
Not before the tool has demonstrably worked in the department you are mandating it in. A mandate produces started sessions rather than used output, and it destroys the abandonment signal that would have told you the tool was failing. Where adoption has plateaued above roughly 70 percent and the remaining gap is habit, a workflow standard with a documented exception route is a reasonable step; a blanket mandate at week four is not.
How much protected time do champions need?
About four hours a month per champion during the first quarter, released as clinical time rather than added to it, and roughly half that afterwards. The exact number matters less than the fact that it is written down and honoured. Champion programmes that rely on goodwill reliably decay within about six weeks, and the decay is usually blamed on the champion rather than the design.
What do we do about a department that refuses to use it?
Investigate before escalating. In most cases refusal is a fit problem: visit type, room environment, speaker count or vocabulary. Sit in three clinics, read ten drafts and you will normally have the answer within a week. If the tool genuinely does not fit, release the seats and redeploy them, which is a better outcome than a compliance campaign that produces started sessions and discarded drafts.
How do we know whether the tool is actually saving time?
Only by comparing to a baseline you captured before go live, using EHR event log measures such as after hours time and time in notes rather than self report alone. Perception and measured time frequently diverge, in both directions, and vendor supplied dashboards usually measure activity inside their own product rather than total documentation burden. Ask your EHR vendor for the clinician efficiency metrics before you launch.
Who should own adoption after the project team stands down?
A named operational owner in the clinical organisation, with the monthly segmented report as their standing artifact, plus a designated recipient for vendor release notes. Ownership that reverts to IT tends to reduce adoption work to ticket handling, and ownership that reverts to nobody produces a quiet decline that surfaces at renewal.
Know what changed before your vendor tells you
A monthly regulatory and vendor intelligence note for people who have to sign off on this. What moved in HIPAA, ONC and state AI rules, and which vendor claims stopped being true.
Book an evaluation call at any point. No obligation.