Analysis

Pajama Time: What the Documentation Burden Evidence Actually Says

ByClunic Research Team13 min read

Free tool

AI Scribe ROI Calculator

Put your own numbers in and see the range, not one flattering figure.

Need it signed off?

Thirty free minutes with an analyst on the vendor, the workflow and the rule you are unsure about.

Book an evaluation call

What do the foundational studies actually show?

Nearly every claim about documentation burden traces back to a small number of papers, and it is worth knowing exactly what they measured, because the numbers get repeated far outside their original scope.

StudyDesign and sampleHeadline findingWhat it does not establish
Sinsky and colleagues, Annals of Internal Medicine, 2016Direct observation time and motion study, 57 physicians in family medicine, internal medicine, cardiology and orthopedics across four states, 430 hours observed, 21 physicians also kept after hours diaries27.0 percent of the office day on direct clinical face time and 49.2 percent on EHR and desk work; in the examination room, 52.9 percent face time and 37.0 percent EHR and desk work; diarists reported 1 to 2 hours of after hours work each night, mostly EHR tasksNot a national sample, not a random one, and four specialties only. The after hours figure rests on 21 self reported diaries.
Arndt and colleagues, Annals of Family Medicine, 2017Retrospective analysis of Epic event log data for 142 family medicine physicians in one southern Wisconsin system over three years, validated against direct observation355 minutes, or 5.9 hours, of EHR time in an 11.4 hour workday: 4.5 hours during clinic and 1.4 hours after clinic. Clerical and administrative work accounted for 157 minutes, or 44.2 percent of EHR time, and inbox management for 85 minutes, or 23.7 percentOne system, one specialty, one EHR. Event log time attribution has known measurement caveats.
Holmgren and colleagues, JAMA Internal Medicine, 2020Comparison of EHR metadata across US and non US health systems using vendor supplied usage dataUS clinicians spent substantially more time in the EHR per day than clinicians in comparable non US systems, and received far more inbox messagesDescriptive. It shows the US pattern is unusual, not why, and does not isolate the effect of any single policy.

Two things follow. First, the widely quoted figures are real, well conducted and specific. Second, they describe ambulatory practice in the mid 2010s, mostly primary care, before ambient documentation existed. Using them as the before number in a 2026 business case assumes nothing has changed in a decade, which is a claim in itself.

What does pajama time actually measure?

The term describes EHR work done outside scheduled clinic hours, and in practice it is defined by whatever your EHR vendor's efficiency reporting counts as outside hours. That definitional looseness matters more than it sounds.

Most vendor signal is built on event log timestamps. A clinician logged into the chart at nine in the evening registers as after hours EHR time. But event logs cannot distinguish a physician writing five notes from one who left the chart open on a second monitor while doing something else, and different systems apply different idle timeouts to that problem. Two organisations reporting different after hours numbers may simply be applying different timeouts.

The scheduled hours boundary is equally soft. A clinician whose contracted session ends at five but whose clinic reliably runs to six thirty generates after hours time that is really clinic overrun, not documentation debt. Part time clinicians and those with non standard schedules are systematically misclassified by definitions built around a nine to five template.

None of this makes the measure useless. It makes it a within organisation trend measure rather than a cross organisation benchmark. Your after hours minutes this quarter against your after hours minutes last quarter, using the same definition and the same population, is a real number. Your after hours minutes against a figure in a vendor deck is not a comparison, and treating it as one is how business cases end up defending a baseline nobody measured.

The same caution applies to the burden measures that matter clinically. Arndt and colleagues found that inbox management alone consumed 85 minutes of the average day, nearly a quarter of total EHR time, and a documentation tool that does not touch the inbox cannot move that quarter. Matching the tool to the component of burden you actually have is the whole exercise, which is why inbox triage and ambient documentation are separate procurements with separate cases.

How much of the burden is the inbox?

More than most documentation business cases assume, and growing faster than the note writing component.

Arndt's 2017 breakdown put inbox management at 23.7 percent of EHR time in family medicine. The volume driving that number has not been stable since. Portal messaging rose sharply through the pandemic period and did not return to its previous level, and Holmgren and colleagues have published repeatedly on both the volume trend and the effects of interventions such as billing messages as e-visits.

The distributional point is the one that gets lost. A 2026 Health Affairs analysis by Holmgren and colleagues examined how portal messages are distributed across patients and physicians and found the load highly uneven: a minority of physicians carry a multiple of the median volume, and a minority of patients generate a large share of the messages. An organisation reasoning from its average inbox time will systematically under serve the clinicians who are actually drowning, and those clinicians are the ones whose attrition costs the most.

There is also an opportunity cost that rarely appears in a business case. Holmgren and colleagues reported in Health Affairs in 2024 that documentation burden crowds out health information exchange use, with each additional hour of documentation associated with a 7.1 percent reduction in review of outside patient records. Documentation time is not only unpleasant, it displaces the reading of records from elsewhere, which is a quality argument rather than a wellbeing one and tends to carry further with a board.

The practical implication for procurement is that if your clinicians' complaints are about the inbox, an ambient documentation tool addresses a different problem well. Both may be worth buying. They are not substitutes, and the ROI models we publish, including the AI scribe ROI calculator, are deliberately scoped to one at a time for that reason.

Is the burden getting better or worse?

The honest answer is that the wellbeing measures have improved for three consecutive years while the underlying burden measures have moved much less, and it is not clear the two are causally linked.

On wellbeing, the American Medical Association's national physician burnout survey reported that 43.2 percent of physicians experienced at least one symptom of burnout in 2024, down from 48.2 percent in 2023, and the AMA has since reported a further decline to 41.9 percent in 2025. That is a genuine, sustained improvement from a peak above 60 percent in 2021, and it is measured consistently across a large multi organisation sample.

It is also not evidence that AI reduced burnout. The decline began before ambient documentation reached meaningful scale, it coincides with the recovery from an extraordinary period of pandemic pressure, and the survey does not attribute causes. Vendors that place their adoption curve next to this trend line are inviting an inference the data does not support. Anyone building an internal case should say so explicitly, because a board that later notices the confound will discount everything else in the deck.

On the burden itself, the more useful reading is that the components have shifted rather than shrunk. Note writing appears more tractable to automation than the inbox, and inbox volume has risen. An organisation can plausibly reduce time in notes while total EHR time stays flat, which is a real improvement in the quality of the work but not the number most business cases promised.

What does the evidence say AI scribes actually deliver?

The strongest available evidence is more modest and more interesting than the marketing, and it is improving quickly in quality as multi site studies replace single site pilots.

SourceWhat was measuredFinding
Peterson Health Technology Institute, March 2025 report on AI adoption in health systemsInterviews and data from health systems deploying ambient scribesConsistent reductions in reported burnout and cognitive load, but no clear financial return demonstrated. Time savings feedback was mixed: some systems reported reduced after hours documentation, others reported no measurable difference or no reliable way to tell.
Health system pilots reported in the same PHTI workSelf reported burnout before and after, short pilot windowsMass General Brigham reported a 40 percent relative reduction in reported burnout over a six week pilot survey; MultiCare reported a 63 percent reduction among clinicians surveyed after its pilot. Both are self reported, short window and uncontrolled.
Multi site study by Holmgren and colleagues with the ACDC Collaborative, JAMA, 2026EHR event log time before and after AI scribe adoption across a multi site collaborativeAdoption was associated with 13.4 fewer minutes of total EHR time per eight scheduled patient hours.
Study by Holmgren and colleagues, JAMA Internal Medicine, 2024Team based documentation support, a non AI comparatorVisit volume rose 6.0 percent while EHR documentation time fell 9.1 percent, a useful benchmark for what a well run human intervention achieves.

Put the two strongest numbers next to each other. Roughly thirteen minutes saved per eight scheduled patient hours is real, measured from event logs rather than surveys, and worth having. It is also about a quarter of an hour a day, not the two hours that appears in sales conversations. Meanwhile the burnout and cognitive load effects are large, consistent across sites and reported by clinicians themselves. The most defensible summary as of mid 2026 is that ambient documentation reliably makes the work feel better, modestly reduces measured EHR time, and has not yet demonstrated a clear financial return at the system level.

That summary is a perfectly good reason to buy. It is a bad reason to promise a finance director a headcount reduction.

Where do vendors overreach the evidence?

Five patterns, all of which can be checked in a demo.

Borrowing the 2016 baseline. The Sinsky figures describe observed ambulatory practice in four specialties in 2016. Presenting them as your current state, then subtracting a product effect from them, produces a saving that was never measured in your organisation. Ask for the vendor's own before and after data from a comparable site.

Reporting time in notes as time saved. Time spent in the note editor can fall while total EHR time is unchanged, because the work moved rather than disappeared. Insist that any claim is stated as total EHR time or after hours time, using the same definition before and after.

Extrapolating from enthusiast pilots. Early adopters in a six week pilot are the most favourable population and the most favourable window available. Retention at ninety days across a full department is the number that predicts the second year, and it is almost never the number in the deck.

Converting saved minutes into revenue automatically. Saved documentation time becomes revenue only if it is converted into visits, and that conversion requires schedule changes, demand and clinician willingness. The team based documentation study above achieved a 6.0 percent visit volume increase alongside its time saving, which shows conversion is possible and shows roughly what scale is realistic. A model that assumes full conversion of every saved minute is not a forecast.

Presenting the burnout trend as an AI effect. The national decline in burnout began before ambient documentation scaled. Correlation with an adoption curve is not evidence.

None of this means the products do not work. It means the honest case is narrower than the pitch, and an honest case survives its first internal audit. We apply the same tests to products in the AI medical scribe comparison and to pricing claims in the scribe pricing comparison.

How should you measure your own baseline?

Published studies tell you what is plausible. Only your own data tells you what is true in your organisation, and capturing it costs a few weeks of attention before go live rather than an argument afterwards.

  1. Pull EHR vendor efficiency metrics for at least eight weeks before deployment. Total EHR time per scheduled hour, time outside scheduled hours, time in notes, time in inbox. Fix the definitions in writing, including the idle timeout, and never change them mid study.
  2. Segment by specialty, visit type and clinician full time equivalent. Aggregates hide the distribution, and the distribution is where the decisions are.
  3. Capture note turnaround. Time from encounter close to signature, which is the measure clinicians feel most directly and which moves earlier than total time.
  4. Run a short validated wellbeing measure at baseline. A single item burnout measure administered before and after is worth more than a bespoke satisfaction survey afterwards, because it is comparable to published benchmarks.
  5. Keep a concurrent comparison group. Clinicians in the same specialty who are not yet enabled. Wave rollouts give you this for free if you plan for it, and without it you cannot separate the tool's effect from seasonal variation.
  6. Read fifty notes. Quality is not in the event log. Sample before and after, against a rubric, with a clinician reviewer.

The baseline is the single most common omission in healthcare AI deployments, and its absence is what turns a year two review into a debate about impressions. Building it into the plan before go live is part of what a deployment roadmap is for, and it is the measurement half of the adoption work described in our clinician adoption playbook.

What do the numbers actually support?

A short list of claims we would be willing to defend in front of a board as of mid 2026, and the ones we would not.

Supported. Ambulatory physicians spend a large share of the working day on EHR and desk work, with roughly two hours of it for every hour of direct patient contact in the observed sample. A meaningful amount of that work falls outside scheduled hours. Inbox management is a substantial and growing share of the total, distributed very unevenly across clinicians. Ambient documentation is associated with a modest measured reduction in total EHR time and with consistent, sizeable improvements in reported burnout and cognitive load.

Not supported. That AI scribes save two hours a day. That the national burnout decline is attributable to AI adoption. That documentation time savings convert automatically into visit volume or revenue. That findings from four specialties in 2016 describe your organisation in 2026. That any of these effects hold uniformly across specialties, and specifically that they hold in procedural specialties, in non English consultations or in noisy multi speaker environments, where the evidence is thin to absent.

The gap between those two lists is not an argument against buying. It is an argument for buying with a measured baseline, a specific target and a stated definition of failure, which is the difference between a deployment that survives its second year review and one that quietly becomes an unused licence. That is the same failure pattern we set out in why AI pilots fail in healthcare.

If you would rather establish the baseline and the target before signing anything, that is precisely the scope of an AI readiness audit, and it is considerably cheaper than discovering at renewal that nobody measured the before.

Sources

Primary material behind the claims above. Read the source before acting on any summary of it.

Questions we get asked

How many hours a day do physicians spend on the EHR?

The most cited measured figure is from Arndt and colleagues in 2017: 5.9 hours of EHR time in an 11.4 hour workday among 142 family medicine physicians, of which 1.4 hours fell after clinic hours. That is one specialty, one system and one EHR, so treat it as an illustration of scale rather than a benchmark. Your own EHR vendor's efficiency metrics are the only figure that describes you.

What is pajama time?

EHR work performed outside scheduled clinic hours, typically evenings and weekends. It is measured from EHR event log timestamps, which means the number depends on how your system defines scheduled hours and how long an idle session counts as active. It is a reliable trend measure within one organisation using one definition, and an unreliable benchmark between organisations.

Do AI scribes actually save time?

The best current evidence says yes, modestly. A 2026 JAMA study by Holmgren and colleagues with the ACDC Collaborative found AI scribe adoption associated with 13.4 fewer minutes of total EHR time per eight scheduled patient hours. The Peterson Health Technology Institute's 2025 review found consistent reductions in burnout and cognitive load but mixed evidence on time savings and no clear financial return. Two hours a day is not a figure the published literature supports.

Has physician burnout gone down because of AI?

Burnout has gone down; the attribution to AI is not established. The AMA reported 43.2 percent of physicians with at least one burnout symptom in 2024, down from 48.2 percent in 2023 and from a peak above 60 percent in 2021, with a further decline to 41.9 percent reported for 2025. That trend began before ambient documentation reached scale and coincides with recovery from pandemic conditions, so a causal claim would need study designs that these surveys do not provide.

Which measure should we use to justify an AI documentation purchase?

Total EHR time per scheduled patient hour and time outside scheduled hours, both segmented by specialty, both captured for at least eight weeks before deployment, both defined identically before and after. Add note turnaround time and a single item burnout measure. Avoid time in notes on its own, because it can fall while total time is unchanged, which is the most common way a favourable result turns out to be an artefact.

Does the documentation burden evidence apply outside primary care?

Much less than it is assumed to. The foundational studies concentrate on ambulatory primary care and a small number of other specialties, and the effects of ambient tools vary widely by visit length, structure and environment. Procedural specialties, non English consultations and noisy multi speaker settings are all under studied. If your deployment depends on results in one of those settings, treat the published evidence as inapplicable and measure locally.