Use case

Medical Coding Automation with AI Agents

Last updated / Reviewed by Clunic Research Team

Quick answer

AI medical coding tools read the clinical note and propose ICD-10, CPT and HCC codes. Assisted products suggest codes a certified coder accepts or rejects. Autonomous products submit a defined subset of charts with no human review. The accountability line sits at that choice, not at the accuracy figure a vendor quotes, because the provider signs the claim either way.

Free tool

Charge Capture Leakage Calculator

A missed charge never shows up as a denial.

Need it signed off?

Thirty free minutes with an analyst on the vendor, the workflow and the rule you are unsure about.

Book an evaluation call

The numbers

Estimated Medicare fee for service improper payment rate, FY 2024, of which insufficient documentation is consistently the largest single cause
7.66%CMS: Improper payment rates and additional data, Comprehensive Error Rate Testing, CMS (opens in a new tab)
Increase in Medicare inpatient stays billed at the highest severity level between FY 2014 and FY 2019, while stays at lower severity levels fell
Almost 20%HHS: Trend Toward More Expensive Inpatient Hospital Stays in Medicare Emerged Before COVID-19 and Warrants Further Scrutiny, OEI-02-18-00380, HHS OIG, February 2021 (opens in a new tab)
Risk adjusted payments in one year driven by diagnoses reported only through chart reviews and health risk assessments, with no other record of the beneficiary receiving care for that condition
$9.2 billionHHS: Some Medicare Advantage Companies Leveraged Chart Reviews and Health Risk Assessments To Disproportionately Drive Payments, OEI-03-17-00474, HHS OIG, September 2021 (opens in a new tab)
Independently validated, peer reviewed accuracy studies of autonomous coding engines that we could locate as of mid 2026. Every accuracy figure in this market is currently vendor reported
None found

What is AI medical coding automation?

Coding turns a clinical encounter into the codes a payer will adjudicate: ICD-10-CM for diagnoses, CPT and HCPCS for procedures and services, and, where risk adjustment applies, the hierarchical condition categories that drive capitated payment. An AI coding tool reads the documentation and proposes those codes.

Three quite different products are sold under that one description, and the difference is not a feature difference. It is a question of who is accountable when the claim is wrong.

  • Computer assisted coding. The oldest of the three. Natural language processing and terminology matching surface candidate codes from the chart. A coder still builds the claim. Most large hospitals have had some version of this for a decade.
  • Assisted or augmented AI coding. A language model proposes a full code set with a rationale and a pointer to the supporting text in the note. A certified coder accepts, edits or rejects each line. Throughput rises, accountability does not move.
  • Autonomous coding. A defined slice of charts, usually the most repetitive and least ambiguous specialties, is coded and released to billing with no human review. Everything below a confidence threshold falls out to a coder.

Only the third one changes the operating model, and it is the only one where a vendor demo tells you very little. What matters is which charts fall into the autonomous lane, who set that boundary, and how often anybody checks it. That question sits inside a wider programme, which we set out on the revenue cycle automation page.

Autonomous or assisted coding: which should you buy?

The honest answer for most organisations is assisted first, autonomous later and only in specialties where the coding logic is genuinely narrow. Radiology, pathology, ophthalmology and some outpatient procedural work are the usual candidates because the input document is structured and the code set is small. Inpatient DRG assignment is not a candidate, and any vendor that says it is should be asked for the audit results.

DimensionComputer assisted codingAssisted AI codingAutonomous coding
Human in the loopYes, coder builds the claimYes, coder approves every lineNo, for charts above the confidence threshold
Typical scopeInpatient and outpatientAny specialtyNarrow, document driven specialties
Where the accuracy risk landsCoder fatigue, alert blindnessAutomation bias in reviewThreshold setting and drift
Effect on coder headcountSmallThroughput per coder risesRole shifts to audit and exception work
What you must monitorSuggestion acceptance rateOverride rate by coder and by codeFallout rate, and a blind audit sample of released charts
Realistic time to productionAlready installed in most hospitalsThree to six monthsNine to eighteen months, specialty by specialty

The most common procurement mistake is buying the autonomous tier and then operating it as an assisted tier because nobody was willing to sign off the threshold. You pay the premium and get none of the labour change. Deciding that in advance, in writing, is part of what we do during vendor selection.

How well does it handle E/M leveling?

Evaluation and management leveling is the hardest thing to automate honestly, because since the 2021 office visit revisions the level turns on medical decision making or total time, not on how many bullet points appear in the history and exam. Both of those inputs are judgements about the encounter rather than facts extractable from the text.

Total time is the easier of the two and the one most often documented badly. If the clinician does not record time, the model cannot infer it, and a tool that estimates time from note length is manufacturing a billing input. Ask any vendor whether it ever proposes a time based level without a documented time statement. The answer should be no.

Medical decision making is harder. The number and complexity of problems addressed, the data reviewed and the risk of the option chosen are all read from a narrative written for clinical purposes rather than billing ones. Models spot the elements that were documented. They are poor at knowing what was considered and not written down, which is exactly where undercoding lives.

This is the strongest argument for pairing coding automation with better documentation at the source rather than deploying it on top of thin notes. An ambient scribe that captures decision making in the note improves coding accuracy more reliably than a coding engine reading a two line assessment, and several documentation vendors on our AI medical scribe shortlist now market coding support on that basis.

What does it do for HCC capture and risk adjustment?

Risk adjustment is where AI coding tools generate the most enthusiasm and the most exposure, because the same capability that finds a legitimately documented chronic condition also finds a condition that was mentioned once, three years ago, in a discharge summary.

Two capabilities are worth having. The first is retrospective suggestion: reading the chart before or after the visit and flagging conditions that are documented and clinically active but were not coded on the encounter. The second is prospective prompting: surfacing likely conditions to the clinician at the point of care so that the assessment either documents and addresses them or explicitly does not.

The second is safer and slower. The first is faster and is the one regulators have been looking at. In September 2021 the HHS Office of Inspector General reported that diagnoses reported only through chart reviews and health risk assessments, with no other service record for the condition, drove 9.2 billion US dollars of risk adjusted payments in a single year, and identified twenty Medicare Advantage companies whose share of those payments was disproportionate to their enrolment.

The lesson for a provider group in a risk contract is not that HCC capture is illegitimate. It is that a code with no accompanying care is a finding waiting to happen. Set the rule before the tool arrives: a suggested condition is only coded if the encounter documents that it was assessed, and that rule belongs in the governance pack we build during an AI governance engagement.

Does coding automation increase or decrease audit risk?

Both, and which one depends entirely on how the tool is tuned and who tuned it.

Upcoding risk. A model optimised on historical claims learns the billing behaviour of the organisation that trained it, including any drift already present. If your baseline leans high, the tool will industrialise that lean and apply it evenly across every chart, which is precisely the pattern that shows up in a data driven audit. The OIG has already published on this shape of problem in the inpatient setting: it found that the number of Medicare stays billed at the highest severity level rose almost 20 percent between fiscal 2014 and 2019 while lower severity stays fell, that nearly a third of highest severity stays were unusually short, and that more than half were supported by only a single qualifying diagnosis.

Undercoding risk. Less discussed and more common in practices that have been burned once. A conservative threshold, or a coder review culture that rejects anything unfamiliar, produces a book of claims that is defensible and short. Nobody audits for revenue you did not bill, so this failure is invisible unless you measure it deliberately.

The practical control is the same in both directions: a blind audit sample. Pull a random monthly sample of finished claims, have a coder who did not touch them recode from the documentation alone, and record the variance in both directions. Do this before the tool arrives so you have a baseline, and keep doing it afterwards. Organisations that only start auditing after go live cannot tell the difference between a tool problem and a pre-existing one.

If your denial pattern already includes coding related rejections, the same sample tells you something about downstream cost too. We connect the two views on the denial management page.

How should you read a vendor accuracy claim?

Sceptically, and with four questions. As of mid 2026 we have not been able to locate an independently run, peer reviewed accuracy study of an autonomous coding engine in a US provider setting. Every headline number in this market is vendor reported, computed on a dataset the vendor chose, against a ground truth the vendor defined. That does not make the numbers false. It makes them uncomparable.

  1. Accuracy against what? Agreement with the organisation's existing coders is not the same as agreement with an independent auditor. If the reference standard is your own coding, the tool is being measured on its ability to reproduce your current error rate.
  2. Accuracy on which charts? A 95 percent figure quoted across all charts, when 40 percent fell out to humans, is really a claim about the easy 60 percent. Ask for accuracy on released charts only, plus the fallout rate.
  3. Which unit? Per code, per claim or per line. Per code accuracy always looks better, because most codes on most claims are trivial.
  4. Measured over what period? Coding rules change every October, payer edits change constantly, and clinical documentation style drifts. A figure from a 2024 evaluation is a historical fact, not a forecast.

Then ask the question that actually settles it: will the vendor accept a paid, blind bake off on your own charts, scored by an independent certified auditor, before you sign. Vendors confident in their numbers usually will. That test is the centrepiece of how we run a vendor selection in this category, and it is worth more than every case study on a website.

What happens to the coding team?

The work changes shape before it changes size, and organisations that plan only for the second part get the transition wrong.

In an assisted deployment, a coder stops building claims from scratch and starts adjudicating suggestions. That sounds easier and is in fact harder to do well for eight hours, because the failure mode is automation bias: after a few hundred correct suggestions, review becomes acknowledgement. Practices that hold accuracy through this transition tend to rotate coders between review work and audit work rather than leaving anyone on approval duty all day.

In an autonomous deployment, the residual human work is threshold governance, fallout coding, denial root cause analysis and audit. That is a more senior job than production coding and should be paid and titled as one. The people who understand why a payer rejected a code combination are the ones you need to tune the tool, and the ones most likely to leave if the messaging is that software replaced them.

The honest planning position is that the coding function becomes smaller and more senior over several years, not next quarter, and that most of the reduction comes from vacancies not being refilled. Say that plainly to the team at the start. For hospital scale programmes we set out how this sequences on the hospitals page, and the retraining component belongs in staff training rather than in a vendor onboarding call.

What does it cost, and how do you measure the return?

Published pricing is not available from most vendors in this category. The pricing shapes we see quoted are per chart, per encounter, a percentage of the coded charge, or a per coder subscription for assisted products. Percentage of charge pricing deserves care: it aligns the vendor's revenue with a higher code, which is the one incentive you least want in a coding tool.

Measure four things, and measure them before anything is installed.

  • Coding lag, in days from encounter to coded and billed. This is usually the first number to move and the easiest to defend.
  • Charts per coder per day, split by specialty, because a blended figure hides everything.
  • Blind audit variance, in both directions, expressed as a percentage of claims with any code change and the net revenue effect of those changes.
  • Coding related denial rate, which is the downstream check that the codes were not just fast but adjudicable. Our denial rate benchmark gives you a reference point for that one.

Note what is missing from that list: revenue lift. It is the number vendors lead with and the hardest to attribute honestly, because case mix, payer mix, contract terms and volume all move at the same time. If you do claim it, claim it against a documented baseline period and say what else changed.

How do you pilot coding automation safely?

Pick one specialty, run in shadow mode first, and do not let the pilot touch a released claim for at least a month.

Shadow mode means the tool codes every chart in parallel with your coders and neither sees the other's output. At the end of the month you have three things: an agreement rate, a list of every disagreement, and a sample of disagreements adjudicated by an independent auditor. That third artefact is what tells you whether the differences are the tool being wrong or your coders being conservative. It is the only part of the pilot that cannot be reconstructed later.

Only then decide what, if anything, moves to autonomous release, and write the threshold down with a named owner and a review date. Add a standing monthly blind audit and a defined circuit breaker: a variance level at which autonomous release is suspended without a committee meeting.

Two compliance items run in parallel. The vendor is a business associate, so the agreement, the retention terms and the position on model training all need settling before charts leave your estate, which is covered on the HIPAA and AI compliance page. And if the suggestions surface inside a certified EHR module, the transparency obligations that certified decision support carries are worth understanding before the configuration is set, which is the subject of the ONC HTI-1 page.

If you are not sure whether coding is the right first agent at all, it usually is not. Denials and prior authorisation are more measurable and less exposed, and the sequencing question is the first thing we work through in an AI readiness audit.

Questions we get asked

Is autonomous medical coding safe to use?

It is safe in narrow, document driven specialties where the code set is small and the threshold is conservative, and it is not safe as a general replacement for coders. The controls that make it defensible are a documented confidence threshold with a named owner, a monthly blind audit of released claims, and a suspension rule that triggers on variance without needing a committee. The provider remains accountable for the claim regardless of what the software did.

Can AI coding tools cause upcoding?

Yes, and the mechanism is usually inherited rather than invented. A model tuned on an organisation's historical claims reproduces whatever billing lean already existed and applies it consistently, which makes the pattern easier for an auditor to detect, not harder. HHS OIG has published on exactly this shape of drift in Medicare inpatient severity billing. A blind audit sample measured before deployment is the only way to know which direction you started from.

Does AI coding help with HCC and risk adjustment?

It can, and it is the highest value and highest exposure application at once. Prospective prompting, which surfaces likely chronic conditions to the clinician during the encounter, is safer than retrospective chart mining because the condition gets assessed and documented rather than only coded. Set a firm rule that a suggested condition is coded only where the encounter shows it was addressed.

How accurate is AI medical coding?

Vendors publish accuracy figures in the mid to high nineties, but as of mid 2026 we could not find an independently run, peer reviewed study validating any of them in a US provider setting. The figures are not comparable to each other because each vendor chooses its own dataset, ground truth and unit of measurement. Ask for accuracy on released charts only, alongside the fallout rate, and run a blind bake off on your own charts before signing.

Will AI replace medical coders?

It changes the job before it changes the headcount. Production coding shrinks, while threshold governance, fallout work, denial root cause analysis and audit grow, and those are more senior tasks. Most organisations reach a smaller and more senior team over several years, largely by not refilling vacancies. Plan the retraining before the tool arrives, because the coders who understand payer edits are the people who make the tool work.

What does AI medical coding software cost?

Published pricing is not available from most vendors in this category. Quoted structures include per chart, per encounter, a percentage of coded charge and a per coder subscription. Treat percentage of charge pricing with caution, because it pays the vendor more when the code is higher. Budget separately for integration, the audit function and the internal owner, which are usually larger than the licence in year one.