Quick answer
AI vendor selection is a five week independent evaluation that turns a crowded market into a scored shortlist for one specific workflow. It defines your requirements, tests the vendors against them, verifies the security and compliance claims rather than accepting them, and leaves you with a written decision record. We take no commissions from any vendor.
What is an AI vendor selection engagement?
It is a structured evaluation that reduces a category to a scored shortlist for one specific workflow in your organisation. It has three parts: agreeing what you actually need, testing the vendors against that rather than against their own feature lists, and verifying the claims that carry compliance risk.
The third part is the one buyers most often skip and most often regret. A vendor page saying HIPAA compliant is a marketing statement, not a status any regulator confers. What matters is whether the vendor signs a business associate agreement, which subprocessors sit behind it, what the contract says about using your data to train models, and whether the audit logging is detailed enough to reconstruct an incident. Those are checkable, and checking them changes shortlists.
The output is a scored matrix plus a written decision record. The decision record matters as much as the score, because in two years the person reviewing the contract will not remember why the second placed vendor lost, and an undocumented decision is renegotiated from scratch every time somebody new arrives.
What problem does this actually solve?
Healthcare AI is a market where the marketing has outrun the disclosure. Published pricing is rare, published accuracy figures are usually measured by the vendor on data the vendor selected, and integration claims range from a certified application in the EHR's own marketplace to a screen scraping workaround that breaks at the next upgrade. Comparing those on a website is not possible.
Two failure modes follow. The first is deciding on demo quality, which measures the presenter rather than the product. The second is a committee that cannot separate two finalists, escalates the decision, and loses six months, by which point the original business case has aged out and the process restarts.
Both are solved the same way: decide what matters and how much it is worth before you look at anyone, then test against it consistently. Most of the value here is in the sequence, not the expertise. If your evaluation has stalled, the reason is usually that requirements were written after the demos rather than before.
Who is it for, and who is it not for?
It is for organisations with a defined workflow, three or more plausible vendors, and a decision that has to be defensible to somebody who was not in the room. It fits naturally where procurement requires a scored rationale, and where the contract term is long enough that being wrong is expensive.
It is not for organisations that have already chosen and want validation. We will give you an honest answer, and where the honest answer is that the incumbent choice is wrong, it is unwelcome and expensive to ignore. It is also not proportionate for a small practice picking an ambient scribe: that decision can be made in an afternoon from the comparison page and a two week trial, and the money is better spent elsewhere.
If the harder question is which workflow to automate at all rather than which product to buy, the AI readiness audit is the right engagement, and vendor selection follows it. If you have already chosen and now need the rollout planned, that is the deployment roadmap.
What happens week by week?
Week one, requirements. We work from the workflow as it is actually run rather than as it is documented, and turn it into weighted requirements. Weighting happens here, before anyone has seen a demo, because a matrix weighted afterwards will quietly favour whoever demonstrated best.
Week two, market scan and longlist. We assemble the candidate set, including vendors you have not heard of and, where relevant, the module your existing EHR vendor already sells you. Each candidate gets a first pass screen on the requirements that are absolute, which typically removes half of them.
Week three, structured demonstrations. Every vendor is shown the same scenario script, including the edge cases: the difficult accent, the multi payer claim, the patient who says something the workflow did not anticipate. We attend and score against the matrix during the session rather than from memory afterwards.
Week four, verification. Security documentation, business associate agreement terms, subprocessor lists, data residency, model training clauses, audit logging detail, incident history and integration evidence. This is where claims either survive contact with paperwork or do not. Reference calls run in the same week, using a guide rather than an open conversation.
Week five, decision record. The scored matrix, the total cost of ownership model and the written rationale, presented to the group that has to approve it. We state what we could not verify, because a matrix with no gaps in it is a matrix that stopped asking.
What is in the decision record?
The full deliverable list is above. Three parts do the heavy lifting.
The scored matrix shows every vendor against every weighted requirement with the evidence for each score, so a reviewer can disagree with a specific judgement rather than with the conclusion as a whole. Scores that rest on a vendor assertion are marked as such and never carry the same weight as scores that rest on a document or a demonstration.
The verified compliance position records, per vendor: whether a business associate agreement is offered and on whose paper, the named subprocessors, where data is processed and retained, what the contract permits regarding model training on your data, the granularity of audit logs, and whether the product makes any claim that would place it in scope for FDA device regulation. Where a vendor declines to answer, that is recorded as a decline rather than as a gap.
The total cost of ownership model separates licence from implementation, integration and internal effort, and is handed over editable. A licence quote compared against zero is not a business case, and the internal effort line is the one that most often turns a cheap option into an expensive one.
How does this fit the CARE method?
CARE is our four step method: Chart, Architect, Run, Evaluate. Vendor selection sits inside Architect, alongside integration and governance design, because in practice those three decisions constrain one another.
Chart produces the ranked workflow shortlist and the baseline, usually through the readiness audit. Selection without that step tends to evaluate products against a workflow nobody has measured.
Architect is where the vendor choice, the integration design and the control set are decided together. A vendor that scores well on function and poorly on your EHR's interface model is not a good choice, and the only way to see that is to design all three at once. The deployment roadmap is the other half of this step.
Run is the bounded pilot, which is why the contract checklist insists on pilot terms with a real exit. Evaluate compares the result against the baseline and the thresholds. The decision record is what makes Evaluate honest, because it recorded in advance what you were buying the product to do.
How do you score AI vendors without rigging the matrix?
You can run this yourself. Most bad evaluations are bad for structural reasons rather than for lack of expertise, and four rules remove most of the damage.
Weight before you look. Write and sign off the weights before the first demo. A matrix weighted afterwards is a rationalisation, and everyone in the room will know it.
Separate absolute requirements from scored ones. An absolute requirement is a pass or fail gate: will they sign a business associate agreement, do they support your EHR's interface, will they commit contractually not to train on your data. Anything absolute should never appear in the scored section, because a high score elsewhere must not be allowed to buy its way past a gate.
Score evidence, not assertions. Use three tiers: demonstrated in front of us, documented in something contractual, or asserted. Asserted claims can still score, but they should never score the same as demonstrated ones, and the tier should be visible in the matrix.
Run the same script for everyone, including the edge cases. Vendor led demos show the happy path. Ask each vendor to run your scenarios, including the ones that go wrong, and watch what the product does when it is unsure. An agent that fails loudly is safer than one that fails smoothly, and you will only see the difference if you ask.
A workable scoring structure looks like this.
| Category | Typical weight | What actually earns the score |
|---|---|---|
| Workflow fit | High | Your scenarios run end to end in the demo, including the exceptions |
| Integration | High | A named production reference on your EHR at your scale, not a roadmap item |
| Security and compliance | Gate plus score | Signed business associate agreement terms, subprocessor list, audit log detail |
| Accuracy and oversight | Medium | How the product behaves when uncertain, and whether a human can see why |
| Total cost of ownership | Medium | Licence plus implementation plus integration plus your own staff time |
| Viability and support | Medium | Funding position, named support model, and how upgrades are handled |
| Exit | Gate | Data export format, notice period, and what happens to your records |
What should you ask in an AI vendor security review?
These are the questions that change answers, and they are all reasonable to ask before signing anything.
- Will you sign a business associate agreement, and on whose paper? A vendor handling protected health information on your behalf is a business associate under HIPAA, and HHS publishes sample provisions you can compare against. A vendor that will only sign its own version with liability capped at one month of fees has told you something.
- Who are your subprocessors, and does that list include a model provider? Most healthcare AI products call a third party model. That provider is in your data flow, and you need it named, contracted and covered.
- Is our data used to train or improve your models, and can we opt out contractually? Ask for the clause, not the answer. A verbal no with no contractual language is not a no.
- What exactly is in the audit log, and for how long? Enough to reconstruct who saw what, what the agent produced, what a human changed and when. Logging that records only successful transactions cannot support an incident investigation.
- What happens when the model is updated? Ask whether you are notified, whether you can stay on a prior version, and whether behaviour has changed in ways that would invalidate your validation work. Silent model updates are a genuine governance problem and most contracts are silent about them.
- What is your incident history and your notification commitment? Ask for the breach notification timeline in the contract, and compare it against what your own obligations require.
- On exit, in what format do we get our data, and what do you keep? Get this before signature. Nobody negotiates a good exit during a bad one.
Where a product also makes a diagnostic or treatment claim, the regulatory question is separate and prior: the FDA page sets out when software becomes a regulated device. Where the workflow touches utilization review or patient facing communication, state rules increasingly apply on top, which the California page covers in detail.
What does it cost, and why do we take no commission?
The engagement is a fixed fee stated in a written proposal before work begins, scaled to the size of the organisation and the number of vendors in scope rather than billed hourly. The strategy call that produces the proposal is free.
We take no vendor commissions, no referral fees and no paid placements, and we resell nothing. In this category that is unusual enough to be worth stating plainly, because a large share of healthcare AI advice is paid for by the vendors being recommended. The business model is the reason our answer can be that none of the shortlist is good enough yet, or that the module your EHR vendor already sells you is adequate and free of a new integration.
Where the reasoning behind our published views can be checked, it is published. The comparison pages record what we verified and what we could not, and the ROI calculator exposes its assumptions on the page. If the shortlist is settled and the next question is how to get it live, that is the deployment roadmap, and the governance work that has to run alongside it is AI governance and compliance.
What you are left holding
The engagement is finished when these are true, not when the calendar says so.
- A scored shortlist with the reasoning attached, in a form your procurement process can accept
- A verified compliance position per vendor, distinguishing what the vendor documents from what it merely asserts
- A decision record that states plainly what could not be verified rather than smoothing over it
- A contract checklist you can take into negotiation, whichever vendor you choose
Questions we get asked
How much does an AI vendor selection engagement cost?
A fixed fee, agreed in a written proposal before the work starts and scaled to the organisation and the number of vendors in scope. We quote after a short free call. There is no commission on whatever you subsequently buy, which is the point of the model.
Do you have relationships with the vendors you evaluate?
We speak to vendors, attend their briefings and ask them for documentation, which is how verification works. We take no fees, referrals, commissions or paid placements from any of them, we resell nothing, and no vendor pays for coverage on this site. Where a vendor has declined to answer something, we record the decline rather than filling the gap.
How many vendors do you evaluate?
Typically a longlist of eight to twelve reduced to three or four for structured demonstrations and full verification. Carrying more than four into deep evaluation usually costs more in committee time than the extra option is worth, and carrying fewer than three tends to mean the market scan was too narrow.
Can you evaluate a vendor we have already shortlisted?
Yes, and we will also tell you if the shortlist itself has a gap. A single vendor verification is a smaller piece of work and we will scope it as one rather than inflating it into a full selection.
What if no vendor is good enough?
Then the decision record says so, and says what would have to change for that to stop being true. In fast moving parts of this market that is a real outcome, and waiting two quarters is often cheaper than a three year contract with a product that is not ready.
Do you handle the contract negotiation?
No. We produce the checklist of terms worth negotiating and the evidence behind each one, and your legal team or counsel negotiates from it. We are not lawyers and nothing we produce is legal advice.
Make it a formal evaluation
Everything we publish is free to read and free to argue with. When the decision has to be signed, dated and defended to a board, we run the evaluation against your own estate. We take no vendor commissions.
- A 30 minute evaluation call with an analyst, no pitch deck.
- A read on the vendors and the rules in play, and the use cases we would not touch yet.
- A written proposal with scope, sequence and a fixed fee.
- No obligation
- Direct with an analyst, not a sales rep
- BAA available before any PHI discussion