Not long ago I read a vendor pitch for an AI documentation tool. The selling point, stated almost as a victory: the system listened to the visit, settled on a diagnosis, and produced a ready-to-submit ICD-10 code. Off it went.
In the world of billing, that’s a finished transaction. In the world of treating a patient, it’s nearly meaningless.
That gap, between a code that satisfies a payer and the information a physician needs to treat a person, is what nearly every healthcare AI tool is racing past. As we push these systems toward precision medicine and genomics, that gap stops being an annoyance and becomes a patient safety issue.
A billing code is not a diagnosis
ICD-10 and the coding systems around it were built to do one thing well: classify an encounter so a claim can be paid. The trouble starts when we ask them to stand in for clinical reality, because a diagnostic category and a clinical diagnosis are not the same thing.
Consider how codes are chosen. An AI tool (or a hurried clinician) often selects whatever sits first in a list, or whatever is specific enough to clear the claim. “Other chronic pulmonary disease” will get a COPD encounter paid, but it tells you nearly nothing about what the patient actually has. “Malignant neoplasm of brain, unspecified” will get the bill paid, but it doesn’t tell you the patient has a glioblastoma, which demands a completely different treatment path and prognosis than the more favorable tumors sharing that same vague code.
For everyday internal medicine, that imprecision is already a problem. An unspecified code lands on the problem list and stays there, influencing decisions for years. But the stakes climb sharply the moment you move toward genomics and targeted therapy, where the treatment decision can hinge on a phenotype, the observable expression of a patient’s genetic makeup, and the precise variant underneath it.
Where precision pays off
Consider a real-world example. Charcot-Marie-Tooth disease, or CMT, is a hereditary neuropathy I’ve encountered exactly once in my career. A patient presents with a cluster of findings: frequent falls, trouble writing, muscle cramps and atrophy, weakening grip, and a family history of the same. Individually, none is decisive. Together, they form a phenotype that should point toward CMT, and then toward genetic screening to identify the specific variant.
Here is why that specificity matters now in a way it didn’t a few years ago. For a long time, there was no real treatment for CMT, just over-the-counter pain relief. Today, targeted therapies are in development. The only way a patient reaches one is through the right diagnosis, with the right level of specificity, that points to the drug likely to work. If the diagnosis is vague, you never start down that path.
Now contrast two related conditions. CMT and a closely associated hereditary sensory and motor neuropathy can look similar at the bedside, yet carry different genetic underpinnings, and patients respond differently because they are different diseases. The ICD-10 code for one of them resolves to “hereditary motor and sensory neuropathy.” That phrase tells a treating clinician essentially nothing. It’s like being told a lot of cars are green. Great, but what does that help me do?
I’ve lived this from the other side, too, watching a patient carry a generic ‘polyarthritis’ label for years before anyone reached the real answer, whether rheumatoid or, harder still, psoriatic arthritis, where the deciding clue was a single patch of skin nobody thought to look for. The label that finally fits is the one that unlocks the right treatment. The broad category just delays it.
Ambient AI hears the words. It doesn’t know what to ask next.
Ambient documentation tools fall short with a subtle failure that’s easy to miss. These systems listen to a visit, infer a diagnosis, and attach a billing code. What they don’t do is prompt the clinician for the additional findings that would refine or correct that diagnosis. They provide a best guess from what was said aloud, but can’t tell you that asking about one more symptom, exam finding, or piece of family history might yield a better answer.
A clinical system built on a structured knowledge foundation works the other way around. When a clinician documents toward a diagnosis, the system can surface a full set of associated findings at once, the elements that distinguish one diagnosis from a near neighbor, rather than forcing the clinician to recall every attribute of a condition they may have seen once in twenty years. No busy clinician has time for that, and asking them to is a strange answer to the burnout problem.
The order matters. In much of today’s ambient AI, the billing code comes first and the diagnosis is reverse engineered to fit it. Clinically, that’s backwards. The clinical picture should come first; the code should fall out of it. That’s the difference between an expert clinical system and a billing system, and between data that a downstream tool can reason with and data that quietly carries every ambiguity forward.
Specificity is the prerequisite, not the bonus
This is also why structured, clinically-specific data is the precondition for precision medicine. A rich phenotype, the detailed history and physical findings captured at the point of care, is what makes a genetic workup meaningful. An accurate phenotype can help determine which therapies a patient is likely to respond to and which they won’t.
Drug-knowledge partners are already using that kind of rich clinical input to flag, at the point of care, whether a patient is likely to respond to a specific therapy. None of it works on a vague ICD-10 code. A broad diagnostic category cannot be reliably connected to a genetic variant, a biomarker, or a targeted pathway, no matter how efficiently it travels between systems. The specificity must be present in the data before it moves; it cannot be added in transit.
Slow down and get it right
I’ll end with the concern I keep coming back to. The pressure to adopt these tools is enormous, and the pitch is always time saved, less documentation burden. Those are real benefits. But saving a clinician time is not the same as delivering better care, and we are not asking the second question often enough.
The concern is that clinicians are leaning on these systems before fully understanding their limits, trusting AI to pull the relevant information out of a note and, too often, getting it wrong. That doesn’t produce smarter medicine. It produces clinicians who have stopped pairing the tool with their own judgment. That’s how mistakes multiply.
The fix isn’t to abandon AI. It’s to give it a clinical knowledge foundation that reflects how medicine actually reasons. If the technology prompts for what’s missing and distinguishes the diagnosis that matters from the one that merely pays, it can hand the clinician a more complete picture rather than a faster guess. The granularity the industry once resisted is now the floor. Building above it is the only way these tools earn the trust we’re already extending them.
Photo: cat-scape, Getty Images
Dr. Jay Anders is Chief Medical Officer of Medicomp Systems. Dr. Anders supports product development, serving as a representative and voice for the physician and healthcare community that Medicomp’s products serve. Prior to joining Medicomp, Dr. Anders served as Chief Medical Officer for McKesson Business Performance Services, where he was responsible for supporting development of clinical information systems for the organization. He was also instrumental in leading the first integration of Medicomp’s Quippe Physician Documentation into an EHR. Dr. Anders spearheads Medicomp’s clinical advisory board, working closely with doctors and nurses to ensure that all Medicomp products are developed based on user needs and preferences to enhance usability.
This post appears through the MedCity Influencers program. Anyone can publish their perspective on business and innovation in healthcare on MedCity News through MedCity Influencers. Click here to find out how.
