Data challenges with AI: Omissions, bloated problem lists and unnecessary token burn

Data challenges with AI: Omissions, bloated problem lists and unnecessary token burn

Hallucination is not the biggest risk in clinical AI. Omission is — and it is far harder to detect.


John Laursen, SVP at IMO Health, has spent his career on the layer of healthcare AI that gets the least attention: clinical terminology and the semantic data infrastructure underneath every model deployed in a hospital. IMO Health's terminology has been built and curated since 1994 and now sits behind roughly 12 billion terminology search transactions a year across US provider organisations and every major EHR.


In this interview with Tjaša Zajc, Laursen makes the case that structured data was necessary but is no longer sufficient. AI reasoning across a thirty-year patient chart needs semantic continuity — an understanding that clinical language recorded in the 1990s and language recorded today can mean the same thing. Without it, health systems are investing in models that cannot reliably interpret their own records.


The conversation also covers what happens when ambient AI scribes get it wrong, why accumulated clinical data has become a computational cost rather than an asset, and why clinician trust is the constraint that determines how fast clinical AI can move.


Guest:

John Laursen — Senior Vice President, IMO Health (Chicago, US)


Host:

Tjaša Zajc — Faces of Digital Health


What the conversation covers:

- Why omissions, not hallucinations, are the underrated risk in clinical AI

- What a semantic layer does that structured data alone cannot

- How clinical terminology maps to SNOMED CT and ICD-10 — and why those code sets were built for different purposes

- Ambient AI scribes: what happens when a model mishears or over-infers a diagnosis

- The billing and clinical consequences of an error entering the patient record

- Why problem lists hundreds of entries long now cost money in token burn

- Patient-generated and AI-generated content entering the EHR, and why health systems resist it

- Translating lay language into clinical terminology without losing specificity

- Ambient documentation, billing intensity and friction with payers

- How data quality expectations differ between the US, the NHS, the Gulf states and Singapore

- Who governs clinical data as coding complexity increases

- Why AI performance breaks down on rare disease and the difficult 20% of cases

- Knowledge graphs as a grounding source for clinical AI models

- What health systems should require from AI vendors before clinical deployment


Chapters:


02:20 Why the data layer decides what clinical AI can do

03:27 Inside IMO Health: 12 billion terminology searches a year

05:36 Keeping terminology current: SNOMED, ICD-10 and clinical governance

07:35 The semantic bridge: why structured data alone is not enough

10:17 Patient language versus clinical language in the record

12:23 When an ambient scribe mishears: clinical and billing consequences

14:53 Omissions, bloated problem lists and unnecessary token burn

19:12 Outside the US: the NHS, the Gulf, Singapore and coding complexity

20:46 Who governs clinical data as complexity increases

23:35 Patient-side AI recorders and resistance to external data

26:08 Ambient documentation, billing intensity and payer friction

29:21 The last 20%: rare disease, model limits and AI governance

33:38 Grounding, clinician trust and the cost of misfiring


Faces of Digital Health:

Website: https://www.facesofdigitalhealth.com

Newsletter: https://fodh.substack.com

Spotify: https://open.spotify.com/show/4cElKJHrauyP6QJQaCkvdY

Apple Podcasts: https://podcasts.apple.com/gb/podcast/faces-of-digital-health/id1194284040

LinkedIn: https://www.linkedin.com/company/faces-of-digital-health


#digitalhealth #healthcareAI #clinicalinformatics #EHR #ambientAI #interoperability #healthdata #SNOMED #healthIT #medicalcoding