Welcome to AI Health Uncut, a brutally honest newsletter on AI, innovation, and the state of the healthcare market. If you’d like to sign up to receive issues over email, you can do so here.
TL;DR:
On August 12, my co-host Alex Koshykov and I interviewed Tom Kelly, co-founder and CEO of Heidi Health, and Simon Kos, Heidi’s Global Chief Medical Officer and formerly Microsoft’s Chief Medical Officer, on Digital Health Inside Out.
Kelly’s most surprising claim:
Frontier models stopped getting better on Heidi’s core tasks somewhere around Claude Sonnet 3.5. After that, new releases were “noise” on their tasks, and the models “get different, not better.” So Heidi went back to building its own models, and claims the result is both cheaper and better.
That is the mirror image of the argument I made in June in “OpenEvidence Goes Hippocratic AI”: the Nature Medicine study showed frontier models beating specialized clinical AI on every benchmark. I’ll show you where the two claims collide, where they quietly agree, and why the real disagreement is about data moats and economics, not intelligence.
Kelly also cited an unpublished randomized trial claiming only 1 in 10 doctors catches a dangerous hallucination in AI notes, predicted 90% of AI scribes will die, and said “yes, for sure” to becoming the first FDA-regulated AI scribe.
In June 2025 I gave Heidi a neutral rating in my “Top 50 AI Scribes: The Brutal Truth“ ranking (verbose notes, editing burden, weak integration, no independent validation). Inside: what aged well, what I got wrong, and a rapid-fire tour of the rest of the episode.
We finally got Heidi on the show

Quick context. Heidi Health is the Melbourne-born company that grew into one of the most widely used AI scribes on the planet. Kelly is a former vascular surgery resident with a computer science and math background who started building clinical tools in 2019 and first raised for Heidi in 2021. Kos spent years at InterSystems and Cerner, then 15 years at Microsoft, before joining Heidi in October 2025, the same month Heidi announced a $65M Series B led by Steve Cohen’s Point72 at a $465M valuation. The volume claims keep climbing: roughly a million visits a week in early 2025, 2 million at the Series B, and on our show Kelly casually said 3 million.
One fun fact. Heidi began as a history-taking tool that suggested a diagnosis. History into diagnosis: “Hi to Di.” Everyone heard “high to die,” nobody could spell it, and Heidi was born. 😉 Meanwhile, the YouTube transcription bot that processed our own episode renamed me “Sergey Politico” and the show “Digital Halt Inside Out.” I choose to read that as supporting evidence for everything that follows. 😂
Tom Kelly’s claim: the frontier stopped improving on Heidi’s tasks

When GPT-3.5 landed, Kelly and his co-founder Yu Liu had the same reaction at the same time: “Oh no, we’re probably cooked.” In 2023, Heidi stopped trying to out-model the labs and bet on the application layer. Then early 2025 brought scale, and with it a lottery: “I’ve sent the same request one second apart and it’s literally like toddler and PhD.” Kelly even claims performance sags with API load, accusing providers of quietly quantizing models during busy hours. But his core claim was bigger:
“The frontier wasn’t getting better on our tasks. It was kind of logarithmic. Probably Sonnet 3.5 is when it stopped getting better. The next one, on our task, would be like noise.”
—Tom Kelly, CEO of Heidi
And the kicker:
“The models get different, not better. As they get more intelligent, they stop following rules.”
—Tom Kelly, CEO of Heidi
Every builder in this space knows that feeling: a doctor spends hours tuning templates, and a smarter model that improvises is a worse model for that doctor.
So Heidi went back to its deep-tech roots and started fine-tuning its own models. The constraint makes it interesting: no training on patient-identifiable data, and outside the US not even on de-identified data. So they generate synthetic transcripts with frontier models as supervisors, then reinforce against real clinician preference signal (pick output A or B) at 3 million visits a week. Kelly’s conclusion: “We actually just get a data set that the frontiers don’t have,” and the result is “both cheaper and better on our tasks.” The cheaper part comes with numbers: speech-to-text on their own rented hardware is, he says, 100 times cheaper than APIs, and running it through Gemini Flash would cost 2 to 3 million dollars a month while hallucinating more medication names. The strategic logic, verbatim: “They’ll always put 40% to 50% gross margin on cloud, and that’s our margin to eat.”
Two months ago I argued the exact opposite
Now you understand why I was squirming in my chair, in the good way.
On June 12, Nature Medicine published the NYU Langone study showing general-purpose frontier models beating OpenEvidence and UpToDate Expert AI on all three benchmarks, with the specialized tools performing no better than Google’s free AI Overview on real physician queries. I covered it in “OpenEvidence Goes Hippocratic AI,” and I’d been making the underlying argument since “Painful AI Adoption in Medicine“ in August 2024:
Frontier models are erasing the gap with fine-tuned clinical tools, and specialized fine-tuning as a business moat is disappearing.
People laughed at me in 2024. Nobody was laughing in June.
So when the founder of one of the fastest-growing clinical AI companies, a man in the trenches betting his company on the claim, tells me the frontier stopped getting better on his tasks, I have to take it seriously.
Who is right? We are answering different questions.
The Nature Medicine study measured medical knowledge and answer quality, and on that axis Kelly barely disputes the frontier’s win. His concession on air: “I can only imagine how many clinical questions are landing in Claude and ChatGPT. So they actually do have pretty good products on that task.”
Kelly’s plateau (the word is mine, from my question on air, his words were “stopped getting better” and “noise”) is about a different axis: production reliability on a narrow task. Same transcript, same template, ten thousand times a day, no invented drug names, at survivable cost. On that axis, he argues, raw intelligence stopped being the bottleneck around Claude Sonnet 3.5. What matters now is consistency, rule-following, latency, and cost. Framed that way, both stories are true simultaneously.
The frontier keeps getting smarter, but it stopped getting better for Heidi’s tasks.
The remaining disagreement is defensibility. Kelly’s moat has three legs:
preference data the labs can’t touch (their enterprise agreements prohibit training on customer data),
regulation (in the UK, Heidi is a regulated medical device, so it’s not as simple as “API endpoint, Claude Fable, build me an AI scribe, thanks, I’m done”), and
economics (hyperscalers won’t undercut their own cloud margins, “it’s never a can’t, it’s a focus issue”).
My counter, which I said a version of on air: every leg is rented. Focus changes with one board meeting, and the day Big AI goes all in on clinical, the moat evaporates faster than a Twitter thread. To his credit, Kelly didn’t pretend otherwise:
“If they give us a better, cheaper model that meets our standards and is consistent, we would use it. [...] Maybe Fable 7 distilled is incredible at everything and it’s no longer required to do this. I’m not sure.”
That’s more honesty than OpenEvidence has produced on the record, ever. 😉
The safety numbers nobody publishes
Kelly says Heidi ran a randomized trial, not yet published, showing AI scribes introduce errors doctors overwhelmingly fail to catch: “One in 10 doctors will find a dangerous hallucination in notes, and average models create a large number of these dangerous hallucinations.” He described Heidi’s own worst safety incident: a hallucinated chemotherapy drug name, nearly identical to the real one, wrong molecule, plausible dose, invisible to a tired reviewer. The published literature (NOHARM, the UCSF emergency department study) points the same direction.
Which sets up his most quotable prediction: “90% of AI scribes will die.” The survivors, he says, will be the ones that prove consistent performance on safety benchmarks. For the record, I counted 126 AI scribes over a year ago, most of them copycats. If Kelly is right, roughly 113 are walking dead. I won’t pretend to be sad, because I’ve written ad nauseam about the "AI tourist" / copycat problem in clinical AI.
Skeptics would point out that Alex Lebrun, my friend and the CEO of another leading AI scribe company, Nabla, made similar prediction back in January 2025. So far, the jury is still out… That’s the funny thing about forecasting. You can eventually be proven right, but only after everyone who heard the prediction is dead. 😉
When I asked whether Heidi would become the first FDA-regulated AI scribe in the US, Kelly didn’t hedge: “Yes, for sure,” followed by the line I keep thinking about: “It just freaks me out how little the US systems care about safety.” A vendor asking for more regulation of his own category is either very confident or very good at marketing.
My June 2025 Heidi scorecard, revisited
Time to eat my vegetables. In Part 4 of “Top 50 AI Scribes: The Brutal Truth“ (June 2025), I gave Heidi a neutral rating. My reasons: middling note quality with drafts clinicians called way too verbose, a real editing burden, subpar integration (a standalone app with copy-paste into the EHR), and no independent validation. I credited competitive pricing, broad specialty fit, privacy posture, and building in-house instead of being an AI tourist.
Fourteen months later, the audit of my audit.
What aged well: the validation critique stands fully. The impact reports are self-published, the randomized trial is unpublished, and Kelly himself concedes that on structured, problem-based charting benchmarks Heidi is “maybe not quite as good,” his defense being that a scribe welded to the record is “just a feature of the EHR” and not the game Heidi wants to win.
What aged badly: the integration critique needs a major update, with Heidi now live on Epic via SMART on FHIR at Beth Israel Lahey Health and Franciscan Alliance, on the athenahealth Marketplace and eClinicalWorks via Vim, plus a revenue cycle partnership with R1. And one correction: my table listed Heidi as bootstrapped. Wrong then (Blackbird), wrong now (Point72). I stand corrected. 😉
What changed underneath me: per Computer Weekly, more than 80% of Heidi’s workloads now run on their own models. The product I rated and the product shipping today are not the same product, which is exactly why every scorecard in this market has a shelf life measured in months.
Will Heidi submit to an independent blind bake-off, the way Cleveland Clinic ran one across five scribes before choosing Ambience? We ran out of time before I could ask. Tom, Simon, consider it a standing invitation.
Everything else from the episode, in brief
We covered far more than one thesis in 66 minutes. The rapid-fire version:
The mission. Heidi’s stated goal is to “double the world’s healthcare capacity,” and Kos says that vision, not the scribe, is what pulled him from Microsoft. Kelly’s link from notes to capacity: more than half of weekly active users now ask clinical questions inside Heidi, so evidence gets applied in the room, and every saved hour either sees another patient or keeps a doctor from quitting.
The time-savings evidence. Heidi’s own impact reports across Australia and New Zealand, the UK, and Canada claim 59 minutes returned per 8-hour shift, and in New Zealand, where every hospital with an emergency department runs Heidi, an extra patient seen per ED shift. Kos also framed unpaid documentation overtime as a legal risk, citing class actions by doctors. Patient acceptance of recording, they say, exceeds 99%. All self-reported, so trust accordingly.
Two kinds of scribes. Kelly’s taxonomy: EHR-welded form fillers (”not useful, just a feature of the EHR”) versus open-ended AI surfaces where the real US value lives, in denial responses, claims, and patient-specific summaries. Heidi never built the former, which conveniently explains why it benchmarks poorly on it.
The billing elephant. I brought up Texas Oncology going from 3.0 to 4.1 billed diagnosis codes per patient with ambient AI and Riverside’s 14% jump in documented HCCs. Shiv Rao of Abridge told us on this show that scribes are becoming billing tools, and Kelly’s response was blunt: “I think Shiv’s saying that because he has to say that.” His version: doctors chronically under-document (his orthopedic surgeons wrote “day 3, total knee” no matter what happened to the patient), capturing reality is the job, fabrication is the red line, and the endgame is scribes and payers working from one shared transcript instead of an AI-versus-AI upcoding and denial arms race that raises costs on both sides. Hence his interest in the R1 partnership. I remain more cynical about US incentives, and said so.
The public option aside. Asked about fixing those incentives, Kelly, an Australian, said a government-funded alternative creates a counterbalance that pure private markets can’t: in an emergency “you just have to go to the nearest hospital,” and in an all-private system “you’re bankrupt and you’re in trouble.” He doubts the US will ever do it.
The Epic-doom exception. When Christina Farr came on our show, she called Heidi the one smart exception to her thesis that Epic dooms US health AI startups. Kelly confirmed it was half strategy, half luck: after a 2022 HLTH trip he concluded Heidi could never out-raise the US land grab, so they went free, horizontal (dentists, vets, hospitals), and Commonwealth-first, then followed organic US adoption instead of buying it. The ambition now is to be the chosen AI productivity layer for doctors, closer to OpenEvidence or Doximity than to an Epic checkbox. Systems that want the record filler won’t pick Heidi, “and that’s fine.”
Kos’s two busted myths. Fifteen years in health IT taught him that everything must ship through the EHR and that product-led growth can’t work in healthcare. Heidi, he says, broke both. His frame: the EHR remains the system of record, but an AI “system of work” is forming on top, like a slick banking app running against a COBOL mainframe. In the US, where IT decisions are top-down and policy violations have teeth, the scribe gets flattened into an EHR feature, which is why Heidi’s US play needs enterprise sales muscle on top of viral adoption.
Stickiness and the free tier. Three sessions in, Kelly claims roughly 60% long-term retention, powered by “a million little details” like preserving your manual edits when you switch templates. About 100,000 clinicians use Heidi free every week, which he calls “almost like our marketing budget,” funded by owning the models. His analogy: why does Granola exist when every tool has transcription? “It’s just better.”
Sovereignty. A quiet reason for in-house models: countries like New Zealand want clinical AI running in their own cloud, and frontier APIs at those volumes get “super expensive, like a dollar a query.”
The economics tease. No revenue disclosure, but Kelly claims the business works ex-growth and quoted an investor calling Heidi’s economics “the best they’ve seen in the industry.”
Hardware and a jab at Abridge. Heidi now ships its own microphone, paid for now at subscale, possibly bundled later. When Alex noted Abridge’s foundation-model partnership with NVIDIA, the deadpan came back instantly: “It’s fake news. Partnerships aren’t real.” A joke. Mostly.
Why Heidi’s social media is actually funny. Kelly credits the Australian instinct to not take things too seriously and a deliberate wink at the cutting dark humor that gets clinicians through the day. In a commoditized market, he argues, people pick the vendor they like being around.
The roadmap. Less about the next product, more about renting out capability: 150 of “some of the world’s best AI engineers” doing co-development and new IP with enterprises that can’t hire that talent themselves. Scribe, Evidence, and comms are capabilities, and the enterprise motion is “help us solve this gnarly problem.”
My verdict
The Nature Medicine result and the Kelly plateau are not opposites. Frontier models are winning the knowledge game, which is the game OpenEvidence sells, and that moat is gone, as I wrote in June. The production game (consistency, template obedience, drug-name accuracy, regulatory error bars, unit economics at 3 million visits a week) is a different sport, and Kelly makes the strongest case I’ve heard that it currently favors the specialist. Notice the word currently. His own hedge says it best: the moment a frontier lab ships a cheaper, consistent model that meets his standards, he’ll use it. The moat isn’t the model. It’s the preference data, the regulatory scaffolding, and the willingness to be audited.
So publish that trial, publish the impact report methodology, and walk into an independent bake-off with cameras rolling.
A founder who says “90% of AI scribes will die” is implicitly claiming a seat among the survivors. In June 2025 I wasn’t ready to grant Heidi that seat. In August 2026, they’ve earned a strong maybe, which, from me, is practically a love letter. 😉 We had roughly 25 more questions we never got to, including data retention defaults and what multiple Point72 actually paid. Tom and Simon promised a round two in a year. I intend to collect.
Watch the full episode here:
Like what you’re reading in this newsletter? Want more in-depth investigations and research? Alright then—go tell your friends!
👉👉👉👉👉 Hi! My name is Sergei Polevikov. I’m an AI researcher and a healthcare AI startup founder. In my newsletter ‘AI Health Uncut,’ I combine my knowledge of AI models with my unique skills in analyzing the financial health of digital health companies. Why “Uncut”? Because I never sugarcoat or filter the hard truth. I don’t play games, I don’t work for anyone, and therefore, with your support, I produce the most original, the most unbiased, the most unapologetic research in AI, innovation, and healthcare. Thank you for your support of my work. You’re part of a vibrant community of healthcare AI enthusiasts! Your engagement matters. 🙏🙏🙏🙏🙏





The Nature Medicine article is Open Evidence learning the Bitter Lession (https://en.wikipedia.org/wiki/Bitter_lesson). With regard to AI scribes, ever since Meta released wave2vec in 2020 and then MMS and then the OmniLingual ASR suite into open source, I'm seeing a lot claiming to be privacy-preserving open-source voice transcription apps. It seems a strange space in which to try to be profitable when you can get it for free. But doctors want and need turnkey applications, not something to install with brew or pip. And integration into the EHR is probably a must. That's where the competition will be. Going after dentists and veterinarians is smart, until Epic invades that space, too.
Great piece. The part about "plateaued" depending on whose task list you measure against is the key one. On narrow clinical tasks it can look flat while the frontier keeps moving elsewhere, and the safety data nobody publishes is what actually decides deployment. Thanks for putting the scorecard out there.