Forwarded this by a colleague? Subscribe free and get it every week.
THIS WEEK
A new review found published patient-outcome studies for three of 1,357 FDA-authorized AI devices. Its search could miss unregistered studies and private manufacturer data, but the gap is wide. Everything below lands somewhere on it: tools scaling fast, studies that stop short, and a way to check an AI answer yourself.
In this issue:
MSK's cancer-genomics database moves into OpenEvidence
The VA expands AI scribes without note-quality data
Doctors strike over who controls the AI in their EHR
Spotlight: FDA-cleared isn't the same as tested
How-to: safe prompting, plus a free field guide
LATEST NEWS
MSK's cancer-genomics database is now inside OpenEvidence
Memorial Sloan Kettering and OpenEvidence announced a two-way partnership on Sept. 16 (MSK press release). MSK's OncoKB, a curated database that grades whether a tumor alteration is actionable and which therapies go with it, is being built into OpenEvidence for clinicians outside MSK. OpenEvidence also moves into MSK's own Epic workflow. OncoKB is partially recognized by FDA, and its evidence levels follow AMP/ASCO/CAP consensus. The claim that more than half of U.S. hematologist-oncologists use OpenEvidence is the company's.
Why it matters: When you ask about a specific alteration, check that MSK's evidence level and the literature answer agree before you act on either. The announcement reports no evaluation of the combined answers.
The VA expands AI scribes without note-quality data
The VA is extending ambient scribes from Abridge and Knowtex to all primary care teams and to behavioral health, rehab, and medical and surgical specialists, Christian Robles reports in Nextgov/FCW. More than 986,000 primary care visits have used them since October 2025. Veterans consent verbally and can opt out mid-visit; vendors report fewer than 1% decline. The VA supplied no note-quality data from its rollout. A VHA-funded study the article cites, Reddy et al. in Annals of Internal Medicine, compared notes from 11 AI scribes with notes from 18 human note takers on five standardized primary care visits. Blinded raters scored human notes higher in all five cases on a 50-point quality scale; the widest gap, on a low-back-pain visit, was 43.8 versus 20.3.
Why it matters: The rollout is scaling on adoption and satisfaction numbers. Before you sign a drafted note, check it against what happened in the room, starting with medications and allergies.
Doctors strike, in part, over who controls the AI
About 150 hospitalists, OB-GYNs, psychiatrists, and intensivists at Allina Health's Mercy Hospital in Minnesota struck Sept. 14–17. Jeremy Olson reports in the Minnesota Star Tribune that it is believed to be the first strike by staff physicians at a U.S. private hospital. Pay and sick leave were in dispute. So was AI: physicians described EHR tools nudging billing codes, discharges, and medication choices, and wanted contract language keeping the software a tool, not a decision-maker. Allina said AI adoption is clinically led. A tentative first contract followed Sept. 23, Max Nesterak reports in the Minnesota Reformer.
Why it matters: If your EHR suggests codes, discharges, or orders, find out who turned that feature on and how you override it. The tentative contract includes professional-autonomy language; reports don't establish whether it covers AI.
IN THE LITERATURE
Journal of the American College of Radiology — Goldberg-Stein et al.
Finding: Running in the background on 3,856 consecutive brain CTAs, an FDA-cleared aneurysm tool (Aidoc) found 55 aneurysms radiologists missed. More than half were under 3 mm.
Limitation: Prospective shadow mode with blinded readers, but only disagreements were adjudicated, limiting confidence in the accuracy estimates. Patient outcomes were not measured.
Bottom line: More small aneurysms found; whether that helps patients is unknown.
arXiv preprint (v4) — Wu et al.
Finding: On 1,100 tasks built from real specialist consults, recommendations from the worst of 24 AI systems carried potential for severe harm in 24.6% of tasks; omissions made up more than 80% of severe errors. Clinical tools such as OpenEvidence did better than general chatbots. In a randomized comparison, 101 physicians scored higher with an AI assistant than with conventional resources, but below the AI alone, often because they dropped its useful suggestions.
Limitation: Simulated consults, potential rather than actual harm, not peer reviewed.
Bottom line: Check what the answer left out.
BMJ Digital Health & AI — Seuren et al.
Finding: A narrative review of 27 papers on the risks of ambient scribes; 13 were empirical.
Limitation: Most concerns it catalogs, from deskilling to patients disclosing less, rest on editorials.
Bottom line: The frequency and clinical impact of those risks remain unclear.
SPOTLIGHT
FDA-cleared isn't the same as tested
Abulibdeh, Celi, and colleagues at MIT, Johns Hopkins, Harvard, and elsewhere went through every AI-enabled device FDA had authorized through Dec. 5, 2025: 1,357 of them (PLOS Digital Health). They matched each against ClinicalTrials.gov and PubMed. Thirty-four (2.5%) were linked to a registered prospective trial. Three (0.2%) had measured a patient outcome such as death, complications, or readmission. Industry ran 32 of the 34 trials. Anesthesiology had 22 cleared devices and no registered trials at all.
Most of these devices reached the market through 510(k), which asks for substantial equivalence to an earlier device, not proof that patients do better. The search could miss unregistered studies and private manufacturer data, and it cannot tell us how much evidence is missing. But a clinician can't weigh evidence nobody can see.
Two pieces this month give the gap a shape. In NEJM AI, Nenadic and colleagues lay out five phases of evidence: building the model, internal validation, external validation, a prospective look at how the tool changes decisions, and monitoring after deployment. They argue most published studies stop at internal or external validation. In JACR, Zhang and Hochhegger argue that radiology AI too often clears the bar of beating no tool at all, rather than beating current practice. This week's aneurysm study is a fair example: prospective, carefully done, and still silent on whether patients did better.
Conversational AI sits further back. In Nature Medicine, Schaekermann, Rodman, Goh, and colleagues, most of them at Google, which funded the piece, argue that benchmark scores can't stand in for prospective studies in real clinics. Meanwhile FDA's TEMPO pilot is letting four devices reach Medicare patients without premarket authorization, including an AI voice agent that delivers CBT for depression and anxiety. FDA's own page says their effectiveness hasn't been evaluated; manufacturers must report real-world data instead.
The takeaway: Before using a tool, ask what it was tested to do, in whom, and against what comparator. Then ask whether patient outcomes were measured. If the answer is accuracy on an old dataset, you are the prospective study, so report what goes wrong through your safety system.
HOW-TO
Before you try this: check your organization's approved-tool list, and never put patient information into a tool it hasn't approved.
Ask AI a clinical question without leaking PHI or trusting a bad answer
This issue keeps landing in the same place. The evidence behind most AI tools is thin, and the worst errors are the things an answer leaves out. You can't fix the evidence base. You can change how you ask and how you check. Our free field guide, Safe Prompting for Clinicians, puts that on three pages. Here's the core.
The rule. AI can draft, summarize, and explore options. You still own every clinical decision.
The SAFE method.
Sanitize and set the role. Use an approved account. Give only the details the task needs, and tell the tool it's assisting a clinician. Deleting names and MRNs isn't HIPAA de-identification. Under HIPAA, a vendor that handles PHI for your organization needs a signed BAA first.
Ask one focused question. Name the task, the setting, and the decision you're making.
Flag uncertainty. Ask for sources and for what's unknown. Asking doesn't make the citations real.
Evaluate before use. Open the key sources. Recheck doses, units, and risk scores yourself.
A prompt to copy. Fill the brackets.
You're assisting a licensed clinician. Don't invent facts or fill in missing data.
Task: [one clinical question or drafting task]
Context: [setting, urgency, and only the clinical details this task needs]
Give me:
1. A concise answer
2. The sources that support it
3. Important uncertainties or missing information
End by telling me what might be wrong or might not apply here.When to stop. Never act on AI-generated patient-specific dosing or contraindications without checking a primary reference. If the answer conflicts with the patient or a trusted source, cites something you can't confirm, or adds facts you never gave it, go back to your usual workflow. Look for what's missing, not just what's wrong.
One worked example (fictional case). From the guide, with identifiers left out:
A patient treated for DKA is being considered for urgent surgery. Current values: pH 7.34, bicarbonate 19, anion gap 11, glucose 168, potassium 4.2, beta-hydroxybutyrate 0.7. Hemodynamics are stable.
Are these findings consistent with DKA resolution? Identify remaining concerns, information I should verify, and perioperative considerations. Do not assume missing information is normal. Cite current guidance.My check on whatever comes back: Does it name the resolution criteria it used, and do they match current consensus guidance and your local protocol? Look hard at the ketone value, which sits right at the threshold, and don't accept a clean yes/no that skips it. Did it ask about anything I left out, such as timing of the last insulin dose? Then I decide, not the tool.
Get the full guide. The PDF adds add-on prompts for evidence reviews and patient handouts, a BAA checklist, and five checks before anything reaches the chart.
THAT'S IT FOR THIS WEEK
How was this issue?
Send us a tool or a how-to. Using something in clinic that other clinicians should know about? Submit it here and it may show up in a future issue. We ask before we publish anything and credit you by name only if you want us to.
Know one colleague who’d use this? Forward this email, or send them the link below.
Loading Dose AI is for educational purposes only and is not medical advice. Nothing here should guide the care of a specific patient, and reading it does not create a physician-patient relationship. AI tools help gather and draft each issue; I read every source and edit every word. Opinions are mine and not my employer's. Full disclaimer
