Ideas
How clinical AI gets evaluated.
Notes on what gets measured when an AI-enabled workflow is tested, what gets missed, and what would make a pilot worth trusting. Written for the people who have to decide whether to deploy.
September 2026
Tested on the wrong question
A model can score well on every metric its builders chose and still make care worse. Three distinct failures — the wrong target, the wrong data, the wrong place — with the evidence for each, and the five questions I would ask before a pilot.
Read it
More to come, including notes from the conversation series.