Stanford researchers put nine models through 508 questions about real patient records, checked by clinicians. The best run recalled 78% of the needed facts and the weakest reached 43%. Scores dropped further whenever an answer had to be pieced together from several visits.