{# Structured data, built in the view: hand-written JSON in a template is one unescaped quote away from invalid, and an invalid block is silently ignored by every crawler that reads it. #}
Skoora IELTS Practise free →

AI and the doctor

IELTS Academic Reading — IELTS Practice Originals, Reading Practice Test 9, Passage 3

{# No redirect, deliberately: the text above must be the same thing a reader and a crawler get, and the CTA is how a reader crosses into the timed app. #} Answer the questions on this passage The real 13–14 questions for Passage 3, marked instantly, with the sentence that proves each answer. Start now

The passage

Why software that matches specialists has changed so little of what happens to patients

A Software that classifies medical images now performs, on the benchmark tasks used to evaluate it, at a level comparable to experienced specialists. Systems have been approved by regulators for detecting the eye disease that accompanies diabetes, for flagging strokes on brain scans and for reading screening mammograms, and studies of each have reported accuracy that would have been dismissed as fantasy fifteen years ago. The results are not fraudulent and the technology is not a fashion. What is nevertheless true is that very little of it has changed what happens to patients, and the gap between the two statements is the subject worth examining. It is not a gap peculiar to this technology. Most medical innovations that work in a study fail to survive contact with a clinic, and the reasons are usually organisational; what is unusual here is the confidence with which the intermediate step has been assumed away.

B The first reason is that the tasks on which such systems are tested resemble clinical work only distantly. A benchmark presents a clean image, of a known type, from a known machine, labelled by consensus, and asks a single question with a definite answer. A clinic presents an image of uncertain quality from an unfamiliar scanner, of a patient whose history matters, and asks what should be done. A system that answers the first question superbly may be useless for the second, not because it is inaccurate but because accuracy on a narrow question was never the binding constraint.

C The second reason is a technical failure with a mundane name and serious consequences. Models learn the statistical regularities of the data they were trained on, including regularities that have nothing to do with disease: the make of the scanner, the position of a label, the fact that patients scanned lying down are more likely to be seriously ill. When such a model meets data from a different hospital, its performance degrades, sometimes severely, and the degradation is not announced. This is the crucial property. A model that has stopped working does not report that it has stopped working; it continues to produce outputs of exactly the same form, at exactly the same rate, with exactly the same appearance of authority. A widely deployed system for predicting sepsis, evaluated independently at a large hospital system after being installed at hundreds, was found to miss most of the cases it was supposed to catch while generating alerts on thousands of patients who did not develop the condition.

D That episode illustrates a regulatory difficulty as much as a technical one. Approval processes designed for devices assume a fixed product whose behaviour does not change, whereas a model may be retrained, may be deployed on populations quite unlike those it was tested on, and may degrade quietly as clinical practice around it shifts. Several regulators are now constructing frameworks for continuous monitoring of deployed models, which is plainly the right approach and is not yet in place anywhere at scale.

E The proposed remedy for all of this is to keep a clinician in the loop, and it is less of a remedy than it appears. Automation bias is well documented across aviation, industry and medicine: when a system is usually right, the people supervising it stop supervising, and their independent judgement decays through disuse. A reviewer who has agreed with a machine four hundred times is not an independent check on the four hundred and first. Designing an arrangement in which the human contribution remains real is a problem in the organisation of work rather than in software, and it has received a small fraction of the attention devoted to improving the models.

F None of which argues that the technology is worthless, and the applications that have quietly succeeded suggest where the value lies. Systems that triage a queue so that the urgent scan is read first, that measure something a human measures slowly and imprecisely, or that take over documentation, save time without asking anyone to trust a judgement. They are unglamorous, they do not appear in headlines about machines outperforming doctors, and they are the ones actually in use. The pattern is familiar from every previous wave of automation: the applications that survive are the ones that remove drudgery rather than judgement, and they are adopted quietly because nobody has to be persuaded to give anything up.

G The distribution of the benefit is the question that ought to be asked more often. The strongest argument for automated diagnosis is not that it exceeds a specialist in a wealthy hospital but that it could provide a competent reading where there is no specialist at all, and much of the world has no specialist at all. Whether that happens depends on decisions about who the systems are sold to, what they cost, and whether they work on the equipment available in the places that need them — questions of procurement and design rather than of accuracy. The technical problem is largely solved and the interesting problems are the ones nobody gets a prize for solving.

More reading from Reading Practice Test 9