In my first year of medical school, one of the few lectures I still remember was on Bayes’ Theorem. Perhaps the inclusion of it in our curriculum was because Harvard is still a little embarrassed by the famous 1978 study from its own medical school, where researchers tested how well 60 physicians and medical students understood priors and base-rate probabilities in the context of clinical scenarios.
"If a test to detect a disease whose prevalence is 1/1,000 has a false positive rate of 5 per cent, what is the chance that a person found to have a positive result actually has the disease, assuming that you know nothing about the person's symptoms or signs?"
(Note: As the question did not mention a false-negative rate, a 100% true-positive rate—or perfect sensitivity—was mathematically assumed).
This is your moment - are you smarter than 20 medical school students, 20 residents, and 20 attending physicians from 1978?
The correct answer is ~1.96%. In this study, a dismal 18% of medical professionals got it right. With the advent of the internet and the addition of biostatistics to the MCAT, have we made any progress? In a 2014 replication study, researchers gave an identical diagnostic word problem to a new cohort of medicine’s best and brightest. Only 23% answered correctly. Evolution did not optimize us for math.
Regardless, as a longtime LessWrong lurker and Slate Star Codex reader, I was thrilled to attend that lecture. We were finally going to talk about probabilistic thinking and how one ought to change their mind in clinical care. Experimentation in the practice of medicine. The House approach. Indeed, the session was a strong practical overview.
But once the class ended, my Bayesian dream mostly faded. Medical training is practically about building models of thought and those models were not Bayesian. During my 2.5 years of clinical rotations, I saw it come up only once or twice, strictly in regards to labs and testing, and almost exclusively when used to confirm a physician’s existing beliefs (“Well, sure the test came back positive, but actually, the false positive rate is high enough that I am going to ignore it.”)
A physician’s job, in my view, boils down to three main components:
Providing comfort
Performing procedures
Perpetrating testing
I would argue that for non-procedural doctors (and with much of the real human connection and care tasks being performed by nurses, care coordinators, and patient advocates) roughly 80% of being a physician is just testing patients. Every lab is a test, sure. But less intuitively, every physical exam procedure is also a test. Every question we ask a patient is itself a diagnostic test. One could ask: What is the positive predictive value of me phrasing a question this way versus another? And if a patient answers affirmatively, how should that affect my map of the territory.
A few academic clinicians actually take this seriously; I remember a domestic violence training session where a researcher highlighted how their work showed specific phrasing in patient interviews could significantly increase the PPV of uncovering abuse in a specific population, and thus recommended we adopt a standard phrasing. But that kind of thinking is far from the norm. And don’t even think about actually conditioning your priors on the specific, complex knowns of a patient.
Fundamentally, I believe medicine today is not practiced with the goal of truth-seeking. The dominant model is still the Kahneman-esque “slow thinking and fast thinking” approach (despite its questionable rigor) relying heavily on heuristics, pattern recognition, and rigid algorithms and flowcharts. The gold standard. Too often do we limit our understanding of a patient to that moment, that one encounter (and perhaps a note from the physician before us) and not the actual sea of data we have and should be using to update our priors.
In this reality, humans are no better than LLMs. In fact, they will be predictably worse. Language models think fast and slow and can easily mimic medical training’s System 1 / System 2 framing. Therefore, it is no surprise LLMs are already achieving superhuman performance on the clinical tasks we measure doctors on: discrete exams and the discrete visit. It begs the question: is this all that is needed to master human health?
My answer is no. Superhuman performance on today’s clinical tasks is not enough to fix medicine, let alone the broader healthcare system. If all it took were superhuman doctors, then my friends at HMS and the attendings I’ve worked with - who by all objective measures are superhuman in their dedication and intellect - would have already been enough to fix American healthcare. We wouldn’t be seeing what we are seeing today: outcomes for overall groups and diverse subgroups are falling behind, costs spiraling out of control, and major insurers like Massachusetts Blue Cross are dropping coverage for new therapies such as GLP-1s for real medical indications despite overwhelming clinical evidence.
LLM-based AI doctors, while powerful and useful for patients already, are still the weaker and less ambitious future for medicine. I believe there is a better path, one actually informed by mathematics and the best method we know for truth-seeking in the world. I call it medicine with prescience, and it suggests a new type of medicine that is practiced based on forecasting and medical superintelligence.
What Matters in Health
Professor Isaac Kohane, a close mentor, helped authored a report in 2011 which called for the creation of a new data network that would begin to integrate the influx of data from the newly institutionalized electronic health record, government population data and biostatistics, and genetic repositories into what he called a “new taxonomy of disease.” It was a thoughtful idea that correctly foretold the struggles and opportunities that could arise from the massive amounts of data collected, combined with the massive amounts of compute that I am sure even Professor Kohane would not have predicted we would now have. What it did not foresee was that models themselves could utilize that data far more effectively than humans could.
I believe compute and data will allow models to gain an intuition about the future that stretches far beyond human capacity. This is prescience.

We already see hints of this in a few, remarkable humans. If you ask around a hospital, doctors will always mention that one older colleague who just seems to know when a patient is going to crash or die, before the vitals even shift. It’s an intuition built on a lifetime of subtle inputs and uninterpretable data processing. It is not part of the standard curriculum of medicine. And it is likely not a wide enough skill to scale (unless we find a source of melange).
Until recently, the idea that we could practice medicine in any meaningful way purely based on data as opposed to human structure was ridiculous. But as compute has expanded, as transformers continue to reflect The Bitter Lesson, and as the amount of data we collect on every human life has far exceeded the ability of any human to process, we have created the perfect conditions for machine prescience.
A superintelligent machine can scale prescience to the practice of medicine itself. Instead of attempting to diagnose and treat based on a population-conformed “gold standard” of care, while simultaneously trying to optimize for 15-minute visit times and RVU targets, machines can become pure optimization functions for what actually matters: health. The exact optimization function can be refined, but I would roughly call it a function of quality adjusted life years (QALYs) and patient satisfaction or comfort. And the nice thing about a function over a human is that it can be clearly audited and refined.
What would medicine with prescience look like? I envision a machine continuously trained on an evergrowing amount of chronological, sequential data until it develops prescience. That machine is then taught to find the best sequences out of the countable infinity and take the moves that result in those sequences. A system capable of calculating the true, unvarnished probabilities of a healthier future and optimized for a pure goal. That’s it. Leave the reasoning and the heuristics to the LLMs and the patient care and interaction to the humans, but the cognitive core of medicine and its decision making ought to be left to those who can see further than we can. Lest we begin an evitable conflict. For the patient, this means we feed the machine every bit and every byte of data known about you, and so long as you are human and live in the system, the machine can plot a trajectory for your life that helps you become your own version of Jeanne Calment.
This shift in the architecture of care toward probabilistic prescience brings alongside it many other benefits. We begin to solve the alignment problems paralyzing American healthcare.
Modern medicine is a broken four-player game. Patients, providers, payors, and pharma/producers are trapped together in a labyrinth of fiercely misaligned incentives, fighting over billing codes and prior authorizations while a human being sits waiting in a paper gown. It is equivalent to a prisoner’s dilemma. For decades, our only answer to this systemic friction has been to demand that doctors become superhuman: that they memorize more, work longer, and somehow carry the weight of a deeply flawed architecture on their own shoulders. Doctors must be responsible for health while every incentive pushes against it.
But we cannot win a game of life simply by working harder. Try as I might, I am no Kasparov. Medicine with prescience allows us to do away with the divides and deep suspicions we have fostered between all parties. Instead, payer, providers, builders, and patients can be aligned around their one goal. And while a reinforcement learning algorithm has an incentive when it plays a game, it is auditable. Today’s game players are not.
At its core, medicine has always been a quiet rebellion against humanity’s own fragility. It is man picking a fight with nature. It is the audacious act of standing at a bedside, or in an operating room, and altering the trajectory of a life. For centuries, we have fought that battle armed with little more than heuristics, hunches, and the sheer endurance of exhausted clinicians. We do not need superhuman doctors. We need systems that finally allow doctors to be human again. If we build the machines to watch the dark, we free the physician to do what they do best: to pull up a chair, look a patient in the eye, and care.
—

I love the way you separated superhuman performance on the clinical tasks we measure from actually fixing medicine. I’m less sure whether a machine can ever have the pure goal you describe, because someone still has to decide what healthier means and whose tradeoffs count. How do you imagine setting that objective without recreating the same payer, provider, and patient conflicts inside the model?