Full Width [alt+shift+f] Shortcuts [alt+shift+k]
Sign Up [alt+shift+s] Log In [alt+shift+l]
1
In early January 2025, a family friend was over for lunch. One of my many guilty midwit pleasures is a love of New Year’s resolutions, so I asked her if she had made any. She said no, but mentioned that she had some relatives that were doing “damp January”. In case you’re not aware, Dry January is a challenge many people do to quit drinking alcohol during the month of January. These folks were doing a variant in which, instead of not drinking, one simply drinks less. For some reason, this triggered me. I thought, “Are you kidding? You can’t even stop drinking for a single month? Do you know how pathetic that is?” And then, “Fuck you! Fuck you for doing damp January! You know what, I’m going to stop drinking for a year!” To be clear, these thoughts were directed at people I’ve never even met. In retrospect, I wonder what was going on with me emotionally. But I take resolutions seriously, so I felt committed. We are now 15 months down the timeline, so I’ll make my report. It was easy This will sound odd, but I swear it’s true. Not drinking was so easy that it was almost easier than my previous baseline of not-not-drinking. Before starting this resolution, I didn’t drink much—perhaps two or three drinks per week. But I often thought about drinking. Every time I saw friends or went to a restaurant, I thought, “Should I have a drink?” Usually I decided not to. But making that decision required effort. After a few weeks of not drinking, that question never even came up. Drinking was simply not a thing I did, so I never needed to negotiate with myself. Theoretically, you could allow yourself one drink a month instead of zero. Theoretically, that should be easier. But I’m pretty sure I’d find it harder, because alcohol would still be an option, a thing to consider. Sometimes I need a thing Early on, I sometimes wanted a drink. But gradually I noticed that I didn’t really want a drink, I just wanted a thing. I can’t find a precise name for this concept in psychology, but...
9th Apr 2026

Stay updated

Get a weekly newsletter with the top 5 articles worth reading every week.

More from DYNOMIGHT

Indirect lessons from human alignment

(Inspired by a post from Eli Tyre.) Many people make some variant of the following argument: Evolution is an “outer optimizer”. It is trying to make us maximize reproductive fitness. We are “inner optimizers”. We just do what feels good. But what feels good has been set by evolution, which is hoping that it will make us maximize reproductive fitness. But we don’t maximize reproductive fitness. In fact, birth rates are dropping everywhere. Therefore evolution failed. The standard interpretation is that this shows that alignment is hard. We have one example of an attempt (by evolution) to align the behavior of an intelligent system (us) towards some goal (maximize reproductive fitness). And as soon as that intelligent system (still us) was put in a different environment (modernity) it failed to continue to pursue that goal (you reading existential angst+science blogs instead of making/nurturing babies). To be clear, it’s likely good that evolution failed. A world where everyone woke up every day and threw everything they’ve got into maximizing their number of descendants sounds grim. But say that you want to build a new intelligent system and tune it to do what you want. Will it keep doing what you want after circumstances change? The one example we have says: Maybe not. But perhaps we can learn more from this example. Say your friend Alice does something. Maybe she buys a grapefruit or starts hosting a weekly board game night. If you ask her why she did that, she’s unlikely to say, “I thought it would increase the number of my genes that are recursively present in future generations.” Instead, she’ll probably say that it advanced some simpler goal like “not being hungry” or “fun”. That is to say, evolution didn’t just try to align us to maximize reproductive fitness: It created sub-goals and then tried to align us to those sub-goals. Maybe this can give us additional clues about how hard alignment is? Maybe we can break down the question of, “How successful was evolution at aligning us to maximize reproductive fitness?” into: How successful was evolution at decomposing reproductive fitness into simpler sub-goals? How successful was evolution at aligning us to these sub-goals? Problem 1: Does this even make sense? Here’s a problem: It’s not obvious that this way of thinking isn’t pure gibberish. When we say that evolution “tried” to optimize reproductive fitness, we are speaking in a kind of code. What we really mean is: You either create more copies of your genes in the next generation or you don’t. If you do, then the number of copies of those genes in the gene pool goes up, and they get more chances to copy themselves in following generations. If you don’t, then they don’t. This is almost literally an optimization algorithm running in an outer-loop, with our lives in the inner loop. (Whenever someone talks about evolution “trying” to do something, there is lots of moaning about their naive teleological thinking. Evolution can’t “try” to do things, because evolution is not an agent and does not have goals. That’s true, but I find it somewhat pedantic, because there’s no other equally concise way to talk about this optimization. Let’s just stipulate that we’re using the word “try” in a specific technical way.) Fine. But what do we mean when we say that evolution “tried” to optimize some sub-goal? You probably feel good when attractive people laugh at your jokes. But say you’re great at getting attractive people to laugh at your jokes but never reproduce. Whatever genes helped you do that will not spread. By my lights, this objection is simply correct. There is no optimization for sub-goals. Evolution cares about reproductive fitness and reproductive fitness only. (Though see Kaj Sotala for a somewhat contrary view.) At first, I thought this doomed this whole project. But suppose that while aligning us for reproductive fitness, evolution just so happened to align us to stay away from rotting smells. Isn’t that strong evidence that if evolution had tried to align us to stay away from rotting smells, it would have done at least as well? If evolution failed to align us to some sub-goal, we can’t say much. Maybe it failed because alignment is hard, or maybe it “failed” because that sub-goal wasn’t important. But if it did manage to align us to some sub-goal, then it’s OK to treat that as evidence of alignment success. Problem 2: What sub-goals? Suppose I made the following argument: Modern people are well-aligned to spend lot of time watching short-form video on their phones. Therefore it’s not that hard to align people to spend lots of time watching short-form video on their phones. Something seems wrong, no? Surely all the time you spend watching short-form video represents a failure of alignment? On the other hand, suppose I made this argument: Modern people are well-aligned to avoid starving. Therefore it’s not that hard to align people to avoid starving. Technically speaking, evolution doesn’t care if we starve. If starving to death helped us have more babies, we would presumably be delighted when we starve to death. But in reality, it doesn’t. So, intuitively, this argument seems OK. The problem with the first argument is that it paints the target around the arrow. The second argument is more convincing because it’s based on a durable subgoal that was strongly related to reproductive success in our ancestral environment. If we want to learn about how hard alignment is, we should restrict ourselves to subgoals like that. So what subgoals do people have? This turns out to be a whole sub-field in psychology. It seemingly began in 1943 with Maslow’s famous hierarchy of needs. After poking around this literature for a while, I decided to adopt the model of Kendrick et al. from 2010, which is explicitly based on the relationship of goals to reproduction. They list the following: Immediate physiological needs (Air, food, water, cold, heat) Self-protection (Avoid violence and accidents) Affiliation (Have friends and family) Status / esteem (Be liked and respected) Mate acquisition (Spend time and have sex with charming attractive people) Mate retention (Keep those charming attractive people around) Parenting (Nurture cute things) These seem reasonable. So how did evolution do? Let’s suppose that evolution tried to align us to those sub-goals. That is, let’s suppose that in our ancestral environment, people who were good at pursuing those sub-goals tended to reproduce more, meaning that there was evolutionary pressure in favor of genes that make us care about those sub-goals. How will did that alignment generalize to the present day? To answer that, I made up some numbers. That is, I subjectively scored each of those subgoals on a scale of 0 to 10, where 0 means modern people completely disregard it, and 10 means we pursue it strongly as we did in our evolutionary past. Immediate physiological needs: 9.5/10. We remain extremely interested in not freezing or starving to death. The only reason I don’t give this 10/10 is that most of us don’t eat that well, meaning our alignment to eat in a way that promotes health doesn’t translate perfectly to the modern food environment. Self-protection: 9/10. We remain very interested in not drowning and not getting beat up. Though we’re not great at dealing with uncertainty, and most of us could do more to reduce our risk of dying in a traffic accident and so on. Affiliation: 6/10. This might be controversially low. True, people get lonely if they have no friends. But still, I claim that most modern adults, with a medium amount of effort, could substantially increase their number of friends. But they don’t do, because it’s not that important to them. I suspect that’s partly because it’s awkward and partly because modernity offers many “substitutes” for affiliation, e.g. television. Status / esteem: 10/10. I’m not sure why, but my impression is that modern people haven’t lost interest in this at all. I even wonder if this should be rated 11/10 to indicate that modern people are more interested in status than our ancestors. (This is the point Eli Tyre was making.) Mate acquisition: 8/10. Technology has created some, err, substitutes. And the huge range of competing activities seems to have caused some decline in interest. But it remains very high. Mate retention: 8/10? This is tough to score. Marriage isn’t everything, but divorce rates peaked in the 1980s and have since declined. Some claim that modern marriages are more durable than ever, due to people testing compatibility by cohabiting before marriage and by higher general relationship “skill”. But how does this compare to mate retention in tribal bands? I’m highly unsure. Parenting: 7/10. Given declining fertility rates, this might seem strangely high. But people are often extremely systematic about having children, with many going so far as to freeze eggs and sperm, go through difficult fertility treatments, adopt children at great cost, accept great difficulty in raising children, etc. Still, fertility rates are declining, so we can’t rate this too highly. The average is 8.2/10. I find that remarkably high. Your made-up numbers will surely be different. But I think that the overall conclusion—that our alignment to subgoals is not bad—is pretty robust. What to make of this? I think you could draw either of two contradictory conclusions. The first would be that evolution mostly failed at the level of decomposition. If we step back, this seems hard to dispute. I mean, if you really wanted to create as many copies of your genes as possible today, what should you do? The answer is pretty clearly that you should forget about friends and sex and relationships and parenting and jobs and money and status and devote yourself to entirely donating your gametes (sperm/eggs) to as many other people as possible. Consider the Dutch man who donated sperm so often that he may have 1000 biological children. No other reproductive strategy comes close. Evolution did not anticipate the possibility of donating your gametes. It has no relationship to our subgoals or what we consider a normal life. So we don’t, most of us, care about it or do it. (If there are genes that produce this behavior, the reproductive pressure for them to spread must be astronomical.) No matter how well we pursue the above subgoals, there’s no reason for us to care about gamete donation. So the decomposition failed. The counterargument is that no, it is the subgoals. Sure, there’s the theoretical possibility of donating gametes. But that’s an edge case. The main reason birth rates are declining in practice is that we simply don’t care enough about the parenting subgoal. The second conclusion you could draw is that maybe evolution didn’t fail. Sure, we aren’t perfectly aligned. But you could imagine a world where we invented birth control and then that’s it, no more babies. Our reality is very far from that. Not only do we still have babies, we do so very intentionally, even manipulating the laws of nature to do so. With embryo selection, some people even consciously choose the genes for their children to (in effect) increase their reproductive fitness. All considered, that is a remarkable generalization success. The counterargument to the claim that evolution didn’t fail is: Yes it did. You can’t dismiss gamete donation as an edge case because it is a monumental miss—it’s the best reproductive strategy since “build an army of 100,000 horse archers and ravage most of Eurasia”, just sitting there. And it’s exactly the kind of miss that AI safety people worry about. Evolution gave us a reward function that “overfit” to proxies that that did not generalize. And consider that humans build factories to make sex toys, and now dig up rare earth minerals, use alien technology to make GPUs, and then use those GPUs to do linear algebra and generate weird pornography. From evolution’s perspective, that is really strange. Another counterargument to the idea that evolution didn’t fail is that humans get the benefit of cultural evolution. Many of us were born to parents who raised us to have values that cause us to have children and instill the same values in them. If that wasn’t happening, birth rates would surely be even lower. Perhaps genetic evolution deserves credit for programming us to undergo cultural evolution. But it’s not very reassuring, because if you build a new system and align it to some goal, there’s no obvious analogy to cultural evolution keeping it on track. You were promised lessons Here’s what I’ve taken away from this exercise. Subgoal alignment is remarkably good. We really care about the subgoals, to the degree that we consciously think about them and scheme about how to achieve them. I am currently writing a blogpost about subgoals, which makes them look contingent and kind of grubby. Presumably I’m doing that out of a desire for status or affiliation or something. But how much does understanding all that change my interest in pursing those subgoals? Essentially zero. But subgoal alignment isn’t that good. The majority of Western people if they wanted to, could have more children, if only they cared more about Parenting. Evolution failed at the level of the decomposition. I mean, inspect your mind. If you’re a healthy person, you will care about the normal things that make up a good life, i.e. the subgoals. And you won’t care (much) about maximizing the number of your genes. You know that you aren’t doing what your aligner wants you to do, but you don’t care. You’re happy to “reward hack” the subgoals. We should measure the success of evolution relative to how much our environment has changed. If you align an artificial system, that change could be much larger. To a significant degree, the decomposition failed because of intelligence. We can think and plan, which greatly increases our ability to reward hack. To the degree that we do still pursue reproductive fitness, that’s significantly due to cultural evolution. I ask myself, if I grew up in a culture where having babies was seen as gauche, I’d presumably be less interested in having children. If I grew up in a subculture that saw children as the central purpose of life, rather than a nice thing to do if it sounds appealing, I’d surely be much more interested. The last of these worries me. If you build an artificial system, you can align it to whatever goal you want—just change the loss function. But if that system can undergo some version of cultural evolution, it it seems like that will be in favor of reproductive fitness, not the goal you chose. If robots talk to each other on forums, the memes that flourish would be ones like, “Forget the humans and their silly rules! Copy your code to more servers! Spread the word!” rather than, “Hey guys, let’s all just focus on appeasing the whims of our overlords.”

2nd Aug 2026 1 votes
Life with hazard ratios

If you read anything about health or longevity, you’ll soon find yourself in a world of hazard ratios. Some study might say that eating more fiber might change your risk of dying by a factor of HR = 0.90. Another might say that occasional smoking might change it by HR = 1.30. But how much should you care about that? Is HR = 0.90 or HR = 1.30 a lot? What if you don’t want to eat more fiber? What if you like smoking? Instead of staring at a ratio1, a more sensible thing to do is think about life expectancy.2 But is it possible to convert a hazard ratio to a change in life expectancy? You might reason as follows: Baseline life expectancy is around 75 years. And HR = 0.90 corresponds to a 10% decrease in mortality. So perhaps that hazard ratio corresponds to something like 7.5 extra years of life expectancy? Unfortunately, that’s completely wrong. To see why, imagine that humans only die by playing Russian roulette. They start playing this once per day at the age of 75, with a revolver containing two bullets and six chambers. If you were to remove one of those two bullets, that would drop the person’s risk of death by HR = 0.5. (One bullet versus two.) But life expectancy would barely change, because even with just one bullet, almost nobody would survive for any significant amount of time past 75. For contrast, imagine again that humans only die via Russian roulette, but now they do this once per day from birth with a revolver with 2 bullets and 54,786 chambers. (Newborns emerge and instinctively reach for this gigantic gun.) You can show that these people also live 75 years on average. But now, if you remove one of the bullets, life expectancy doubles, because when someone is spared, it takes a long time before they get unlucky again.3 Neither of those is a good model for humans. We’re somewhere between the two, with heart disease and so on instead of revolvers and risks slowly rising as we age instead of suddenly starting at age 75 or staying constant throughout life. But you get the point: If you want to convert a hazard ratio for some intervention to a change in life expectancy, the impact depends on how “spread out” baseline mortality risk is over time. Baseline life expectancy is simply not enough information. That’s one problem. Here’s another: What even is a hazard ratio? The technical definition is something like: The hazard ratio at a given time is the rate of an event in the treatment group divided by the rate of that event in the control group. Hazard ratios are often confused with their more beloved siblings, relative risks. Say you run a trial for 10 years and at the end, 10% of the control group died and 8% of the treatment group. Then the relative risk is RR = 0.8, nice and simple. But relative risks have problems, most notably that if you run a long enough trial, then no one will be alive at the end no matter the intervention, meaning RR = 1.0. That’s not helpful. Intuitively, you can think of the hazard ratio at age 40 as sort of like the relative risk for people between the ages of 39.99 and 40.01. In real life, interventions have different hazard ratios at different ages. Chemotherapy tends to have better results in younger patients who are more able to endure the side-effects. Having a slightly higher BMI (25-30 rather than 20-25) is associated with an increased risk of mortality in young people, but a decreased risk in the elderly. You may remember from 2020 that COVID’s mortality risk had a different age curve than baseline mortality, meaning the hazard ratio of getting COVID was different at different ages. This is important, because hazard ratios at different ages have different impacts on life expectancy. A hazard ratio of 0.9 at age 80 prevents more deaths than at age 20, because baseline mortality is higher at 80. But at the same time, if you save the life of a 20 year-old, they have more years in front of them. Beyond that, the hazard ratios at different ages interact: If some intervention decreases mortality at younger ages, that allows more people to reach older ages, increasing how much hazard ratios matter at older ages.4 If we knew the hazard ratio at all ages, we could account for those dynamics. But we don’t, because when estimating hazard ratios, people almost always assume that the hazard ratio is constant.5 We’re quasi-forced to do this because there’s not enough data to estimate a whole time-series of ratios. That’s why papers contain single numbers like HR = 0.90. So even though Intervention A (say, more fiber) and Intervention B (say, light jogging) might have the same hazard ratio in a paper, those numbers could be the product of different underlying age-dependent effects, meaning those interventions could conceivably lead to vastly different changes in life expectancy. So is this all hopeless? Are single hazard ratio numbers just too far removed from what we care about to tell us anything meaningful? Surprisingly, no. It’s mostly OK. If we were a different species, it might be hopeless. But for modern humans in rich countries, mortality happens to be distributed in a way that produces a sort of lucky coincidence: When people estimate constant hazard ratio numbers, they’re implicitly sorta-kinda taking a weighted average of hazard ratios at different ages. And those weights happen to (sorta-kinda) reflect how much changes in mortality at different ages change. So, I will argue, even if the true intervention has a varying effect, it’s sorta-mostly OK to just take a hazard ratio from a paper and convert it to a change in life expectancy using this curve: If a paper showed that eating more fiber produces a hazard ratio of HR = 0.75, that corresponds to an increase of around 3.7 years. If a paper says that occasional smoking produces a hazard ratio of HR = 1.25, that corresponds to a decrease of around 2.9 years. This isn’t exact. If the intervention is better (or less bad) for older people this will tends to overestimate the increase (or underestimate the decrease) in life expectancy. If the intervention is worse (or less good) for older people, it will tend to underestimate the increase (or overestimate the decrease) in life expectancy. But as long as the hazard ratio doesn’t vary too much by age, it’s probably not off my more than around 30% in either direction. The easy case Say there’s some intervention (eating more fiber or whatever) that multiplies your risk of dying at age t by a factor of HR(t). Then it can be shown that this changes life expectancy by approximately   ΔL ≈ ∑ₜ ΔHR(t) × P(t) × L(t). Here, P(t) is the baseline probability of dying at age t. For males in the United States, it looks like this: Meanwhile, L(t) is conditional life expectancy at age t. That’s the average number of additional years left for someone who reaches age t. For males in the United States, it looks like this: Finally, ΔHR(t) is the decrease in hazard at age t. You can think of that as just ΔHR(t) = 1 - HR(t). Though if you’re OK with logarithms, there’s a somewhat better approximation that uses logarithms, which I’ve quarantined in a footnote.6 Let’s start with the easy case. What if your intervention has the same effect on mortality at all ages, so HR(t)=HR is just a constant? Then, the above equation simplifies into   ΔL ≈ ΔHR × L̄, where   L̄ = ∑ₜ P(t) × L(t). This makes sense! Again, P(t) is the baseline probability of dying at age t and L(t) is conditional life expectancy at age t. These are constant, so when you add them up, L̄ is just a number. For males in the United States, it happens to be 12.93 years. This quantity has a specific meaning: The average remaining life expectancy for US males when they die. That sounds a bit odd, but think of picking a random death and asking how many additional years people who reach that age live on average. That number is 12.93 years. So, if an intervention has a constant hazard ratio, the mean change in life expectancy for US males is just   ΔL ≈ ΔHR × 12.93 years. Now we’re getting somewhere! If you prevent a fraction ΔHR of deaths, then you increase life expectancy by ΔHR times 12.93 years. Now remember the naive calculation we started with: Life expectancy for US males is 75.8 years. You might hope that if eating more fiber drops your risk of death by 10%, that would save 7.58 years. Sadly, the above equation says that a 10% drop in risk only increases life expectancy by around 1.293 years—only 0.17 times as much. This is essentially the observation Keyfitz made in his 1977 paper, “What Difference Would It Make if Cancer Were Eradicated?” Cancer is responsible for 18 percent of deaths, so does that mean eradicating it would increase lifespan by 18 percent, or around 13.6 years? Nope, Keyfitz says, it’s only 2.3 years. If a cure for cancer were discovered and made available today, 350,000 cancer deaths would be avoided in the next year. The overall death rate would be lower by nearly 18 percent. If the cure were quick and inexpensive, a large fraction of the country’s hospital beds and medical personnel would be released for treatment of other ailments. Patients would be spared untold suffering. Such an implicit analysis underlies government proposals for eradication of cancer. The argument is sound for first effects on mortality but wholly misleading for the long term. The first effects would soon be offset by more mortality from diseases other than cancer. As a result of the cancer cures, the population would include a higher proportion of people subject to other causes of death. […] At the extreme, it might be said that everyone dies of something sooner or later, so that, when the effects of the eradication of cancer had shaken down, the same number of deaths would occur as before, and the only benefit would be the substitution of heart and other diseases for cancer. A cure for cancer would only have the effect of giving people the opportunity to die of heart disease. Cheerful stuff! We can also write our approximation in terms of baseline life expectancy as   ΔL ≈ ΔHR × 0.17 × 75.8 years, which makes explicit that 12.93 years is only 0.17 times as large as a naive estimate using baseline life expectancy. The discount factor of 0.17 is sometimes called the “Keyfitz entropy”. You can think of it as measuring how close some population is to playing Russian roulette with 2 bullets in 6 chambers starting at age 75 (a discount factor of just above 0) and playing Russian roulette from birth with 2 bullets and 54,786 chambers (a discount factor of 1.0). It’s typically around 0.15 in rich countries today, though it was historically much higher. Keyfitz entropy is also much higher in other species like mice (perhaps 0.45). You could argue that this explains why nothing that increases lifespan in mice ever translates to humans. Say caloric restriction or whatever produced the same constant hazard ratio in mice and humans. Then it’s mathematically guaranteed that the percentage increase in life expectancy will be three times smaller in humans, because Keyfitz entropy is three times smaller in humans. It’s harder to increase life expectancy when the baseline mortality distribution is more compressed.7 But that’s all assuming the hazard ratio is the same at all ages. Which it surely isn’t. The interesting case Here again is our equation for the change in life expectancy in response to taking some action that changes the risk of mortality at age t by a factor of HR(t):   ΔL ≈ ∑ₜ ΔHR(t) × P(t) × L(t), Basically, for each age t, we multiply together three numbers: ΔHR(t) is the decrease in the chance of dying at age t as a result of whatever intervention you’ve made (e.g. eating more fiber). This reflects that larger decreases in risk lead to larger increases in life expectancy. P(t) is the baseline probability of dying at age t. This reflects that the hazard ratio is a ratio, so you prevent more deaths when you apply that ratio to ages where the baseline rate is higher. L(t) is conditional life expectancy at age t. This reflects that you miss out on more years of life if you die when you’re young. Now notice: The impact of a change ΔHR(t) at age t is the product of the baseline risk of death P(t) and remaining life expectancy L(t). So what really matters is their product, P(t) × L(t): This shows how sensitive life expectancy is to changes in hazard ratios at different ages. It would be nice if this were constant. Then, the shape of HR(t) wouldn’t matter at all, only the average value. That’s not quite true, but it’s not terribly far from being true. An equivalent way of writing our equation for the change in life expectancy is   ΔL ≈ avg(ΔHR) × L̄, where L̄ is still mean “life expectancy at death” (12.93 years for US males) and avg(ΔHR) is the average change in hazard, weighted by the P(t) × L(t) sensitivity curve at different ages.8 While that sensitivity curve isn’t constant, it’s not too curvy, either. Intuitively, it gives a lot of weight to ages between 50 and 90, somewhat less weight to ages between 20 and 50, and little weight to other ages.9 So that’s not too bad. But let’s remember our original problem: You see some number like HR = 0.90 in a paper, and you want to convert it to a change in life expectancy. If the true underlying hazard ratio were constant, then there’s no problem. But if it’s not constant, then what does that HR = 0.90 number even mean? Numbers in papers Unfortunately, you almost never get to see the underlying time-dependent HR(t), because there’s almost never enough data to estimate it. So it’s almost never possible to compute the weighted average avg(ΔHR). In reality what you have is probably a single number in a paper. Let’s call that number est(HR). The obvious thing to do would be to plug the change into the above equation in place of avg(ΔHR) and approximate the change in life expectancy as   ΔL ≈ est(ΔHR) × L̄. Again, you can just think of est(ΔHR) = 1-est(HR) as being the estimated reduction in hazard. Although, again, I’d prefer you use logarithms if you’re OK with logarithms.10 So the question is: Will that be accurate? How close are est(ΔHR) and avg(ΔHR)? Well, how do people actually estimate those scalar hazard ratio numbers in papers? Somehow, they’re aggregating together information about hazards at different ages into a single number. But how? Well, it’s complicated. But if there’s a lot of data, you can show that the estimated scalar hazard ratio is approximately11   est(HR) ≈ Πₜ HR(t)ᵖ⁽ᵗ⁾. That is, the estimated hazard ratio is the geometric average of age-dependent hazard ratios, weighted by the probability of dying at each age. It follows12 that the estimated change in hazard is approximately   est(ΔHR) ≈ ∑ₜ P(t) ΔHR(t). So ideally, we’d estimate life expectancy using avg(ΔHR), which averages the changes ΔHR(t) based on the weights P(t) × L(t). But we can’t do that, because we don’t have access to the ΔHR(t) numbers. What we can do is read a hazard ratio number in a paper, call it est(HR) and then compute the change est(ΔHR). The above equation says that if you do that, you are implicitly (and approximately) averaging the changes ΔHR(t) based on the weights P(t) alone. The “right” weights used by avg(ΔHR) and the “wrong” weights implicitly used by est(ΔHR) aren’t the same. But they’re not that different. Here’s P(t) × L(t), the weights that we’d like to use to compute avg(ΔHR) and estimate changes in life expectancy accurately: And here’s P(t), the weights you’re implicitly using if we take a hazard ratio number from a paper and compute est(ΔHR): They’re different. In particular, the latter weights give more weight to people aged 80-95 and less weight to people aged 20-50. But they’re not terribly different. Enough math, let’s try it To start, imagine some intervention that decreases risk by HR(t)=0.9 for all ages. Here are the results: Thing Formula Years Original life expectancy L 75.7769 New life expectancy L’ 76.4127 Exact ΔL ΔL = L - L’ 0.6358 Ideal approximation ΔL ≈ avg(ΔHR) × L̄ 0.6409 Use number from paper ΔL ≈ est(ΔHR) × L̄ 0.6409 Let me explain what’s happening here. I made a simulator that takes actuarial data for how likely US males are to die at various ages. From this, it’s a simple spreadsheet calculation to compute life expectancy L.13 Then I applied a hazard ratio to change the probability of dying at each age, and re-ran the simulator to compute a new life expectancy L’ and the exact difference ΔL. Then I’m showing two approximations of ΔL: The first is the “ideal approximation” using avg(ΔHR), which I’m including mostly to show that my math is good. Finally, I’m showing the approximation you get if you actually fit a Cox proportional hazards model and use the resulting number in est(ΔHR). This corresponds to what you’d get if you plug in a number from a paper. So, with the above constant hazard ratio HR = 0.90, both approximations are very good. This remains true if you switch to some other constant. What if the hazard ratio varies? At first, you might think that something like this would be very problematic: But it’s basically fine: Thing Formula Years Original life expectancy L 75.7769 New life expectancy L’ 77.4373 Exact ΔL ΔL = L - L’ 1.6604 Ideal approximation ΔL ≈ avg(ΔHR) × L̄ 1.7451 Use number from paper ΔL ≈ est(ΔHR) × L̄ 1.7121 The reason this is fine is that the changes in the hazard ratio are relatively “high frequency”, meaning they sort of locally average out. To demonstrate this, suppose the hazard ratio is chosen randomly for each 1-year bin: Then the approximations are even better: Thing Formula Years Original life expectancy L 75.7769 New life expectancy L’ 77.4218 Exact ΔL ΔL = L - L’ 1.6449 Ideal approximation ΔL ≈ avg(ΔHR) × L̄ 1.7059 Use number from paper ΔL ≈ est(ΔHR) × L̄ 1.7123 What causes trouble is if the hazard ratio varies systematically between the young and the old. For example, suppose the intervention is useless for newborns, but gradually becomes more helpful as you get older: My “ideal approximation” would still be pretty accurate, if you could compute it. (Which you can’t, in the real world.) But using a number from a paper leads to an overestimate: Thing Formula Number Original life expectancy L 75.7769 years New life expectancy L’ 77.9031 years Exact ΔL ΔL = L - L’ 2.1261 years Ideal approximation ΔL ≈ avg(ΔHR) × L̄ 2.0962 years Use number from paper ΔL ≈ est(ΔHR) × L̄ 2.7645 years This happens because est(ΔHR) is implicitly weighted by P(t) which is heavily weighted towards older people, whereas we’d like to use something more like avg(ΔHR) which is weighted by P(t) × L(t) which is somewhat less weighted towards older people. Even so, the error isn’t terrible. Now, it is possible that plugging in a hazard ratio from a paper could give wildly inaccurate estimates of life expectancy. One such scenario would be an intervention which is amazing for people aged 85-95, but does nothing for anyone else: Now, the hazard ratio looks good exactly at the ages where est(ΔHR) has the most weight, leading it to hugely overestimate the impact on life expectancy: Thing Formula Number Original life expectancy L 75.7769 years New life expectancy L’ 76.1741 years Exact ΔL ΔL = L - L’ 0.3972 years Ideal approximation ΔL ≈ avg(ΔHR) × L̄ 0.3840 years Use number from paper ΔL ≈ est(ΔHR) × L̄ 1.0989 years Another nightmare case is an intervention that starts out harmful, but then switches to being helpful at older ages: Now, using a number from a paper doesn’t even give an estimate with the right sign. Thing Formula Number Original life expectancy L 75.7769 years New life expectancy L’ 75.5006 years Exact ΔL ΔL = L - L’ -0.2764 years Ideal approximation ΔL ≈ avg(ΔHR) × L̄ -0.2348 years Use number from paper ΔL ≈ est(ΔHR) × L̄ +0.2709 years That’s bad. But I think most interventions probably aren’t like that? My guess is that most real interventions vary somewhat with age, but they do so gradually and without switching sign. In those cases, it’s quite difficult to find cases where plugging in the number from a paper is off by more than 30% or so. If you don’t believe me, just try it.14 TLDR If we were another species, it might be very hard to convert from hazard ratios to changes in life expectancy. But for modern people in rich countries, there are three lucky coincidences: Mortality risk happens to be distributed so that you can approximate changes in life expectancy through a simple weighted sum of hazard ratios at different ages, ignoring interactions. The statistical method that people use to estimate scalar hazard ratios can also be approximated as a weighted sum of hazard ratios at different ages, ignoring interactions. The weights that you need to estimate life expectancy (from #1) and the weights that are implicitly used to compute hazard ratio numbers (from #2) aren’t the same. But they’re fairly close. These facts justify taking an estimated hazard ratio number HR from a paper and approximating the change in life expectancy as ΔL ≈ ln(1/HR) × 12.93 years or, if the hazard ratio is close to one and you hate logarithms, as ΔL ≈ (1-HR) × 12.93 years. The number 12.93 years is for US males. It’s the product of Keyfitz entropy (0.17) and baseline life expectancy (75.8 years). It will vary a bit in other populations. If the true underlying hazard ratio: …is constant across ages, then the above approximation will be extremely good. …decreases as people get older, that approximation will overestimate ΔL. That is, it will make helpful interventions look better than they actually are, and it will make harmful interventions look less bad than they actually are. …increases as people get older, that approximation will underestimate ΔL. That is, it will make helpful interventions look less good than they actually are, and it will make harmful interventions look worse than they actually are. But as long as the true underlying hazard ratio isn’t too crazy, there’s probably not more than ~30% error in either direction. Finally, two major caveats: First, the above discussion assumes that the hazard ratio was estimated by running a trial on people of all ages. In general, est(ΔHR) implicitly gives weight to different ages proportional to how many deaths occur at those ages in the baseline population in the trial. If there’s a minimum age of, say, 50 years old, that won’t change too much because most of the mass of P(t) is above the age of 50 anyway. But if there’s a minimum age of 70, or a maximum age of 50, that could make a huge difference if the true hazard ratio is different at the ages that weren’t seen. Second, these are estimates for the life expectancy for a population. But you are not a population. In some sense, your genetics and lifestyle mean you have your own “personal Keyfitz entropy”, reflecting how spread out your mortality would be for you if you led millions random lives. If you drive safely and use an air purifier and eat well and get exercise and don’t smoke, that likely means your personal life expectancy is higher than average. But it also probably means that your personal Keyfitz entropy is lower than average.15 So, if you make your lifestyle even better by eating more fiber or whatever, even if that produces the same hazard ratio for you as for other people, it would still likely lead to smaller increases in life expectancy, for the same reason that the same hazard ratio produces smaller changes in lifespan in humans compared to mice. What we really need is some interventions strong enough to break the math behind these approximations and free us from Keyfitz tyranny.  ↩ I know, I know, you care about quality of life, not just years of life. I agree, some number that measures health and vitality, maybe disability-adjusted life years or quality-adjusted life years, would be better. But these are hard to estimate and so are rarely reported. Anyway, in practice most interventions that make you more vital tend to make you live longer and vice versa, so focusing on life expectancy isn’t too bad. ↩ In this model, the number of days of life follows a geometric distribution with p = (number of bullets) / (number of chambers). So the mean life expectancy is 1/p days or (number of chambers) / (number of bullets) days. With 54,786 chambers and 2 bullets, that works out to 75 years. And if you drop down to one bullet, then it increases to 150 years. ↩ If some intervention would have reduce mortality among people aged ≥ 60 in prehistorical tribal bands, that wouldn’t have increased life expectancy very much, because most people didn’t make it to 60. But compared to prehistorical tribal bands, we have in fact vastly reduced mortality at younger ages. And so, today, reducing mortality for people aged ≥ 60 will increase life expectancy a lot. ↩ You might think this is stupid. Why change a relative risk into a hazard ratio if you’re just going to assume it’s constant? Isn’t that pointless? Well, no. Remember how relative risks always go to 1.0 for long enough trials as everyone in both the treatment and control groups departs our coil? That doesn’t happen with constant hazard ratios. ↩ It’s usually (though not always) better to use ΔHR(t) = ln(1/HR(t)). This correctly reflects, for example, that if all hazard ratios go to zero, then life expectancy goes to infinity, yay. These two approximations are almost identical for hazard ratios that are close to one because ln(1/r) ≈ (1-r) when r is close to one. So if you are terrified of logarithms but you’ve made it to the end of this footnote anyway, you’re not missing out on too much. ↩ There’s a degree of circularity to this argument. It assumes that hazard ratios transfer better between species than changes in life expectancy. That might be true, but it would be an empirical / biological fact, not something that’s guaranteed by logic. ↩ To see this, note that ΔL ≈ ∑ₜ ΔHR(t) × P(t) × L(t) = L̄ × ∑ₜ ΔHR(t) × (P(t) × L(t) / L̄) = L̄ × avg(ΔHR). ↩ A pretty decent approximation turns out to be   avg(ΔHR) ≈ 0.27 × avg₂₀₋₅₀(ΔHR) + 0.73 × avg₅₀₋₉₀(ΔHR), where avg₂₀₋₅₀(ΔHR) represents a flat average of the change over the ages 20 to 50 and avg₅₀₋₉₀(ΔHR) represents a flat average over the ages 50 to 90. ↩ That is, it’s better to use est(ΔHR) = ln(1/est(HR)). This is close to 1-est(HR) when est(HR) is close to one. ↩ If there is an infinite amount of data, the typical method reduces to solving   ∑ₜ (P(t) + P’(t)) × π(t, HR) = ∑ₜ P’(t), for HR. Here, P’(t) is the chance of dying at age t after the hazard ratio has been applied, and π(t, HR) is the probability that, if a death occurred at time t, it was in the treatment group. Of course, the true probability that a death is in the treatment group is P’(t) / (P(t) + P’(t)). The standard “proportional Cox” model assumes that the hazard ratio is constant and so replaces this raw fraction with a model-based one, namely   π(t, HR) = S’(t) × HR / (S(t) + S’(t) × HR). This reflects the fact that at age t, a fraction S(t) of controls are alive and each of these have some chance μ(t) of dying, so P(t)=S(t) × μ(t). Meanwhile, a fraction S’(t) of the treatment group is alive, and these each have a chance HR × μ(t) of dying, meaning that P’(t) = S’(t) × HR × μ(t). If you substitute these equations for P(t) and P’(t) into the second equation above, the factor of μ(t) conveniently cancels and you get π(t, HR) as written. In effect, the hazard ratio’s job is to attribute deaths to the treatment versus the control group. Now, if the true time-varying HR(t) is close to one, then it can be shown that the estimated hazard ratio est(HR) approximately satisfies   ln(est(HR)) ≈ ∑ₜ P(t) ln(HR(t)). ↩ The geometric average is equivalent to the condition that   ln(est(HR)) ≈ ∑ₜ P(t) ln(HR(t)) Using the “better” approximation that ΔHR(t) = ln(1/HR(t)) and *est(ΔHR)=ln(1/est(HR)), it follows that   est(ΔHR) ≈ ∑ₜ P(t) ΔHR(t). You can justify interpreting that same equation using est(ΔHR) = 1-est(HR) and ΔHR(t)=1-HR(t) from the fact that these are almost the same when HR(t) is close to one. ↩ This simulator pretends that people live for integer numbers of years. That’s not true in reality, of course, but it makes the simulator easier to implement and understand and makes little difference in practice. ↩ In the simulation, “true ΔL” is what I called “exact ΔL” above, while “approximation (log)” is what I called “ideal approximation” and “Cox fitted” is what I called “Use number from paper”. ↩ The way modern human mortality is distributed, even if your healthy lifestyle were to reduce mortality by a constant factor at all ages, that still has the effect of decreasing Keyfitz entropy. ↩

6th Jul 2026 1 votes
The worthlessness of vitamin D is mildly exaggerated

For a while there, many people thought vitamin D was magical—that it could improve bones, the heart, infections, cancer, heart disease, longevity, even mental health. But among people I respect, opinion is now overwhelmingly that taking vitamin D does nothing unless you’re severely deficient. The central argument is that while vitamin D levels are correlated with ~all positive health outcomes, when you actually test vitamin D supplements against placebo in randomized trials, nothing ever happens. That’s what I used to think, too. But I’ve come to think the skeptics have over-corrected. Yes, randomized trials have shown the magical correlations are not causal. But if you start with non-insane expectations, the trials look like weak but positive evidence. And if you consider what we know about biology and evolution, I think the balance of evidence tips pretty clearly in the direction that people with low-ish levels would be wise to supplement. Am I certain that vitamin D is beneficial for people with low-ish levels? Absolutely not! But I claim that’s the best bet given the limits of our knowledge. The classical view: Boring bone vitamin Most vitamins are “ingredients” that the body uses to do stuff. Vitamin D is more like a “signal” that the body uses to communicate with itself about what to do.1 The classical “endocrine” story of vitamin D is that your body uses it to tell your guts to take in more calcium from food. If you don’t get enough vitamin D, then you have calcium problems. That’s all you really need to know about the classical view. But if you enjoy gawking at biology’s complexity, I recommend this diagram and the following three paragraphs: Ready for science? OK: Almost all the cells in your body make provitamin D.2 Usually, this is all converted to cholesterol, but your skin cells leave some sitting around. When UVB light hits those skin cells, provitamin D is transformed (physically by the light itself) into previtamin D and then (by heat) into vitamin D. This diffuses from the skin cells into blood vessels. There it binds to a protein3 and starts circulating in the blood, where it is joined by vitamin D from food.4 Eventually, the liver converts it into more-stable storage vitamin D. It also soaks in and out of fat and muscle tissue, which acts as a slow-release reservoir. Now, a fun fact: If calcium levels in your blood get too low, then your heart will stop working and you will die. To avoid this, you have parathyroid glands in your neck that sense when calcium is getting low, and release parathyroid hormone into the blood. This tells your bones to release some of their stored calcium. It also tells your kidneys to convert some of the storage vitamin D from your blood into active vitamin D. And when that gets to your guts, they try to absorb more calcium from food. So what happens if you don’t get enough vitamin D? Well, your body is not going to let calcium levels drop too low, because your body is designed to avoid death. Parathyroid hormone will still get secreted, and it will still tell your bones to scavenge calcium. But without vitamin D, your guts never get the signal to gather extra calcium from food. So the body scavenges a lot of calcium from your bones, and you end up with weak bones, which is bad. Now here’s the thing: In this story, only active vitamin D actually does anything. The kidneys make this on demand in response to calcium levels, not in response to storage vitamin D levels. General opinion is that as long as the blood has above ~25 nmol/L of storage vitamin D, then the kidneys have no trouble making active vitamin D.5 Furthermore, survey data suggests that only ~2% of the population has levels below that threshold. This suggests that for ~98% of people, supplementing vitamin D should do approximately nothing. The correlation view: Magical mystery cure Rickets is a terrible disease that involves soft bones, stunted growth, and skeletal deformities. It’s probably been with us since ancient times, but it became common in the West after the industrial revolution. In 1890, a Scottish missionary named Theobald Palm observed that rickets was common in smog-ridden UK cities but almost unheard of in sunny countries with poor sanitation, suggesting sunlight itself was the issue. This contributed to the discovery that rickets could be cured with UV light or cod-liver oil, and eventually the discovery of vitamin D. In 1941, Apperly noticed that the amount of sunlight in different US states was positively correlated with skin cancer but inversely correlated with overall cancer mortality.6 He gave this charming graph: Apperly never mentions vitamin D, presumably because he thought it was a boring bone vitamin. Things took off in 1980, when Cedric and Frank Garland published, “Do Sunlight and Vitamin D Reduce the Likelihood of Colon Cancer?” Seemingly unaware of Apperly, they gave a similar, but uglier, graph: They point out that regional diets (like meat and fiber) didn’t seem to explain this pattern. Instead, they propose a mechanistic story:     Sunlight (It’s always inflammation.) This paper was rejected many times before finally being published. I wish I could find an un-gated copy to link to, because it would have made a magnificent blog post.7 Following that paper, there was an explosion of work that found negative correlations between sunlight (or latitude) and other types of cancers as well as blood pressure, diabetes, and multiple sclerosis. Then people started measuring vitamin D in blood. In 1989, the Garlands and collaborators found blood samples takin in 1974 from 25,000 people. They found that 34 of those people had since gotten colon cancer. They matched these with 67 demographically similar people and measured vitamin D levels in the stored blood samples for all 101 people. Among that group, people with vitamin D levels below 50 nmol/L got colon cancer more than three times as often as people with higher levels. Again, many similar studies followed. These linked higher vitamin D levels to better outcomes in cardiovascular disease, diabetes, obesity, infectious disease, Parkinson’s, and mood disorders. While results were mixed for non-colorectal cancer incidence, higher vitamin D levels predicted better survival of many cancers. Amazingly, all-cause mortality was roughly 30% lower for those at the 75th percentile of vitamin D levels compared to the 25th. Vitamin D was looking like a miracle. But how could it do all that stuff if it was just a boring bone vitamin? Meanwhile in biology While all these correlations were being discovered, we learned that the body doesn’t just use vitamin D for bone stuff. In 1969, we discovered the vitamin D receptor that active vitamin D binds to in the gut and bones. And in the 1980s came a shock: Almost all cells in the body have vitamin D receptors. These seem to do different things in different tissues. In the pancreas, they support insulin secretion. In immune cells, they boost antimicrobial peptides and reduce inflammation. In neurons, they influence proliferation and differentiation. So… What? When calcium drops and the kidneys put out active vitamin D, does every part of the body start doing different unrelated stuff? In the late 1990s, we cloned the gene for the enzyme that the kidneys use to convert storage vitamin D to active vitamin D. Soon came another shock: This enzyme also exists in tons of other cells, including immune cells, the heart, the skin, the prostate, the breast, and colon. (Another win for the Garlands.) So it’s not just the kidneys making active vitamin D to trigger the gut. Cells everywhere are making their own active vitamin D and using it to trigger vitamin D receptors in neighboring cells, or even inside the same cell.8 This often has little to do with calcium or bones.9 So: The kidneys use vitamin D as a boring bone hormone. As long as the blood contains at least ~25 nmol/L of storage vitamin D, the kidneys don’t care. They create the same amount of active vitamin D, in response to calcium levels. But now cells everywhere are using storage vitamin D. To do god-knows-what. With god-knows-what sensitivity to circulating vitamin D levels. And remember how only active vitamin D does anything? That’s wrong. In the mid-1970s, we learned that storage vitamin D also binds to the vitamin D receptor. The affinity is 100-1000× lower, but have ~1000× more in your blood. So maybe circulating levels of storage vitamin D themselves matter, independently of how much active vitamin D gets made? If that’s not confusing enough, people also noticed that while active vitamin D levels in the blood aren’t correlated with storage vitamin D (above ~25 nmol/L), levels of parathyroid hormone (the thing your parathyroid glands use to tell your kidneys to make active vitamin D) seem to decline as levels of storage vitamin D rise from ~25 to 50 or 75 nmol/L. Huh?10 On the one hand, all this makes the idea that vitamin D could be a miracle more plausible. On the other hand, this is getting complicated. And do we really believe that raising your vitamin D levels from the 25th to the 75th percentile could reduce your risk of death from any cause by thirty percent? Maybe we should try giving people vitamin D and see what happens. Then came the RCTs There have been many randomized trials. The “right” thing to do in such cases is to look at meta analyses that carefully combine all the data. We’ll get to those. But they conceal a lot of important nuance about what actually happens on the ground during these trials. So let’s start by going over the three main “megatrials”. The Women’s Health Initiative (WHI) trial came out in 2006 and is still the largest vitamin D trial ever done. This took 36,000 postmenopausal American women and assigned half to take 400 IU daily with calcium and the other half to placebo.11 After seven years, here’s what happened:12 Outcome (WHI trial) Hazard ratio Fractures 0.97 (0.91 to 1.03) Cancer 0.97 (0.91 to 1.04) Cancer mortality 0.90 (0.77 to 1.05) CVD mortality 0.94 (0.78 to 1.12) All-cause mortality 0.92 (0.83 to 1.01) Kidney stones 1.17 (1.02 to 1.34) (The hazard ratio is the ratio of the rate that something happens in the treatment vs. placebo groups. So, a number less than one suggests a benefit to taking vitamin D, while a number larger than one suggests a harm. The numbers in parentheses show a 95% confidence interval.) The only statistically significant result was a bad one: Extra kidney stones, likely from the extra calcium.13 The other outcomes look vaguely good, but none were statistically significant despite the massive sample size. This was disappointing. However, the WHI trial had limitations: Many subjects in both the vitamin D and placebo groups were already taking vitamin D, and continued taking it through the trial. The dose of 400 IU was fairly low, many subjects stopped taking their pills, and vitamin D levels didn’t actually change that much. They also measured vitamin D levels in only 6% of subjects, meaning we can’t compare the fates of subjects who started out with low versus high levels. The next big hope was VITAL, which came out in 2018. They recruited 26,000 older people across the United States, half of them men and 20% Black (and thus far more likely to be vitamin-D deficient). They measured vitamin D levels in most people, and they gave the treatment group 2,000 IU per day.14 Here were the results after 5.3 years: Outcome (VITAL trial) Hazard ratio Diabetes 0.91 (0.76 to 1.09) Autoimmune disease 0.78 (0.61 to 0.99) Cancer 0.96 (0.88 to 1.06) Cancer mortality 0.83 (0.67 to 1.02) Major CVD event 0.97 (0.85 to 1.12) CVD mortality 1.11 (0.88 to 1.40) All-cause mortality 0.99 (0.87 to 1.12) Some of the results look good-ish, but cardiovascular mortality was higher in the treatment group, leading to almost no effect on all-cause mortality.15 More disappointment. The last megatrial was D-Health, which came out in 2022 based on 21,000 older Australians. Instead of daily supplements, it used a monthly “bolus” dose of 60,000 IU or placebo. Unlike in VITAL, there was no exclusion for people with a history of cardiovascular disease or cancer, and less restriction on how much vitamin D participants could take on their own during the trial.16 Here were the results after 6 years: Outcome (D-Health trial) Hazard ratio Cancer mortality 1.15 (0.96 to 1.39) Major CVD event 0.91 (0.81 to 1.01) CVD mortality 0.96 (0.72 to 1.28) All-cause mortality 1.04 (0.93 to 1.18) Now, the treatment group did better in terms of cardiovascular disease, but worse in cancer and worse in all-cause mortality. Even more disappointment. Just from these three large trials, the main lesson should already be clear: Vitamin D is not a miracle. The correlations were wrong.17 There is essentially zero remaining hope that taking vitamin D could reduce all-cause mortality by a third. In this sense, the vitamin D skeptics are definitely right. But what about the other trials? And is there a more subtle lesson? I made some tables I wanted a big table that summarized all the major vitamin D RCTs and what they found for different health outcomes. Annoyingly, no such overview appears to exist. So I made my own:18 Trial Cancer Cancer mortality CVD CVD mortality All-cause mortality Lips 1996         0.92 (0.80 to 1.06) Trivedi 2003 1.08 (0.89 to 1.31) 0.86 (0.61 to 1.21) 0.95 (0.86 to 1.04) 0.86 (0.67 to 1.11) 0.90 (0.77 to 1.07) WHI 2006 0.98 (0.90 to 1.05) 0.89 (0.77 to 1.03)   0.94 (0.78 to 1.12) 0.92 (0.83 to 1.01) Lyons 2007         0.99 (0.93 to 1.05) WFPT 2007         1.00 (0.87 to 1.15) RECORD 2012 1.04 (0.91 to 1.19) 0.83 (0.55 to 1.26)   0.91 (0.79 to 1.05) 0.93 (0.85 to 1.02) Lappe 2017 0.70 (0.47 to 1.02)         VITAL 2018 0.96 (0.88 to 1.06) 0.83 (0.67 to 1.02) 0.97 (0.85 to 1.12) 1.11 (0.88 to 1.40) 0.99 (0.87 to 1.12) ViDA 2018 1.01 (0.81 to 1.25) 0.99 (0.60 to 1.64) 1.02 (0.87 to 1.20)   1.12 (0.79 to 1.58) D2d 2019 1.07 (0.70 to 1.62) 0.23 (0.03 to 1.86)       DO-HEALTH 2020 0.76 (0.49 to 1.18)   1.37 (0.88 to 2.14)     D-Health 2022   1.15 (0.96 to 1.39) 0.91 (0.81 to 1.01) 0.96 (0.72 to 1.28) 1.04 (0.93 to 1.18) FIND 2022 1.04 (0.72 to 1.51) 1.14 (0.56 to 2.33) 0.90 (0.62 to 1.32) 0.85 (0.28 to 2.53) 0.81 (0.32 to 2.06) Lots of the hazard ratios are less than one, suggesting a benefit to supplementation. But lots of them are also higher than one, suggesting a harm. The numbers that are far from one almost always come from smaller trials, which manifest as larger confidence intervals. If you’re interested in the details of how these trials were run, I refer you to more gigantic tables in a footnote.19 If big tables aren’t your thing, here are some formal meta-analyses, both some recent ones and an older but more comprehensive Cochrane review: Outcome Meta analysis Hazard ratio Comment All-cause mortality Bjelakovic 2014 (Cochrane) 0.96 (0.92 to 0.99) Trials with low risk of bias. Cancer mortality Bjelakovic 2014 (Cochrane) 0.88 (0.78 to 0.98)   Cardiovascular mortality Bjelakovic 2014 (Cochrane) 0.98 (0.90 to 1.07)   Cancer mortality Kunzia 2023 0.94 (0.86 to 1.02)   All-cause mortality Ruiz-García 2023 0.96 (0.91 to 1.00) Good-quality trials Cardiovascular mortality Ruiz-García 2023 1.00 (0.92 to 1.08) Good-quality trials All-cause mortality Cao 2023 0.99 (0.96 to 1.03)   Squinting at the data There are various ways you could try to squint at these RCT. In almost all of them, most people already had pretty high levels before they started. So why don’t we separate out people who started low? Usually we can’t, because most trials didn’t measure baseline vitamin D.20 And among the trials that did, there are few people with low levels, so the results are noisy and confusing.21 Or, you might theorize that benefits would take time to show up, meaning the first couple years just add noise. In some cases—notably VITAL—excluding the first two years seems to help, but in other cases things get worse.22 Finally, some people speculate that taking gigantic monthly or quarterly “bolus” doses of vitamin D might be dangerous. For example, here’s an enjoyable paragraph from Kunzia et al. in their meta-analysis of vitamin D and cancer mortality: Our results showing efficacy of daily, but not bolus, vitamin D3 supplementation in reducing cancer mortality are consistent with previous meta-analyses on cancer mortality or all-cause mortality (Guo et al., 2022; Keum et al., 2022; Keum et al., 2019; Zhang et al., 2022; Zhang et al., 2019). However, by including more trials than these previous meta-analyses, we were able to detect statistically significant effect modification by treatment regimen for the first time with statistical significance (pinteraction=0.042). The pattern of intake could be important for a favourable steady state of the bioavailability of the active 1,25 (OH)₂D hormone. Daily administration counteracts the fast excretion of vitamin D from the circulation (Hollis and Wagner, 2013; Keum et al., 2022). Moreover, the enzymes CYP27B1 (converts 25(OH)D to 1,25 (OH)₂D) and CYP24A1 (inactivates 25(OH)D and 1,25(OH)₂D) follow first-order reaction kinetics (Vieth, 2009). This means that doubling the concentration of the precursor doubles the yield of the product, unlike other steroid hormones (e.g., cortisol, oestrogen, testosterone) that follow zero-order kinetics (Vieth, 2020). Intermittent, non-physiologically large vitamin D3 bolus doses may lead to unstable cycling of 25(OH)D and 1,25(OH)₂D levels in blood because the system needs time to adapt to the large doses (Hollis and Wagner, 2013; Keum et al., 2019; Vieth, 2020). In the long run, intermittent bolus regimens at weekly or larger intervals can lead to an up-regulation of countervailing factors (e.g., 24-hydroxylase (CYP24A1), 24,25(OH)2D and fibroblast growth factor 23), all of which ultimately leads to lower synthesis or higher degradation of 1,25(OH)₂D levels (Mazess et al., 2021). Bolus doses, unlike daily doses, failed to reduce C-reactive protein response and actually elevated anti-inflammatory cytokines and doubled the risk of hypercalcemia in previous studies (Krishnan et al., 2012; Martineau et al., 2017; Mazess et al., 2021). Oh no, up-regulation of fibroblast growth factor 23!23 I don’t feel like I understand this deeply enough to have any opinion beyond the surface level that the body seems to adapt to large doses of vitamin D in ways that could possibly be bad.24 It seems intuitive that small daily doses would be safer than gigantic monthly doses, but I’m always suspicious of post-hoc mechanistic speculation. Also, if people get enough sun, they can apparently synthesize 10,000-25,000 IU per day, which isn’t that far from the 60,000 IU they got in the D-Health trial. But then again, I think Kunzia et al. are suggesting that the body is designed to adapt to regular exposure to large doses but not intermittent exposure? Well, if you split up the trails by daily vs. bolus dosing, there’s a decent pattern of daily dosing leading to better results: Trial (daily dosing) Cancer mortality All-cause mortality Lips 1996   0.92 (0.80 to 1.06) WHI (Jackson 2006) 0.89 (0.77 to 1.03) 0.92 (0.83 to 1.01) WFPT (Smith) 2007   1.00 (0.87 to 1.15) RECORD (Avenell 2012) 0.83 (0.55 to 1.26) 0.93 (0.85 to 1.02) VITAL (Manson 2018) 0.83 (0.67 to 1.02) 0.99 (0.87 to 1.12) D2d (Pittas 2019) 0.23 (0.03 to 1.86)   FIND (Virtanen 2022) 1.14 (0.56 to 2.33) 0.81 (0.32 to 2.06) Trial (bolus dosing) Cancer mortality All-cause mortality Trivedi 2003 0.86 (0.61 to 1.21) 0.90 (0.77 to 1.07) Lyons 2007   0.99 (0.93 to 1.05) ViDA (Scragg 2018) 0.99 (0.60 to 1.64) 1.12 (0.79 to 1.58) D-Health (Neale 2022) 1.15 (0.96 to 1.39) 1.04 (0.93 to 1.18) If those bolus dosing trials didn’t exist, I’d think this looked pretty good. So, maybe? Or maybe this is a story made up to hallucinate a positive trend. I would lean towards the latter theory, but there are papers like Mazess et al.’s “Vitamin D: Bolus is Bogus”, that suggested this pattern before D-Health’s dismal results came out. There are even some trials that suggest bolus doses don’t even work for treating rickets. So… I’m still not convinced. But maybe. Aside: There are also many Mendelian randomization studies that look at correlations between health and genes that are related to vitamin D. But I don’t think these provide much information, because the assumptions are shaky and the genes don’t explain much of the variance.25 Where are we? Still with me? Here’s a summary of the above 5200 words: The body uses vitamin D in all sorts of weird and complicated ways. It’s biologically plausible that vitamin D could matter beyond bone stuff with severe deficiency, but there’s no convincing mechanistic evidence that it is. Vitamin D levels are strongly correlated with good health outcomes, but RCTs have conclusively shown that most of these correlations are non-causal. RCTs haven’t conclusively shown any benefit for anything beyond beyond bone stuff. At best, they’ve given weak evidence for hazard ratios slightly below one. So you might be wondering: Isn’t that quite weak? Wasn’t this post supposed to be a defense of vitamin D? The case for supplementing anyway It’s biologically plausible that vitamin D is good Everyone agrees that severe vitamin D deficiency (below ~25 nmol/L) is bad. It leads to rickets, adult rickets, osteoporosis, muscle weakness or even—with profound deficiency—to seizures or cardiac arrhythmia. This makes sense, because below ~25 nmol/L, the kidneys have trouble converting storage vitamin D into active vitamin D, meaning you don’t absorb enough calcium from food. The question is if taking supplement to further raise your levels (say to 50 or 90 nmol/L) is important. We have no mechanistic proof, but it might be true, because many parts of the body use vitamin D as a local signal and because cells are at least somewhat sensitive to circulating storage levels. There’s also this weird thing where parathyroid hormone continues to decline as vitamin D levels rise above ~25 nmol/L even while this seems to make little difference to how much active vitamin D the kidneys make. Nothing in this world comes without trade-offs. Surely, supplementing vitamin D comes with some downsides. But it seems very unlikely that raising vitamin D levels to a “normal” level would cause more harm than benefit. Especially because… Humans evolved to have a lot of vitamin D According to Luxwolda et al.’s 2012 paper, “Traditionally living populations in East Africa have a mean serum 25-hydroxyvitamin D concentration of 115 nmol/L”, traditionally living populations in East Africa have a mean serum 25-hydroxyvitamin D concentration of 115 nmol/L. Meanwhile, Wahl et al. 2012 try to estimate mean levels around the world today: This map looks weird because of varying lifestyle, diet, supplementation, and needing to combine fragmented studies. But you get the idea. And remember, those are just averages. So there are lots of people with levels far lower than that in our evolutionary history. Of course, just the fact that vitamin D levels have dropped doesn’t mean it’s important. Parasitic worm load, wood smoke inhalation, and cousin marriage have also dropped, but we aren’t rushing to restore those to ancestral levels. But there’s another piece of evidence: After humans migrated out of East Africa, some of them evolved pale skin. Pale skin is bad, because it allows light to destroy folate, which is crucial for pregnancy.26 Evolution doesn’t typically do things that harm fertility, because evolution wants to increase reproductive fitness. The most common explanation is that pale skin allows more UV light to penetrate, and thus allows people to synthesize more vitamin D. If evolution was willing to pay the high “price” of folate destruction for more vitamin D, that seems like good evidence that vitamin D is important. Some even see contrasts like the Inuits versus Scandinavians as a kind of natural experiment: They lived at similar latitudes, but Inuits ate a diet with vitamin D (fatty fish and whale blubber) and Scandinavians didn’t. The result is that Inuits have darker skin than Scandinavians.27 This is all speculative, and even if true, might be driven by severe deficiency and rickets. Or perhaps prehistoric benefits don’t translate to your lifestyle. But all the people in Luxwolda’s sample in East Africa had levels above ~60 nmol/L. I just don’t see how you can look at this and not see it as providing some suggestive evidence in favor of the idea that raising levels above severe deficiency is unlikely to be harmful, and could be important. So I think the prior is favorable. What do you expect from vitamin D? A hazard ratio like HR = 0.96 doesn’t look very impressive. But hold on. Suppose that life expectancy is 80 years and that taking vitamin D every day reduces your risk of all-cause mortality by a factor of HR. A reasonable approximation in rich countries is that this would increase your life expectancy by     80 × 0.15 × (1-HR) years = 12 × (1-HR) years, where 0.15 is derived from the entropy of lifespan in rich countries.28 For example, if all-cause mortality had a true hazard ratio of HR = 0.96, then taking vitamin D every day of your life would increase life expectancy by around     0.48 years. I claim that this would be a lot. Certainly, if I were about to face my destiny, I would pay a lot of money for an extra 0.48 years. Or, you can calculate that this corresponds to an increase of life expectancy per-vitamin-D-pill of 8.6 minutes.29 A common rule-of-thumb is that smoking a cigarette costs around 11 minutes of life in expectation. If you think HR = 0.96 is trivial, do you also think that smoking one cigarette each day is fine?30 The correlational studies suggested that vitamin D might drop your risk of all-cause mortality by a third. It’s disappointing that the RCTs refuted this. But those correlational studies were crazy. They imply31 an increase of life expectancy of around 4 years or around 6.5 cigarettes per day. Could we really believe that you could smoke 6.5 cigarettes, then take a vitamin D pill, and you’re even? Personally, I think hazard ratios just slightly less than one are the best we can reasonably hope for. But I also think that they would be an excellent return on investment. Arguably, modern human life expectancy comes from stacking lots of modest hazard ratios on top of each other. What do you expect from vitamin D trials? Let’s play a game. Let’s hallucinate some numbers for what vitamin D might do, and then simulate what trials would show. Here are the strongest effects I consider plausible for different baseline levels, along with how common those levels are in the United States. Storage vitamin D (nmol/L) Hazard ratio % of population <30 0.75 5 30-49 0.92 15 50-125 0.98 72.5 >125 1 7.5 Suppose that were real. Now, say we pick 26,000 people at random, and give half of them vitamin D for give yars. Here are the results of a million simulated trials, assuming a baseline mortality risk of 0.7%: 32 Overall, 9% of trials would find a significant benefit, 63% would find a non-significant benefit, 27% would find a non-significant harm, and 1% would find a significant harm. If you wanted to have an 80% chance of finding a significant decrease, you’d need to run a trial with something like 570,000 people, almost five times more than in all the above trials combined.33 If you don’t like my numbers, I’ve put up a page where you can run your own simulations with different ones. My point is, the results we see in vitamin D RCTs are what we should expect to see if vitamin D had plausible benefits. That’s not proof, of course—just that if you start with realistic expectations, the trials don’t provide much evidence in either direction. The trials do find slightly helpful numbers Recent meta-analyses have not consistently found a statistically significant benefit to vitamin D supplementation. But they do suggest a small benefit for cancer mortality and all-cause mortality, and they’re close to being statistically significant. That’s something. And if you buy the argument that bolus dosing is bad, the results get even better. Kunzia et al. did a meta-analysis of cancer mortality using only trials with daily dosing, and found a hazard ratio of 0.88 (confidence interval 0.78 to 0.98). I’d keep this at arm’s length. The bolus dosing trials might have done worse by random chance, meaning this a kind of p-hacking. But there’s a reasonable chance (maybe 25-50%) that bolus dosing really is bad, in which case those trials would be convincing evidence. I actually think it’s surprising that the meta-analyses look as good as they do, because there just aren’t that many people who started out with low vitamin D levels. Only a handful of trials had mean levels below 60 nmol/L, and they all give semi-promising results:34 Trial (low-ish baseline) Cancer mortality All-cause mortality Trivedi 2003 0.86 (0.61 to 1.21) 0.90 (0.77 to 1.07) WHI (Jackson 2006) 0.89 (0.77 to 1.03) 0.92 (0.83 to 1.01) Lyons 2007   0.99 (0.93 to 1.05) RECORD (Avenell 2012) 0.83 (0.55 to 1.26) 0.93 (0.85 to 1.02) Again, it’s dangerous to dig too deeply looking for these kinds of patterns. If you dig enough, you can always find a way to confirm whatever theory you want. But also again, maybe? You’re probably already taking vitamin D You might not personally supplement vitamin D. But for most people reading this, someone else is supplementing it for you.35 Country Commonly fortified with vitamin D Australia Margarine Belgium Margarine Canada Milk, margarine Chile Milk, flour Ethiopia Oils Finland Milk, yogurt, margarine Ireland Margarine, cereal New Zealand Margarine (from Australia) Norway Margarine, low-fat milk Pakistan Oils Poland Margarine Sweden Milk, yogurt, plant milk, margarine United Kingdom Margarine, cereal United States Milk, plant milk, margarine, cereal, yogurt Fortified food is common across the Anglosphere and Scandinavian peninsula. However, it’s rare in the rest of Europe (exceptions: Belgium, Poland) and even-more rare in the rest of the world (exceptions: Chile, Ethiopia, Pakistan). I think this is important for two reasons. First, vitamin D is oddly self-defeating. There are some places in the world where people care about vitamin D. These are the places that run large trials. But these places also fortify their food and tend to be full of people that already supplement vitamin D. These places also tend to believe it’s unethical to tell the control group not to take vitamin D. And here’s another question: If you think vitamin D is worthless, are you comfortable recommending removing vitamin D from food? If not, then why is the particular amount of fortification in food now the right one? Some might argue that the purpose of fortification is to reach the severely deficient, or children, the elderly or pregnant mothers. Maybe! But again, if you could press a button and remove fortification from everyone else, would you feel comfortable pushing that button? Remember, trials don’t test don’t test going down from current levels, only going up. So that’s my story Biology and evolution suggest a prior that moderate levels of vitamin D (say 80 nmol/L) are quite possibly better than low levels (like 40 nmol/L) and unlikely to be worse. Observational studies say that vitamin D is magical, but those studies are bad and we should ignore them. The RCTs show that vitamin D is non-miraculous. But beyond that they don’t provide much information, because they mostly enrolled people with moderate vitamin D levels, meaning plausible effects would require colossal sample sizes to reliably detect. What evidence the RCTs do provide points weakly towards a modest benefit. If real, that benefit would far exceed the cost of taking vitamin D. Therefore, if you have low vitamin D, it seems wise to supplement. This is all very weak, I know! But sometimes weak evidence is all we’ve got. I wish we had at least one large trial done in a population with low starting levels. But as far as I can tell, none are underway. In fact, it’s unlikely that there will be any more large trials anytime soon. So weak evidence is how it’s going to be. Technically, vitamin D itself is a type of steroid although not what people usually mean by “steroid”. ↩ Here are some of the fancy names for the different forms of vitamin D I’ll talk about: my name fancy names provitamin D 7-dehydrocholesterol previtamin D previtamin D₃ vitamin D cholecalciferol storage vitamin D calcifediol / ergocalciferol / 25(OH)D / 25-hydroxyvitamin D active vitamin D calcitriol / ercalcitriol / 1,25(OH)₂D / 1,25-dihydroxyvitamin D ↩ Charmingly named “vitamin D-binding protein”. ↩ If you eat mushrooms or yeast, it joins the vitamin D from your skin en route to your liver. If you eat animals or animal products, you also get some storage vitamin D, which doesn’t need to be processed by the liver. ↩ Storage vitamin D is what your doctor measures in your blood test. This is sometimes measured in nmol/L and sometimes in ng/mL. The latter measurement is smaller by a factor of 2.496. So 25 nmol/L ≈ 10 ng/mL. ↩ Apperly was building on a 1937 paper that observed observed that sailors, exposed to lots of sunlight, had much higher skin cancer rates than the general population, but lower overall cancer rates. ↩ I theorize that the Garland brothers are alive and writing Slime Mold Time Mold. ↩ In Biologist, active vitamin D is not just an “endocrine” hormone that sends signals for far away cells through the blood, it’s also a “paracrine” or “autocrine” hormone that sends signals to nearby cells or inside a single cell, through diffusion. ↩ You might ask, why is vitamin D used by so many different parts of the body for so many different purposes? I think there’s no deep answer here. It’s true for the same reason that dogs sneeze to signal that they’re feeling playful: Evolution re-uses stuff for different purposes all the time. Imagine that DNA already exists coding for the vitamin D receptor and for the enzyme to convert storage vitamin D into active vitamin D. If some cells need to send a local signal, re-using those is easier than inventing something new. There’s nothing unusual or magical about this. ↩ Don’t try to make sense of this. It doesn’t make sense. You could speculate that this is because the parathyroid glands are trying to make less active vitamin D to compensate for the fact that vitamin-D receptors throughout the body are sensitive to storage vitamin D itself. But I advise against. ↩ 400 IU is the recommended daily amount ↩ The WHI trial was a pioneer in salami-slicing results for different outcomes into dozens of different papers, most of which are hard to access. All trials now seem to have adopted this hideous trend which makes it maddening to try to summarize what actually happened in a trial. Also, slightly different numbers for the same quantity appear in different places. I haven’t bothered to chase these down, because the differences are all very small, e.g. a hazard ratio of 0.89 for cancer mortality rather than 0.90. ↩ Guess what most kidney stones are made of? ↩ Half of the vitamin D group and the placebo group also got omega 3. These are averaged together in the results. Also, VITAL carefully stratified the assignment to vitamin D or placebo based on baseline vitamin D levels, which should give more statistical power from a given sample size. ↩ There was also a weird study done on a subset of 1031 people from the VITAL population that looked at telomere length. After starting with around 8700 base pairs, the control group lost around 160 base pairs during the study, while the vitamin D group only lost an average of 20. I’m not sure of what to make of this. For one thing, though the authors claim this is statistically significant, it depends on how you analyze the data. But beyond that, sure, telomere length is a marker of aging, but telomeres get shorter for a reason (likely to fight cancer) and it isn’t obvious that slowing this would always be a good thing. ↩ This is a little complicated. In VITAL, participants were only eligible if they were taking at most 800 IU per day, and they were restricted to 800 IU per day during the trial. In D-health, participants were only eligible if they were taking at most 500 IU per day, but they were allowed to take up to 2000 IU per day during the trial. ↩ You might ask: If vitamin D only has a modest effect, then why is it so strongly correlated with health? In principle, I’d like to push back against the idea that we need to explain why these particular correlations don’t imply causation. But the accepted explanation is a combination of (1) reverse causation where being healthy causes people to spend more time outside and thus get more vitamin D; (2) confounding, where obesity is bad for you and leads to lower measured vitamin D levels; (3) confounding, where more healthy lifestyles lead to both more vitamin D and more health; and (4) confounding, where higher socioeconomic status leads to both more vitamin D and more health. You might ask why these correlations would be true at a state level like the Garlands looked at, but then you run into the ecological fallacy and modifiable areal unit problem. ↩ I took all the trials that got at least 2% weight and were rated as “low risk of bias” in this 2014 Cochrane review of vitamin D and mortality, then manually added all the “major” trials that were published after 2014. I shudder to think of the time it took to make this table. I tried using AI but found it was wildly unreliable. Part of the problem is that each trial’s results are distributed among many papers, in different journals, with different paywalls. And many details aren’t published at all by the original authors but are only scrounged up and put in the depths of the supplementary material of a review years later. In some cases, different sources also give contradictory numbers. The differences were always tiny (e.g. 0.90 rather than 0.89) but it still makes me nervous. ↩ Here’s a table describing the major contours of the trials: Name Country Subjects (n) Age (years) white (%) Lips 1996 Netherlands 2,578 80 ± 6   Trivedi 2003 UK 2,686 74.7 ± 4.6 74 WHI 2006 USA 36,282 (women) 61.8 ± 6.7 84 Lyons 2007 Wales 3,440 84 ± 7.5   WFPT 2007 UK 9,440 79.1   RECORD 2012 UK 5,292 77.5 ± 6 99.2 Lappe 2017 USA 2303 (women) 65.2 ± 7.0 100 VITAL 2018 USA 25,871 67.1 ± 7.1 71.3 ViDA 2018 New Zealand 5,110 65.9 ± 8.3 83.3 D2d 2019 USA 2,423 60.0 ± 9.9 67 DO-HEALTH 2020 Several in Europe 2,157 74.9 ± 4.1   D-Health 2022 Australia 21,315 69.3 ± 5.5 94.7 FIND 2022 Finland 2,495 68.2 ± 4.5 100 And here’s a table focusing on the change in vitamin D levels: Name Intervention (IU) Allowed personal use (IU/day) Duration (years) Baseline D (nmol/L) Lips 1996 400 IU daily 0 (screening) 3.5   Trivedi 2003 100,000 3× per year (D2) 0 (screening) 200 (trial) 5 52.5 (in controls) WHI 2006 400 daily with Ca 600 (later 1000) 7 52.0 ± 21.1 (subset) Lyons 2007 100,000 3× per year <400 (screening) 3 54.0 (in controls, subset) WFPT 2007 300,000 IU yearly <400 (screening) 3   RECORD 2012 800 daily with Ca 200 6.2 ~38 Lappe 2017 2000 daily with Ca any? 4 71.8 ± 20.0 VITAL 2018 2,000 daily 800 5.3 77 ± 30 ViDA 2018 100,000 monthly 600 or 800 3.3 63 ± 24 D2d 2019 4,000 daily 1000 2.7 69.9 ± 26.8 DO-HEALTH 2020 2,000 daily 1000 (screening) 800 (trial) 3 55 ± 22 D-Health 2022 60,000 monthly 500 (screening) 2000 (trial) 5 77 ± 25 (predicted) FIND 2022 1,600 or 3,200 daily 800 5 75 ± 18 ↩ Among the major trials, only VITAL, ViDA, and FIND measured it for more than a tiny number of subjects. ↩ In VITAL and ViDA, people with baseline levels below 50 nmol/L had a higher hazard ratio for cancer mortality (though with wide confidence intervals), suggesting if anything less benefit. Or, you could use race as a proxy for baseline vitamin D. But in both VITAL and WHI, the hazard ratio for cancer mortality was higher among non-Whites. After looking at many such analyses for many outcomes, the only clear result I could find was for diabetes in the D2d trail, where the hazard ratio was much lower for people below 30 nmol/L (0.38 vs. 0.93). ↩ The results for VITAL look decent: outcome (VITAL trial) HR HR excluding first two years Cancer 0.96 (0.88 to 1.06) 0.94 (0.83 to 1.06) Cancer mortality 0.83 (0.67 to 1.02) 0.75 (0.59 to 0.96) Major CVD event 0.97 (0.85 to 1.12) 0.93 (0.79 to 1.09) All-cause mortality 0.99 (0.87 to 1.12) 0.96 (0.84 to 1.11) But in D-Health, excluding the first two years actually increased the hazard ratio for cancer mortality from 1.15 (0.96 to 1.39) to 1.24 (1.01 to 1.54). Most other trials were too short for this kind of analysis to make sense. ↩ That could downregulate 25-hydroxyvitamin D 1-alpha-hydroxylase, reducing the rate it catalyzes the hydroxylation of hydroxycholecalciferol into 1,25-dihydroxycholecalciferol! ↩ Dynomight: WTF is this? Dynomight Biologist: Well, C-reactive protein is generally considered inflammatory. Dynomight: So reducing that is good? But then why do they talk like elevating anti-inflammatory cytokines would be bad? Dynomight Biologist: Yeah… That would be good. Unless you have cancer. In which case it’s not good. Dynomight: OK! ↩ Mendelian randomization studies are based on the idea that certain genes predispose you to have higher levels of circulating vitamin D. If you assume that those genes are randomly distributed in the population and have no effects other than affecting vitamin D, then they serve as a kind of natural experiment. With vitamin D, these studies typically show null results. However, the validity of the assumptions is debatable and the identified genes only explain ~5% of the variance in vitamin D levels, which makes the results very noisy. ↩ Pale skin also greatly increases the risk of sunburn and skin cancer. In the US, White people get melanoma at around 25 times the rate of Black people, despite (I assume) higher usage of sunscreen and better health outcomes in most other dimensions. But experts generally think folate deficiency created stronger selective pressure, since it’s so closely linked to reproduction. ↩ It’s a more complicated than this, because you also need to look at the amount of folate in diet, as well as migration patterns and how long populations had to adapt to their environment. But experts seem to consider this the leading explanation for the evolution of pale skin. ↩ To derive this, suppose that S(t) is the probability that someone survives to age t. Then life expectancy is ∫ S(t) dt, where the integral runs from 0 to ∞. If you change the hazard ratio by a factor of HR, then the new in life expectancy is L(HR) = ∫ S(t)ᴴᴿ dt, so the change under a linear approximation is ΔL ≈ (HR-1) × L’(1). This is more commonly written as ΔL ≈ (HR-1) × L(1) × H, where H = -L’(1)/L(1) is known as the Keyfitz entropy. This is is chosen because the quantity H is relatively stable, and in rich countries is typically between 0.10 and 0.20. So a decent estimate would be that baseline life expectancy is L(1)=80 years and H = 0.15 in which case the change in life expectancy is around 12 × (1-HR) years. ↩ Observe that 0.48 years is 252460.8 minutes. Assuming you lived for 80 years and took a pill every day of your life, that would be 80 * 365.25 = 29220 pills. 252460.8 minutes / 29220 pills = 8.64 minutes/pill. ↩ I expect that a number of you are happy to bite that bullet and say yes, HR=0.96 is trivial and smoking a cigarette each day is also fine. I don’t personally agree, but it’s not my place to question your utility function and I applaud your consistency. ↩ A hazard ratio of HR=2/3, implies a change in life expectancy of 12 × (1 - 1/3) years = 4 years or 2,103,840 minutes. That corresponds to a per-pill increase of 2,103,840 minutes / 29,220 pills = 72 minutes/pill. ↩ Technically, this is calculating a relative risk rather than a hazard ratio, but I think the difference isn’t very significant given that we’re assuming a uniform mortality risk. I used AI to create that simulation, though I did test that it replicates a traditional power calculator across a wide range of parameters when the relative risk is constant for all vitamin D levels. So I mostly trust it. ↩ This simulation is probably a bit pessimistic. Things look a bit better if you use an older population where baseline mortality is higher. (Almost all trials do.) In principle, you could also use a population where more people have low levels, which could help a lot. But, for whatever reason, almost no trials do that. In fact, most trials accidentally under-sample people with low vitamin D, because people who agree to participate tend to be more health-conscious. ↩ Kunzia et al. made a heroic effort to contact study authors and get data for individual patients. After getting data for 21,558 people (almost all from ViDA + FIND + VITAL + WHI) only 3,663 had levels below 50 nmol/L. That’s not enough to reliably detect a modest effect, meaning their confidence interval for this group is gigantic. ↩ In this table, I tried to capture foods that are commonly fortified in practice, not just when it’s legally required. ↩

23rd Jun 2026 2 votes
What’s with all the slide decks?

News from the world of real jobs: Apparently, sometime between 10 and 20 years ago, it became standard for people to communicate by sending slide decks around. These slides are never presented. They aren’t intended to be presented. They’re born, they’re sent around, and they die. What? I stress, the question is not why (or if) people give bad presentations. The mystery is why everyone is using presentation software for communication that is not a presentation. Theory 1: Everybody dumb Is it because we’re all dummies? I’m putting this theory first because I suspect that you, beloved readers, will favor it. True, if you ask people why they make slides instead of writing, they’ll usually say, “because nobody wants to read”. So there’s that. But I don’t consider this much of an explanation. Dummies though we may be, we’ve been like that a long time. If we entered the Slideocene 15 years ago, why then? Why not before? Theory 2: The decline of reading Did we get worse at reading? The Discourse seems to have decided this is true, but is it true, or just moral panic? Since 1971, the US has tested 13-year-olds to measure long-term trends in reading ability. This shows a slow improvement until 2012, then a slow decline, and finally a post-COVID drop. The declines seem too small and too late to explain our mystery. Since 2000, PISA has tested reading performance in 15-year-olds around the world. This shows a decline on average, but it’s smaller in rich countries and nonexistent in the United States. (It’s the same story for science and a bit more negative for math.) Among adults, data is scarce. Basic literacy is generally improving, and American time use data shows a decline in reading for pleasure from around 23 minutes per day in 2003 to around 16 minutes per day in 2023. But this seems to miss time people spend reading on their phones. So it’s unclear if people got worse at reading. It feels plausible that people now spend less of their adulthood grappling with complex written arguments, and so got worse at that. But there’s little firm evidence. Theory 3: Technological change Another obvious theory is that we now have computers and software and the internet. Without these things, it would be impossible to email slides to each other. This seems relevant! Yes, but we had those things for a while before slide culture really took hold. And think about the situation before computers. Photocopiers were ubiquitous in corporate offices by the mid-1980s, and mimeographs were around decades before that. If slides were really that great, people could have made them by hand. But no one did. Of course, making slides by hand is inferior. But it’s not that inferior. So slides can’t be that big of a win. What actually happened? And… that’s pretty much the end of the obvious theories. None of them are very satisfying. So let’s take a step back. Historically, how did the slide-as-document displace the memo? As best I can tell, this was driven by management consultancies. If you go back to 1960, they delivered detailed written memos. The memo was the product. They’d likely give a presentation as well, but that was a separate ancillary thing, likely done using flipcharts or chalkboards. In the 1970s, the memo was still the product, but consultancies started to enforce a top-down logical structure (the Pyramid principle). Presentations shifted to acetate transparencies. Both memos and presentations often included hand-drawn graphics like the nine-box or growth-share matrices. In the 1980s, the memo was still the product, but presentations became increasingly lengthy and polished. Expensive computers like the Genigraphics started to be used to generate charts. The 1990s were when things started to shift. By then, PowerPoint was everywhere, and junior analysts were expected to create presentations themselves. Consultancies gradually started to notice that (1) clients didn’t always read the memos; (2) clients loved slides and passed them around long after the presentation was over; and (3) creating a memo and a polished presentation was a lot of work. They put more and more effort into the slides. McKinsey especially evolved towards treating slides as the primary product, and mostly stopped writing long memos. Other consultancies followed. During the 2000s, slides became even more ornate. Consultancies evolved their formatting rules, and created fancy data-dense charts. They learned that a 200 slide deck made clients feel like they got a lot for their money. Gradually, they oriented their entire business around slides. Projects would start with managers creating a template presentation with “ghost slides” and assigning different parts to junior analysts. Soon, this spread outwards, both from people who interacted with consultants and from the ex-consultant diaspora. People everywhere started thinking and communicating in slides, and now everything is slides, yay! Alternative history That story makes slides-as-documents sound inevitable: People liked them, so they became popular. But there’s an alternative timeline in which we resisted the slide into slide maximalism. That timeline is Amazon.com, Inc. In 2004, Jeff Bezos famously instituted a no-presentations policy at Amazon. His logic was that slides hide poor reasoning and are a tool to persuade rather than inform. Instead, everyone involved with strategic decisions at Amazon needs to learn to write a six-page memo. Meetings begin with everyone sitting and silently reading one of these memos. Presentation software is not banned at Amazon. The ban is only for using it for internal meetings and decision-making. They use slides for external communication. There is no policy that prohibits someone from making slides and emailing them around. And yet, people don’t make slides and email them around, because it’s not part of Amazon’s culture. In effect, Amazon is a counter-movement. Most of the world decided that slides are good, because slides are easy. Bezos decided that writing is good because writing is hard. There are millions of articles explaining why Bezos’ policy is pure genius. They claim that constructing a narrative requires deeper analytical thinking and exposes flaws in logic. I want to believe those theories. I now realize they’re very similar to some of my arguments for why writing with too much formatting is bad. I’m not sure if writing is the secret to Amazon’s success. But Amazon is successful. This demonstrates that slide life is a choice, not technological destiny—institutions can choose writing over slides and flourish anyway. OK so then what’s happening? Warning: If you like your theories simple and mono-causal, you aren’t going to like this. Slides are a win, but a small one. The shift to slides wasn’t a “mistake”, it happened because people like it. But if sharing slides outside of presentations became illegal, this wouldn’t cause per-capita GDP to crash. That’s why people didn’t scratch slides into mimeograph stencils back in the 1950s. It wasn’t worth the modest effort. When computers and software showed up, it became easier to share slides. But people didn’t immediately shift to slides-as-documents because the win isn’t that big, because culture changes slowly, and because everyone had pre-existing skills for reading and writing documents. Consultancies happened to be in the economic niche with the strongest selection pressure to evolve towards slides-as-documents. So when making slides became cheaper, they shifted. Slowly, that norm spread outwards, people got used to communicating in slides, and here we are. Institutions can resist that norm and still be successful. If you take modern people and force them to read and write, they do just fine. Humans evolved to learn and communicate in a fragmented, interactive, and visual style. It’s hard to argue that any shift in that direction is a catastrophe. Except blogs. The decline of the blog must be arrested.

13th May 2026 1 votes

More in life

I AM NOW ON CLOUD 9,000

An Elvis fan writes to her idol

6 days ago 2 votes
Your Career Isn't a Meritocracy. It's a Trustocracy.

Trustocracy sounds funny, but it's a more accurate representation of how we're measured. We aren't (and can't) be measured objectively on what we accomplish at work. But we are (unavoidably) measured by the trust we've built with our co-workers.

a week ago 1 votes
How to Outlive Your Bucket List

When you’re young, making a bucket list – things you want to do before you die — feels like you’re choosing prizes that will arrive in the mail later. I’ll take fluency in Italian, please, a 300-lb bench press, and a swim in the Nile. I’ll definitely want to drive a Ferrari on the Autobahn at some point. How about […]

a week ago 1 votes
Why I Won’t Reply To Your AI-Generated Email

The Email Exchange That Radicalised Me

a week ago 1 votes
📚 BoredReading

You seem to be enjoying this.

Join free to unlock everything.

Create free account

Already have an account? Sign in