More from DYNOMIGHT
(Inspired by a post from Eli Tyre.) Many people make some variant of the following argument: Evolution is an “outer optimizer”. It is trying to make us maximize reproductive fitness. We are “inner optimizers”. We just do what feels good. But what feels good has been set by evolution, which is hoping that it will make us maximize reproductive fitness. But we don’t maximize reproductive fitness. In fact, birth rates are dropping everywhere. Therefore evolution failed. The standard interpretation is that this shows that alignment is hard. We have one example of an attempt (by evolution) to align the behavior of an intelligent system (us) towards some goal (maximize reproductive fitness). And as soon as that intelligent system (still us) was put in a different environment (modernity) it failed to continue to pursue that goal (you reading existential angst+science blogs instead of making/nurturing babies). To be clear, it’s likely good that evolution failed. A world where everyone woke up every day and threw everything they’ve got into maximizing their number of descendants sounds grim. But say that you want to build a new intelligent system and tune it to do what you want. Will it keep doing what you want after circumstances change? The one example we have says: Maybe not. But perhaps we can learn more from this example. Say your friend Alice does something. Maybe she buys a grapefruit or starts hosting a weekly board game night. If you ask her why she did that, she’s unlikely to say, “I thought it would increase the number of my genes that are recursively present in future generations.” Instead, she’ll probably say that it advanced some simpler goal like “not being hungry” or “fun”. That is to say, evolution didn’t just try to align us to maximize reproductive fitness: It created sub-goals and then tried to align us to those sub-goals. Maybe this can give us additional clues about how hard alignment is? Maybe we can break down the question of, “How successful was evolution at aligning us to maximize reproductive fitness?” into: How successful was evolution at decomposing reproductive fitness into simpler sub-goals? How successful was evolution at aligning us to these sub-goals? Problem 1: Does this even make sense? Here’s a problem: It’s not obvious that this way of thinking isn’t pure gibberish. When we say that evolution “tried” to optimize reproductive fitness, we are speaking in a kind of code. What we really mean is: You either create more copies of your genes in the next generation or you don’t. If you do, then the number of copies of those genes in the gene pool goes up, and they get more chances to copy themselves in following generations. If you don’t, then they don’t. This is almost literally an optimization algorithm running in an outer-loop, with our lives in the inner loop. (Whenever someone talks about evolution “trying” to do something, there is lots of moaning about their naive teleological thinking. Evolution can’t “try” to do things, because evolution is not an agent and does not have goals. That’s true, but I find it somewhat pedantic, because there’s no other equally concise way to talk about this optimization. Let’s just stipulate that we’re using the word “try” in a specific technical way.) Fine. But what do we mean when we say that evolution “tried” to optimize some sub-goal? You probably feel good when attractive people laugh at your jokes. But say you’re great at getting attractive people to laugh at your jokes but never reproduce. Whatever genes helped you do that will not spread. By my lights, this objection is simply correct. There is no optimization for sub-goals. Evolution cares about reproductive fitness and reproductive fitness only. (Though see Kaj Sotala for a somewhat contrary view.) At first, I thought this doomed this whole project. But suppose that while aligning us for reproductive fitness, evolution just so happened to align us to stay away from rotting smells. Isn’t that strong evidence that if evolution had tried to align us to stay away from rotting smells, it would have done at least as well? If evolution failed to align us to some sub-goal, we can’t say much. Maybe it failed because alignment is hard, or maybe it “failed” because that sub-goal wasn’t important. But if it did manage to align us to some sub-goal, then it’s OK to treat that as evidence of alignment success. Problem 2: What sub-goals? Suppose I made the following argument: Modern people are well-aligned to spend lot of time watching short-form video on their phones. Therefore it’s not that hard to align people to spend lots of time watching short-form video on their phones. Something seems wrong, no? Surely all the time you spend watching short-form video represents a failure of alignment? On the other hand, suppose I made this argument: Modern people are well-aligned to avoid starving. Therefore it’s not that hard to align people to avoid starving. Technically speaking, evolution doesn’t care if we starve. If starving to death helped us have more babies, we would presumably be delighted when we starve to death. But in reality, it doesn’t. So, intuitively, this argument seems OK. The problem with the first argument is that it paints the target around the arrow. The second argument is more convincing because it’s based on a durable subgoal that was strongly related to reproductive success in our ancestral environment. If we want to learn about how hard alignment is, we should restrict ourselves to subgoals like that. So what subgoals do people have? This turns out to be a whole sub-field in psychology. It seemingly began in 1943 with Maslow’s famous hierarchy of needs. After poking around this literature for a while, I decided to adopt the model of Kendrick et al. from 2010, which is explicitly based on the relationship of goals to reproduction. They list the following: Immediate physiological needs (Air, food, water, cold, heat) Self-protection (Avoid violence and accidents) Affiliation (Have friends and family) Status / esteem (Be liked and respected) Mate acquisition (Spend time and have sex with charming attractive people) Mate retention (Keep those charming attractive people around) Parenting (Nurture cute things) These seem reasonable. So how did evolution do? Let’s suppose that evolution tried to align us to those sub-goals. That is, let’s suppose that in our ancestral environment, people who were good at pursuing those sub-goals tended to reproduce more, meaning that there was evolutionary pressure in favor of genes that make us care about those sub-goals. How will did that alignment generalize to the present day? To answer that, I made up some numbers. That is, I subjectively scored each of those subgoals on a scale of 0 to 10, where 0 means modern people completely disregard it, and 10 means we pursue it strongly as we did in our evolutionary past. Immediate physiological needs: 9.5/10. We remain extremely interested in not freezing or starving to death. The only reason I don’t give this 10/10 is that most of us don’t eat that well, meaning our alignment to eat in a way that promotes health doesn’t translate perfectly to the modern food environment. Self-protection: 9/10. We remain very interested in not drowning and not getting beat up. Though we’re not great at dealing with uncertainty, and most of us could do more to reduce our risk of dying in a traffic accident and so on. Affiliation: 6/10. This might be controversially low. True, people get lonely if they have no friends. But still, I claim that most modern adults, with a medium amount of effort, could substantially increase their number of friends. But they don’t do, because it’s not that important to them. I suspect that’s partly because it’s awkward and partly because modernity offers many “substitutes” for affiliation, e.g. television. Status / esteem: 10/10. I’m not sure why, but my impression is that modern people haven’t lost interest in this at all. I even wonder if this should be rated 11/10 to indicate that modern people are more interested in status than our ancestors. (This is the point Eli Tyre was making.) Mate acquisition: 8/10. Technology has created some, err, substitutes. And the huge range of competing activities seems to have caused some decline in interest. But it remains very high. Mate retention: 8/10? This is tough to score. Marriage isn’t everything, but divorce rates peaked in the 1980s and have since declined. Some claim that modern marriages are more durable than ever, due to people testing compatibility by cohabiting before marriage and by higher general relationship “skill”. But how does this compare to mate retention in tribal bands? I’m highly unsure. Parenting: 7/10. Given declining fertility rates, this might seem strangely high. But people are often extremely systematic about having children, with many going so far as to freeze eggs and sperm, go through difficult fertility treatments, adopt children at great cost, accept great difficulty in raising children, etc. Still, fertility rates are declining, so we can’t rate this too highly. The average is 8.2/10. I find that remarkably high. Your made-up numbers will surely be different. But I think that the overall conclusion—that our alignment to subgoals is not bad—is pretty robust. What to make of this? I think you could draw either of two contradictory conclusions. The first would be that evolution mostly failed at the level of decomposition. If we step back, this seems hard to dispute. I mean, if you really wanted to create as many copies of your genes as possible today, what should you do? The answer is pretty clearly that you should forget about friends and sex and relationships and parenting and jobs and money and status and devote yourself to entirely donating your gametes (sperm/eggs) to as many other people as possible. Consider the Dutch man who donated sperm so often that he may have 1000 biological children. No other reproductive strategy comes close. Evolution did not anticipate the possibility of donating your gametes. It has no relationship to our subgoals or what we consider a normal life. So we don’t, most of us, care about it or do it. (If there are genes that produce this behavior, the reproductive pressure for them to spread must be astronomical.) No matter how well we pursue the above subgoals, there’s no reason for us to care about gamete donation. So the decomposition failed. The counterargument is that no, it is the subgoals. Sure, there’s the theoretical possibility of donating gametes. But that’s an edge case. The main reason birth rates are declining in practice is that we simply don’t care enough about the parenting subgoal. The second conclusion you could draw is that maybe evolution didn’t fail. Sure, we aren’t perfectly aligned. But you could imagine a world where we invented birth control and then that’s it, no more babies. Our reality is very far from that. Not only do we still have babies, we do so very intentionally, even manipulating the laws of nature to do so. With embryo selection, some people even consciously choose the genes for their children to (in effect) increase their reproductive fitness. All considered, that is a remarkable generalization success. The counterargument to the claim that evolution didn’t fail is: Yes it did. You can’t dismiss gamete donation as an edge case because it is a monumental miss—it’s the best reproductive strategy since “build an army of 100,000 horse archers and ravage most of Eurasia”, just sitting there. And it’s exactly the kind of miss that AI safety people worry about. Evolution gave us a reward function that “overfit” to proxies that that did not generalize. And consider that humans build factories to make sex toys, and now dig up rare earth minerals, use alien technology to make GPUs, and then use those GPUs to do linear algebra and generate weird pornography. From evolution’s perspective, that is really strange. Another counterargument to the idea that evolution didn’t fail is that humans get the benefit of cultural evolution. Many of us were born to parents who raised us to have values that cause us to have children and instill the same values in them. If that wasn’t happening, birth rates would surely be even lower. Perhaps genetic evolution deserves credit for programming us to undergo cultural evolution. But it’s not very reassuring, because if you build a new system and align it to some goal, there’s no obvious analogy to cultural evolution keeping it on track. You were promised lessons Here’s what I’ve taken away from this exercise. Subgoal alignment is remarkably good. We really care about the subgoals, to the degree that we consciously think about them and scheme about how to achieve them. I am currently writing a blogpost about subgoals, which makes them look contingent and kind of grubby. Presumably I’m doing that out of a desire for status or affiliation or something. But how much does understanding all that change my interest in pursing those subgoals? Essentially zero. But subgoal alignment isn’t that good. The majority of Western people if they wanted to, could have more children, if only they cared more about Parenting. Evolution failed at the level of the decomposition. I mean, inspect your mind. If you’re a healthy person, you will care about the normal things that make up a good life, i.e. the subgoals. And you won’t care (much) about maximizing the number of your genes. You know that you aren’t doing what your aligner wants you to do, but you don’t care. You’re happy to “reward hack” the subgoals. We should measure the success of evolution relative to how much our environment has changed. If you align an artificial system, that change could be much larger. To a significant degree, the decomposition failed because of intelligence. We can think and plan, which greatly increases our ability to reward hack. To the degree that we do still pursue reproductive fitness, that’s significantly due to cultural evolution. I ask myself, if I grew up in a culture where having babies was seen as gauche, I’d presumably be less interested in having children. If I grew up in a subculture that saw children as the central purpose of life, rather than a nice thing to do if it sounds appealing, I’d surely be much more interested. The last of these worries me. If you build an artificial system, you can align it to whatever goal you want—just change the loss function. But if that system can undergo some version of cultural evolution, it it seems like that will be in favor of reproductive fitness, not the goal you chose. If robots talk to each other on forums, the memes that flourish would be ones like, “Forget the humans and their silly rules! Copy your code to more servers! Spread the word!” rather than, “Hey guys, let’s all just focus on appeasing the whims of our overlords.”
For a while there, many people thought vitamin D was magical—that it could improve bones, the heart, infections, cancer, heart disease, longevity, even mental health. But among people I respect, opinion is now overwhelmingly that taking vitamin D does nothing unless you’re severely deficient. The central argument is that while vitamin D levels are correlated with ~all positive health outcomes, when you actually test vitamin D supplements against placebo in randomized trials, nothing ever happens. That’s what I used to think, too. But I’ve come to think the skeptics have over-corrected. Yes, randomized trials have shown the magical correlations are not causal. But if you start with non-insane expectations, the trials look like weak but positive evidence. And if you consider what we know about biology and evolution, I think the balance of evidence tips pretty clearly in the direction that people with low-ish levels would be wise to supplement. Am I certain that vitamin D is beneficial for people with low-ish levels? Absolutely not! But I claim that’s the best bet given the limits of our knowledge. The classical view: Boring bone vitamin Most vitamins are “ingredients” that the body uses to do stuff. Vitamin D is more like a “signal” that the body uses to communicate with itself about what to do.1 The classical “endocrine” story of vitamin D is that your body uses it to tell your guts to take in more calcium from food. If you don’t get enough vitamin D, then you have calcium problems. That’s all you really need to know about the classical view. But if you enjoy gawking at biology’s complexity, I recommend this diagram and the following three paragraphs: Ready for science? OK: Almost all the cells in your body make provitamin D.2 Usually, this is all converted to cholesterol, but your skin cells leave some sitting around. When UVB light hits those skin cells, provitamin D is transformed (physically by the light itself) into previtamin D and then (by heat) into vitamin D. This diffuses from the skin cells into blood vessels. There it binds to a protein3 and starts circulating in the blood, where it is joined by vitamin D from food.4 Eventually, the liver converts it into more-stable storage vitamin D. It also soaks in and out of fat and muscle tissue, which acts as a slow-release reservoir. Now, a fun fact: If calcium levels in your blood get too low, then your heart will stop working and you will die. To avoid this, you have parathyroid glands in your neck that sense when calcium is getting low, and release parathyroid hormone into the blood. This tells your bones to release some of their stored calcium. It also tells your kidneys to convert some of the storage vitamin D from your blood into active vitamin D. And when that gets to your guts, they try to absorb more calcium from food. So what happens if you don’t get enough vitamin D? Well, your body is not going to let calcium levels drop too low, because your body is designed to avoid death. Parathyroid hormone will still get secreted, and it will still tell your bones to scavenge calcium. But without vitamin D, your guts never get the signal to gather extra calcium from food. So the body scavenges a lot of calcium from your bones, and you end up with weak bones, which is bad. Now here’s the thing: In this story, only active vitamin D actually does anything. The kidneys make this on demand in response to calcium levels, not in response to storage vitamin D levels. General opinion is that as long as the blood has above ~25 nmol/L of storage vitamin D, then the kidneys have no trouble making active vitamin D.5 Furthermore, survey data suggests that only ~2% of the population has levels below that threshold. This suggests that for ~98% of people, supplementing vitamin D should do approximately nothing. The correlation view: Magical mystery cure Rickets is a terrible disease that involves soft bones, stunted growth, and skeletal deformities. It’s probably been with us since ancient times, but it became common in the West after the industrial revolution. In 1890, a Scottish missionary named Theobald Palm observed that rickets was common in smog-ridden UK cities but almost unheard of in sunny countries with poor sanitation, suggesting sunlight itself was the issue. This contributed to the discovery that rickets could be cured with UV light or cod-liver oil, and eventually the discovery of vitamin D. In 1941, Apperly noticed that the amount of sunlight in different US states was positively correlated with skin cancer but inversely correlated with overall cancer mortality.6 He gave this charming graph: Apperly never mentions vitamin D, presumably because he thought it was a boring bone vitamin. Things took off in 1980, when Cedric and Frank Garland published, “Do Sunlight and Vitamin D Reduce the Likelihood of Colon Cancer?” Seemingly unaware of Apperly, they gave a similar, but uglier, graph: They point out that regional diets (like meat and fiber) didn’t seem to explain this pattern. Instead, they propose a mechanistic story: Sunlight (It’s always inflammation.) This paper was rejected many times before finally being published. I wish I could find an un-gated copy to link to, because it would have made a magnificent blog post.7 Following that paper, there was an explosion of work that found negative correlations between sunlight (or latitude) and other types of cancers as well as blood pressure, diabetes, and multiple sclerosis. Then people started measuring vitamin D in blood. In 1989, the Garlands and collaborators found blood samples takin in 1974 from 25,000 people. They found that 34 of those people had since gotten colon cancer. They matched these with 67 demographically similar people and measured vitamin D levels in the stored blood samples for all 101 people. Among that group, people with vitamin D levels below 50 nmol/L got colon cancer more than three times as often as people with higher levels. Again, many similar studies followed. These linked higher vitamin D levels to better outcomes in cardiovascular disease, diabetes, obesity, infectious disease, Parkinson’s, and mood disorders. While results were mixed for non-colorectal cancer incidence, higher vitamin D levels predicted better survival of many cancers. Amazingly, all-cause mortality was roughly 30% lower for those at the 75th percentile of vitamin D levels compared to the 25th. Vitamin D was looking like a miracle. But how could it do all that stuff if it was just a boring bone vitamin? Meanwhile in biology While all these correlations were being discovered, we learned that the body doesn’t just use vitamin D for bone stuff. In 1969, we discovered the vitamin D receptor that active vitamin D binds to in the gut and bones. And in the 1980s came a shock: Almost all cells in the body have vitamin D receptors. These seem to do different things in different tissues. In the pancreas, they support insulin secretion. In immune cells, they boost antimicrobial peptides and reduce inflammation. In neurons, they influence proliferation and differentiation. So… What? When calcium drops and the kidneys put out active vitamin D, does every part of the body start doing different unrelated stuff? In the late 1990s, we cloned the gene for the enzyme that the kidneys use to convert storage vitamin D to active vitamin D. Soon came another shock: This enzyme also exists in tons of other cells, including immune cells, the heart, the skin, the prostate, the breast, and colon. (Another win for the Garlands.) So it’s not just the kidneys making active vitamin D to trigger the gut. Cells everywhere are making their own active vitamin D and using it to trigger vitamin D receptors in neighboring cells, or even inside the same cell.8 This often has little to do with calcium or bones.9 So: The kidneys use vitamin D as a boring bone hormone. As long as the blood contains at least ~25 nmol/L of storage vitamin D, the kidneys don’t care. They create the same amount of active vitamin D, in response to calcium levels. But now cells everywhere are using storage vitamin D. To do god-knows-what. With god-knows-what sensitivity to circulating vitamin D levels. And remember how only active vitamin D does anything? That’s wrong. In the mid-1970s, we learned that storage vitamin D also binds to the vitamin D receptor. The affinity is 100-1000× lower, but have ~1000× more in your blood. So maybe circulating levels of storage vitamin D themselves matter, independently of how much active vitamin D gets made? If that’s not confusing enough, people also noticed that while active vitamin D levels in the blood aren’t correlated with storage vitamin D (above ~25 nmol/L), levels of parathyroid hormone (the thing your parathyroid glands use to tell your kidneys to make active vitamin D) seem to decline as levels of storage vitamin D rise from ~25 to 50 or 75 nmol/L. Huh?10 On the one hand, all this makes the idea that vitamin D could be a miracle more plausible. On the other hand, this is getting complicated. And do we really believe that raising your vitamin D levels from the 25th to the 75th percentile could reduce your risk of death from any cause by thirty percent? Maybe we should try giving people vitamin D and see what happens. Then came the RCTs There have been many randomized trials. The “right” thing to do in such cases is to look at meta analyses that carefully combine all the data. We’ll get to those. But they conceal a lot of important nuance about what actually happens on the ground during these trials. So let’s start by going over the three main “megatrials”. The Women’s Health Initiative (WHI) trial came out in 2006 and is still the largest vitamin D trial ever done. This took 36,000 postmenopausal American women and assigned half to take 400 IU daily with calcium and the other half to placebo.11 After seven years, here’s what happened:12 Outcome (WHI trial) Hazard ratio Fractures 0.97 (0.91 to 1.03) Cancer 0.97 (0.91 to 1.04) Cancer mortality 0.90 (0.77 to 1.05) CVD mortality 0.94 (0.78 to 1.12) All-cause mortality 0.92 (0.83 to 1.01) Kidney stones 1.17 (1.02 to 1.34) (The hazard ratio is the ratio of the rate that something happens in the treatment vs. placebo groups. So, a number less than one suggests a benefit to taking vitamin D, while a number larger than one suggests a harm. The numbers in parentheses show a 95% confidence interval.) The only statistically significant result was a bad one: Extra kidney stones, likely from the extra calcium.13 The other outcomes look vaguely good, but none were statistically significant despite the massive sample size. This was disappointing. However, the WHI trial had limitations: Many subjects in both the vitamin D and placebo groups were already taking vitamin D, and continued taking it through the trial. The dose of 400 IU was fairly low, many subjects stopped taking their pills, and vitamin D levels didn’t actually change that much. They also measured vitamin D levels in only 6% of subjects, meaning we can’t compare the fates of subjects who started out with low versus high levels. The next big hope was VITAL, which came out in 2018. They recruited 26,000 older people across the United States, half of them men and 20% Black (and thus far more likely to be vitamin-D deficient). They measured vitamin D levels in most people, and they gave the treatment group 2,000 IU per day.14 Here were the results after 5.3 years: Outcome (VITAL trial) Hazard ratio Diabetes 0.91 (0.76 to 1.09) Autoimmune disease 0.78 (0.61 to 0.99) Cancer 0.96 (0.88 to 1.06) Cancer mortality 0.83 (0.67 to 1.02) Major CVD event 0.97 (0.85 to 1.12) CVD mortality 1.11 (0.88 to 1.40) All-cause mortality 0.99 (0.87 to 1.12) Some of the results look good-ish, but cardiovascular mortality was higher in the treatment group, leading to almost no effect on all-cause mortality.15 More disappointment. The last megatrial was D-Health, which came out in 2022 based on 21,000 older Australians. Instead of daily supplements, it used a monthly “bolus” dose of 60,000 IU or placebo. Unlike in VITAL, there was no exclusion for people with a history of cardiovascular disease or cancer, and less restriction on how much vitamin D participants could take on their own during the trial.16 Here were the results after 6 years: Outcome (D-Health trial) Hazard ratio Cancer mortality 1.15 (0.96 to 1.39) Major CVD event 0.91 (0.81 to 1.01) CVD mortality 0.96 (0.72 to 1.28) All-cause mortality 1.04 (0.93 to 1.18) Now, the treatment group did better in terms of cardiovascular disease, but worse in cancer and worse in all-cause mortality. Even more disappointment. Just from these three large trials, the main lesson should already be clear: Vitamin D is not a miracle. The correlations were wrong.17 There is essentially zero remaining hope that taking vitamin D could reduce all-cause mortality by a third. In this sense, the vitamin D skeptics are definitely right. But what about the other trials? And is there a more subtle lesson? I made some tables I wanted a big table that summarized all the major vitamin D RCTs and what they found for different health outcomes. Annoyingly, no such overview appears to exist. So I made my own:18 Trial Cancer Cancer mortality CVD CVD mortality All-cause mortality Lips 1996 0.92 (0.80 to 1.06) Trivedi 2003 1.08 (0.89 to 1.31) 0.86 (0.61 to 1.21) 0.95 (0.86 to 1.04) 0.86 (0.67 to 1.11) 0.90 (0.77 to 1.07) WHI 2006 0.98 (0.90 to 1.05) 0.89 (0.77 to 1.03) 0.94 (0.78 to 1.12) 0.92 (0.83 to 1.01) Lyons 2007 0.99 (0.93 to 1.05) WFPT 2007 1.00 (0.87 to 1.15) RECORD 2012 1.04 (0.91 to 1.19) 0.83 (0.55 to 1.26) 0.91 (0.79 to 1.05) 0.93 (0.85 to 1.02) Lappe 2017 0.70 (0.47 to 1.02) VITAL 2018 0.96 (0.88 to 1.06) 0.83 (0.67 to 1.02) 0.97 (0.85 to 1.12) 1.11 (0.88 to 1.40) 0.99 (0.87 to 1.12) ViDA 2018 1.01 (0.81 to 1.25) 0.99 (0.60 to 1.64) 1.02 (0.87 to 1.20) 1.12 (0.79 to 1.58) D2d 2019 1.07 (0.70 to 1.62) 0.23 (0.03 to 1.86) DO-HEALTH 2020 0.76 (0.49 to 1.18) 1.37 (0.88 to 2.14) D-Health 2022 1.15 (0.96 to 1.39) 0.91 (0.81 to 1.01) 0.96 (0.72 to 1.28) 1.04 (0.93 to 1.18) FIND 2022 1.04 (0.72 to 1.51) 1.14 (0.56 to 2.33) 0.90 (0.62 to 1.32) 0.85 (0.28 to 2.53) 0.81 (0.32 to 2.06) Lots of the hazard ratios are less than one, suggesting a benefit to supplementation. But lots of them are also higher than one, suggesting a harm. The numbers that are far from one almost always come from smaller trials, which manifest as larger confidence intervals. If you’re interested in the details of how these trials were run, I refer you to more gigantic tables in a footnote.19 If big tables aren’t your thing, here are some formal meta-analyses, both some recent ones and an older but more comprehensive Cochrane review: Outcome Meta analysis Hazard ratio Comment All-cause mortality Bjelakovic 2014 (Cochrane) 0.96 (0.92 to 0.99) Trials with low risk of bias. Cancer mortality Bjelakovic 2014 (Cochrane) 0.88 (0.78 to 0.98) Cardiovascular mortality Bjelakovic 2014 (Cochrane) 0.98 (0.90 to 1.07) Cancer mortality Kunzia 2023 0.94 (0.86 to 1.02) All-cause mortality Ruiz-García 2023 0.96 (0.91 to 1.00) Good-quality trials Cardiovascular mortality Ruiz-García 2023 1.00 (0.92 to 1.08) Good-quality trials All-cause mortality Cao 2023 0.99 (0.96 to 1.03) Squinting at the data There are various ways you could try to squint at these RCT. In almost all of them, most people already had pretty high levels before they started. So why don’t we separate out people who started low? Usually we can’t, because most trials didn’t measure baseline vitamin D.20 And among the trials that did, there are few people with low levels, so the results are noisy and confusing.21 Or, you might theorize that benefits would take time to show up, meaning the first couple years just add noise. In some cases—notably VITAL—excluding the first two years seems to help, but in other cases things get worse.22 Finally, some people speculate that taking gigantic monthly or quarterly “bolus” doses of vitamin D might be dangerous. For example, here’s an enjoyable paragraph from Kunzia et al. in their meta-analysis of vitamin D and cancer mortality: Our results showing efficacy of daily, but not bolus, vitamin D3 supplementation in reducing cancer mortality are consistent with previous meta-analyses on cancer mortality or all-cause mortality (Guo et al., 2022; Keum et al., 2022; Keum et al., 2019; Zhang et al., 2022; Zhang et al., 2019). However, by including more trials than these previous meta-analyses, we were able to detect statistically significant effect modification by treatment regimen for the first time with statistical significance (pinteraction=0.042). The pattern of intake could be important for a favourable steady state of the bioavailability of the active 1,25 (OH)₂D hormone. Daily administration counteracts the fast excretion of vitamin D from the circulation (Hollis and Wagner, 2013; Keum et al., 2022). Moreover, the enzymes CYP27B1 (converts 25(OH)D to 1,25 (OH)₂D) and CYP24A1 (inactivates 25(OH)D and 1,25(OH)₂D) follow first-order reaction kinetics (Vieth, 2009). This means that doubling the concentration of the precursor doubles the yield of the product, unlike other steroid hormones (e.g., cortisol, oestrogen, testosterone) that follow zero-order kinetics (Vieth, 2020). Intermittent, non-physiologically large vitamin D3 bolus doses may lead to unstable cycling of 25(OH)D and 1,25(OH)₂D levels in blood because the system needs time to adapt to the large doses (Hollis and Wagner, 2013; Keum et al., 2019; Vieth, 2020). In the long run, intermittent bolus regimens at weekly or larger intervals can lead to an up-regulation of countervailing factors (e.g., 24-hydroxylase (CYP24A1), 24,25(OH)2D and fibroblast growth factor 23), all of which ultimately leads to lower synthesis or higher degradation of 1,25(OH)₂D levels (Mazess et al., 2021). Bolus doses, unlike daily doses, failed to reduce C-reactive protein response and actually elevated anti-inflammatory cytokines and doubled the risk of hypercalcemia in previous studies (Krishnan et al., 2012; Martineau et al., 2017; Mazess et al., 2021). Oh no, up-regulation of fibroblast growth factor 23!23 I don’t feel like I understand this deeply enough to have any opinion beyond the surface level that the body seems to adapt to large doses of vitamin D in ways that could possibly be bad.24 It seems intuitive that small daily doses would be safer than gigantic monthly doses, but I’m always suspicious of post-hoc mechanistic speculation. Also, if people get enough sun, they can apparently synthesize 10,000-25,000 IU per day, which isn’t that far from the 60,000 IU they got in the D-Health trial. But then again, I think Kunzia et al. are suggesting that the body is designed to adapt to regular exposure to large doses but not intermittent exposure? Well, if you split up the trails by daily vs. bolus dosing, there’s a decent pattern of daily dosing leading to better results: Trial (daily dosing) Cancer mortality All-cause mortality Lips 1996 0.92 (0.80 to 1.06) WHI (Jackson 2006) 0.89 (0.77 to 1.03) 0.92 (0.83 to 1.01) WFPT (Smith) 2007 1.00 (0.87 to 1.15) RECORD (Avenell 2012) 0.83 (0.55 to 1.26) 0.93 (0.85 to 1.02) VITAL (Manson 2018) 0.83 (0.67 to 1.02) 0.99 (0.87 to 1.12) D2d (Pittas 2019) 0.23 (0.03 to 1.86) FIND (Virtanen 2022) 1.14 (0.56 to 2.33) 0.81 (0.32 to 2.06) Trial (bolus dosing) Cancer mortality All-cause mortality Trivedi 2003 0.86 (0.61 to 1.21) 0.90 (0.77 to 1.07) Lyons 2007 0.99 (0.93 to 1.05) ViDA (Scragg 2018) 0.99 (0.60 to 1.64) 1.12 (0.79 to 1.58) D-Health (Neale 2022) 1.15 (0.96 to 1.39) 1.04 (0.93 to 1.18) If those bolus dosing trials didn’t exist, I’d think this looked pretty good. So, maybe? Or maybe this is a story made up to hallucinate a positive trend. I would lean towards the latter theory, but there are papers like Mazess et al.’s “Vitamin D: Bolus is Bogus”, that suggested this pattern before D-Health’s dismal results came out. There are even some trials that suggest bolus doses don’t even work for treating rickets. So… I’m still not convinced. But maybe. Aside: There are also many Mendelian randomization studies that look at correlations between health and genes that are related to vitamin D. But I don’t think these provide much information, because the assumptions are shaky and the genes don’t explain much of the variance.25 Where are we? Still with me? Here’s a summary of the above 5200 words: The body uses vitamin D in all sorts of weird and complicated ways. It’s biologically plausible that vitamin D could matter beyond bone stuff with severe deficiency, but there’s no convincing mechanistic evidence that it is. Vitamin D levels are strongly correlated with good health outcomes, but RCTs have conclusively shown that most of these correlations are non-causal. RCTs haven’t conclusively shown any benefit for anything beyond beyond bone stuff. At best, they’ve given weak evidence for hazard ratios slightly below one. So you might be wondering: Isn’t that quite weak? Wasn’t this post supposed to be a defense of vitamin D? The case for supplementing anyway It’s biologically plausible that vitamin D is good Everyone agrees that severe vitamin D deficiency (below ~25 nmol/L) is bad. It leads to rickets, adult rickets, osteoporosis, muscle weakness or even—with profound deficiency—to seizures or cardiac arrhythmia. This makes sense, because below ~25 nmol/L, the kidneys have trouble converting storage vitamin D into active vitamin D, meaning you don’t absorb enough calcium from food. The question is if taking supplement to further raise your levels (say to 50 or 90 nmol/L) is important. We have no mechanistic proof, but it might be true, because many parts of the body use vitamin D as a local signal and because cells are at least somewhat sensitive to circulating storage levels. There’s also this weird thing where parathyroid hormone continues to decline as vitamin D levels rise above ~25 nmol/L even while this seems to make little difference to how much active vitamin D the kidneys make. Nothing in this world comes without trade-offs. Surely, supplementing vitamin D comes with some downsides. But it seems very unlikely that raising vitamin D levels to a “normal” level would cause more harm than benefit. Especially because… Humans evolved to have a lot of vitamin D According to Luxwolda et al.’s 2012 paper, “Traditionally living populations in East Africa have a mean serum 25-hydroxyvitamin D concentration of 115 nmol/L”, traditionally living populations in East Africa have a mean serum 25-hydroxyvitamin D concentration of 115 nmol/L. Meanwhile, Wahl et al. 2012 try to estimate mean levels around the world today: This map looks weird because of varying lifestyle, diet, supplementation, and needing to combine fragmented studies. But you get the idea. And remember, those are just averages. So there are lots of people with levels far lower than that in our evolutionary history. Of course, just the fact that vitamin D levels have dropped doesn’t mean it’s important. Parasitic worm load, wood smoke inhalation, and cousin marriage have also dropped, but we aren’t rushing to restore those to ancestral levels. But there’s another piece of evidence: After humans migrated out of East Africa, some of them evolved pale skin. Pale skin is bad, because it allows light to destroy folate, which is crucial for pregnancy.26 Evolution doesn’t typically do things that harm fertility, because evolution wants to increase reproductive fitness. The most common explanation is that pale skin allows more UV light to penetrate, and thus allows people to synthesize more vitamin D. If evolution was willing to pay the high “price” of folate destruction for more vitamin D, that seems like good evidence that vitamin D is important. Some even see contrasts like the Inuits versus Scandinavians as a kind of natural experiment: They lived at similar latitudes, but Inuits ate a diet with vitamin D (fatty fish and whale blubber) and Scandinavians didn’t. The result is that Inuits have darker skin than Scandinavians.27 This is all speculative, and even if true, might be driven by severe deficiency and rickets. Or perhaps prehistoric benefits don’t translate to your lifestyle. But all the people in Luxwolda’s sample in East Africa had levels above ~60 nmol/L. I just don’t see how you can look at this and not see it as providing some suggestive evidence in favor of the idea that raising levels above severe deficiency is unlikely to be harmful, and could be important. So I think the prior is favorable. What do you expect from vitamin D? A hazard ratio like HR = 0.96 doesn’t look very impressive. But hold on. Suppose that life expectancy is 80 years and that taking vitamin D every day reduces your risk of all-cause mortality by a factor of HR. A reasonable approximation in rich countries is that this would increase your life expectancy by 80 × 0.15 × (1-HR) years = 12 × (1-HR) years, where 0.15 is derived from the entropy of lifespan in rich countries.28 For example, if all-cause mortality had a true hazard ratio of HR = 0.96, then taking vitamin D every day of your life would increase life expectancy by around 0.48 years. I claim that this would be a lot. Certainly, if I were about to face my destiny, I would pay a lot of money for an extra 0.48 years. Or, you can calculate that this corresponds to an increase of life expectancy per-vitamin-D-pill of 8.6 minutes.29 A common rule-of-thumb is that smoking a cigarette costs around 11 minutes of life in expectation. If you think HR = 0.96 is trivial, do you also think that smoking one cigarette each day is fine?30 The correlational studies suggested that vitamin D might drop your risk of all-cause mortality by a third. It’s disappointing that the RCTs refuted this. But those correlational studies were crazy. They imply31 an increase of life expectancy of around 4 years or around 6.5 cigarettes per day. Could we really believe that you could smoke 6.5 cigarettes, then take a vitamin D pill, and you’re even? Personally, I think hazard ratios just slightly less than one are the best we can reasonably hope for. But I also think that they would be an excellent return on investment. Arguably, modern human life expectancy comes from stacking lots of modest hazard ratios on top of each other. What do you expect from vitamin D trials? Let’s play a game. Let’s hallucinate some numbers for what vitamin D might do, and then simulate what trials would show. Here are the strongest effects I consider plausible for different baseline levels, along with how common those levels are in the United States. Storage vitamin D (nmol/L) Hazard ratio % of population <30 0.75 5 30-49 0.92 15 50-125 0.98 72.5 >125 1 7.5 Suppose that were real. Now, say we pick 26,000 people at random, and give half of them vitamin D for give yars. Here are the results of a million simulated trials, assuming a baseline mortality risk of 0.7%: 32 Overall, 9% of trials would find a significant benefit, 63% would find a non-significant benefit, 27% would find a non-significant harm, and 1% would find a significant harm. If you wanted to have an 80% chance of finding a significant decrease, you’d need to run a trial with something like 570,000 people, almost five times more than in all the above trials combined.33 If you don’t like my numbers, I’ve put up a page where you can run your own simulations with different ones. My point is, the results we see in vitamin D RCTs are what we should expect to see if vitamin D had plausible benefits. That’s not proof, of course—just that if you start with realistic expectations, the trials don’t provide much evidence in either direction. The trials do find slightly helpful numbers Recent meta-analyses have not consistently found a statistically significant benefit to vitamin D supplementation. But they do suggest a small benefit for cancer mortality and all-cause mortality, and they’re close to being statistically significant. That’s something. And if you buy the argument that bolus dosing is bad, the results get even better. Kunzia et al. did a meta-analysis of cancer mortality using only trials with daily dosing, and found a hazard ratio of 0.88 (confidence interval 0.78 to 0.98). I’d keep this at arm’s length. The bolus dosing trials might have done worse by random chance, meaning this a kind of p-hacking. But there’s a reasonable chance (maybe 25-50%) that bolus dosing really is bad, in which case those trials would be convincing evidence. I actually think it’s surprising that the meta-analyses look as good as they do, because there just aren’t that many people who started out with low vitamin D levels. Only a handful of trials had mean levels below 60 nmol/L, and they all give semi-promising results:34 Trial (low-ish baseline) Cancer mortality All-cause mortality Trivedi 2003 0.86 (0.61 to 1.21) 0.90 (0.77 to 1.07) WHI (Jackson 2006) 0.89 (0.77 to 1.03) 0.92 (0.83 to 1.01) Lyons 2007 0.99 (0.93 to 1.05) RECORD (Avenell 2012) 0.83 (0.55 to 1.26) 0.93 (0.85 to 1.02) Again, it’s dangerous to dig too deeply looking for these kinds of patterns. If you dig enough, you can always find a way to confirm whatever theory you want. But also again, maybe? You’re probably already taking vitamin D You might not personally supplement vitamin D. But for most people reading this, someone else is supplementing it for you.35 Country Commonly fortified with vitamin D Australia Margarine Belgium Margarine Canada Milk, margarine Chile Milk, flour Ethiopia Oils Finland Milk, yogurt, margarine Ireland Margarine, cereal New Zealand Margarine (from Australia) Norway Margarine, low-fat milk Pakistan Oils Poland Margarine Sweden Milk, yogurt, plant milk, margarine United Kingdom Margarine, cereal United States Milk, plant milk, margarine, cereal, yogurt Fortified food is common across the Anglosphere and Scandinavian peninsula. However, it’s rare in the rest of Europe (exceptions: Belgium, Poland) and even-more rare in the rest of the world (exceptions: Chile, Ethiopia, Pakistan). I think this is important for two reasons. First, vitamin D is oddly self-defeating. There are some places in the world where people care about vitamin D. These are the places that run large trials. But these places also fortify their food and tend to be full of people that already supplement vitamin D. These places also tend to believe it’s unethical to tell the control group not to take vitamin D. And here’s another question: If you think vitamin D is worthless, are you comfortable recommending removing vitamin D from food? If not, then why is the particular amount of fortification in food now the right one? Some might argue that the purpose of fortification is to reach the severely deficient, or children, the elderly or pregnant mothers. Maybe! But again, if you could press a button and remove fortification from everyone else, would you feel comfortable pushing that button? Remember, trials don’t test don’t test going down from current levels, only going up. So that’s my story Biology and evolution suggest a prior that moderate levels of vitamin D (say 80 nmol/L) are quite possibly better than low levels (like 40 nmol/L) and unlikely to be worse. Observational studies say that vitamin D is magical, but those studies are bad and we should ignore them. The RCTs show that vitamin D is non-miraculous. But beyond that they don’t provide much information, because they mostly enrolled people with moderate vitamin D levels, meaning plausible effects would require colossal sample sizes to reliably detect. What evidence the RCTs do provide points weakly towards a modest benefit. If real, that benefit would far exceed the cost of taking vitamin D. Therefore, if you have low vitamin D, it seems wise to supplement. This is all very weak, I know! But sometimes weak evidence is all we’ve got. I wish we had at least one large trial done in a population with low starting levels. But as far as I can tell, none are underway. In fact, it’s unlikely that there will be any more large trials anytime soon. So weak evidence is how it’s going to be. Technically, vitamin D itself is a type of steroid although not what people usually mean by “steroid”. ↩ Here are some of the fancy names for the different forms of vitamin D I’ll talk about: my name fancy names provitamin D 7-dehydrocholesterol previtamin D previtamin D₃ vitamin D cholecalciferol storage vitamin D calcifediol / ergocalciferol / 25(OH)D / 25-hydroxyvitamin D active vitamin D calcitriol / ercalcitriol / 1,25(OH)₂D / 1,25-dihydroxyvitamin D ↩ Charmingly named “vitamin D-binding protein”. ↩ If you eat mushrooms or yeast, it joins the vitamin D from your skin en route to your liver. If you eat animals or animal products, you also get some storage vitamin D, which doesn’t need to be processed by the liver. ↩ Storage vitamin D is what your doctor measures in your blood test. This is sometimes measured in nmol/L and sometimes in ng/mL. The latter measurement is smaller by a factor of 2.496. So 25 nmol/L ≈ 10 ng/mL. ↩ Apperly was building on a 1937 paper that observed observed that sailors, exposed to lots of sunlight, had much higher skin cancer rates than the general population, but lower overall cancer rates. ↩ I theorize that the Garland brothers are alive and writing Slime Mold Time Mold. ↩ In Biologist, active vitamin D is not just an “endocrine” hormone that sends signals for far away cells through the blood, it’s also a “paracrine” or “autocrine” hormone that sends signals to nearby cells or inside a single cell, through diffusion. ↩ You might ask, why is vitamin D used by so many different parts of the body for so many different purposes? I think there’s no deep answer here. It’s true for the same reason that dogs sneeze to signal that they’re feeling playful: Evolution re-uses stuff for different purposes all the time. Imagine that DNA already exists coding for the vitamin D receptor and for the enzyme to convert storage vitamin D into active vitamin D. If some cells need to send a local signal, re-using those is easier than inventing something new. There’s nothing unusual or magical about this. ↩ Don’t try to make sense of this. It doesn’t make sense. You could speculate that this is because the parathyroid glands are trying to make less active vitamin D to compensate for the fact that vitamin-D receptors throughout the body are sensitive to storage vitamin D itself. But I advise against. ↩ 400 IU is the recommended daily amount ↩ The WHI trial was a pioneer in salami-slicing results for different outcomes into dozens of different papers, most of which are hard to access. All trials now seem to have adopted this hideous trend which makes it maddening to try to summarize what actually happened in a trial. Also, slightly different numbers for the same quantity appear in different places. I haven’t bothered to chase these down, because the differences are all very small, e.g. a hazard ratio of 0.89 for cancer mortality rather than 0.90. ↩ Guess what most kidney stones are made of? ↩ Half of the vitamin D group and the placebo group also got omega 3. These are averaged together in the results. Also, VITAL carefully stratified the assignment to vitamin D or placebo based on baseline vitamin D levels, which should give more statistical power from a given sample size. ↩ There was also a weird study done on a subset of 1031 people from the VITAL population that looked at telomere length. After starting with around 8700 base pairs, the control group lost around 160 base pairs during the study, while the vitamin D group only lost an average of 20. I’m not sure of what to make of this. For one thing, though the authors claim this is statistically significant, it depends on how you analyze the data. But beyond that, sure, telomere length is a marker of aging, but telomeres get shorter for a reason (likely to fight cancer) and it isn’t obvious that slowing this would always be a good thing. ↩ This is a little complicated. In VITAL, participants were only eligible if they were taking at most 800 IU per day, and they were restricted to 800 IU per day during the trial. In D-health, participants were only eligible if they were taking at most 500 IU per day, but they were allowed to take up to 2000 IU per day during the trial. ↩ You might ask: If vitamin D only has a modest effect, then why is it so strongly correlated with health? In principle, I’d like to push back against the idea that we need to explain why these particular correlations don’t imply causation. But the accepted explanation is a combination of (1) reverse causation where being healthy causes people to spend more time outside and thus get more vitamin D; (2) confounding, where obesity is bad for you and leads to lower measured vitamin D levels; (3) confounding, where more healthy lifestyles lead to both more vitamin D and more health; and (4) confounding, where higher socioeconomic status leads to both more vitamin D and more health. You might ask why these correlations would be true at a state level like the Garlands looked at, but then you run into the ecological fallacy and modifiable areal unit problem. ↩ I took all the trials that got at least 2% weight and were rated as “low risk of bias” in this 2014 Cochrane review of vitamin D and mortality, then manually added all the “major” trials that were published after 2014. I shudder to think of the time it took to make this table. I tried using AI but found it was wildly unreliable. Part of the problem is that each trial’s results are distributed among many papers, in different journals, with different paywalls. And many details aren’t published at all by the original authors but are only scrounged up and put in the depths of the supplementary material of a review years later. In some cases, different sources also give contradictory numbers. The differences were always tiny (e.g. 0.90 rather than 0.89) but it still makes me nervous. ↩ Here’s a table describing the major contours of the trials: Name Country Subjects (n) Age (years) white (%) Lips 1996 Netherlands 2,578 80 ± 6 Trivedi 2003 UK 2,686 74.7 ± 4.6 74 WHI 2006 USA 36,282 (women) 61.8 ± 6.7 84 Lyons 2007 Wales 3,440 84 ± 7.5 WFPT 2007 UK 9,440 79.1 RECORD 2012 UK 5,292 77.5 ± 6 99.2 Lappe 2017 USA 2303 (women) 65.2 ± 7.0 100 VITAL 2018 USA 25,871 67.1 ± 7.1 71.3 ViDA 2018 New Zealand 5,110 65.9 ± 8.3 83.3 D2d 2019 USA 2,423 60.0 ± 9.9 67 DO-HEALTH 2020 Several in Europe 2,157 74.9 ± 4.1 D-Health 2022 Australia 21,315 69.3 ± 5.5 94.7 FIND 2022 Finland 2,495 68.2 ± 4.5 100 And here’s a table focusing on the change in vitamin D levels: Name Intervention (IU) Allowed personal use (IU/day) Duration (years) Baseline D (nmol/L) Lips 1996 400 IU daily 0 (screening) 3.5 Trivedi 2003 100,000 3× per year (D2) 0 (screening) 200 (trial) 5 52.5 (in controls) WHI 2006 400 daily with Ca 600 (later 1000) 7 52.0 ± 21.1 (subset) Lyons 2007 100,000 3× per year <400 (screening) 3 54.0 (in controls, subset) WFPT 2007 300,000 IU yearly <400 (screening) 3 RECORD 2012 800 daily with Ca 200 6.2 ~38 Lappe 2017 2000 daily with Ca any? 4 71.8 ± 20.0 VITAL 2018 2,000 daily 800 5.3 77 ± 30 ViDA 2018 100,000 monthly 600 or 800 3.3 63 ± 24 D2d 2019 4,000 daily 1000 2.7 69.9 ± 26.8 DO-HEALTH 2020 2,000 daily 1000 (screening) 800 (trial) 3 55 ± 22 D-Health 2022 60,000 monthly 500 (screening) 2000 (trial) 5 77 ± 25 (predicted) FIND 2022 1,600 or 3,200 daily 800 5 75 ± 18 ↩ Among the major trials, only VITAL, ViDA, and FIND measured it for more than a tiny number of subjects. ↩ In VITAL and ViDA, people with baseline levels below 50 nmol/L had a higher hazard ratio for cancer mortality (though with wide confidence intervals), suggesting if anything less benefit. Or, you could use race as a proxy for baseline vitamin D. But in both VITAL and WHI, the hazard ratio for cancer mortality was higher among non-Whites. After looking at many such analyses for many outcomes, the only clear result I could find was for diabetes in the D2d trail, where the hazard ratio was much lower for people below 30 nmol/L (0.38 vs. 0.93). ↩ The results for VITAL look decent: outcome (VITAL trial) HR HR excluding first two years Cancer 0.96 (0.88 to 1.06) 0.94 (0.83 to 1.06) Cancer mortality 0.83 (0.67 to 1.02) 0.75 (0.59 to 0.96) Major CVD event 0.97 (0.85 to 1.12) 0.93 (0.79 to 1.09) All-cause mortality 0.99 (0.87 to 1.12) 0.96 (0.84 to 1.11) But in D-Health, excluding the first two years actually increased the hazard ratio for cancer mortality from 1.15 (0.96 to 1.39) to 1.24 (1.01 to 1.54). Most other trials were too short for this kind of analysis to make sense. ↩ That could downregulate 25-hydroxyvitamin D 1-alpha-hydroxylase, reducing the rate it catalyzes the hydroxylation of hydroxycholecalciferol into 1,25-dihydroxycholecalciferol! ↩ Dynomight: WTF is this? Dynomight Biologist: Well, C-reactive protein is generally considered inflammatory. Dynomight: So reducing that is good? But then why do they talk like elevating anti-inflammatory cytokines would be bad? Dynomight Biologist: Yeah… That would be good. Unless you have cancer. In which case it’s not good. Dynomight: OK! ↩ Mendelian randomization studies are based on the idea that certain genes predispose you to have higher levels of circulating vitamin D. If you assume that those genes are randomly distributed in the population and have no effects other than affecting vitamin D, then they serve as a kind of natural experiment. With vitamin D, these studies typically show null results. However, the validity of the assumptions is debatable and the identified genes only explain ~5% of the variance in vitamin D levels, which makes the results very noisy. ↩ Pale skin also greatly increases the risk of sunburn and skin cancer. In the US, White people get melanoma at around 25 times the rate of Black people, despite (I assume) higher usage of sunscreen and better health outcomes in most other dimensions. But experts generally think folate deficiency created stronger selective pressure, since it’s so closely linked to reproduction. ↩ It’s a more complicated than this, because you also need to look at the amount of folate in diet, as well as migration patterns and how long populations had to adapt to their environment. But experts seem to consider this the leading explanation for the evolution of pale skin. ↩ To derive this, suppose that S(t) is the probability that someone survives to age t. Then life expectancy is ∫ S(t) dt, where the integral runs from 0 to ∞. If you change the hazard ratio by a factor of HR, then the new in life expectancy is L(HR) = ∫ S(t)ᴴᴿ dt, so the change under a linear approximation is ΔL ≈ (HR-1) × L’(1). This is more commonly written as ΔL ≈ (HR-1) × L(1) × H, where H = -L’(1)/L(1) is known as the Keyfitz entropy. This is is chosen because the quantity H is relatively stable, and in rich countries is typically between 0.10 and 0.20. So a decent estimate would be that baseline life expectancy is L(1)=80 years and H = 0.15 in which case the change in life expectancy is around 12 × (1-HR) years. ↩ Observe that 0.48 years is 252460.8 minutes. Assuming you lived for 80 years and took a pill every day of your life, that would be 80 * 365.25 = 29220 pills. 252460.8 minutes / 29220 pills = 8.64 minutes/pill. ↩ I expect that a number of you are happy to bite that bullet and say yes, HR=0.96 is trivial and smoking a cigarette each day is also fine. I don’t personally agree, but it’s not my place to question your utility function and I applaud your consistency. ↩ A hazard ratio of HR=2/3, implies a change in life expectancy of 12 × (1 - 1/3) years = 4 years or 2,103,840 minutes. That corresponds to a per-pill increase of 2,103,840 minutes / 29,220 pills = 72 minutes/pill. ↩ Technically, this is calculating a relative risk rather than a hazard ratio, but I think the difference isn’t very significant given that we’re assuming a uniform mortality risk. I used AI to create that simulation, though I did test that it replicates a traditional power calculator across a wide range of parameters when the relative risk is constant for all vitamin D levels. So I mostly trust it. ↩ This simulation is probably a bit pessimistic. Things look a bit better if you use an older population where baseline mortality is higher. (Almost all trials do.) In principle, you could also use a population where more people have low levels, which could help a lot. But, for whatever reason, almost no trials do that. In fact, most trials accidentally under-sample people with low vitamin D, because people who agree to participate tend to be more health-conscious. ↩ Kunzia et al. made a heroic effort to contact study authors and get data for individual patients. After getting data for 21,558 people (almost all from ViDA + FIND + VITAL + WHI) only 3,663 had levels below 50 nmol/L. That’s not enough to reliably detect a modest effect, meaning their confidence interval for this group is gigantic. ↩ In this table, I tried to capture foods that are commonly fortified in practice, not just when it’s legally required. ↩
News from the world of real jobs: Apparently, sometime between 10 and 20 years ago, it became standard for people to communicate by sending slide decks around. These slides are never presented. They aren’t intended to be presented. They’re born, they’re sent around, and they die. What? I stress, the question is not why (or if) people give bad presentations. The mystery is why everyone is using presentation software for communication that is not a presentation. Theory 1: Everybody dumb Is it because we’re all dummies? I’m putting this theory first because I suspect that you, beloved readers, will favor it. True, if you ask people why they make slides instead of writing, they’ll usually say, “because nobody wants to read”. So there’s that. But I don’t consider this much of an explanation. Dummies though we may be, we’ve been like that a long time. If we entered the Slideocene 15 years ago, why then? Why not before? Theory 2: The decline of reading Did we get worse at reading? The Discourse seems to have decided this is true, but is it true, or just moral panic? Since 1971, the US has tested 13-year-olds to measure long-term trends in reading ability. This shows a slow improvement until 2012, then a slow decline, and finally a post-COVID drop. The declines seem too small and too late to explain our mystery. Since 2000, PISA has tested reading performance in 15-year-olds around the world. This shows a decline on average, but it’s smaller in rich countries and nonexistent in the United States. (It’s the same story for science and a bit more negative for math.) Among adults, data is scarce. Basic literacy is generally improving, and American time use data shows a decline in reading for pleasure from around 23 minutes per day in 2003 to around 16 minutes per day in 2023. But this seems to miss time people spend reading on their phones. So it’s unclear if people got worse at reading. It feels plausible that people now spend less of their adulthood grappling with complex written arguments, and so got worse at that. But there’s little firm evidence. Theory 3: Technological change Another obvious theory is that we now have computers and software and the internet. Without these things, it would be impossible to email slides to each other. This seems relevant! Yes, but we had those things for a while before slide culture really took hold. And think about the situation before computers. Photocopiers were ubiquitous in corporate offices by the mid-1980s, and mimeographs were around decades before that. If slides were really that great, people could have made them by hand. But no one did. Of course, making slides by hand is inferior. But it’s not that inferior. So slides can’t be that big of a win. What actually happened? And… that’s pretty much the end of the obvious theories. None of them are very satisfying. So let’s take a step back. Historically, how did the slide-as-document displace the memo? As best I can tell, this was driven by management consultancies. If you go back to 1960, they delivered detailed written memos. The memo was the product. They’d likely give a presentation as well, but that was a separate ancillary thing, likely done using flipcharts or chalkboards. In the 1970s, the memo was still the product, but consultancies started to enforce a top-down logical structure (the Pyramid principle). Presentations shifted to acetate transparencies. Both memos and presentations often included hand-drawn graphics like the nine-box or growth-share matrices. In the 1980s, the memo was still the product, but presentations became increasingly lengthy and polished. Expensive computers like the Genigraphics started to be used to generate charts. The 1990s were when things started to shift. By then, PowerPoint was everywhere, and junior analysts were expected to create presentations themselves. Consultancies gradually started to notice that (1) clients didn’t always read the memos; (2) clients loved slides and passed them around long after the presentation was over; and (3) creating a memo and a polished presentation was a lot of work. They put more and more effort into the slides. McKinsey especially evolved towards treating slides as the primary product, and mostly stopped writing long memos. Other consultancies followed. During the 2000s, slides became even more ornate. Consultancies evolved their formatting rules, and created fancy data-dense charts. They learned that a 200 slide deck made clients feel like they got a lot for their money. Gradually, they oriented their entire business around slides. Projects would start with managers creating a template presentation with “ghost slides” and assigning different parts to junior analysts. Soon, this spread outwards, both from people who interacted with consultants and from the ex-consultant diaspora. People everywhere started thinking and communicating in slides, and now everything is slides, yay! Alternative history That story makes slides-as-documents sound inevitable: People liked them, so they became popular. But there’s an alternative timeline in which we resisted the slide into slide maximalism. That timeline is Amazon.com, Inc. In 2004, Jeff Bezos famously instituted a no-presentations policy at Amazon. His logic was that slides hide poor reasoning and are a tool to persuade rather than inform. Instead, everyone involved with strategic decisions at Amazon needs to learn to write a six-page memo. Meetings begin with everyone sitting and silently reading one of these memos. Presentation software is not banned at Amazon. The ban is only for using it for internal meetings and decision-making. They use slides for external communication. There is no policy that prohibits someone from making slides and emailing them around. And yet, people don’t make slides and email them around, because it’s not part of Amazon’s culture. In effect, Amazon is a counter-movement. Most of the world decided that slides are good, because slides are easy. Bezos decided that writing is good because writing is hard. There are millions of articles explaining why Bezos’ policy is pure genius. They claim that constructing a narrative requires deeper analytical thinking and exposes flaws in logic. I want to believe those theories. I now realize they’re very similar to some of my arguments for why writing with too much formatting is bad. I’m not sure if writing is the secret to Amazon’s success. But Amazon is successful. This demonstrates that slide life is a choice, not technological destiny—institutions can choose writing over slides and flourish anyway. OK so then what’s happening? Warning: If you like your theories simple and mono-causal, you aren’t going to like this. Slides are a win, but a small one. The shift to slides wasn’t a “mistake”, it happened because people like it. But if sharing slides outside of presentations became illegal, this wouldn’t cause per-capita GDP to crash. That’s why people didn’t scratch slides into mimeograph stencils back in the 1950s. It wasn’t worth the modest effort. When computers and software showed up, it became easier to share slides. But people didn’t immediately shift to slides-as-documents because the win isn’t that big, because culture changes slowly, and because everyone had pre-existing skills for reading and writing documents. Consultancies happened to be in the economic niche with the strongest selection pressure to evolve towards slides-as-documents. So when making slides became cheaper, they shifted. Slowly, that norm spread outwards, people got used to communicating in slides, and here we are. Institutions can resist that norm and still be successful. If you take modern people and force them to read and write, they do just fine. Humans evolved to learn and communicate in a fragmented, interactive, and visual style. It’s hard to argue that any shift in that direction is a catastrophe. Except blogs. The decline of the blog must be arrested.
In early January 2025, a family friend was over for lunch. One of my many guilty midwit pleasures is a love of New Year’s resolutions, so I asked her if she had made any. She said no, but mentioned that she had some relatives that were doing “damp January”. In case you’re not aware, Dry January is a challenge many people do to quit drinking alcohol during the month of January. These folks were doing a variant in which, instead of not drinking, one simply drinks less. For some reason, this triggered me. I thought, “Are you kidding? You can’t even stop drinking for a single month? Do you know how pathetic that is?” And then, “Fuck you! Fuck you for doing damp January! You know what, I’m going to stop drinking for a year!” To be clear, these thoughts were directed at people I’ve never even met. In retrospect, I wonder what was going on with me emotionally. But I take resolutions seriously, so I felt committed. We are now 15 months down the timeline, so I’ll make my report. It was easy This will sound odd, but I swear it’s true. Not drinking was so easy that it was almost easier than my previous baseline of not-not-drinking. Before starting this resolution, I didn’t drink much—perhaps two or three drinks per week. But I often thought about drinking. Every time I saw friends or went to a restaurant, I thought, “Should I have a drink?” Usually I decided not to. But making that decision required effort. After a few weeks of not drinking, that question never even came up. Drinking was simply not a thing I did, so I never needed to negotiate with myself. Theoretically, you could allow yourself one drink a month instead of zero. Theoretically, that should be easier. But I’m pretty sure I’d find it harder, because alcohol would still be an option, a thing to consider. Sometimes I need a thing Early on, I sometimes wanted a drink. But gradually I noticed that I didn’t really want a drink, I just wanted a thing. I can’t find a precise name for this concept in psychology, but often, some deep part of my brain seems to scream, “I WANT A THING.” It could be alcohol, but I found dessert worked just as well. I suspect that a new shirt or meeting a new dog would also work. I was not able to stop my brain from doing this. When it demanded a thing, I gave it a thing. I just substituted a non-alcohol thing. So, over the year, I became interested in desserts and even-more interested in tea. The struggle was The Chocolates. Shortly after I made this resolution, my mother gave me a bag of chocolates that each contained a bit of whiskey. In general, I don’t keep chocolate at home. If anyone gives me chocolate, I immediately eat all of it and then text the giver, “Thanks for the chocolate, I ate it instead of dinner, it’s all gone, this is what will always happen if you give me chocolate.” But I couldn’t eat the Chocolates, because they contained alcohol. I managed to get guests to eat a few. A couple of times I came close to draining out the alcohol and eating the chocolate container. I even considered throwing them away, but that felt wrong. So instead I spent a year glaring at them and waiting for them to apologize for the anguish they were causing me. This represented half the difficulty of this resolution. I do not recommend it. Keep your things separate. Alcohol is bad for sleep Have you heard that alcohol is bad for sleep? Because alcohol is bad for sleep. I’ve always known that was true, abstractly. But sleep is variable. If I didn’t sleep well on an individual night, I was never sure: Was that because of the alcohol, or was it random variation? After a year without alcohol, I am very confident that yes indeed, alcohol is bad for sleep, because my sleep during 2025 was much better than in previous years. Sure, like anyone else, I still sometimes wake up and start thinking about oblivion rushing towards me, and how everything I love will vanish into time, and how all that was once future and hope inevitably becomes static and dust, and how the plague of bluetooth speakers continues to spread across the globe. But now: less! I wish there was a drug I could take that would give me energy and improve my mood and make me physically healthier and smarter, all without side-effects. I don’t think such a drug exists. But we do have the opposite! So, sadly, I’ve come to believe that alcohol is basically the perfect anti-nootropic. That’s not because it makes you dumb while you’re drunk. (True, but who cares?) Rather, that’s because it is bad for sleep, and therefore makes you worse across all dimensions the next day. Alcohol is good for socializing sometimes I did find not drinking to have one clear downside: It’s just not that much fun to hang out with people who are drinking if you are not drinking yourself. To be clear, this is a limited effect. It’s only an issue at bars or certain parties where people are there to drink. I don’t go to many such gatherings, but when I did, I felt it was less fun. It’s not that I missed alcohol. Instead, my theory is that drinking parties are a sort of joint role-playing exercise: “Let’s all get together and collectively reduce our inhibitions and see what happens.” It’s fun not (just) because everyone is taking a recreational drug, but because it’s a joint social experience. If you don’t drink, then you aren’t fully participating. It seems like it should be possible to reproduce this effect without alcohol. You could imagine other ways to push the social equilibrium out of balance. Like… Masks? Or weird environments? Or mutual disclosure games? Should people get together and do a group cold plunge? Unfortunately, all these are complicated and/or carry some kind of social stigma. So until we figure something better out, this is a real cost. It was minor for me, but it probably depends a lot on where you are in life. Other effects All other effects were minor. I guess I saved money at restaurants. I actually lost a bit of weight over the year, despite all the extra desserts, though I can’t say for sure if alcohol was the cause. Otherwise, once I stopped thinking of alcohol as an option, I rarely thought about the resolution at all, except when I saw those damn chocolates. Aftermath Towards the end of the year, I started wondering if I should quit drinking forever. But I never came to a conclusion, because I rarely thought about alcohol. I considered having a drink at midnight on New Year’s eve, but I happened to be on a plane that crossed the international date line, and thus skipped New Year’s eve. And then… for the first few months of 2026, I still didn’t drink. That wasn’t because of any decision. It just never seemed appealing because (a) sleep and (b) I’d broken the mental link between want thing and drink alcohol. Eventually, I ate the chocolates, and I had a glass of wine when visiting some friends. If I can continue rarely drinking while almost never thinking about drinking, I’ll probably do that. If I slowly slide back into always thinking of alcohol as a live option and always negotiating with myself, I might just resolve to quit forever. So that’s my story. Obviously, it’s heavily colored by my own idiosyncrasies, so it’s hard to say if it offers any general lesson. I do think people underrate the long-term health impact of drinking. The effect on heart disease is debated, but everyone agrees that any alcohol increases the risk of cancer. Still, the long-term effects from occasional light drinking probably aren’t huge. What’s really underrated is the short-term effects, via worse sleep. If I had to give advice, it would be this: If you drink, and you think you might be better off not drinking, why not try it? Maybe you’ll find that champagne is essential to your happiness and drink it every night, to hell with the costs. Maybe you’ll find a different baseline, or maybe you’ll quit forever. Whatever you decide, you’ll have full information.
More in life
Laetitia@Work #104
Trustocracy sounds funny, but it's a more accurate representation of how we're measured. We aren't (and can't) be measured objectively on what we accomplish at work. But we are (unavoidably) measured by the trust we've built with our co-workers.
When you’re young, making a bucket list – things you want to do before you die — feels like you’re choosing prizes that will arrive in the mail later. I’ll take fluency in Italian, please, a 300-lb bench press, and a swim in the Nile. I’ll definitely want to drive a Ferrari on the Autobahn at some point. How about […]
The Email Exchange That Radicalised Me