More from Rachel Thomas, PhD
As a former mathematician, I was used to nobody reading what I wrote. So when I first began blogging in 2015, I never expected that several of my blog posts would go viral or to have multiple journalists contact me (including from NPR, Wired, and Fortune), make the front page of Hacker News (over 10 times), receive conference keynote invitations, and be interviewed on podcasts. I do not consider myself a “natural” writer. In college, I tried to avoid classes that required essays, because writing was a struggle for me. It wasn’t until I was 30 that I set out to practice writing more. I share tips I use for blogging here, which include being willing to put a lot of time into a single post, incorporating high quality information, and having a clear idea of my intended audience. I have selected some of my most popular and impactful posts below. Several of these were originally posted on Medium or fast.ai (the two sites where my writing used to live). They are grouped into clusters based on theme. I hope you might enjoy reading these if you haven’t seen them before! Challenging Conventional Wisdom Questioning widely-held assumptions about tech culture, education, and health has been the basis for several popular posts. If you think women in tech is just a pipeline problem, you haven’t been paying attention (2015) In 2015, I felt burnt out and disillusioned by my experiences working in tech. I was frustrated with how much the popular conversation was still focused on “the pipeline problem”: training young girls to code while ignoring all the adult women being driven out of the tech industry by mistreatment. I spent 9 months researching and writing this post. It went viral and remains my most popular essay. This, together with my other posts, led to me being interviewed and quoted in Wired several times regarding diversity in tech (as well as other AI topics). My first post Trends to Avoid When Founding a Startup (2018) The dominant narrative for Bay Area tech startups is to try to raise venture capital, achieve exponential hypergrowth, and hire lots of computer science PhDs. I argued that these approaches not only harm employees, but lead to weaker companies and worse products. My family’s unlikely homeschooling journey (2022) Many people hold a stereotyped and outdated view of homeschooling, not realizing the explosion of innovative, non-traditional education options available in recent years. My husband and I never planned to homeschool, but we unexpectedly found that our child thrives with this approach. Your Immune System is Not a Muscle (2024) The misleadingly named “Hygiene Hypothesis” is often used to justify the misconception that all microbes are good for us. However, this theory is more accurately reframed as the “Old friends hypothesis”: humans co-evolved with friendly bacteria and some parasites. We did not co-evolve with the crowd infections of mega-cities and 100,000 global flights per day. AI Beyond Elite Institutions Machine learning isn’t just for those at billion dollar companies. These posts highlight unconventional practitioners and offer practical guidance for people in varied domains. Deep Learning: Not Just for Silicon Valley (2017) Our goal at fast.ai is making AI accessible to people outside of elite institutions, who are tackling meaningful problems in low-resource areas. This post introduced some of our earliest international fellows and the diverse range of problems they were working on. I always enjoyed writing about fascinating use cases from our deep learning community. How (and why) to create a good validation set (2017) An all-too-common scenario: a seemingly impressive machine learning model is a complete failure when implemented in production. Advice on one common culprit of this, and how to avoid it. In the early years of fast.ai, I wrote numerous posts with practical advice for machine learning. This article asked the question, “Can A.I. conquer its Excel problem? An Introduction to Deep Learning for Tabular Data (2018) Deep learning is not just for images and text. Companies such as Pinterest and Instacart are also applying it to tabular data, the type of data you might normally put in a spreadsheet. This post caught the attention of a reporter with Fortune, who ended up interviewing me and writing about the topic here. Debunking AI Hype & Holding Tech Accountable The narratives about AI put forth by major tech companies are often misleading about what is necessary, what values matter, and what types of harms can result. Google’s AutoML: Cutting Through the Hype (2018) In a 3-part series, I countered claims that all data scientists need customized, bespoke neural network architectures. While I was nervous about disagreeing with both Google’s CEO and head of AI, my posts led to an invitation to keynote the prestigous ICML AutoML workshop. Seven years later, my critiques have been proved valid, with transfer learning a cornerstone of ML and automated neural network search not commonly used. By 2023, we were supposed to all be using AutoML neural architecture search Five Things That Scare Me About AI (2019) AI ethics is not just a theoretical topic. I was (and still am) alarmed about the harms already being caused to human beings by AI systems irresponsibly applied to healthcare, employment decisions, policing, and more. The Problem with Metrics is a Big Problem for AI (2019) Overemphasizing metrics leads to a variety of real-world harms, including manipulation, gaming, and a myopic focus on the short-term. AI is metric optimization on steroids. I later turned this blog post into an academic paper, together with David Uminsky. Two disturbing case studies I keep returning to are how computerized algorithms have been used to cut healthcare and to fire teachers Deep learning gets the glory, deep fact checking gets ignored (2025) A microbiologist discovered hundreds of errors in a paper that used AI to classify enzymes. This is a case study of how challenging it can be to evaluate AI claims outside our area of expertise, as well as of the misaligned incentives that reward flashy results, but not diligent fact-checking. This has been by far my most popular post on LinkedIn. Immunology & Science Decoding T cells with AI (2024) T cells are a crucial component of the adaptive immune system. Accurately pedicting what they can bind to would impact a range of treatments. Numerous algorithms have been developed for this question, but the problem is far from solved. The surface of a T cell. I’ve enjoyed exploring how AI is being applied to immunology Scientists Just Connected the Dots Between Viruses and… Everything (2025) For a long time, catching frequent viruses was considered both inevitable and harmless. But it turns out that common, seemingly-mild viruses have disturbing long-term health impacts. A thread about viruses Thanks for joining me on this walk through the past! Also, you can subscribe to be notified of new blog posts by submitting your email below: I look forward to reading your responses. Create a free GitHub account to comment below.
Most people catch many viruses in their lives– for example, over 90% of adults have Epstein-Barr virus, and adults catch the flu about once every 5 years. For a long time, catching frequent viruses was considered both inevitable and harmless. But it turns out that common, seemingly-mild viruses have disturbing long-term health impacts. Common respiratory viruses increase the risk of heart attacks and strokes. Viruses are linked to dementia and Alzheimer’s Disease. They can re-awaken cancer cells in patients whose cancer was previously in remission. Persistent infections accelerate aging and undermine longevity. Viruses can be the trigger that kicks off life-long autoimmune diseases. New studies come out each week confirming that viruses can harm the health of your heart, blood vessels, brain, nervous system, and gut. Please pause and let this sink in. If we were to truly internalize this information, there would be massive shifts in the practice of medicine, scientific research, and public policy. Pathogens accelerate aging in many ways. Proal and VanElzakker, 2025 What you can avoid (infections) may be just as important as what you seek out (exercise, healthy foods). This news seems depressing. It’s too late, viruses are everywhere, everyone has already caught them– what can be done? There actually is a lot we can do. First of all, developing new anti-viral therapies and treatments should be a top priority. Second, regardless of what infections you’ve already had, preventing or reducing future infections will have a positive impact. There is exciting work happening towards both of these goals, including AI-assisted drug design, patient-led biomedical research, initiatives to improve indoor air quality and new technologies for cleaning the air. How can viruses cause all these bad outcomes when some people who catch them are fine? Human health is complicated. Disease development involves a complex interplay of factors: infections, underlying genetics, environment, the microbiome, and more. Let’s return to the example of Epstein-Barr Virus (EBV). EBV has been strongly linked to Multiple Sclerosis, prolonged fatigue, and 6 different types of cancer. Given that almost everyone has had EBV, even though “only” a percentage of people develop these lasting impacts, this is a major cause of suffering. Viruses tilt the probabilities against you A new world and a powerful idea You may wonder why so many of these health issues are on the rise, when viruses are nothing new. Our world has changed drastically in recent decades compared to most of human evolution. We live in a hyperconnected age of global mega-cities and record numbers of large international flights now. We spend our time indoors in crowded, poorly ventilated buildings. These factors have allowed viruses to travel faster and farther than ever before. The misleadingly named “Hygiene Hypothesis” is often used to justify the misconception that all microbes are good for us. However, this theory is more accurately reframed as the “Old friends hypothesis”: humans co-evolved with friendly bacteria and some parasites. Viruses are not our friends, but rather enemies. We did not co-evolve with these crowd infections of mass travel, mega-cities, and indoor confines. Not all infections are the same! Modern crowd infections are causing huge harm. Figure from Rook, 2014 The idea that viruses are contributing so much to human suffering and long-term disease is powerful. It will transform how we approach medicine, health, and aging, if we let it. This revelation is one of the key reasons that I decided to make a mid-life career pivot, stepping back from fulfiling work in AI to return to graduate school in Microbiology-Immunology, a journey I have been chronicling here on my blog. I hope to spend the next few decades applying my machine learning skills to problems at the intersection of infections, multi-omic data sets, the microbiome, and chronic disease. Below, I will share some of what has captured my attention and upended my old views on disease and medicine. Viruses have many ways to wreak havoc Viruses have evolved to evade, outmatch, commandeer, and otherwise hurt our immune systems. Here is an incomplete and overlapping list of ways that viruses can harm us: 1. Persistence Some viruses quietly stick around for years or decades after our initial illness. They may re-awaken later to cause more problems, or they may spawn surprising issues that we don’t recognize as part of our initial infection. When they persist in our cells, viruses can impact gene expression, hijacking processes our cells need to gain nutrition and energy. Dr. Amy Proal, a researcher in this area, says that treating persistent infections will be necessary to combat aging and extend healthspan. 2. Autoimmunity During an infection, sometimes our immune cells get confused into attacking our own tissue that may “look” similar to the virus (this process is known as molecular mimicry). Once it has mistakenly learned to attack self-tissue, the immune system may continue to do so, even after the virus has been defeated. This is just one of several ways by which viruses can trigger autoimmune diseases such as Lupus, Multiple Sclerosis, Rheumatoid Arthritis, or Type 1 Diabetes. A confused antibody decides to attack a pathogen, as well as the similar-looking myelin covering of the nerves, causing Guillain-Barré syndrome (Comic from Creative Med Doses) 3. Microbiome changes You might expect a stomach bug like norovirus to change the gut microbiome for the worse. Surprisingly, respiratory viruses such as Influenza, RSV, and Covid all harm the gut microbiome too. This is bad news, since the gut microbiome helps to regulate the immune system and produces neurotransmitters for our brain. 4. Immune Dysregulation There are a bunch of ways that the immune system can malfunction (including the ones listed above). Measles can cause immune amnesia, where the immune system forgets previous infections it had learned to fight, leading people to catch the exact same diseases again. There is growing evidence that covid has a negative impact on the immune system as well. 5. Reactivation of other pathogens Infection with a new virus can wake up old infections that were sleeping quietly in your cells. It is unfair, but sometimes viruses will gang up on you, re-activating other viruses (or bacteria) that weren’t bothering you before. 6. Cardiac damage Chickenpox/Shingles, Influenza, and Covid all raise the risk of heart attacks and strokes. Viruses have many ways of harming our cardiac systems: inflammation, damage to the blood vessles, increased blood clots, and damage to the heart. From a meta-analysis of 48 studies about respiratory viruses triggering heart attacks & strokes (Nguyen, et al, 2025) 7. Cancer Cancer involves a failure of the immune system to kill cells that have gone rogue and turned over to the dark side. In 2008, it was estimated that viral infections contribute to 15-20% of human cancer cases. Additional research further linking viruses and cancer has come out since then, so the percentage may be higher now. Both flu and covid infections can reawaken “sleeping” cancer cells that had previously been in remission or cause cancer to spread. An article from MD Anderson on 8 Viruses that Cause Cancer 8. Cumulative impacts You might hope that you could catch a virus, get it over, and be done with it. Unfortunately, that is often not the case. A young college student was fine after having covid twice, but then struggled to walk short distances after her 3rd infection. A Colorado newspaper columnist was skiing, biking, mountain climbing, and running half-marathons up until his 5th covid infection. At this point, he developed pain, fatigue, and migraines that prevent him from doing the activities he loves most. These are not just isolated anecdotes, research confirms the cumulative dangers of repeat infections. In children, a second covid infection is more likely to cause Long Covid than the first infection. Whatever your previous history, reducing risk of future infections is a worthwhile goal. The above mechanisms are not exclusive. For example, some microbiome changes can make it easier for pathogens to pass from the gut into the bloodstream and provoke an autoimmune reaction (a process I talked about in this 5-minute video) The Paradigm Shift For most viruses, people focus on just a few weeks of initial symptoms. This is the wrong way to think about infections. Viral meningitis or EBV increases your risk of Alzheimer’s or dementia, 5-15 years later. Chicken pox (varicella zoster virus) can reactivate decades afterwards as shingles, which itself then leads to increased risk of stroke for at least the following year. We need to radically change how we think about viruses. There is much we still don’t know about the immune system. Early during the covid pandemic, many experts made definitive statements about the risks of covid, assuming that those who didn’t die in the first few weeks must be completely fine. However, perturbations from infections that initially seem minor can have far-reaching, long-lasting, and time-delayed impacts. There is a ton that is still unknown. Trying to figure out how viruses hijack cell processes, alter microbiomes, and dysregulate the immune system are complex questions. Researching these areas with curiosity, determination, and an open mind will reveal a lot. Reasons for Hope It can be gloomy to think about all the damage viruses can cause. The good news is that we don’t have to resign ourselves to these outcomes. Facing the disturbing reality that many viruses are worse than we thought is just the first step towards coming up with creative new solutions. There are some bright, curious, and determined people focused on these problems, although we need even more hands and brains to get involved. The breadth and depth of the harms caused by viruses can focus biomedical research in new directions. Most viruses do not have effective anti-viral treatments. This creates a huge need. Scientific inquiry can fail catastrophically when those closest to the problem are not included. Patient-led research gives me hope, because it is centered on the expertise of those closest to the problem. I am also optimistic about the use of AI for discovering new drugs and designing immune therapies. On the prevention side, reducing how frequently people get sick will have a big impact. Different viruses spread in different ways. In recent years, we have learned that many infections are airborne. Healthy indoor air is a human right, like access to clean drinking water. The UN recently held a high-level event focused on the right to clean air. There are many measures we can take to reduce transmission of airborne diseases, such as improved ventilation, air purification, and far-UVC technologies. Parliament houses, venues for elites, and barns for pigs have already received these air quality upgrades. We need children in schools, employees in workplaces, and patients in hospitals to get the same protections. Hopefully, we are on the cusp of a clean air revolution, with more people and organizations recognizing that healthy indoor air is essential. N95 masks offer an immediate way to significantly reduce how often you get sick. Thankfully, the N95s available currently are more comfortable and more effective than the surgical or cloth masks that many of us wore back in 2020. On the brink of an indoor air quality revolution Conclusion Viruses can harm our cardiac health and cognition, and increase our chances of cancer. If this revelation is fully realized, it will change how the field of medicine operates, priorities in research funding, and public policy on everything from indoor air quality standards to paid sick leave and school attendance. I believe we are on the threshold of what could be a drastic shift in better understanding, preventing, and treating viruses, thus unlocking longer and healthier lives. Related posts you may also be interested in: 5 Devious Tricks Pathogens Use Against Us Viruses are weirder, worse, & more preventable than you realise Viruses: The Silent Triggers of Autoimmune and Neurodegenerative Diseases Your Immune System is Not a Muscle If you enjoy my posts, please subscribe to be notified of new posts via email: I look forward to reading your responses. Create a free GitHub account to comment below.
Deep learning is glamorous and highly rewarded. If you train and evaluate a Transformer (a state-of-the-art language model) on a dataset of 22 million enzymes and then use it to predict the function of 450 unknown enzymes, you can publish your results in Nature Communications (a very well-regarded publication). Your paper will be viewed 22,000 times and will be in the top 5% of all research outputs scored by Altmetric (a rating of how much attention online articles receive). However, if you do the painstaking work of combing through someone else’s published work, and discovering that they are riddled with serious errors, including hundreds of incorrect predictions, you can post a pre-print to bioRxiv that will not receive even a fraction of the citations or views of the original. In fact, this is exactly what happened in the case of these two papers: Functional annotation of enzyme-encoding genes using deep learning with transformer layers | Nature Communications Limitations of Current Machine-Learning Models in Predicting Enzymatic Functions for Uncharacterized Proteins | bioRxiv A Tale of two Altmetric Scores This pair of papers on enzyme function prediction make for a fascinating case study on the limits of AI in biology and the harms of current publishing incentives. I will walk through some of the details below, although I encourage you to read the papers for yourself. This contrast is a stark reminder of how hard it can be to evaluate the legitimacy of AI results without deep domain expertise. The Problem of Determining Enzyme Function Enzymes are what catalyze reactions, so they are crucial for making things happen in living organisms. Enzyme Commission (EC) numbers provide a hierarchical classification system for thousands of different functions. Given a sequence of amino acids (the building blocks of all proteins, including enzymes), can you predict what the EC number (and thus, the function) is? This seems like a problem that is custom-made for machine learning, with clearly defined inputs and outputs. Moreover, there is a rich dataset available, with over 22 million enzymes and their EC numbers listed in the online database UniProt. An Approach with Transformers (AI model) A research paper used a transformer deep learning model to predict the functions of enzymes with previously unknown functions. It seemed like a good paper! The authors used a reasonable, well-regarded neural network architecture (two transformer encoders, two convolutional layers, and a linear layer) that had been adopted from BERT. They looked at regions with high attention to confirm that these were biologically significant, which suggests that the model had learned underlying meaning and provided interpretability. They used a standard training, validation, and test split on a dataset with millions of entries. The researchers then applied the model to a dataset where no “ground truth” was known to make ~450 novel predictions. For these novel predictions, they randomly selected three to test in vitro and confirmed that the predictions were accurate. A transformer model, shown on the left, was used to predict Enzyme Commission numbers for uncharacterized enzymes in E. coli. Three of these were tested in vitro (Fig 1a and Fig 4 from Kim, et al.) The Errors The Transformer model in the Nature Communications paper made hundreds of “novel” predictions that are almost certainly erroneous. The paper had followed a standard methodology of evaluating performance on a held-out test set, and did quite well on that (although later investigation suggests there may have been data leakage). The results claimed for enzymes where no ground truth is known were full of errors. For instance, the gene E. coli YjhQ was predicted to be a mycothiol synthase, but mycothiol is not synthesized by E. coli at all! The gene yciO, which evolved from the gene TsaC, had already been shown a decade earlier in vivo to not have the same function as TsaC, yet the Nature Communications paper concluded it did have the same function. Of the 450 “novel” results given in the paper, 135 of these results were not novel at all; they were already listed in the online database UniProt. Another 148 showed unreasonably high levels of repetition, with the same very specific enzyme functions reappearing up to 12 times for genes of E. coli, which biologically implausible. Most of the “novel” results from the transformer paper were either not novel, unusually repetitious, or incorrect paralogs (Fig 5 from de Crecy, et al.) The Microbiology Detective How did these errors come to light? After the model had been trained, validated, and evaluated on a dataset involving millions of entries, it was used to make ~450 novel predictions, and three of these were tested in vitro. It just so happens that one of the enzymes selected for in vitro testing, yciO, had already been studied extensively over a decade earlier by Dr. de Crécy-Lagard. When Dr. de Crécy-Lagard read that deep learning had predicted that yciO had the same function of another gene, TsaC, she knew from her long years in the lab that this was incorrect. Her previous research had shown that the TsaC gene is essential in E. coli even if yciO is present in the same genome and even when yciO gene is overexpressed. Moreover, the yciO activity reported by Kim et al. is more than four orders of magnitude (i.e. 10,000 times) weaker than that of TsaC. All this suggests that yciO does NOT serve the same key function as TsaC. Two enzymes with a common evolutionary ancestor, but different functions (Fig 7 from de Crecy, et al.) YciO and TsaC do have structural similarities, and YciO evolved from an ancestor of TsaC. Decades of research on protein and enzyme evolution have shown that new functions often evolve via duplication of an existing gene, followed by diversification of its function. This poses a common pitfall in determining enzyme function, because the genes will have many similarities with the ones they duplicated and then diversified from. Thus, looking at structural similarities is only one type of evidence for considering enzyme function. It is also crucial to look at other types of evidence, such as neighborhood context of the genes, substrate docking, gene co-occurrence in metabolic pathways, and other features of the enzymes. It is important to look at multiple types of evidence when classifying enzyme function (Fig 2 from de Crecy, et al.) Hundreds of Likely Erroneous Results Spotting this one error inspired de Crécy-Lagard and her co-authors to take a closer look at all of the enzymes found to have novel results in the Kim, et al, paper. They found that 135 of these results were already listed in the online database used to build the training set and thus not actually novel. An additional 148 of the results contained a very high level of repetition, with the same highly specific functions reappearing up to 12 times. Biases, data imbalance, lack of relevant features, architectural limitations, or poor uncertainty calibration can all lead models to “force” the most common labels from the training data. Other examples were proven wrong via biological context or a literature search. For instance, the gene YjhQ was predicted to be a mycothiol synthase but mycothiol is not synthesized by E. coli. YrhB was predicted to synthesize a particular compound, which was already predicted to be synthesized by the enzyme QueD. A form of E. coli with a QueD mutant was unable to synthesize the compound, showing that this is not in fact the function of YrhB. Rethinking Enzyme Classification and “True Unknowns” Identifying enzyme function actually consists of two quite different problems which are commonly conflated: propagating known function labels to enzymes in the same functional family discovering truly unknown functions The authors of the second paper observe, “By design, supervised ML-models cannot be used to predict the function of true unknowns.” While machine learning can be useful for propagating known functions to additional enzymes, there are many types of errors that can occur: including failing to propagate labels when they should, propagating labels when they should not, curation mistakes, and experimental mistakes. Unfortunately, erroneous functions are being entered into key online databases such as UniProt, and this incorrect data may be further propagated if it is used to train prediction models. This is a problem that increases over time. Need for Domain Expertise It is not news that AI work will be more highly rewarded and supported than work that closely inspects the underlying data and integrates deep domain knowledge. The aptly titled “Everyone Wants to do the Model Work, not the Data Work” paper involving dozens of machine learning practitioners working on high-stakes AI projects and found that inadequate-application domain expertise was one of a few key causes of catastrophic failures. Sources of cascading failures in machine learning systems (Fig 1 from Sambasivan, et al.) These papers also serve as a reminder of how challenging (or even impossible) it can be to evaluate AI claims in work outside our own area of expertise. I am not a domain expert in the enzyme functions of E. coli. And for most deep learning papers I read, domain experts have not gone through the results with a fine-tooth comb inspecting the quality of the output. How many other seemingly-impressive papers would not stand up to scrutiny? The work of checking hundreds of enzyme predictions is less glamorous than the work of building the AI model that generated them, yet it is even more important. How can we better incentivize this type of error-checking research? At a time when funding is being slashed, I believe we should be doing the opposite and investing even more into a range of scientific and biomedical research, from a variety of angles. And we need to push back on an incentive system that is disproportionately focused on flashy AI solutions at the expense of quality results. Related Reading: The problem with metrics is a big problem for AI Gaps and Risks of AI in the Life Sciences “AI will cure cancer” misunderstands both AI and medicine You can subscribe to be notified of new blog posts by submitting your email below: I look forward to reading your responses. Create a free GitHub account to comment below.
DNA sequencing hasn’t lived up to the hype Twenty to thirty years ago, politicians, scientific leaders, journalists, and even Nobel laureates predicted that sequencing the human genome would revolutionize how we treat disease. And while the advances in DNA sequencing that have occurred since then have improved recognition and treatment for some cancers and rare diseases, on the whole the field has not lived up to earlier hype. Time Magazine covers from 1994 and 1999 about genetics In an article titled “Why sequencing the human genome failed to produce big breakthroughs in disease”, a biology professor highlights that most common diseases are not caused by a single gene. In fact, common diseases are often linked to hundreds of gene variants, and even collectively, these variants still account for only a small fraction of disease variance. Here, I want to focus on two other key limitations of DNA sequencing, and how they are now being addressed with new approaches. What DNA can’t tell us First, DNA can’t answer many questions about how cells and organisms work in practice. A neuron in the brain has the same DNA as a liver cell, yet the two have completely different functions. This is because different segments of DNA are turned off or on in different cells. To understand how cells are actually working, you need to know about proteins and RNA (RNA is the intermediary which translates DNA into protein). Proteins are what build the structure of cells, catalyze chemical reactions within the cell, and allow communication between cells. Healthy and unhealthy cells in the same organism will usually have the same DNA. For instance, if some regions of the intestines are experiencing an IBD (irritable bowel disease) flare and others aren’t, they would all have the same DNA, yet likely different RNA and protein levels. A second big problem is that many key sequencing techniques destroy spatial information. You essentially may have to put tissue or cells into a blender in order to get rich information about DNA or RNA sequences. While this data is informative, it turns out that locations of cells within a piece of tissue, and locations of regions within a cell, are also very important! Again, considering the case in which some regions of the intestine are inflamed due to IBD, yet others aren’t, mixing them all together in a blender will lose or distort useful information. Look at the difference between crypts in a healthy segment of the colon (on the left) compared to inflamed crypts (on the right). Differences between a healthy bowel (on the left) and an inflamed bowel (on the right). Source: mypathologyreport.ca Recently, we have seen a rise in breakthroughs that allow us to obtain data about location. Spatial techniques are a necessary and exciting step beyond DNA sequencing. The power of spatial information showed up as a major theme at a conference I attended last year, and spatial techniques have been recognized by Nature Methods as “Method of the year” twice in the last 5 years. A Few Major Areas of Innovation What is genetic sequencing anyway? There are a number of different types of sequencing that have been invented in the last 30 years. Walking through a brief history will illustrate what these technologies are, and what they can and cannot do. To make it concrete, let’s look at the example of how they have been applied to cancer treatment. Sanger Sequencing: This is an older technology dating back to the 1970s, and which was the main way of sequencing DNA up until 2005. Sanger sequencing was one of the methods used in the mid-90s to identify the genes BRCA1 and BRCA2 as key genetic risk factors for breast cancer. The process of discovering BRCA1/BRCA2 involved scientists slowly zeroing in on their chromosomal locations over a period of years. Several other cancer genes were discovered during this time period as well. While Sanger sequencing is effective on smaller amounts of DNA, it can be quite slow to deal with larger volumes. It took over 10 years to sequence the first copy of the human genome using Sanger Sequencing. It is still used today as a simple and reliable way to test for known mutations (such as BRCA1/BRCA2) or for smaller tasks. High-Throughput Sequencing: New technology released in the mid-2000s allowed DNA to be cut into lots of short pieces and for millions of pieces to be sequenced in parallel at once. This approach, called high-throughput sequencing, was significantly faster than Sanger sequencing. High-throughput sequencing has many applications, including to cancer treatment, by making it cheaper and faster to sequence DNA to identify particular mutations which can influence treatment decisions. Method of the Year | Nature Methods Long-read sequencing: High-throughput sequencing has the advantage of high accuracy, but the downside of short sequence lengths. In the 2010s, technologies were released with the opposite set of strengths and weaknesses. Long-read sequencing provides the advantage of long sequence lengths, although the downside of lower accuracy. To compare, high-throughput sequencing uses DNA strands that are a few hundred base pairs long, whereas long-read sequencing uses DNA strands that are tens of thousands base pairs long. Both technologies have different strengths and are widely used today. Single-Cell Sequencing: High-throughput sequencing involves sequencing the DNA of many cells at once, but sometimes it is useful to sequence individual cells. In a tumor, different cells can have different mutations. It is possible that a small subset of the mutations may drive metastasis (the spread of cancer to other areas) or resistance to treatment. Identifying these driver mutations can guide treatment decisions, since particular driver mutations can predict the effectiveness of various drugs. Key mutations may be drowned out in the average if you sequence the entire tumor. This is one reason why it is useful to be able to sequence single cells, and not just obtain the average of many cells. Single-cell sequencing was selected as Nature Methods method of the year in 2013. A Lego interpretation of bulk RNA-seq; single-cell RNA-seq; spatial transcriptomics; and the original organ. Source: Bo Xia, @BoXia7 Multi-Omics: The study of DNA is genomics. DNA alone gives us an incomplete picture of an organism. Epigenomics can provide information about which regions of DNA are active or silenced. To understand how different cells function, as well as cells in different states of disease or health, you also need to know about their RNA and proteins. This data is contained in the field of transcriptomics (transcripts are strands of RNA transcribed from DNA) and proteomics (the proteins in a cell). Metabolomics looks at small molecules (such as sugars, amino acids, and vitamins) within the body and exposomics includes all sorts of environmental exposures. Collectively, these fields are known as -omics or multi-omics. It is valuable to combine multiple types of -omics together for richer sources of information, since each has different strengths, limitations, and insights to offer. Multi-Omics is an exciting area that draws on lots of data, with applications to cancer, infectious disease and immunology. Spatial: Spatial information lets us see all the variation within a section of the body– such as a segment of the intestines, the liver, or a cancerous tumor. This variation can often be significant for understanding disease and treatment prognosis. To better understand why spatial techniques are useful, let us dive into some background about cancer. Tumors aren’t just lumps of bad cells Cancer is defined as excess cell division. I used to think that tumors were just clumps of “bad cells”, where “good” and “bad” were binary states. This is incorrect. Tumors are not uniform, and within the category of “bad” there is a great deal of variation and heterogeneity. Different cells within a tumor may have different mutations from one another. And our immune systems sometimes build complex defense structures within tumors in attempts to more effectively fight them. How close a cancerous cell is to one of these immune structures impacts how likely the body is to destroy the cancerous cell. Notice all the variation within this tumor! Immunologist Dr. Angela Ferguson, who studies head and neck cancers, describes tumors as having a “physical landscape”. She has shown how the organization and structure within a tumor can predict and guide treatment outcomes. Her work found that cancer progression is “landscape-dependent”, where landscape refers to the locations of immune cells and structures within a tumor. It is not enough to study cancer cells in isolation. We need to understand their layout. Methods that effectively put tumor cells into a blender in order to sequence them, disrupting their spatial information, are insufficient on their own. Mapping that spatial information can hold the keys for more effective treatment. Single cell sequencing approaches allow a greater number of genes to be measured, whereas spatial approaches measure fewer genes but also provide location information. Combining these two approaches can prove powerful. Programming Libraries Applied to Spatial -Omics We are living at a time when multi-omics, spatial information, and user-friendly programming libraries are converging for easier exploration and discovery. For instance, below is an image I created using the common Python programming libraries pandas and matplotlib of data the NIH has shared about a rare liver disease. The image shows a slice of liver tissue, with gene expression overlaid in a color scale ranging from purple (low expression) to yellow (highest expression). Using Matplotlib to display a cross-section of liver with gene expression The NIH dataset contain images of slices of liver and expression of many different genes, from both healthy patients and those with a rare liver disease. https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE240429. The plot shows expression of albumin, a protein which helps transport other molecules around the body, overlaid on top of the liver, illustrating which regions produce more or less of this protein. Data science tools are invaluable for transforming and plotting such data. In addition to being able to use standard Python libraries (such as Pandas and Matplotlib) for visualing this data, there are many specialist libraries as well. The Scverse (Sc = single cell) includes a number of Python libraries focused on single cell analysis. The SC Verse includes libaries for single cell sequencing analysis There are also popular R libraries for single cell analysis and multi-omics, such as Harmony, Seurat, and mixOmics. With all of these tools and technologies, it is an exciting time to be working at the intersection of data science and microbiology. The causes of most diseases are complex and multi-factorial, although new approaches of spatial multi-omics provide unprecedented types of useful information. Related Reading: What AI can tell us about microscope slides Gaps and Risks of AI in the Life Sciences AI and Immunology You can subscribe to be notified of new blog posts by submitting your email below: I look forward to reading your responses. Create a free GitHub account to comment below.
The lavender images below show breast tissue. There are many questions doctors could want to answer using these images: They could want to know whether there are tumors present or not. If there is a tumor, doctors would want to classify its stage, make predictions about how likely the patient is to respond to treatment, and to detect whether the tumor has spread from another organ. All of these are questions which people are now tackling with machine learning. They fall within the area of computational pathology, often abbreviated CPath. In the past year, two CPath AI models were released which achieved state-of-the-art results. Here I will discuss an introduction to this field, what these models do, and what some key challenges are going forward. Breast tissue images from the BACH: Grand challenge on breast cancer histology CPath foundation models There is a powerful idea about how to make more accurate CPath models. Rather than train a model on a single type of tissue and a single task (e.g. identifying cancer in breast tissue), train a model on images of tissue from many different organs (breasts, lymph nodes, lungs, prostate, heart,…) and on multiple different tasks (recognizing cancer, determining the stage and subtype of the cancer, segmenting cells, and predicting treatment outcomes). Patterns learned from one dataset or one task are likely to generalize to others. Such models are known as CPath foundation models. In general, a foundation model is a machine learning model which is trained on a sufficiently diverse large dataset which can then be adapted for a range of downstream tasks. This idea is commonly used in the area of language models such as Chat-GPT and Claude.ai. Language foundation models are trained on many types of language tasks and intended to generalize across different corpuses of text (e.g. wikipedia, reddit posts, academic papers, online conversations, news articles, and more). ImageNet models trained to recognize a huge variety of different pictures often serve as foundation models for images. The success of foundation models within the areas of language and more general images is a key reason why we might expect pathology foundation models to be useful too. Tissues are groups of cells with similar structure and function. Different types of tissue within the human body include nervous, muscle, connective, and epithelial tissue. Image: Wikimedia Two notable CPath foundation models were released in 2024: Prov-GigaPath and UNI. Both models achieved state-of-the-art performance on dozens of pathology tasks (although they were not directly compared to one another). Another relevant paper (from Kaiko.ai) studied the impact of dataset size and model size on CPath model performance. Learning the Vocab Medicine is full of jargon and specialized vocabulary. Pathology refers to the study of disease. It is a broad field, and can include everything from dissecting dead bodies to analyzing blood samples. One key focus of computational pathology is analyzing and interpreting whole slide images (WSIs) and in some cases combined with accompanying meta-data about a patient. Whole slide images refers to the complete microscope slide, although in many cases the region of interest (such as particular cancerous or inflamed cells) may be much smaller, just occupying a subset of the slide. Machine learning (ML) is a subfield of Artificial intelligence (AI) which involves learning from past data, and is increasingly being used with great success in pathology. The focus of most computational pathology ML models is on images of tissue, on microscope slides. That is what we will focus on in this post as well. So Many Tasks! There are many different benchmarks that CPath models can be tested on. These involve numerous datasets: related to different areas of the body, with different sizes, and with different purposes. They also involve a variety of tasks, including binary classification, image segmentation, and outcome prediction. Prov-GigaPath attained state-of-the-art performance on 25 out of the 26 tasks it was evaluated on and UNI attained state-of-the-art performance on 34 different tasks. Here I will give examples of just 3 of these tasks. Task: prostate cancer cell grading In the 1960s, the pathologist Dr. Donald Gleason came up with a grading scale for rating cells as they progressed from normal to prostate cancer. The Gleason Grading system is still widely used and is considered a powerful predictor of how prostate cancer patients will fare. A major medical image conference (MICCAI) held a competition in 2022 for researchers to create algorithms to determine the Gleason grades when given images of prostate tissue. Examples of UNI predictions of Gleason grades for a section of prostate tissue. Figure 3b from the UNI paper The prostate tissue is shown in pink, and segments have been colored in blocks based on where they fall on the Gleason scale. Task: identifying early signs of rejection after a heart transplant Rejection is the main cause of mortality in patients who have received a heart transplant. Since the early stages of rejection can be asymptomatic, it is standard for patients to receive frequent biopsies for 1-2 years following a transplant. These are known as endomyocardial biopsies (EMB), since they remove a small sample of tissue from the inner lining (endo) of the heart (cardial) muscle (myo). Accurately interpreting the results of these biopsies is a key question. Underestimating the chance of rejection could lead to dangerous delays in treatment, but overestimating could lead to alarm and unnecessary follow-ups or treatment. Assessment of the sampled tissue by experienced pathologists has higher variability than many other tasks, such as cancer diagnosis. Deep learning is being used to tackle this task, in models such as Cardiac Rejection Assessment Neural Estimator (CRANE) and the CPath foundation model UNI. Each row shows a different sample of cardiac tissue, with a different medical issue. On the far left are the whole slide images, then zoomed in at higher resolution on a key Region of Interest (ROI). On the far right is a heat map for the most zoomed in area showing which features the algorithm has identified as significant. Figure 3 from the CRANE paper. Task: Genetic Mutations in Cancer For several common genetic mutations in tumors, there are specific drugs known to target those mutations. This has a direct application for clinical treatment. Since genetic mutations can change the form and function of cells, it is reasonable to expect that this information could be deduced from images of the cancer cells. Deep learning models have been built to identify genetic mutations from tissue slides. The benefits of using a computational approach are that it can be scaled as an increasing number of relevant genetic mutations and molecular biomarkers are being discovered. Task-specific models have been built for this, and this is one of the tasks that foundation models can be tested on. Different types of cancer listed along the y-axis and 20 common genetic mutations listed on the y-axis. Figure 1D from Kather, 2020. We need more data One key challenge in the area of CPath foundation models is gathering enough training data. The Cancer Genome Atlas (TCGA) was an ambitious project launched in 2006 by the National Cancer Institute in the USA. Over a 12 year period, samples were collected from over 11,000 patients of 33 different cancer types, and all this data was made publicly available. While this is a rich dataset and a useful resource, all 3 papers we’ve looked at concluded that TCGA is not large enough for effective foundation models. In addition to limited data size, TCGA also has limited diversity, consisting mostly of slides from the primary site of cancer, but not metastasized cancers or different types of tissues. Researchers at Kaiko.ai tested the impact of scaling both the size of their model and the size of the training dataset. While they found limited need to scale model size beyond a certain point, they found that larger datasets continued to lead to increased performance. They concluded that TCGA was likely not large enough and shared their plans to build a larger training set, and are now partnering with cancer centers across Europe to create a dataset for their model. The researchers behind two other CPath foundation models reached the same conclusion about data set size, and gathered massive datasets to train their models. This required partnering with healthcare centers. Prov-GigaPath, a model created by Microsoft Research and Providence Genomics involved data from 30,000 patients across 28 cancer centers (which are part of Providence Healthcare company). UNI, a cPath model created by a team at Harvard, MIT, and the Broad Institute, involved the creation of the Mass-100K: a dataset with over 100K whole slide images across 20 tissue types collected from Mass General Hospital, Brigham & Women’s Hospital, and Genotype-Tissue Expression (GTEx) consortium. These partnerships and curation of training datasets are currently a crucial component of building CPath foundation models. Curating datasets carefully poses many challenges as well. Combining data from different sources, which often use different protocols for how slides are sampled and prepared, can introduce significant biases. Different scales CPath foundation models face the difficulty of capturing both local patterns (that show up in a small tile within a slide) and global patterns across the whole slide. Many tiny tiles are found within a slide. Some models, such as the Hierarchical Image Pyramid Transformer (from several of the same authors as UNI), use hierarchical approaches to deal with these multiple scales. Hierarchical Structure of Whole-Slide Images, Figure 1 from Chen, et al, 2020 Other models, such as Prov-GigaPath, treat the tiles as tokens, encoding both the tiles and the slide as a whole as model inputs. Prov-GigaPath uses both a slide encoder and a tile encoder to take into account these two different scales. Treating slides as tokens, Figure 1a from the Prov-GigaPath paper In pathology clinics, diagnosis and treatment decisions are often made at the patient level, whereas CPath models are often highly focused on regions of interest. Accommodating the multiple relevant scales (small tiles, whole slides, and patient-level) for pathology is a consideration that CPath models need to balance. Going Forward It is still early in the world of CPath and there are many growth opportunities, including the continued need for large and diverse datasets, ways to further optimize model training, tasks which have previously received less focus, and the difficulties of integrating models into clinical work. As the authors of the kaiko.ai paper wrote, “We are still at the very beginning of developing a truly foundational pathology foundation model.” It is a hopeful sign that these models achieve state-of-the-art results on dozens of benchmarks, but it still remains to be seen when and how they will be used in clinical settings. Related Reading: The Most Common and Useful Neural Nets Using AI to Discover New Antibiotics AI and Immunology You can subscribe to be notified of new blog posts by submitting your email below: I look forward to reading your responses. Create a free GitHub account to comment below.
More in AI
When I finished learning how to build an LLM from scratch, I was left with a mystery: my own models were not as good as OpenAI's original GPT-2 models, despite being based on the same architecture. My models all had 163M parameters, and followed the design from Sebastian Raschka's book "Build a Large Language Model (from Scratch)". That meant that they were pretty much the same as the setup for the OpenAI GPT-2 "small" instance, except that they did not use weight-tying or bias on the QKV matrices. Weight-tying means that you re-use the initial embedding matrix as the output head at the end, and using it means that GPT-2 small saved quite a few parameters -- it was 124M rather than 163M -- at, at least in my own experiments, a cost in quality; similarly, while I found that QKV bias made a tiny improvement in loss terms, I'd felt it was likely within the noise. But GPT-2 small consistently beat my models on an instruction fine-tuning (IFT) task -- also adapted from Raschka's book. That test fine-tunes the model on a subset of the Alpaca dataset, until validation loss starts rising, and then runs a test set through the resulting model. The responses to the test set questions are stored, and then I run all of the responses from all of the models under test past GPT 5.5 in one go to get an aggregate score; more details here. GPT-2 small always did better than any of my models on this. Additionally, it did surprisingly well on a simpler eval -- one that just measured the cross entropy loss it got on a test set. It scored close to my own best models, and better than many of them. What made this result particularly interesting was that the test set in question was a split of my own training data; my models would not have seen it when training (at least, in theory), but it seems likely that it would be much more similar to their own training data than it was to OpenAI's. I've checked two things while probing this mystery: It seems very likely that the GPT-2 models were overtrained by modern standards; would overtraining my own models get them closer? It turned out that no, it probably didn't help with the IFT eval (though there might have been some signal there). It did help quite a lot with the test loss eval, though. The way I was handling dropout in the IFT test might have been unduly benefiting some models while working against others. I decided to standardise on not using dropout during this eval, as (counter-intuitively for me) it seemed to harm the results of most models, even those that had been pre-trained with dropout. In particular, the OpenAI weights were harmed by using dropout, and making a change that benefited them (along with some of my own models) seemed the most conservative approach to take in investigating this. The next thing I wanted to look into was the training data. The exact dataset that the various GPT-2 models were trained on has never been released; all we know about it is from the paper, where they say: [W]e created a new web scrape which emphasizes document quality. To do this we only scraped web pages which have been curated/filtered by humans. Manually filtering a full web scrape would be exceptionally expensive so as a starting point, we scraped all outbound links from Reddit, a social media platform, which received at least 3 karma. This can be thought of as a heuristic indicator for whether other users found the link interesting, educational, or just funny. They called it "WebText". There is an OpenWebText that tries to replicate it, but although they tried to follow the same procedure as the original, there's no guarantee that it is all that similar. By comparison, I'd normally been training against FineWeb. While this is a general web-scraping dataset, without the "curation" provided by using only stuff that was linked from upvoted Reddit posts, it has been refined to remove any obvious junk. I had felt that it was pretty much equivalent. But what if I were wrong about that? I decided to see if I could get better models by using better data. The starting point Here's a table of all of the models I've been comparing to date. The "Test loss" column shows how well the model in question did on that held-back cross entropy loss evaluation. The "IFT epochs" column shows how many epochs of fine-tuning the model needed before its validation loss started rising, the "IFT score" the score that GPT 5.5 gave the model's responses to the test set of my Alpaca data, and the "IFT rank" the model's rank in terms of that score. The OpenAI small model is in there in bold, and I've also included the OpenAI medium model for comparison purposes. Test loss IFT epochs IFT score IFT rank OpenAI weights: medium 3.231442 2 43.75 1 JAX, overtrained one long epoch 3.324953 3 19.77 4 JAX, overtrained two normal epochs 3.326482 4 19.72 5 JAX, with MHA bias, no dropout 3.418784 4 18.69 6 JAX, no MHA bias, no dropout 3.420089 5 21.46 3 JAX, no MHA bias, with dropout 3.476802 5 13.22 15 OpenAI weights: small 3.499677 2 26.00 2 1xrtx3090-stacked-interventions 3.538161 4 13.77 14 8xa100m40-stacked-interventions-1 3.577761 4 10.76 18 Cloud FineWeb, 8x A100 40 GiB 3.673623 3 17.72 7 1xrtx3090-baseline 3.683835 4 15.74 8 8xa100m40-baseline 3.691526 3 14.19 13 Cloud FineWeb, 8x H100 80 GiB 3.724507 4 14.33 12 Cloud FineWeb, 8x A100 80 GiB 3.729900 3 11.34 17 Cloud FineWeb, 8x B200 160 GiB 3.771478 4 14.67 11 Local FineWeb train 3.943522 5 12.31 16 Local FineWeb-Edu extended train 4.134991 5 15.04 9 Local FineWeb-Edu train 4.166892 5 14.99 10 You can see that the OpenAI small model did pretty well in terms of the test loss, when you consider that it has 39M fewer weights than my models and was being tested against a dataset that differs more from its likely training data than it does from my own models'. Additionally, the specific models that did better than OpenAI's small one were all trained with JAX rather than PyTorch -- my hypothesis for that is that it's a result of the JAX ones getting better initial weights by pure chance. But the big difference was in the IFT score. In the specific run that gave the results in this table, the OpenAI small model got 26.00 -- the closest of my own models was more than 4.5 points lower, at 21.46. This difference was consistent over all of my other test runs. The GPT-2 small model was always ahead of mine. (GPT-2 medium, of course, beat GPT-2 small and all of my models, but given that it is twice the size of mine, that's not a big surprise.) Now, quite some time ago, I had tried looking into data quality as a lever to pull for model performance. At the bottom of the table, with the worst test loss of all models, you can see two models: "Local FineWeb-Edu train" "Local FineWeb-Edu extended train" These two were (as you might guess from the names) trained on the FineWeb-Edu dataset, which includes just the most "educational" data from FineWeb. They scored very badly on the test loss score. Given that the test dataset is from FineWeb, that's not a big surprise -- as I've written previously: If you train a model on Jane Austen and then evaluate against Chuck Tingle, then you're not going to get amazing results. But again, GPT-2 had the same issue, and did perfectly well on the test loss eval. On the other hand, while these FineWeb-Edu models' performance on the IFT eval wasn't stellar -- there are plenty of my other models ahead of them -- they did seem to punch above their weight. Consistently across all of the IFT evals I've done, they have scored higher than many of the others -- despite their poor loss on the test eval. Additionally: they were amongst the first models that I trained, before I'd spent time learning about how to optimise my hyperparameters and training loop. They did not use gradient clipping, they did use dropout, their batch size was just "whatever I could squeeze into the GPU", and I didn't set the learning rate to the right kind of value or schedule it over the course of the training run. So maybe a new training run on FineWeb-Edu plus my training improvements would help? And maybe some other tweaks to the training data would be worth looking into? The plan I decided to see what would happen if I trained some models with better-quality data. Specifically, I would train models with my current optimised loop and hyperparameters on four different datasets: FineWeb-Edu -- essentially the same as "Local FineWeb-Edu train" but with a better training setup. This would test the "more educational -> better" hypothesis. A 50:50 split of FineWeb and FineWeb-Edu. I've read that LLMs can be helped by having a decent amount of lower-quality data in their training loop, as it helps them to generalise. Perhaps having some FineWeb in there in addition to the FineWeb-Edu stuff would improve that test loss score while also helping the IFT test? A "curated" dataset containing 45% of its contents from FineWeb, 45% from FineWeb-Edu, and 10% from the Simple English Wikipedia. The full Wikipedia is huge, and full of obscure facts -- while the Simple English one is small and hopefully richer in useful information on a per-token basis. And conveniently, Answer.ai have made a snapshot of it available on Hugging Face Hub. Might deliberately putting a bunch of encyclopaedic data into the training set make the model better at the IFT eval (which has lots of factual questions in it, like "who wrote Pride and Prejudice")? OpenWebText. Even though I was unsure how well it matched the original WebText, given that it was there, it seemed silly to not try training something on it and see how it matched up. I would train each model on 3.2B tokens of the chosen dataset; that's the Chinchilla-optimal amount for my 163M-parameter models. If there were any interesting results, then I might consider doing overtrained models later on. I decided to be at least vaguely scientific about this, and to pre-register some predictions: The FineWeb-Edu-only model would do pretty badly on the test loss, but better than my older FineWeb-Edu models (90%). It would also punch above its weight on the IFT eval (90%). The 50:50 split: I expected it to do worse on the test eval than my JAX FineWeb-only models (70%), but better than the FineWeb-Edu one (90%). I wasn't sure about how it would do on the IFT eval, but thought it might be somewhere in between the two groups (60%). The curated dataset I had high hopes for in terms of the IFT eval -- let's say 80% chance of it being the best of all of my models. For the test loss eval, I expected it to do about as well as the 50:50 split, maybe a little bit worse (70%). I had no idea how the OpenWebText eval would do! Could be worse, could be better. Here's how things turned out. The FineWeb-Edu model I already had a dataset based on FineWeb-Edu ready to go, from when I trained those two original models. It is just the 10B-token sample of the original dataset at the time I generated it last December, formatted appropriately for my training script (details on the dataset card). I kicked off a training run with my JAX code (which I've been using for the other posts in this series): giles@poppy:~/Dev/jax-gpt2-from-scratch (main)$ XLA_PYTHON_CLIENT_MEM_FRACTION=0.95 uv run train.py full-llm-full-train-with-mha-output-bias-fineweb-edu datasets/ 2026-09-11 18:11:47.991583 Downloading dataset Fetching 4 files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 4/4 [00:00<00:00, 1772.93it/s] Download complete: : 0.00B [00:00, ?B/s] | 0/4 [00:00<?, ?it/s] 2026-09-11 18:11:48.226273 Loading dataset into RAM Download complete: : 0.00B [00:00, ?B/s] 2026-09-11 18:16:29.507646 Creating model 2026-09-11 18:16:33.042509 Creating optimizer 2026-09-11 18:16:34.138990 Start train 0%| | 0/33165 [00:00<?, ?it/s] 2026-09-11 18:17:38.486288 Saving checkpoint 1%|▌ | 173/33165 [13:22<39:17:03, 4.29s/it, loss=6.897, tps=21,201] ...and just less than 40 hours later, I had a model: Training complete in 142,912.226 seconds 2026-09-13 09:58:26.437276 Tokens seen: 3,260,252,160 2026-09-13 09:58:26.437284 Throughput: 22,813 tokens/second 2026-09-13 09:58:26.437302 Final train loss: 3.342 2026-09-13 09:58:26.437309 Done I converted the saved JAX safetensors file from the last checkpoint into a format that would be compatible with my PyTorch eval code, and ran my smoke test: how would it complete the sentence "Every effort moves you"? Every effort moves you closer to God’s Kingdom, and even closer to Him. As we can see in That was nice and coherent -- if unusually religious! -- so that was promising. I ran the test eval: giles@perry:~/Dev/ddp-base-model-from-scratch (main)$ uv run test_loss.py datasets/ ../jax-gpt2-from-scratch/runs/full-llm-full-train-with-mha-output-bias-fineweb-edu/model.json ../jax-gpt2-from-scratch/runs/full-llm-full-train-with-mha-output-bias-fineweb-edu/checkpoints/latest/pytorch-model.safetensors Fetching 4 files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 4/4 [00:00<00:00, 2758.50it/s] 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 3200/3200 [03:52<00:00, 13.74it/s] Loss against our test dataset: 3.632900 That was pretty good, putting it at a better test loss than all of the models I had trained without optimised hyperparameters, and worse than all of the ones I had trained on FineWeb with optimised hyperparameters. So that fit in with my prediction that it would be better than the old FineWeb-Edu models; the fact that it was also better than the non-optimised training runs with FineWeb seemed sensible enough that I felt silly for not having predicted that it would have fallen exactly there :-) I decided to leave the IFT eval until the end so that I could check all of the models from these experiments together, so it was time to upload this one to Hugging Face, and move on to the next model. 50:50 FineWeb to FineWeb-Edu I put together a new repo with a script to prepare datasets specifically for my training setup. You provide it with config that specifies some source datasets along with information about how to process them and how to mix them together, and it uploads a new dataset to Hugging Face Hub with the required characteristics. For example, for the 50:50 FineWeb to FineWeb-Edu split, the config looked like this: { "seed": 42, "tokens_desired": 10000000000, "upload_dataset_name": "gpjt/fw-fwedu-5050-gpt2-tokens", "sources": [ { "name": "FineWeb", "hf_id": "HuggingFaceFW/fineweb", "hf_name": "sample-10BT", "hf_split": "train", "item_field": "text", "weight": 50 }, { "name": "FineWeb-Edu", "hf_id": "HuggingFaceFW/fineweb-edu", "hf_name": "sample-10BT", "hf_split": "train", "item_field": "text", "weight": 50 } ] } The way the script works is pretty simple: it works out (based on those weights and the tokens_desired) how many tokens it wants from each source dataset, shuffles the items in the sources, then it loops until it has the desired number of tokens or more stored in an output. In the loop, it works out which source is currently most under-represented, grabs an item from it, tokenises it, and adds it to the output. Running it with that 50:50 config seemed to work fine: giles@perry:~/Dev/prepare-llm-training-dataset (main)$ uv run prepare-dataset.py runs/fw-fwedu-5050/ Resolving data files: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████| 27468/27468 [00:00<00:00, 89875.56it/s] Loading dataset shards: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████| 102/102 [00:00<00:00, 133.75it/s] Resolving data files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 2410/2410 [00:00<00:00, 87461.48it/s] Loading dataset shards: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████| 98/98 [00:00<00:00, 200.09it/s] 2026-09-13 20:13:22.000187: Generating dataset; per-source counts 2026-09-13 20:13:22.000217: FineWeb: 5,000,000,000 2026-09-13 20:13:22.000221: FineWeb-Edu: 5,000,000,000 FineWeb: 100%|████████████████████████████████████████████████████████████████████████████████████████████████▉| 4999999705/5000000000 [1:01:33<00:00, 1353639.33token/s] FineWeb-Edu: 5000000363token [1:01:33, 1353639.47token/s] 2026-09-13 21:14:55.747239: Done generating tokens 2026-09-13 21:14:55.748480: FineWeb: 4,999,999,705 / 5,000,000,000 (1.000, 1 iterators) 2026-09-13 21:14:55.748487: FineWeb-Edu: 5,000,000,363 / 5,000,000,000 (1.000, 1 iterators) 2026-09-13 21:14:55.748489: Total: 10,000,000,068 2026-09-13 21:14:55.748491: Catting... 2026-09-13 21:16:29.565152: Catted into a tensor of shape torch.Size([10000000068]) 2026-09-13 21:16:29.566663: Saving... 2026-09-13 21:16:36.006267: Saved 2026-09-13 21:16:36.009413: Uploading to gpjt/fw-fwedu-5050-gpt2-tokens Processing Files (1 / 1) : 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 20.0GB / 20.0GB, 117MB/s New Data Upload : 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 14.6GB / 14.6GB, 98.1MB/s ...du-5050/train.safetensors: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 20.0GB / 20.0GB 2026-09-13 21:17:59.545875: Done So we had almost-perfect 50:50 balance between the datasets, and it saved this dataset on Hugging Face. I ran a script to double-check that it looked sane, and it did, so it was time to spin up a training run: giles@perry:~/Dev/jax-gpt2-from-scratch (main)$ XLA_PYTHON_CLIENT_MEM_FRACTION=0.90 uv run train.py full-llm-full-train-with-mha-output-bias-fw-fwedu-5050 datasets/ 2026-09-13 21:20:59.880918 Downloading dataset Fetching 2 files: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 2/2 [01:13<00:00, 36.70s/it] Download complete: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████| 20.0G/20.0G [01:13<00:00, 1.24GB/s] 2026-09-13 21:22:13.521745 Loading dataset into RAM Download complete: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████████| 20.0G/20.0G [01:13<00:00, 272MB/s] 2026-09-13 21:22:33.787720 Creating model 2026-09-13 21:22:35.501063 Creating optimizer 2026-09-13 21:22:36.043837 Start train 0%| | 0/33165 [00:00<?, ?it/s] 2026-09-13 21:23:11.437206 Saving checkpoint 0%| | 26/33165 [02:20<38:07:05, 4.14s/it, loss=9.308, tps=18,246] That was running on perry, my normal workstation, and I kicked it off in parallel with the "curated" model training run below on poppy my training box, but I'll keep the runs separate for the purposes of this writeup. When this had been running for an hour or so, our power went out. My guess is that having the tumble dryer running, the car charging, the kettle boiling, the electric hob switched on, and two machines doing training runs is a bit too much for our electrics... which might be a problem in the future, especially if (as planned) I make poppy a multi-GPU machine. However, as things stand, I was able to kick it off again after switching the circuit breaker back on, and things held up. Again, about 40 hours later: Training complete in 136,060.457 seconds 2026-09-15 12:05:26.432638 Tokens seen: 3,227,516,928 2026-09-15 12:05:26.432642 Throughput: 23,721 tokens/second 2026-09-15 12:05:26.432650 Final train loss: 3.793 2026-09-15 12:05:26.432653 Done (Note that the numbers reported at the end of a restarted run like this only include what happened after the restart.) I converted it to PyTorch-compatible tensors, and did the smoke test: Every effort moves you on to other options—in fact, it’s not even worth that effort. Just make Looking good! Time for the loss test: giles@perry:~/Dev/ddp-base-model-from-scratch (main)$ uv run test_loss.py datasets/ ../jax-gpt2-from-scratch/runs/full-llm-full-train-with-mha-output-bias-fw-fwedu-5050/model.json ../jax-gpt2-from-scratch/runs/full-llm-full-train-with-mha-output-bias-fw-fwedu-5050/checkpoints/latest/pytorch-model.safetensors Fetching 4 files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 4/4 [00:00<00:00, 1192.07it/s] 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 3200/3200 [03:53<00:00, 13.72it/s] Loss against our test dataset: 3.462454 That was almost in keeping with my prediction that it would do worse than the JAX FineWeb-only models, except that it was better than the worst of those, "JAX, no MHA bias, with dropout": it was actually better than I predicted. So, a promising model. Time to upload it to Hugging Face -- and now let's move on to the next one. The "curated" dataset With my dataset-preparation script, this was easy enough to set up: { "seed": 42, "tokens_desired": 10000000000, "upload_dataset_name": "gpjt/fw-fwedu-simplewiki-gpt2-tokens", "sources": [ { "name": "FineWeb", "hf_id": "HuggingFaceFW/fineweb", "hf_name": "sample-10BT", "hf_split": "train", "item_field": "text", "weight": 45 }, { "name": "FineWeb-Edu", "hf_id": "HuggingFaceFW/fineweb-edu", "hf_name": "sample-10BT", "hf_split": "train", "item_field": "text", "weight": 45 }, { "name": "Simple English Wikipedia", "hf_id": "answerdotai/simplewiki", "hf_name": "articles", "hf_split": "train", "item_field": "md", "weight": 10 } ] } Running that worked nicely: giles@perry:~/Dev/prepare-llm-training-dataset (main)$ uv run prepare-dataset.py runs/fw-fwedu-simplewiki/ Resolving data files: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████| 27468/27468 [00:00<00:00, 90196.13it/s] Loading dataset shards: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████| 102/102 [00:00<00:00, 358.90it/s] Resolving data files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 2410/2410 [00:00<00:00, 88254.11it/s] Loading dataset shards: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████| 98/98 [00:00<00:00, 589.23it/s] 2026-09-13 18:59:04.106327: Generating dataset; per-source counts 2026-09-13 18:59:04.106387: FineWeb: 4,500,000,000 2026-09-13 18:59:04.106407: FineWeb-Edu: 4,500,000,000 2026-09-13 18:59:04.106422: Simple English Wikipedia: 1,000,000,000 FineWeb: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████▉| 4499997964/4500000000 [59:41<00:00, 1256362.56token/s] FineWeb-Edu: 4500000607token [59:41, 1256363.31token/s] Simple English Wikipedia: 1000002889token [59:41, 279192.58token/s] 2026-09-13 19:58:45.874744: Done generating tokens 2026-09-13 19:58:45.876043: FineWeb: 4,499,997,964 / 4,500,000,000 (1.000, 1 iterators) 2026-09-13 19:58:45.876048: FineWeb-Edu: 4,500,000,607 / 4,500,000,000 (1.000, 1 iterators) 2026-09-13 19:58:45.876052: Simple English Wikipedia: 1,000,002,889 / 1,000,000,000 (1.000, 6 iterators) 2026-09-13 19:58:45.876054: Total: 10,000,001,460 2026-09-13 19:58:45.876056: Catting... 2026-09-13 20:00:18.811748: Catted into a tensor of shape torch.Size([10000001460]) 2026-09-13 20:00:18.813169: Saving... 2026-09-13 20:00:22.773873: Saved 2026-09-13 20:00:22.773936: Uploading to gpjt/fw-fwedu-simplewiki-gpt2-tokens Processing Files (1 / 1) : 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 20.0GB / 20.0GB, 143MB/s New Data Upload : 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 19.8GB / 19.8GB, 142MB/s ...plewiki/train.safetensors: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 20.0GB / 20.0GB 2026-09-13 20:01:59.270021: Done One thing that is worth noting in that output is the "6 iterators" for the Simple English Wikipedia. If a source dataset runs out of items while we're building up the results in this script, we start iterating over it again (with a different seed for the shuffle so that the ordering is different). The "6 iterators" means that it needed to do that 6 times -- the original creation of the iterator at the start of the script, and five more. So that means that the Simple English Wikipedia is repeated (oversampled) somewhere between five and six times in the dataset. That's not a bad thing! From what I've read, it's actually quite standard to oversample highly educational content in LLM training datasets. And anyway, the dataset the script generated was 10B tokens, of which we're only using 3.2B for the training run in this post, so it would only appear somewhere between one and two times. The repetition would likely only really cut in if and when we did an overtrained model on the dataset. Anyway, I ran my check against the uploaded dataset -- the first few items were clearly from FineWeb, FineWeb-Edu, and the Simple English Wikipedia. It was time to kick off a training run: giles@poppy:~/Dev/jax-gpt2-from-scratch (main)$ XLA_PYTHON_CLIENT_MEM_FRACTION=0.95 uv run train.py full-llm-full-train-with-mha-output-bias-fw-fwedu-simplewiki datasets/ 2026-09-13 20:24:48.037024 Downloading dataset Downloading (incomplete total...): 0.00B [00:00, ?B/s] Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads. | 0/2 [00:00<?, ?it/s] WARNING:huggingface_hub.utils._http:Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads. Fetching 2 files: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 2/2 [02:51<00:00, 85.85s/it] Download complete: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████████| 20.0G/20.0G [02:51<00:00, 435MB/s] 2026-09-13 20:27:39.934884 Loading dataset into RAM Download complete: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████████| 20.0G/20.0G [02:51<00:00, 116MB/s] 2026-09-13 20:31:20.492877 Creating model 2026-09-13 20:31:24.054143 Creating optimizer 2026-09-13 20:31:25.100832 Start train 0%| | 0/33165 [00:00<?, ?it/s] 2026-09-13 20:32:29.650379 Saving checkpoint 0%|▎ | 107/33165 [08:38<39:05:39, 4.26s/it, loss=7.631, tps=20,293] Again, this was interrupted by the power outage that hit the 50:50 training run, but I was able to restart from a checkpoint. After another 22 hours, it crashed with an error that I've seen before: jax.errors.JaxRuntimeError: INTERNAL: CUDA error: Failed to end stream capture: CUDA_ERROR_STREAM_CAPTURE_INVALIDATED: operation failed due to a previous error during capture [executable_name='jit_train_step'] I put it aside as a one-off oddity when I hit it last time, but this time I dug in a bit more. I noted that it had not ever happened on perry, but seemed to be an issue on poppy, and that poppy had an older version of CUDA and the Nvidia drivers -- might that be the cause? I decided to upgrade those before kicking off the next run, but for now just restarted the run from the most recent checkpoint. (Note for anyone who is hitting the same error: it has not occurred since the upgrade, so that's worth trying.) This time it completed OK: Training complete in 59,564.515 seconds 2026-09-15 15:56:52.909888 Tokens seen: 1,367,212,032 2026-09-15 15:56:52.909894 Throughput: 22,953 tokens/second 2026-09-15 15:56:52.909912 Final train loss: 3.332 2026-09-15 15:56:52.909959 Done Again, these numbers just show what happened after the most recent restart. I copied it over to perry, converted it into a format that was compatible with my PyTorch code, and ran the smoke test: Every effort moves you by the air, for it will make you a better athlete, so your body becomes bigger and stronger Coherent enough -- time for the loss eval: giles@perry:~/Dev/ddp-base-model-from-scratch (main)$ uv run test_loss.py datasets/ ~/Dev/jax-gpt2-from-scratch/runs/full-llm-full-train-with-mha-output-bias-fw-fwedu-simplewiki/model.json ~/Dev/jax-gpt2-from-scratch/runs/full-llm-full-train-with-mha-output-bias-fw-fwedu-simplewiki/checkpoints/latest/pytorch-model.safetensors Fetching 4 files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 4/4 [00:00<00:00, 1007.64it/s] 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 3200/3200 [03:57<00:00, 13.48it/s] Loss against our test dataset: 3.542460 Again, in line with my predictions -- worse than the JAX FineWeb-only models, and indeed than the very best PyTorch one, 1xrtx3090-stacked-interventions, and also worse than the 50:50 split, but better than the FineWeb-Edu one. I uploaded it to Hugging Face, and it was time to move on to what was meant to be the final model for this set of experiments. The OpenWebText run Again, this was a simple enough config to set up: { "seed": 42, "tokens_desired": 10000000000, "upload_dataset_name": "gpjt/openwebtext-gpt2-tokens", "sources": [ { "name": "OpenWebText", "hf_id": "Skylion007/openwebtext", "hf_name": "plain_text", "hf_split": "train", "item_field": "text", "weight": 50 } ] } ...and the build and upload process worked well (and took much less time -- for some reason, sampling randomly from a single dataset is faster than sampling from two or three): giles@perry:~/Dev/prepare-llm-training-dataset (main)$ uv run prepare-dataset.py runs/openwebtext/ Resolving data files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████| 80/80 [00:00<00:00, 32723.26it/s] Resolving data files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████| 80/80 [00:00<00:00, 97940.55it/s] Loading dataset shards: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████| 80/80 [00:00<00:00, 1200.13it/s] 2026-09-15 13:16:47.622617: Generating dataset; per-source counts 2026-09-15 13:16:47.622645: OpenWebText: 10,000,000,000 Resolving data files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████| 80/80 [00:00<00:00, 45602.65it/s] Resolving data files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████| 80/80 [00:00<00:00, 67650.06it/s] Loading dataset shards: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████| 80/80 [00:00<00:00, 307.11it/s] OpenWebText: 10000000024token [31:46, 5246208.64token/s] 2026-09-15 13:48:33.761350: Done generating tokens 2026-09-15 13:48:33.762021: OpenWebText: 10,000,000,024 / 10,000,000,000 (1.000, 2 iterators) 2026-09-15 13:48:33.762026: Total: 10,000,000,024 2026-09-15 13:48:33.762028: Catting... 2026-09-15 13:49:33.115508: Catted into a tensor of shape torch.Size([10000000024]) 2026-09-15 13:49:33.115923: Saving... 2026-09-15 13:49:36.365978: Saved 2026-09-15 13:49:36.366027: Uploading to gpjt/openwebtext-gpt2-tokens Processing Files (0 / 1) : 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████▉| 20.0GB / 20.0GB, 147MB/s New Data Upload : 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 19.9GB / 19.9GB, 147MB/s ...webtext/train.safetensors: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████▉| 20.0GB / 20.0GB 2026-09-15 13:51:16.202890: Done Note that it needed to oversample -- that "2 iterators". OpenWebText is about 40 GiB uncompressed, and so that's about 10B GPT-2 tokens -- presumably just a little bit less. Again, given that I was planning to use just the first 3.2B tokens of the dataset, I didn't feel that it would matter. I ran the check script on the newly-uploaded Hugging Face dataset and all looked well, so that was all set for the training run. I upgraded poppy first with a sudo pacman -Syu to see if that helped with the weird error that I got in the previous run (which, as I said, it looks like it did), then kicked it off: giles@poppy:~/Dev/jax-gpt2-from-scratch (main)$ XLA_PYTHON_CLIENT_MEM_FRACTION=0.95 uv run train.py full-llm-full-train-with-mha-output-bias-openwebtext datasets/ 2026-09-15 16:42:32.606185 Downloading dataset Fetching 2 files: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 2/2 [00:00<00:00, 941.38it/s] Download complete: : 0.00B [00:00, ?B/s] | 0/2 [00:00<?, ?it/s] 2026-09-15 16:42:32.879987 Loading dataset into RAM Download complete: : 0.00B [00:00, ?B/s] 2026-09-15 16:45:40.438791 Creating model 2026-09-15 16:45:43.840269 Creating optimizer 2026-09-15 16:45:44.848351 Start train 0%| | 0/33165 [00:00<?, ?it/s] 2026-09-15 16:46:50.632075 Saving checkpoint 1%|█ | 332/33165 [24:33<38:45:54, 4.25s/it, loss=6.623, tps=22,154] About 31 hours in, it crashed again, but this time it was my own dumb fault: poppy has a relatively small disk and I ran out of space. I fixed that and kicked it off again from the most recent checkpoint, and this time it completed: Training complete in 33,927.995 seconds 2026-09-17 11:25:10.835989 Tokens seen: 779,747,328 2026-09-17 11:25:10.835994 Throughput: 22,982 tokens/second 2026-09-17 11:25:10.836012 Final train loss: 3.165 2026-09-17 11:25:10.836018 Done I converted it to PyTorch for the smoke test: Every effort moves you through each phase, so it's not a complete picture. I'm sure your story was ...which looked solid, so it was time for the test loss eval: giles@perry:~/Dev/ddp-base-model-from-scratch (main)$ uv run test_loss.py datasets/ ~/Dev/jax-gpt2-from-scratch/runs/full-llm-full-train-with-mha-output-bias-openwebtext/model.json ~/Dev/jax-gpt2-from-scratch/runs/full-llm-full-train-with-mha-output-bias-openwebtext/checkpoints/latest/pytorch-model.safetensors Fetching 4 files: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 4/4 [00:00<00:00, 674.76it/s] 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 3200/3200 [03:59<00:00, 13.37it/s] Loss against our test dataset: 4.045255 Our worst score yet in this experiment! Worse than any of my models so far, apart from the two FineWeb-Edu ones I did without optimised hyperparameters. Now, the first draft of this post went straight to the results from here, but the story wasn't quite over yet... Test set contamination GPT-6 Astra is relentless. Before I publish any of these posts, I run them past an editorial board of LLMs to look for issues. GPT-6 Astra not only checked the text, it also visited the code I'd linked to to check that out too, and spotted something problematic. It's obvious in retrospect, but my code to build the new datasets had a high risk of including the contents of the -- in theory held-back -- test set. The way that the test set was generated was that I downloaded the 10B sample of FineWeb back in December, splitting it into 99% training data and 1% "validation". That validation split was about 100M tokens, and I was only using the first 19M or so for actual validation runs during training, so I (somewhat arbitrarily) designated about 19M other tokens starting at position 50M in there as my test set. Now, my new dataset-generation code was just sampling randomly from the complete 10B sample of FineWeb. So there was nothing stopping it from pulling in data that was in that old validation split! That meant that it was quite likely that my new "curated" and "50:50" datasets contained at least some of the test set that was meant to have been held back from the models during training. On reflection, the problem was potentially even worse. FineWeb-Edu is a subset of FineWeb; my existing FineWeb-Edu dataset came from the 10B sample of the Hugging Face original, and so it also could potentially contain documents that I'd put into the test set. The first thing to do was to establish the size of the problem. I wrote a script to take in a "forbidden" dataset and split; this was assumed to be formatted as one big tensor of GPT-2 tokens, which is what all of my datasets are. It would then split it by end-of-text tokens, and generate a hash and a token count for each resulting "document". Optionally, you could restrict it to only considering a subset -- the n tokens starting at position p -- and it would then generate hashes/lengths for the documents inside that slice, or that overlapped it at the start or the end. I ran that to generate a list of hashes for the entire validation set -- the validation split of gpjt/fineweb-gpt2-tokens -- and then used a second script to check my various training sets (and the validation set itself) to see how much of a contamination problem there was. I got these results: Dataset Split Contamination with validation set gpjt/fineweb-gpt2-tokens validation 102163003 out of 102163003 tokens (100.00%) gpjt/fineweb-gpt2-tokens train 636166 out of 102163003 tokens (0.62%) gpjt/fineweb-edu-gpt2-tokens train 672189 out of 102163003 tokens (0.66%) gpjt/fw-fwedu-5050-gpt2-tokens train 49224580 out of 102163003 tokens (48.18%) gpjt/fw-fwedu-simplewiki-gpt2-tokens train 44233824 out of 102163003 tokens (43.30%) gpjt/openwebtext-gpt2-tokens train 212 out of 102163003 tokens (0.00%) So: The validation set was 100% "contaminated" with itself, which was a useful sanity check. The training set of gpjt/fineweb-gpt2-tokens had what I felt was a small level of contamination. It was interesting that there was any at all -- I think that must mean that there are some repeated documents in the original dataset, and some of them wound up with copies in both my training and validation splits. The gpjt/fineweb-edu-gpt2-tokens dataset also had what felt like a reassuringly low level of contamination. Both gpjt/fw-fwedu-5050-gpt2-tokens and gpjt/fw-fwedu-simplewiki-gpt2-tokens, however, looked problematic. In both cases, the training datasets had more than 40% of the validation/test set in them. gpjt/openwebtext-gpt2-tokens was, as you'd expect, almost completely uncontaminated. It looks like maybe one document happened to have been picked up by both the OpenWebText and the FineWeb crawls and then included in the bit of FineWeb I was using for validation. However, these numbers -- while scary, at least for the 50:50 and the curated datasets -- were not quite the ones to use. They showed how much of the full validation set showed up in the full training set; what I actually cared about was how much of the test set -- those 19M tokens starting at position 50M in the validation split -- was in the actual subset of the training datasets that I actually trained on -- the first ~3.2B of them. I re-ran the script to generate hashes for just the test set, and then re-ran the contamination-checking script, telling it just to look at the appropriate subset of the training tokens, and got this: Dataset (first 3.2B tokens only) Split Contamination with test set gpjt/fineweb-gpt2-tokens train 26557 out of 19632681 tokens (0.14%) gpjt/fineweb-edu-gpt2-tokens train 32079 out of 19632681 tokens (0.16%) gpjt/fw-fwedu-5050-gpt2-tokens train 2986889 out of 19632681 tokens (15.21%) gpjt/fw-fwedu-simplewiki-gpt2-tokens train 2682430 out of 19632681 tokens (13.66%) gpjt/openwebtext-gpt2-tokens train None It was clear that there was a problem -- certainly with gpjt/fw-fwedu-5050-gpt2-tokens and gpjt/fw-fwedu-simplewiki-gpt2-tokens. They'd seen what felt like a significant amount of the test set while training, so their results on the test loss eval were dubious at best. I decided to train those two models afresh, and see what the result was in terms of loss. If the difference was huge, I'd look into the risks of the (much smaller) contamination of gpjt/fineweb-gpt2-tokens and gpjt/fineweb-edu-gpt2-tokens. But if it was pretty small, I'd not worry about that too much. I extended the script that prepared datasets so that the config file could specify a forbidden_dataset. Any documents in the source datasets that matched forbidden ones would be excluded from the output. I then updated the config for gpjt/fw-fwedu-5050-gpt2-tokens and gpjt/fw-fwedu-simplewiki-gpt2-tokens so that the whole validation split of gpjt/fineweb-gpt2-tokens was forbidden, and re-generated them. You can see the updated datasets here and here. Running the contamination-checker script against them showed that they were clear. I then re-did the full training runs for those models; the uncontaminated version of the 50:50 split model is here, and the curated one is here. And the good news: both of them actually did very slightly better at the test loss eval than their equivalents that had been trained on the contaminated data: Model Contaminated Test loss JAX, FineWeb/FineWeb-Edu 50:50 No 3.449257 JAX, FineWeb/FineWeb-Edu 50:50 Yes 3.462454 JAX, curated No 3.534068 JAX, curated Yes 3.542460 There are a number of possibilities that come to mind; perhaps learning from the test set just doesn't happen with tiny 163M models like this, or perhaps while the contaminated models were learning, the benefit they got from that was outweighed by the data that they got instead of the test set data being in some way better for training purposes, at least in terms of the loss eval. But anyway, I felt that if the effect of seeing more than 10% of the test set data during training was so tiny, then the effect of seeing less than 0.2% -- which is what the FineWeb-Edu model in this set of training runs had, as did all of my other FineWeb-only models from previous experiments -- would be even smaller and I'd disregard it. That was excellent news! I didn't need to start all of my experiments from scratch. For the rest of this post, I will include the numbers and results for the contaminated models as well as the uncontaminated ones -- they're interesting for several reasons -- but for future posts I'll skip the contaminated ones. So -- finally! -- let's start digging into the final results. Results Firstly, I think it's worth taking a look at all of the test loss results in context. Here they are in a table, with the new models in bold: Test loss OpenAI weights: medium 3.231442 JAX, overtrained one long epoch 3.324953 JAX, overtrained two normal epochs 3.326482 JAX, with MHA bias, no dropout 3.418784 JAX, no MHA bias, no dropout 3.420089 JAX, FineWeb/FineWeb-Edu 50:50 (uncontaminated) 3.449257 JAX, FineWeb/FineWeb-Edu 50:50 (contaminated) 3.462454 JAX, no MHA bias, with dropout 3.476802 OpenAI weights: small 3.499677 JAX, curated (uncontaminated) 3.534068 1xrtx3090-stacked-interventions 3.538161 JAX, curated (contaminated) 3.542460 8xa100m40-stacked-interventions-1 3.577761 JAX, FineWeb-Edu 3.632900 Cloud FineWeb, 8x A100 40 GiB 3.673623 1xrtx3090-baseline 3.683835 8xa100m40-baseline 3.691526 Cloud FineWeb, 8x H100 80 GiB 3.724507 Cloud FineWeb, 8x A100 80 GiB 3.729900 Cloud FineWeb, 8x B200 160 GiB 3.771478 Local FineWeb train 3.943522 JAX, openwebtext 4.045255 Local FineWeb-Edu extended train 4.134991 Local FineWeb-Edu train 4.166892 I think there's something very clear here: with the new models, the more FineWeb that was in the training mix, the better the model did on this eval. I think I might have been subconsciously expecting that in the predictions I did before running these experiments, but in retrospect it's so incredibly obvious that I feel silly for not mentioning it explicitly! But that tells us something interesting. From the description in the paper, whatever OpenAI did the GPT-2 training run on, it was not like FineWeb. It was probably more similar to OpenWebText -- and yet, that model was the one that performed the worst on this test eval, so if it is more like OpenWebText, there must be some other factor involved. But moving on for now: how about the IFT test -- the one that kicked off all of this work in the first place? I generated a set of IFT responses for all of the new models, and then ran them (plus responses for all of the other models on that table above) past GPT 5.5, and found that one of my new models was getting quite close to the original GPT-2 small weights! So I did four more runs, so that I could get an average. Here are the results -- the "IFT score" is the average across all five runs of the judge, and the "IFT rank" is based on that. The "IFT epochs" was from the original result-generation script. Test loss IFT epochs IFT score IFT rank OpenAI weights: medium 3.231442 2 42.36 1 JAX, overtrained one long epoch 3.324953 3 18.67 7 JAX, overtrained two normal epochs 3.326482 4 18.71 6 JAX, with MHA bias, no dropout 3.418784 4 17.90 8 JAX, no MHA bias, no dropout 3.420089 5 20.50 4 JAX, FineWeb/FineWeb-Edu 50:50 (uncontaminated) 3.449257 4 17.69 9 JAX, FineWeb/FineWeb-Edu 50:50 (contaminated) 3.462454 4 19.30 5 JAX, no MHA bias, with dropout 3.476802 5 13.02 21 OpenAI weights: small 3.499677 2 25.19 2 JAX, curated (uncontaminated) 3.534068 4 16.63 10 1xrtx3090-stacked-interventions 3.538161 4 13.51 19 JAX, curated (contaminated) 3.542460 4 13.58 18 8xa100m40-stacked-interventions-1 3.577761 4 10.19 24 JAX, FineWeb-Edu 3.632900 4 24.56 3 Cloud FineWeb, 8x A100 40 GiB 3.673623 3 16.59 11 1xrtx3090-baseline 3.683835 4 15.15 12 8xa100m40-baseline 3.691526 3 13.64 16 Cloud FineWeb, 8x H100 80 GiB 3.724507 4 13.59 17 Cloud FineWeb, 8x A100 80 GiB 3.729900 3 10.79 23 Cloud FineWeb, 8x B200 160 GiB 3.771478 4 13.70 15 Local FineWeb train 3.943522 5 11.87 22 JAX, openwebtext 4.045255 4 13.28 20 Local FineWeb-Edu extended train 4.134991 5 14.29 14 Local FineWeb-Edu train 4.166892 5 14.69 13 If you want to see the full numbers, they're below. The number that initially surprised me, and made me decide to do multiple LLM-judge runs was the one for the "JAX, FineWeb-Edu" model. In my first run it came in at 24.35 vs the OpenAI small weights' 24.93 -- so close that I wondered if it might even beat them on a re-run. However, in the further four runs its score was consistently lower than the OpenAI model's, and the gap extended a bit in some. So, was FineWeb-Edu the clear winner here? Perhaps. If you look at the contaminated/uncontaminated pairs, something interesting pops out. For the 50:50 mix, the model trained with the contaminated dataset got 19.30, and the one trained on the uncontaminated one got 17.69 -- a difference of 1.61. For the "curated" dataset, the situation was even more interesting: uncontaminated got 16.63, while contaminated got 13.58, a delta of 3.05 points. Remember, the contamination issue is about whether or not the model saw the held-back test set during training. It was an issue for the test loss that is based on that test set, but is entirely orthogonal to the IFT test. From the IFT perspective, both contaminated and uncontaminated models in each case saw training data that was -- in theory, at least -- essentially the same in terms of quality. Indeed, the uncontaminated run saw almost the same data in the same order as the contaminated one, except that some items were omitted, and then extra ones were added to the end. The purpose of this set of experiments was to see how data quality affected the results on the IFT test set. But in the case of the curated model, something that should be unrelated to data quality changed the results by 3.05 points! If something as simple as changing which data of the same quality the model is trained with can affect the IFT score so drastically, it makes it a bit harder to be certain as to whether or not data quality really had the effect we were looking for. On the other hand, the FineWeb-Edu model came in at 24.56, which is 4.06 points better than the 20.50 that the closest other model got -- more than the 3.05 points we see in difference between the two curated dataset models. And it's worth noting that the model with 20.50 is "JAX, no MHA bias, no dropout", which has a subtly different architecture -- no bias on the output projection of the multi-head attention blocks. A better comparison might be "JAX, with MHA bias, no dropout", which got a score of 17.90, for a whacking great difference of 6.66 points. I think that without doing a very large number of training runs on different datasets with different mixes, each one created with a different seed, it would be hard to work out exactly what is in the noise here and what is not. However, that would cost a lot in terms of time. I think that the best thing here is to chalk this up as a fairly decent indication that FineWeb-Edu improves matters for the IFT eval, but far from a certainty. But it's certainly worth noting that whatever the noise is, it has a range of at least 3.05 points -- and the FineWeb-Edu model is just 0.63 points short of GPT-2 small! So there could well be something there. Of course, we don't know whether that model got (by chance) the best possible balance of FineWeb-Edu tokens, and could never win -- or whether it got a bad balance and would actually beat GPT-2 with a better one. So that's certainly worth keeping in mind. As an aside, the result for the curated dataset really surprised me. I had expected that it would be the best one, simply because it almost certainly contained more facts. I took a look at its answers to the questions -- one possibility that came to mind might be that it would get better responses to questions like "What is the chemical symbol for chlorine" or "Who wrote Pride and Prejudice" than the others, but would fail on less knowledge-based tasks. But it was terrible at fact-based questions too: Name the author of 'Pride and Prejudice'. What is the periodic symbol for chlorine? As I understand it, many real-world training runs do include (often oversampled) amounts of highly educational training data like this model's dataset did. But perhaps the models that I'm training are just too small to be able to make use of the data they gained that way -- maybe doing things this way and expecting good results is like asking six-year-old children to memorise stuff before they've learned enough to be able to make use of it 1. It's worth noting that the GPT-2 small model also failed on those factual questions. Well, anyway: I think we have some useful results here, so let's work out what that means for next steps. Conclusion The results we got in these experiments point in two interesting directions. The perfect connection between the amount of FineWeb in the training set and the result on the (FineWeb-based) test loss eval, while perfectly obvious in retrospect, really does highlight how mysterious it is that the OpenAI small weights do so well on that test. The fact that FineWeb-Edu did well on the IFT test tells us that there does seem to be value in using richer training data -- though the less-spectacular results of the 50:50 mix and the curated one weaken that a bit, as does the indicator of what the noise due to data selection from equivalently high-quality datasets might be. The OpenWebText result I think I'll ignore, given that -- while in theory it should be similar to what OpenAI trained on -- there are no guarantees, and it might differ in non-obvious ways for non-obvious reasons. I think that the right direction to take this going forward is to separate these two angles. I should chase a higher IFT score, and then once I have nailed that down, I should see what (if anything) might allow me to get the resulting model to improve its test score. But I will need to make sure that whatever dataset I use, I use various "mixes" of it -- versions created with different random seeds. In my earlier experiments with overtraining, I did find that it didn't seem to improve the IFT results -- but it did improve the test loss. So perhaps identifying the right combination of other factors to boost the IFT score, then overtraining the result, might help? Of course, my overtraining tests were with FineWeb, so the connection might not hold up as well if the starting model (as seems likely) was trained on a different dataset. Also, while working through the results here, I've come to the conclusion that the set of models I'm using is a bit confusing -- there are now different hyperparameter settings, small architectural differences (the MHA bias thing), dropout settings during the pre-training, and now datasets. I think that's OK for now; I should see this part of this series as more ideation than actually running the proper experiments. But at the end, when I have some solid hypotheses with a reasonable amount of backup, I should start from scratch: a baseline model, then staged interventions to build up to what (hopefully) will be a model as good as GPT-2 small. Anyway, I'll wrap this one up here. I think that the next lever to pull is (perhaps surprisingly) going to be weight tying. I had previously kind of disregarded that as a possibility, but while I was working on this post, something popped into my mind. The OpenAI models were originally trained with weight tying. My codebase does actually support doing it -- but because I got the OpenAI weights I'm using from the code in "Build a Large Language Model (from Scratch)", when I'm running the IFT test, the weights are not actually tied! We load up a model that has separate but identical embedding and output head matrices, and then we fine-tune that. So those two matrices can vary independently during fine-tuning -- to put it another way, while GPT-2 small was pre-trained with 124M parameters, the IFT test is being done on a 163M-parameter version. Does that give them some non-obvious advantage? And would adding weight-tying to my own models help, either with or without the output heads being independent at fine-tuning time? Stay tuned :-) Appendix: all IFT judge runs Here are the numbers for all of the IFT judge runs, included for completeness. You can see that the LLM judge ranks models very consistently between runs, but there is variation -- that is, on some runs it's in what I think of as a "better mood" than others, and if that's the case, it will give better scores -- but it will give them almost consistently between models, so all of the models do better. Note that (unlike the table above) this one is sorted by the average IFT score rather than the test loss. Model Run 1 Run 2 Run 3 Run 4 Run 5 Average OpenAI weights: medium 42.24 42.16 42.95 41.83 42.61 42.36 OpenAI weights: small 24.93 24.96 25.39 25.01 25.66 25.19 JAX, FineWeb-Edu 24.35 24.55 24.3 24.68 24.9 24.56 JAX, no MHA bias, no dropout 20.5 19.9 20.76 21.25 20.07 20.50 JAX, FineWeb/FineWeb-Edu 50:50 (contaminated) 19.16 18.86 19.61 19.17 19.7 19.30 JAX, overtrained two normal epochs 18.47 18.29 19.17 18.69 18.91 18.71 JAX, overtrained one long epoch 18.04 18.71 19.62 18.41 18.57 18.67 JAX, with MHA bias, no dropout 17.49 17.35 18.33 17.73 18.62 17.90 JAX, FineWeb/FineWeb-Edu 50:50 (uncontaminated) 17.37 17.73 17.53 18.01 17.83 17.69 JAX, curated (uncontaminated) 16.77 16.03 17.3 16.08 16.96 16.63 Cloud FineWeb, 8x A100 40 GiB 16.44 16.23 17.14 16.62 16.54 16.59 1xrtx3090-baseline 14.85 15.07 15.19 15.14 15.51 15.15 Local FineWeb-Edu train 14.37 14.23 15.08 14.79 15 14.69 Local FineWeb-Edu extended train 14.4 14.07 13.82 14.56 14.61 14.29 Cloud FineWeb, 8x B200 160 GiB 13.37 13.05 13.85 13.67 14.57 13.70 8xa100m40-baseline 13.64 13.36 13.9 13.32 13.97 13.64 Cloud FineWeb, 8x H100 80 GiB 13.45 13.32 13.6 13.51 14.07 13.59 JAX, curated (contaminated) 13.09 13.48 13.95 13.23 14.15 13.58 1xrtx3090-stacked-interventions 13.37 13.11 14.04 13.84 13.17 13.51 JAX, openwebtext 12.88 12.7 13.74 13.53 13.53 13.28 JAX, no MHA bias, with dropout 13.19 12.86 12.98 12.85 13.24 13.02 Local FineWeb train 11.75 11.75 12.21 11.46 12.19 11.87 Cloud FineWeb, 8x A100 80 GiB 10.68 10.2 11.03 10.55 11.49 10.79 8xa100m40-stacked-interventions-1 9.44 9.79 10.84 10.2 10.66 10.19 A small boy asleep on his right side, the right arm stuck out, the right hand hanging limp over the edge of the bed. Through a round grating in the side of a box a voice speaks softly. "The Nile is the longest river in Africa and the second in length of all the rivers of the globe. Although falling short of the length of the Mississippi-Missouri, the Nile is at the head of all rivers as regards the length of its basin, which extends through 35 degrees of latitude …" At breakfast the next morning, "Tommy," some one says, "do you know which is the longest river in Africa?" A shaking of the head. "But don't you remember something that begins: The Nile is the …" "The - Nile - is - the - longest - river - in - Africa - and - the - second - in - length - of - all - the - rivers - of - the - globe …" The words come rushing out. "Although - falling - short - of …" "Well now, which is the longest river in Africa?" The eyes are blank. "I don't know." "But the Nile, Tommy." "The - Nile - is - the - longest - river - in - Africa - and - second …" "Then which river is the longest, Tommy?" Tommy burst into tears. "I don't know," he howls. Brave New World, Aldous Huxley ↩
Today's links Voting is to politics as shopping is to boycotts: The big P only matters if the small p is in play. Hey look at this: Delights to delectate. Object permanence: Gilberto Gil v WIPO; Censored Apple wifi hacker talk; Wells Fargo crime-spree started in 1998; Stencils "may not be reproduced"; DVD Jon v Apple DRM; Unpaid diplomatic parking tickets as index of corruption; Tortured Canadian was not a terrorist; Decarbonization at a distance. Upcoming appearances: Brighton, Virtual, South Bend, Hudson, Calgary, Winnipeg, Paris, Vancouver, Victoria, Ottawa, Kilkenny, Montreal. Recent appearances: Where I've been. Latest books: You keep readin' em, I'll keep writin' 'em. Upcoming books: Like I said, I'll keep writin' 'em. Colophon: All the rest. Voting is to politics as shopping is to boycotts (permalink) Here's a funny thing about the right to vote: it wasn't won by voting. From the Magna Carta to the US Constitution to the Emancipation Proclamation to 19th Amendment, voting rights (what you might call "Big P" Politics) were always downstream of protests, riots, petitions, mass movements, strikes and good, old fashioned community organizing (that is, "small p" politics). Which is to say, Big P politics matter, but to make them matter, we need a lot of small p politics. That means that democracy isn't something you do every couple of years with a ballot paper (though that's an important aspect of the process). Democracy is continuous. If you've ever wondered why your vote seems to accomplish so little, I think you can blame the near-abolition of small p politics by Big P politicians of every stripe. Indeed, Obama's genius was summoning up an army of door-knocking, phone-banking small p political activists and then euthanizing that organization after he won the election: https://newrepublic.com/article/140245/obamas-lost-army-inside-fall-grassroots-machine For Obama, the grassroots were useful for one thing: getting out the vote. The last thing he wanted was for millions of activated voters to turn into activists who'd flame him and harangue him and picket him if they didn't like his compromises. Boy, did Obama ever compromise. He let the bank executives who created the Great Financial Crisis off the hook and encouraged them to foreclose on the homes of millions of Americans, the very same public that had bailed them out: https://theweek.com/articles/624777/obamas-biggest-failure He shielded the CIA's torturers from scrutiny and prosecution: https://journals.law.harvard.edu/ilj/2009/04/obama-publishes-torture-memos-immunizes-cia-staff/ He reneged on his promise to shut down Gitmo: https://www.pbs.org/newshour/show/obama-failed-close-guantanamo And his promise to hold the phone companies to account for their complicity in the NSA's mass domestic surveillance: https://www.pbs.org/wgbh/frontline/article/obama-on-mass-government-surveillance-then-and-now/ He stepped up secret drone warfare: https://www.cfr.org/articles/obamas-final-drone-strike-data And unconstitutional domestic surveillance: https://www.eff.org/deeplinks/2017/01/obama-expands-surveillance-powers-his-way-out Whenever I raise this, Obama's apologists come out of the woodwork to tell me that "the president isn't the Green Lantern," and that Obama couldn't act without help from Congress and the Senate, who wouldn't back his plays. I think that Trump's presidency has shown us how much power the president really has even when the legislature won't play ball. But even if you accept the Green Lantern apologetics, the fact remains that Obama could have had a clamoring army of ardent supporters in the streets, defending his agenda against recalcitrants in his own party and wreckers in the GOP. He chose not to have that army. He sent that army home. It's like Obama heard the story about post-election FDR telling civil rights leaders, "I want to do it, now make me do it," and concluded, "I don't want to do it, so I'd better not let anyone make me do it": https://www.quora.com/Did-Franklin-Roosevelt-ever-say-I-agree-with-you-I-want-to-do-it-now-make-me-do-it Of course, Trump is doing everything he can to extinguish both small p politics and Big P Politics. It's not just his wildly illegal voter suppression tactics. He's banning and prosecuting political groups, invoking anti-terror laws (which Obama supported and promised would only be used proportionately and wisely) to chase his grassroots opposition underground: https://www.whitehouse.gov/presidential-actions/2025/09/designating-antifa-as-a-domestic-terrorist-organization/ Liberals are often contemptuous of grassroots movements (cf "basket of deplorables," "Green Lantern" scolding), but the right is terrified of them. The right's political leadership is terrified of its own grassroots, and rightly so, because those people are maniacs, and they're the reason the GOP has been pushed into its most extreme positions. The right's grassroots, meanwhile, are afraid of the left's grassroots. The last thing they want is a militant, organized, mobilized base pushing Dem politicians to take the stands that are wildly and widely popular in America, from Medicare for All to an end to ICE – the Mamdani agenda, in other words. Mamdani is the anti-Obama. He shows what happens when a progressive candidate nurtures and co-governs with their base after the election, using millions of passionate, committed, everyday people to steamroller anyone who gets in the way of his agenda: https://www.nyc.gov/content/100days/pages/ Of course the downside of this is that when Mamdani reneges on his pledges, he is loudly and furiously held to account for it: https://www.thecityreporter.nyc/2026/02/19/mamdani-budget-parks-libraries/ Mamdani understood that he would be corralled into compromises if he won the mayoralty and that when he made those compromises, his base would come after him with the unmistakable fury of betrayed idealists. He also understood that any comfort he enjoyed by sidelining his base while in office would come at a price far higher than being yelled at by his supporters: it would cost him the ability to get anything done. Voting for Mamdani was important. It got him elected. But staying organized – in unions, neighborhood clubs, affinity groups, DSA chapters and mutual aid groups – is what's letting him get stuff done, and stopping him from bailing on his promises as politically infeasible. In other words, voting only matters if it's the final stage of a sustained campaign to build and mobilize popular power. Without that, voting will get you precious little. The right's leadership understands this very well, which is why they've spent years attacking unions, community organizers like Acorn, and activist institutions like Planned Parenthood. We must defend voting rights – Big P Politics – to the bitter end, but we need to defend organizing – small p politics – just as ferociously. The reduction of politics to voting is part of the 50 year neoliberal project whose foremost goal is to make you think of yourself as an atomized individual and not as a member of a polity. Turning "politics" into "voting" is absolutely in line with Margaret Thatcher's dictum that "there is no such thing as society." It's the same move that convinced workers that the answer to bad working conditions is looking your boss in the eye and threatening to change jobs (not forming a union and striking). It's also the same move that transformed "boycotts" into "shopping." Boycotts are a collective enterprise. Before a boycott takes place, small-p political groups hold meetings, organize alternatives and communicate their demands. During a boycott, organizers work to insulate participants from reprisals, like the Montgomery Bus Boycott organizers who reasoned and remonstrated with employers who disciplined workers whose participation made them late for work. And yes, as part of a boycott, you make some consumption choices. You buy X instead of Y. But "shopping" by itself isn't a boycott. You can't "vote with your wallet" (especially not when billionaires get to vote against you with their wallets): https://pluralistic.net/2025/09/13/consumption-choices/#marginal-benefits Shopping isn't politics, and while voting is Politics (Big P), it's also not politics (small p). A boycott, on the other hand, is politics. What's more, "shopping" has the same relationship to "boycotts" that "voting" has to "politics." It's a step you take, after you've laid a lot of groundwork with other people, as part of a mass movement. I understand why shopping and voting are more attractive than boycotts and politics. Meetings suck. Hell is other people: https://locusmag.com/feature/commentary-cory-doctorow-hell-is-other-people/ But changing the system requires systemic work. Hell is other people because other people are great but it's so hard to get them to do things your way. That takes time and understanding and togetherness and arguing and forgiving. Not everyone has time or capacity for that, and at any given time, we don't all have to be doing that work. We can take turns, spelling each other off at times in our lives when we have more or less slack. But lots of us have to be in the fight, or all of us will get screwed. There aren't enough of us doing politics right now. We can tell, because our politicians are so contemptuous of the grassroots that they will sell us out without a moment's hesitation, smugly certain that they will face no consequences for doing so: https://pluralistic.net/2026/09/22/happy-chudmas/#baloney-in-our-slacks Oligarchs have it easy. Where we have to convince people to fight, they can pay or threaten people to bring them into line. But oligarchs' power is wearing thin. The data-center uprising shows how much fury there is out there, looking for a productive outlet: https://www.bloodinthemachine.com/p/with-the-backlash-to-data-centers Data centers are very bad and very visible, so they make for good targets. But data centers are only the physical extrusion of a vast, brutal, extractive system. The most important way to fight data centers is to take everyone you meet protesting one and organize with them to scare the shit out of "your" politicians so they don't dare compromise on anything. Hey look at this (permalink) The Facebook Fake-out https://www.anildash.com/2026/09/29/facebook-fake-out/ Anatomy Unzipped: John of Arderne’s Sweden Scroll (ca. 1425–35) https://publicdomainreview.org/collection/arderne-scroll/ Inside McDonald’s push to have AI price your Big Mac https://www.reuters.com/business/inside-mcdonalds-push-have-ai-price-your-big-mac-2026-09-29/ what is going on with ceiling fans https://mcmansionhell.com/post/829127919552151552/what-is-going-on-with-ceiling-fans From Shitpost to Bullshit https://www.unpopularfront.news/p/from-shitpost-to-bullshit Object permanence (permalink) #25yrsago GWB's press secretary to media: "watch what you do, watch what you say" https://web.archive.org/web/20010926223602/https://www.whitehouse.gov/news/releases/2001/09/20010926-5.html#BillMaher-Comments#BillMaher-Comments #20yrsago Stencils kit “may not be reproduced in any form” https://web.archive.org/web/20061022000842/http://www.fairuseday.com/index.php/2006/10/01/copyright-is-broken/ #20yrsago DVD Jon selling Apple DRM to Apple’s competitors https://web.archive.org/web/20061004191106/https://featured.gigaom.com/2006/10/02/dvd-jon-fairplays-apple/ #20yrsago Unpaid diplomatic parking tickets as index of national corruption https://web.archive.org/web/20130719065306/https://www.theatlantic.com/magazine/archive/2006/10/primary-sources/305203/ #20yrsago Canadian deported to Syria for torture is cleared https://www.theguardian.com/world/2006/oct/02/worlddispatch #20yrsago Gilberto Gil slams WIPO https://fromgeneva.blogspot.com/2006/09/wipo-general-assembly-impressions-from.html #20yrsago Speech given by censored Apple WiFi hacker at ToorCon https://craphound.com/cache_toorcon_2006.txt #10yrsago Company suspected of blame in Office of Personnel Management breach will help run new clearance agency https://www.reuters.com/article/us-usa-security-background-idUSKCN1202M6/ #10yrsago Wells Fargo started demanding fraud of its employees in 1998; Illinois cuts Wells off from state business https://www.citizen.org/wp-content/uploads/wells-fargo-king-of-cross-sell.pdf #10yrsago Google: if you support Amazon’s Echo, you’re cut off from Google Home and Chromecast https://variety.com/2016/digital/news/google-home-amazon-echo-chromecast-1201874125/ #5yrsago How the IMF loan-sharks the global south https://pluralistic.net/2021/10/02/debt-trap/#global-arm-breakers #1yrago Decarbonization at a distance https://pluralistic.net/2025/10/02/there-goes-the-sun/#carbon-shifting Upcoming appearances (permalink) https://www.epl.ca/blogs/post/elbows-up-with-cory-doctorow/ Brighton: Digital Sovereignty and the Post-American Internet (Green Party Conference), Oct 3 https://www.openrightsgroup.org/events/digital-sovereignty-and-the-post-american-internet/ Virtual: How to govern technology in a multipolar digital world (Connecting Current), Oct 6 https://connectingcurrent.tech/how-to-govern-technology-a-multipolar-digital-world/ South Bend: An Evening With Cory Doctorow (Notre Dame), Oct 6 https://franco.nd.edu/events/2026/10/06/an-evening-with-cory-doctorow/ Hudson, OH: Hudson Library, Oct 7 https://engagedpatrons.org/EventsExtended.cfm?SiteID=3850&EventID=596952&PK= Calgary: Wordfest, Oct 8 https://wordfest.com/2026/show/wordfest-presents-cory-doctorow-2026/ Winnipeg: McNally Robinson, Oct 9 https://www.mcnallyrobinson.com/event-18991/An-Evening-with-Cory-Doctorow Paris: Slow Tech Summit, Oct 15 https://slowtechsummit.com/ Vancouver: Read, Resist, Repair, Rejoice (Vancouver Writers Festival), Oct 19 https://writersfest.bc.ca/festival-event-2026/01 Victoria: Munro's Books, Oct 20 https://www.munrobooks.com/events/6113620261020 Vancouver: Life After AI (Vancouver Writers Festival), Oct 22 https://writersfest.bc.ca/festival-event-2026/46 Ottawa: Life After AI (Ottawa Writers Festival), Oct 24 https://writersfestival.org/event/life-after-ai Kilkenny (Kilkenomics), Nov 6-8 https://kilkenomics.com/ Vancouver: Enshittification (Sid Williams Theatre Society), Nov 10 https://www.sidwilliamstheatre.com/events/cory-doctorow-talks-enshittification/ Vancouver: BC Policy Solutions Gala, Nov 12 https://bcpolicy.ca/gala/ Montreal: World Science Fiction Convention, Sep 2-6 https://montreal2027.ca/en Recent appearances (permalink) Terms of Service with Clare Duffy (CNN) https://www.cnn.com/audio/podcasts/terms-of-service-with-clare-duffy/episodes/458ce968-af5d-11f0-b539-13ed2afe25f8 AI, Work, and Power (Software Engineering Daily) AI, Work, and Power https://softwareengineeringdaily.com/podcasts/cory-doctorow-on-ai-work-and-power/ AI, Corporate Power, and the Fight for Worker Control (Plutopia) https://plutopia.io/cory-doctorow-ai-corporate-power-and-the-fight-for-worker-control/ How to Think About AI—Before It’s Too Late (Daniel Solove) https://www.youtube.com/watch?v=_0xR3uEgGcc Could Tech Bosses Destroy Life As We Know It? (Politics JOE) https://www.youtube.com/watch?v=PL4VktU0SgY Latest books (permalink) "The Reverse-Centaur's Guide to AI," a short book about being a better AI critic, Farrar, Straus and Giroux, June 2026 https://us.macmillan.com/books/9780374621568/thereversecentaursguidetolifeafterai/ "Canny Valley": A limited edition collection of the collages I create for Pluralistic, self-published, September 2025 https://pluralistic.net/2025/09/04/illustrious/#chairman-bruce "Enshittification: Why Everything Suddenly Got Worse and What to Do About It," Farrar, Straus, Giroux, October 7 2025 https://us.macmillan.com/books/9780374619329/enshittification/ "Picks and Shovels": a sequel to "Red Team Blues," about the heroic era of the PC, Tor Books (US), Head of Zeus (UK), February 2025 (https://us.macmillan.com/books/9781250865908/picksandshovels). "The Bezzle": a sequel to "Red Team Blues," about prison-tech and other grifts, Tor Books (US), Head of Zeus (UK), February 2024 (thebezzle.org). "The Lost Cause:" a solarpunk novel of hope in the climate emergency, Tor Books (US), Head of Zeus (UK), November 2023 (http://lost-cause.org). "The Internet Con": A nonfiction book about interoperability and Big Tech (Verso) September 2023 (http://seizethemeansofcomputation.org). Signed copies at Book Soup (https://www.booksoup.com/book/9781804291245). "Red Team Blues": "A grabby, compulsive thriller that will leave you knowing more about how the world works than you did before." Tor Books http://redteamblues.com. "Chokepoint Capitalism: How to Beat Big Tech, Tame Big Content, and Get Artists Paid, with Rebecca Giblin", on how to unrig the markets for creative labor, Beacon Press/Scribe 2022 https://chokepointcapitalism.com Upcoming books (permalink) "The Post-American Internet," a geopolitical sequel of sorts to Enshittification, Farrar, Straus and Giroux, 2027 "Unauthorized Bread": a middle-grades graphic novel adapted from my novella about refugees, toasters and DRM, FirstSecond, April 20, 2027 "Enshittification, Why Everything Suddenly Got Worse and What to Do About It" (the graphic novel), Firstsecond, 2027 "The Memex Method," Farrar, Straus, Giroux, 2027 Colophon (permalink) Today's top sources: Currently writing: “Once Is Enemy Action,” a science fiction novel about the origins of modern technofascism. Today's words: 509 (20770 total). "The Post-American Internet," a sequel to "Enshittification," about the better world the rest of us get to have now that Trump has torched America. Fourth draft completed. Submitted to editor. A Little Brother short story about DIY insulin PLANNING This work – excluding any serialized fiction – is licensed under a Creative Commons Attribution 4.0 license. That means you can use it any way you like, including commercially, provided that you attribute it to me, Cory Doctorow, and include a link to pluralistic.net. https://creativecommons.org/licenses/by/4.0/ Quotations and images are not included in this license; they are included either under a limitation or exception to copyright, or on the basis of a separate license. Please exercise caution. How to get Pluralistic: Blog (no ads, tracking, or data-collection): Pluralistic.net Newsletter (no ads, tracking, or data-collection): https://pluralistic.net/plura-list Mastodon (no ads, tracking, or data-collection): https://mamot.fr/@pluralistic Bluesky (no ads, possible tracking and data-collection): https://bsky.app/profile/doctorow.pluralistic.net Medium (no ads, paywalled): https://doctorow.medium.com/ Tumblr (mass-scale, unrestricted, third-party surveillance and advertising): https://mostlysignssomeportents.tumblr.com/tagged/pluralistic "When life gives you SARS, you make sarsaparilla" -Joey "Accordion Guy" DeVilla READ CAREFULLY: By reading this, you agree, on behalf of your employer, to release me from all obligations and waivers arising from any and all NON-NEGOTIATED agreements, licenses, terms-of-service, shrinkwrap, clickwrap, browsewrap, confidentiality, non-disclosure, non-compete and acceptable use policies ("BOGUS AGREEMENTS") that I have entered into with your employer, its partners, licensors, agents and assigns, in perpetuity, without prejudice to my ongoing rights and privileges. You further represent that you have the authority to release me from any BOGUS AGREEMENTS on behalf of your employer. ISSN: 3066-764X
The unstated disagreement that underpins safety debates