More from Rachel Thomas, PhD
As a former mathematician, I was used to nobody reading what I wrote. So when I first began blogging in 2015, I never expected that several of my blog posts would go viral or to have multiple journalists contact me (including from NPR, Wired, and Fortune), make the front page of Hacker News (over 10 times), receive conference keynote invitations, and be interviewed on podcasts. I do not consider myself a “natural” writer. In college, I tried to avoid classes that required essays, because writing was a struggle for me. It wasn’t until I was 30 that I set out to practice writing more. I share tips I use for blogging here, which include being willing to put a lot of time into a single post, incorporating high quality information, and having a clear idea of my intended audience. I have selected some of my most popular and impactful posts below. Several of these were originally posted on Medium or fast.ai (the two sites where my writing used to live). They are grouped into clusters based on theme. I hope you might enjoy reading these if you haven’t seen them before! Challenging Conventional Wisdom Questioning widely-held assumptions about tech culture, education, and health has been the basis for several popular posts. If you think women in tech is just a pipeline problem, you haven’t been paying attention (2015) In 2015, I felt burnt out and disillusioned by my experiences working in tech. I was frustrated with how much the popular conversation was still focused on “the pipeline problem”: training young girls to code while ignoring all the adult women being driven out of the tech industry by mistreatment. I spent 9 months researching and writing this post. It went viral and remains my most popular essay. This, together with my other posts, led to me being interviewed and quoted in Wired several times regarding diversity in tech (as well as other AI topics). My first post Trends to Avoid When Founding a Startup (2018) The dominant narrative for Bay Area tech startups is to try to raise venture capital, achieve exponential hypergrowth, and hire lots of computer science PhDs. I argued that these approaches not only harm employees, but lead to weaker companies and worse products. My family’s unlikely homeschooling journey (2022) Many people hold a stereotyped and outdated view of homeschooling, not realizing the explosion of innovative, non-traditional education options available in recent years. My husband and I never planned to homeschool, but we unexpectedly found that our child thrives with this approach. Your Immune System is Not a Muscle (2024) The misleadingly named “Hygiene Hypothesis” is often used to justify the misconception that all microbes are good for us. However, this theory is more accurately reframed as the “Old friends hypothesis”: humans co-evolved with friendly bacteria and some parasites. We did not co-evolve with the crowd infections of mega-cities and 100,000 global flights per day. AI Beyond Elite Institutions Machine learning isn’t just for those at billion dollar companies. These posts highlight unconventional practitioners and offer practical guidance for people in varied domains. Deep Learning: Not Just for Silicon Valley (2017) Our goal at fast.ai is making AI accessible to people outside of elite institutions, who are tackling meaningful problems in low-resource areas. This post introduced some of our earliest international fellows and the diverse range of problems they were working on. I always enjoyed writing about fascinating use cases from our deep learning community. How (and why) to create a good validation set (2017) An all-too-common scenario: a seemingly impressive machine learning model is a complete failure when implemented in production. Advice on one common culprit of this, and how to avoid it. In the early years of fast.ai, I wrote numerous posts with practical advice for machine learning. This article asked the question, “Can A.I. conquer its Excel problem? An Introduction to Deep Learning for Tabular Data (2018) Deep learning is not just for images and text. Companies such as Pinterest and Instacart are also applying it to tabular data, the type of data you might normally put in a spreadsheet. This post caught the attention of a reporter with Fortune, who ended up interviewing me and writing about the topic here. Debunking AI Hype & Holding Tech Accountable The narratives about AI put forth by major tech companies are often misleading about what is necessary, what values matter, and what types of harms can result. Google’s AutoML: Cutting Through the Hype (2018) In a 3-part series, I countered claims that all data scientists need customized, bespoke neural network architectures. While I was nervous about disagreeing with both Google’s CEO and head of AI, my posts led to an invitation to keynote the prestigous ICML AutoML workshop. Seven years later, my critiques have been proved valid, with transfer learning a cornerstone of ML and automated neural network search not commonly used. By 2023, we were supposed to all be using AutoML neural architecture search Five Things That Scare Me About AI (2019) AI ethics is not just a theoretical topic. I was (and still am) alarmed about the harms already being caused to human beings by AI systems irresponsibly applied to healthcare, employment decisions, policing, and more. The Problem with Metrics is a Big Problem for AI (2019) Overemphasizing metrics leads to a variety of real-world harms, including manipulation, gaming, and a myopic focus on the short-term. AI is metric optimization on steroids. I later turned this blog post into an academic paper, together with David Uminsky. Two disturbing case studies I keep returning to are how computerized algorithms have been used to cut healthcare and to fire teachers Deep learning gets the glory, deep fact checking gets ignored (2025) A microbiologist discovered hundreds of errors in a paper that used AI to classify enzymes. This is a case study of how challenging it can be to evaluate AI claims outside our area of expertise, as well as of the misaligned incentives that reward flashy results, but not diligent fact-checking. This has been by far my most popular post on LinkedIn. Immunology & Science Decoding T cells with AI (2024) T cells are a crucial component of the adaptive immune system. Accurately pedicting what they can bind to would impact a range of treatments. Numerous algorithms have been developed for this question, but the problem is far from solved. The surface of a T cell. I’ve enjoyed exploring how AI is being applied to immunology Scientists Just Connected the Dots Between Viruses and… Everything (2025) For a long time, catching frequent viruses was considered both inevitable and harmless. But it turns out that common, seemingly-mild viruses have disturbing long-term health impacts. A thread about viruses Thanks for joining me on this walk through the past! Also, you can subscribe to be notified of new blog posts by submitting your email below: I look forward to reading your responses. Create a free GitHub account to comment below.
Most people catch many viruses in their lives– for example, over 90% of adults have Epstein-Barr virus, and adults catch the flu about once every 5 years. For a long time, catching frequent viruses was considered both inevitable and harmless. But it turns out that common, seemingly-mild viruses have disturbing long-term health impacts. Common respiratory viruses increase the risk of heart attacks and strokes. Viruses are linked to dementia and Alzheimer’s Disease. They can re-awaken cancer cells in patients whose cancer was previously in remission. Persistent infections accelerate aging and undermine longevity. Viruses can be the trigger that kicks off life-long autoimmune diseases. New studies come out each week confirming that viruses can harm the health of your heart, blood vessels, brain, nervous system, and gut. Please pause and let this sink in. If we were to truly internalize this information, there would be massive shifts in the practice of medicine, scientific research, and public policy. Pathogens accelerate aging in many ways. Proal and VanElzakker, 2025 What you can avoid (infections) may be just as important as what you seek out (exercise, healthy foods). This news seems depressing. It’s too late, viruses are everywhere, everyone has already caught them– what can be done? There actually is a lot we can do. First of all, developing new anti-viral therapies and treatments should be a top priority. Second, regardless of what infections you’ve already had, preventing or reducing future infections will have a positive impact. There is exciting work happening towards both of these goals, including AI-assisted drug design, patient-led biomedical research, initiatives to improve indoor air quality and new technologies for cleaning the air. How can viruses cause all these bad outcomes when some people who catch them are fine? Human health is complicated. Disease development involves a complex interplay of factors: infections, underlying genetics, environment, the microbiome, and more. Let’s return to the example of Epstein-Barr Virus (EBV). EBV has been strongly linked to Multiple Sclerosis, prolonged fatigue, and 6 different types of cancer. Given that almost everyone has had EBV, even though “only” a percentage of people develop these lasting impacts, this is a major cause of suffering. Viruses tilt the probabilities against you A new world and a powerful idea You may wonder why so many of these health issues are on the rise, when viruses are nothing new. Our world has changed drastically in recent decades compared to most of human evolution. We live in a hyperconnected age of global mega-cities and record numbers of large international flights now. We spend our time indoors in crowded, poorly ventilated buildings. These factors have allowed viruses to travel faster and farther than ever before. The misleadingly named “Hygiene Hypothesis” is often used to justify the misconception that all microbes are good for us. However, this theory is more accurately reframed as the “Old friends hypothesis”: humans co-evolved with friendly bacteria and some parasites. Viruses are not our friends, but rather enemies. We did not co-evolve with these crowd infections of mass travel, mega-cities, and indoor confines. Not all infections are the same! Modern crowd infections are causing huge harm. Figure from Rook, 2014 The idea that viruses are contributing so much to human suffering and long-term disease is powerful. It will transform how we approach medicine, health, and aging, if we let it. This revelation is one of the key reasons that I decided to make a mid-life career pivot, stepping back from fulfiling work in AI to return to graduate school in Microbiology-Immunology, a journey I have been chronicling here on my blog. I hope to spend the next few decades applying my machine learning skills to problems at the intersection of infections, multi-omic data sets, the microbiome, and chronic disease. Below, I will share some of what has captured my attention and upended my old views on disease and medicine. Viruses have many ways to wreak havoc Viruses have evolved to evade, outmatch, commandeer, and otherwise hurt our immune systems. Here is an incomplete and overlapping list of ways that viruses can harm us: 1. Persistence Some viruses quietly stick around for years or decades after our initial illness. They may re-awaken later to cause more problems, or they may spawn surprising issues that we don’t recognize as part of our initial infection. When they persist in our cells, viruses can impact gene expression, hijacking processes our cells need to gain nutrition and energy. Dr. Amy Proal, a researcher in this area, says that treating persistent infections will be necessary to combat aging and extend healthspan. 2. Autoimmunity During an infection, sometimes our immune cells get confused into attacking our own tissue that may “look” similar to the virus (this process is known as molecular mimicry). Once it has mistakenly learned to attack self-tissue, the immune system may continue to do so, even after the virus has been defeated. This is just one of several ways by which viruses can trigger autoimmune diseases such as Lupus, Multiple Sclerosis, Rheumatoid Arthritis, or Type 1 Diabetes. A confused antibody decides to attack a pathogen, as well as the similar-looking myelin covering of the nerves, causing Guillain-Barré syndrome (Comic from Creative Med Doses) 3. Microbiome changes You might expect a stomach bug like norovirus to change the gut microbiome for the worse. Surprisingly, respiratory viruses such as Influenza, RSV, and Covid all harm the gut microbiome too. This is bad news, since the gut microbiome helps to regulate the immune system and produces neurotransmitters for our brain. 4. Immune Dysregulation There are a bunch of ways that the immune system can malfunction (including the ones listed above). Measles can cause immune amnesia, where the immune system forgets previous infections it had learned to fight, leading people to catch the exact same diseases again. There is growing evidence that covid has a negative impact on the immune system as well. 5. Reactivation of other pathogens Infection with a new virus can wake up old infections that were sleeping quietly in your cells. It is unfair, but sometimes viruses will gang up on you, re-activating other viruses (or bacteria) that weren’t bothering you before. 6. Cardiac damage Chickenpox/Shingles, Influenza, and Covid all raise the risk of heart attacks and strokes. Viruses have many ways of harming our cardiac systems: inflammation, damage to the blood vessles, increased blood clots, and damage to the heart. From a meta-analysis of 48 studies about respiratory viruses triggering heart attacks & strokes (Nguyen, et al, 2025) 7. Cancer Cancer involves a failure of the immune system to kill cells that have gone rogue and turned over to the dark side. In 2008, it was estimated that viral infections contribute to 15-20% of human cancer cases. Additional research further linking viruses and cancer has come out since then, so the percentage may be higher now. Both flu and covid infections can reawaken “sleeping” cancer cells that had previously been in remission or cause cancer to spread. An article from MD Anderson on 8 Viruses that Cause Cancer 8. Cumulative impacts You might hope that you could catch a virus, get it over, and be done with it. Unfortunately, that is often not the case. A young college student was fine after having covid twice, but then struggled to walk short distances after her 3rd infection. A Colorado newspaper columnist was skiing, biking, mountain climbing, and running half-marathons up until his 5th covid infection. At this point, he developed pain, fatigue, and migraines that prevent him from doing the activities he loves most. These are not just isolated anecdotes, research confirms the cumulative dangers of repeat infections. In children, a second covid infection is more likely to cause Long Covid than the first infection. Whatever your previous history, reducing risk of future infections is a worthwhile goal. The above mechanisms are not exclusive. For example, some microbiome changes can make it easier for pathogens to pass from the gut into the bloodstream and provoke an autoimmune reaction (a process I talked about in this 5-minute video) The Paradigm Shift For most viruses, people focus on just a few weeks of initial symptoms. This is the wrong way to think about infections. Viral meningitis or EBV increases your risk of Alzheimer’s or dementia, 5-15 years later. Chicken pox (varicella zoster virus) can reactivate decades afterwards as shingles, which itself then leads to increased risk of stroke for at least the following year. We need to radically change how we think about viruses. There is much we still don’t know about the immune system. Early during the covid pandemic, many experts made definitive statements about the risks of covid, assuming that those who didn’t die in the first few weeks must be completely fine. However, perturbations from infections that initially seem minor can have far-reaching, long-lasting, and time-delayed impacts. There is a ton that is still unknown. Trying to figure out how viruses hijack cell processes, alter microbiomes, and dysregulate the immune system are complex questions. Researching these areas with curiosity, determination, and an open mind will reveal a lot. Reasons for Hope It can be gloomy to think about all the damage viruses can cause. The good news is that we don’t have to resign ourselves to these outcomes. Facing the disturbing reality that many viruses are worse than we thought is just the first step towards coming up with creative new solutions. There are some bright, curious, and determined people focused on these problems, although we need even more hands and brains to get involved. The breadth and depth of the harms caused by viruses can focus biomedical research in new directions. Most viruses do not have effective anti-viral treatments. This creates a huge need. Scientific inquiry can fail catastrophically when those closest to the problem are not included. Patient-led research gives me hope, because it is centered on the expertise of those closest to the problem. I am also optimistic about the use of AI for discovering new drugs and designing immune therapies. On the prevention side, reducing how frequently people get sick will have a big impact. Different viruses spread in different ways. In recent years, we have learned that many infections are airborne. Healthy indoor air is a human right, like access to clean drinking water. The UN recently held a high-level event focused on the right to clean air. There are many measures we can take to reduce transmission of airborne diseases, such as improved ventilation, air purification, and far-UVC technologies. Parliament houses, venues for elites, and barns for pigs have already received these air quality upgrades. We need children in schools, employees in workplaces, and patients in hospitals to get the same protections. Hopefully, we are on the cusp of a clean air revolution, with more people and organizations recognizing that healthy indoor air is essential. N95 masks offer an immediate way to significantly reduce how often you get sick. Thankfully, the N95s available currently are more comfortable and more effective than the surgical or cloth masks that many of us wore back in 2020. On the brink of an indoor air quality revolution Conclusion Viruses can harm our cardiac health and cognition, and increase our chances of cancer. If this revelation is fully realized, it will change how the field of medicine operates, priorities in research funding, and public policy on everything from indoor air quality standards to paid sick leave and school attendance. I believe we are on the threshold of what could be a drastic shift in better understanding, preventing, and treating viruses, thus unlocking longer and healthier lives. Related posts you may also be interested in: 5 Devious Tricks Pathogens Use Against Us Viruses are weirder, worse, & more preventable than you realise Viruses: The Silent Triggers of Autoimmune and Neurodegenerative Diseases Your Immune System is Not a Muscle If you enjoy my posts, please subscribe to be notified of new posts via email: I look forward to reading your responses. Create a free GitHub account to comment below.
DNA sequencing hasn’t lived up to the hype Twenty to thirty years ago, politicians, scientific leaders, journalists, and even Nobel laureates predicted that sequencing the human genome would revolutionize how we treat disease. And while the advances in DNA sequencing that have occurred since then have improved recognition and treatment for some cancers and rare diseases, on the whole the field has not lived up to earlier hype. Time Magazine covers from 1994 and 1999 about genetics In an article titled “Why sequencing the human genome failed to produce big breakthroughs in disease”, a biology professor highlights that most common diseases are not caused by a single gene. In fact, common diseases are often linked to hundreds of gene variants, and even collectively, these variants still account for only a small fraction of disease variance. Here, I want to focus on two other key limitations of DNA sequencing, and how they are now being addressed with new approaches. What DNA can’t tell us First, DNA can’t answer many questions about how cells and organisms work in practice. A neuron in the brain has the same DNA as a liver cell, yet the two have completely different functions. This is because different segments of DNA are turned off or on in different cells. To understand how cells are actually working, you need to know about proteins and RNA (RNA is the intermediary which translates DNA into protein). Proteins are what build the structure of cells, catalyze chemical reactions within the cell, and allow communication between cells. Healthy and unhealthy cells in the same organism will usually have the same DNA. For instance, if some regions of the intestines are experiencing an IBD (irritable bowel disease) flare and others aren’t, they would all have the same DNA, yet likely different RNA and protein levels. A second big problem is that many key sequencing techniques destroy spatial information. You essentially may have to put tissue or cells into a blender in order to get rich information about DNA or RNA sequences. While this data is informative, it turns out that locations of cells within a piece of tissue, and locations of regions within a cell, are also very important! Again, considering the case in which some regions of the intestine are inflamed due to IBD, yet others aren’t, mixing them all together in a blender will lose or distort useful information. Look at the difference between crypts in a healthy segment of the colon (on the left) compared to inflamed crypts (on the right). Differences between a healthy bowel (on the left) and an inflamed bowel (on the right). Source: mypathologyreport.ca Recently, we have seen a rise in breakthroughs that allow us to obtain data about location. Spatial techniques are a necessary and exciting step beyond DNA sequencing. The power of spatial information showed up as a major theme at a conference I attended last year, and spatial techniques have been recognized by Nature Methods as “Method of the year” twice in the last 5 years. A Few Major Areas of Innovation What is genetic sequencing anyway? There are a number of different types of sequencing that have been invented in the last 30 years. Walking through a brief history will illustrate what these technologies are, and what they can and cannot do. To make it concrete, let’s look at the example of how they have been applied to cancer treatment. Sanger Sequencing: This is an older technology dating back to the 1970s, and which was the main way of sequencing DNA up until 2005. Sanger sequencing was one of the methods used in the mid-90s to identify the genes BRCA1 and BRCA2 as key genetic risk factors for breast cancer. The process of discovering BRCA1/BRCA2 involved scientists slowly zeroing in on their chromosomal locations over a period of years. Several other cancer genes were discovered during this time period as well. While Sanger sequencing is effective on smaller amounts of DNA, it can be quite slow to deal with larger volumes. It took over 10 years to sequence the first copy of the human genome using Sanger Sequencing. It is still used today as a simple and reliable way to test for known mutations (such as BRCA1/BRCA2) or for smaller tasks. High-Throughput Sequencing: New technology released in the mid-2000s allowed DNA to be cut into lots of short pieces and for millions of pieces to be sequenced in parallel at once. This approach, called high-throughput sequencing, was significantly faster than Sanger sequencing. High-throughput sequencing has many applications, including to cancer treatment, by making it cheaper and faster to sequence DNA to identify particular mutations which can influence treatment decisions. Method of the Year | Nature Methods Long-read sequencing: High-throughput sequencing has the advantage of high accuracy, but the downside of short sequence lengths. In the 2010s, technologies were released with the opposite set of strengths and weaknesses. Long-read sequencing provides the advantage of long sequence lengths, although the downside of lower accuracy. To compare, high-throughput sequencing uses DNA strands that are a few hundred base pairs long, whereas long-read sequencing uses DNA strands that are tens of thousands base pairs long. Both technologies have different strengths and are widely used today. Single-Cell Sequencing: High-throughput sequencing involves sequencing the DNA of many cells at once, but sometimes it is useful to sequence individual cells. In a tumor, different cells can have different mutations. It is possible that a small subset of the mutations may drive metastasis (the spread of cancer to other areas) or resistance to treatment. Identifying these driver mutations can guide treatment decisions, since particular driver mutations can predict the effectiveness of various drugs. Key mutations may be drowned out in the average if you sequence the entire tumor. This is one reason why it is useful to be able to sequence single cells, and not just obtain the average of many cells. Single-cell sequencing was selected as Nature Methods method of the year in 2013. A Lego interpretation of bulk RNA-seq; single-cell RNA-seq; spatial transcriptomics; and the original organ. Source: Bo Xia, @BoXia7 Multi-Omics: The study of DNA is genomics. DNA alone gives us an incomplete picture of an organism. Epigenomics can provide information about which regions of DNA are active or silenced. To understand how different cells function, as well as cells in different states of disease or health, you also need to know about their RNA and proteins. This data is contained in the field of transcriptomics (transcripts are strands of RNA transcribed from DNA) and proteomics (the proteins in a cell). Metabolomics looks at small molecules (such as sugars, amino acids, and vitamins) within the body and exposomics includes all sorts of environmental exposures. Collectively, these fields are known as -omics or multi-omics. It is valuable to combine multiple types of -omics together for richer sources of information, since each has different strengths, limitations, and insights to offer. Multi-Omics is an exciting area that draws on lots of data, with applications to cancer, infectious disease and immunology. Spatial: Spatial information lets us see all the variation within a section of the body– such as a segment of the intestines, the liver, or a cancerous tumor. This variation can often be significant for understanding disease and treatment prognosis. To better understand why spatial techniques are useful, let us dive into some background about cancer. Tumors aren’t just lumps of bad cells Cancer is defined as excess cell division. I used to think that tumors were just clumps of “bad cells”, where “good” and “bad” were binary states. This is incorrect. Tumors are not uniform, and within the category of “bad” there is a great deal of variation and heterogeneity. Different cells within a tumor may have different mutations from one another. And our immune systems sometimes build complex defense structures within tumors in attempts to more effectively fight them. How close a cancerous cell is to one of these immune structures impacts how likely the body is to destroy the cancerous cell. Notice all the variation within this tumor! Immunologist Dr. Angela Ferguson, who studies head and neck cancers, describes tumors as having a “physical landscape”. She has shown how the organization and structure within a tumor can predict and guide treatment outcomes. Her work found that cancer progression is “landscape-dependent”, where landscape refers to the locations of immune cells and structures within a tumor. It is not enough to study cancer cells in isolation. We need to understand their layout. Methods that effectively put tumor cells into a blender in order to sequence them, disrupting their spatial information, are insufficient on their own. Mapping that spatial information can hold the keys for more effective treatment. Single cell sequencing approaches allow a greater number of genes to be measured, whereas spatial approaches measure fewer genes but also provide location information. Combining these two approaches can prove powerful. Programming Libraries Applied to Spatial -Omics We are living at a time when multi-omics, spatial information, and user-friendly programming libraries are converging for easier exploration and discovery. For instance, below is an image I created using the common Python programming libraries pandas and matplotlib of data the NIH has shared about a rare liver disease. The image shows a slice of liver tissue, with gene expression overlaid in a color scale ranging from purple (low expression) to yellow (highest expression). Using Matplotlib to display a cross-section of liver with gene expression The NIH dataset contain images of slices of liver and expression of many different genes, from both healthy patients and those with a rare liver disease. https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE240429. The plot shows expression of albumin, a protein which helps transport other molecules around the body, overlaid on top of the liver, illustrating which regions produce more or less of this protein. Data science tools are invaluable for transforming and plotting such data. In addition to being able to use standard Python libraries (such as Pandas and Matplotlib) for visualing this data, there are many specialist libraries as well. The Scverse (Sc = single cell) includes a number of Python libraries focused on single cell analysis. The SC Verse includes libaries for single cell sequencing analysis There are also popular R libraries for single cell analysis and multi-omics, such as Harmony, Seurat, and mixOmics. With all of these tools and technologies, it is an exciting time to be working at the intersection of data science and microbiology. The causes of most diseases are complex and multi-factorial, although new approaches of spatial multi-omics provide unprecedented types of useful information. Related Reading: What AI can tell us about microscope slides Gaps and Risks of AI in the Life Sciences AI and Immunology You can subscribe to be notified of new blog posts by submitting your email below: I look forward to reading your responses. Create a free GitHub account to comment below.
The lavender images below show breast tissue. There are many questions doctors could want to answer using these images: They could want to know whether there are tumors present or not. If there is a tumor, doctors would want to classify its stage, make predictions about how likely the patient is to respond to treatment, and to detect whether the tumor has spread from another organ. All of these are questions which people are now tackling with machine learning. They fall within the area of computational pathology, often abbreviated CPath. In the past year, two CPath AI models were released which achieved state-of-the-art results. Here I will discuss an introduction to this field, what these models do, and what some key challenges are going forward. Breast tissue images from the BACH: Grand challenge on breast cancer histology CPath foundation models There is a powerful idea about how to make more accurate CPath models. Rather than train a model on a single type of tissue and a single task (e.g. identifying cancer in breast tissue), train a model on images of tissue from many different organs (breasts, lymph nodes, lungs, prostate, heart,…) and on multiple different tasks (recognizing cancer, determining the stage and subtype of the cancer, segmenting cells, and predicting treatment outcomes). Patterns learned from one dataset or one task are likely to generalize to others. Such models are known as CPath foundation models. In general, a foundation model is a machine learning model which is trained on a sufficiently diverse large dataset which can then be adapted for a range of downstream tasks. This idea is commonly used in the area of language models such as Chat-GPT and Claude.ai. Language foundation models are trained on many types of language tasks and intended to generalize across different corpuses of text (e.g. wikipedia, reddit posts, academic papers, online conversations, news articles, and more). ImageNet models trained to recognize a huge variety of different pictures often serve as foundation models for images. The success of foundation models within the areas of language and more general images is a key reason why we might expect pathology foundation models to be useful too. Tissues are groups of cells with similar structure and function. Different types of tissue within the human body include nervous, muscle, connective, and epithelial tissue. Image: Wikimedia Two notable CPath foundation models were released in 2024: Prov-GigaPath and UNI. Both models achieved state-of-the-art performance on dozens of pathology tasks (although they were not directly compared to one another). Another relevant paper (from Kaiko.ai) studied the impact of dataset size and model size on CPath model performance. Learning the Vocab Medicine is full of jargon and specialized vocabulary. Pathology refers to the study of disease. It is a broad field, and can include everything from dissecting dead bodies to analyzing blood samples. One key focus of computational pathology is analyzing and interpreting whole slide images (WSIs) and in some cases combined with accompanying meta-data about a patient. Whole slide images refers to the complete microscope slide, although in many cases the region of interest (such as particular cancerous or inflamed cells) may be much smaller, just occupying a subset of the slide. Machine learning (ML) is a subfield of Artificial intelligence (AI) which involves learning from past data, and is increasingly being used with great success in pathology. The focus of most computational pathology ML models is on images of tissue, on microscope slides. That is what we will focus on in this post as well. So Many Tasks! There are many different benchmarks that CPath models can be tested on. These involve numerous datasets: related to different areas of the body, with different sizes, and with different purposes. They also involve a variety of tasks, including binary classification, image segmentation, and outcome prediction. Prov-GigaPath attained state-of-the-art performance on 25 out of the 26 tasks it was evaluated on and UNI attained state-of-the-art performance on 34 different tasks. Here I will give examples of just 3 of these tasks. Task: prostate cancer cell grading In the 1960s, the pathologist Dr. Donald Gleason came up with a grading scale for rating cells as they progressed from normal to prostate cancer. The Gleason Grading system is still widely used and is considered a powerful predictor of how prostate cancer patients will fare. A major medical image conference (MICCAI) held a competition in 2022 for researchers to create algorithms to determine the Gleason grades when given images of prostate tissue. Examples of UNI predictions of Gleason grades for a section of prostate tissue. Figure 3b from the UNI paper The prostate tissue is shown in pink, and segments have been colored in blocks based on where they fall on the Gleason scale. Task: identifying early signs of rejection after a heart transplant Rejection is the main cause of mortality in patients who have received a heart transplant. Since the early stages of rejection can be asymptomatic, it is standard for patients to receive frequent biopsies for 1-2 years following a transplant. These are known as endomyocardial biopsies (EMB), since they remove a small sample of tissue from the inner lining (endo) of the heart (cardial) muscle (myo). Accurately interpreting the results of these biopsies is a key question. Underestimating the chance of rejection could lead to dangerous delays in treatment, but overestimating could lead to alarm and unnecessary follow-ups or treatment. Assessment of the sampled tissue by experienced pathologists has higher variability than many other tasks, such as cancer diagnosis. Deep learning is being used to tackle this task, in models such as Cardiac Rejection Assessment Neural Estimator (CRANE) and the CPath foundation model UNI. Each row shows a different sample of cardiac tissue, with a different medical issue. On the far left are the whole slide images, then zoomed in at higher resolution on a key Region of Interest (ROI). On the far right is a heat map for the most zoomed in area showing which features the algorithm has identified as significant. Figure 3 from the CRANE paper. Task: Genetic Mutations in Cancer For several common genetic mutations in tumors, there are specific drugs known to target those mutations. This has a direct application for clinical treatment. Since genetic mutations can change the form and function of cells, it is reasonable to expect that this information could be deduced from images of the cancer cells. Deep learning models have been built to identify genetic mutations from tissue slides. The benefits of using a computational approach are that it can be scaled as an increasing number of relevant genetic mutations and molecular biomarkers are being discovered. Task-specific models have been built for this, and this is one of the tasks that foundation models can be tested on. Different types of cancer listed along the y-axis and 20 common genetic mutations listed on the y-axis. Figure 1D from Kather, 2020. We need more data One key challenge in the area of CPath foundation models is gathering enough training data. The Cancer Genome Atlas (TCGA) was an ambitious project launched in 2006 by the National Cancer Institute in the USA. Over a 12 year period, samples were collected from over 11,000 patients of 33 different cancer types, and all this data was made publicly available. While this is a rich dataset and a useful resource, all 3 papers we’ve looked at concluded that TCGA is not large enough for effective foundation models. In addition to limited data size, TCGA also has limited diversity, consisting mostly of slides from the primary site of cancer, but not metastasized cancers or different types of tissues. Researchers at Kaiko.ai tested the impact of scaling both the size of their model and the size of the training dataset. While they found limited need to scale model size beyond a certain point, they found that larger datasets continued to lead to increased performance. They concluded that TCGA was likely not large enough and shared their plans to build a larger training set, and are now partnering with cancer centers across Europe to create a dataset for their model. The researchers behind two other CPath foundation models reached the same conclusion about data set size, and gathered massive datasets to train their models. This required partnering with healthcare centers. Prov-GigaPath, a model created by Microsoft Research and Providence Genomics involved data from 30,000 patients across 28 cancer centers (which are part of Providence Healthcare company). UNI, a cPath model created by a team at Harvard, MIT, and the Broad Institute, involved the creation of the Mass-100K: a dataset with over 100K whole slide images across 20 tissue types collected from Mass General Hospital, Brigham & Women’s Hospital, and Genotype-Tissue Expression (GTEx) consortium. These partnerships and curation of training datasets are currently a crucial component of building CPath foundation models. Curating datasets carefully poses many challenges as well. Combining data from different sources, which often use different protocols for how slides are sampled and prepared, can introduce significant biases. Different scales CPath foundation models face the difficulty of capturing both local patterns (that show up in a small tile within a slide) and global patterns across the whole slide. Many tiny tiles are found within a slide. Some models, such as the Hierarchical Image Pyramid Transformer (from several of the same authors as UNI), use hierarchical approaches to deal with these multiple scales. Hierarchical Structure of Whole-Slide Images, Figure 1 from Chen, et al, 2020 Other models, such as Prov-GigaPath, treat the tiles as tokens, encoding both the tiles and the slide as a whole as model inputs. Prov-GigaPath uses both a slide encoder and a tile encoder to take into account these two different scales. Treating slides as tokens, Figure 1a from the Prov-GigaPath paper In pathology clinics, diagnosis and treatment decisions are often made at the patient level, whereas CPath models are often highly focused on regions of interest. Accommodating the multiple relevant scales (small tiles, whole slides, and patient-level) for pathology is a consideration that CPath models need to balance. Going Forward It is still early in the world of CPath and there are many growth opportunities, including the continued need for large and diverse datasets, ways to further optimize model training, tasks which have previously received less focus, and the difficulties of integrating models into clinical work. As the authors of the kaiko.ai paper wrote, “We are still at the very beginning of developing a truly foundational pathology foundation model.” It is a hopeful sign that these models achieve state-of-the-art results on dozens of benchmarks, but it still remains to be seen when and how they will be used in clinical settings. Related Reading: The Most Common and Useful Neural Nets Using AI to Discover New Antibiotics AI and Immunology You can subscribe to be notified of new blog posts by submitting your email below: I look forward to reading your responses. Create a free GitHub account to comment below.
More in AI
There's a lot of polarising discourse right now about the threat AI poses to humanity. Some think it's a farce and others think we face extinction. Here are my thoughts.
A middle ground between the cybersecurity and AI safety communities
An overview of the current state of the engineering market and the AI skills that are in demand
Wendell Berry died last week at 92 at his home in Port Royal, Kentucky, where he farmed his land using traditional techniques and wrote with ... Read more The post Wendell Berry and the Promise of the Deep Life appeared first on Cal Newport.