Full Width [alt+shift+f] Shortcuts [alt+shift+k]
Sign Up [alt+shift+s] Log In [alt+shift+l]
1
DNA sequencing hasn’t lived up to the hype Twenty to thirty years ago, politicians, scientific leaders, journalists, and even Nobel laureates predicted that sequencing the human genome would revolutionize how we treat disease. And while the advances in DNA sequencing that have occurred since then have improved recognition and treatment for some cancers and rare diseases, on the whole the field has not lived up to earlier hype. Time Magazine covers from 1994 and 1999 about genetics In an article titled “Why sequencing the human genome failed to produce big breakthroughs in disease”, a biology professor highlights that most common diseases are not caused by a single gene. In fact, common diseases are often linked to hundreds of gene variants, and even collectively, these variants still account for only a small fraction of disease variance. Here, I want to focus on two other key limitations of DNA sequencing, and how they are now being addressed with new approaches. What DNA can’t tell us First, DNA can’t answer many questions about how cells and organisms work in practice. A neuron in the brain has the same DNA as a liver cell, yet the two have completely different functions. This is because different segments of DNA are turned off or on in different cells. To understand how cells are actually working, you need to know about proteins and RNA (RNA is the intermediary which translates DNA into protein). Proteins are what build the structure of cells, catalyze chemical reactions within the cell, and allow communication between cells. Healthy and unhealthy cells in the same organism will usually have the same DNA. For instance, if some regions of the intestines are experiencing an IBD (irritable bowel disease) flare and others aren’t, they would all have the same DNA, yet likely different RNA and protein levels. A second big problem is that many key sequencing techniques destroy spatial information. You essentially may have to put tissue or cells into a blender in...
12th May 2025

Stay updated

Get a weekly newsletter with the top 5 articles worth reading every week.

More from Rachel Thomas, PhD

A Decade of Writing Stuff That People, To My Surprise, Actually Read

As a former mathematician, I was used to nobody reading what I wrote. So when I first began blogging in 2015, I never expected that several of my blog posts would go viral or to have multiple journalists contact me (including from NPR, Wired, and Fortune), make the front page of Hacker News (over 10 times), receive conference keynote invitations, and be interviewed on podcasts. I do not consider myself a “natural” writer. In college, I tried to avoid classes that required essays, because writing was a struggle for me. It wasn’t until I was 30 that I set out to practice writing more. I share tips I use for blogging here, which include being willing to put a lot of time into a single post, incorporating high quality information, and having a clear idea of my intended audience. I have selected some of my most popular and impactful posts below. Several of these were originally posted on Medium or fast.ai (the two sites where my writing used to live). They are grouped into clusters based on theme. I hope you might enjoy reading these if you haven’t seen them before! Challenging Conventional Wisdom Questioning widely-held assumptions about tech culture, education, and health has been the basis for several popular posts. If you think women in tech is just a pipeline problem, you haven’t been paying attention (2015) In 2015, I felt burnt out and disillusioned by my experiences working in tech. I was frustrated with how much the popular conversation was still focused on “the pipeline problem”: training young girls to code while ignoring all the adult women being driven out of the tech industry by mistreatment. I spent 9 months researching and writing this post. It went viral and remains my most popular essay. This, together with my other posts, led to me being interviewed and quoted in Wired several times regarding diversity in tech (as well as other AI topics). My first post Trends to Avoid When Founding a Startup (2018) The dominant narrative for Bay Area tech startups is to try to raise venture capital, achieve exponential hypergrowth, and hire lots of computer science PhDs. I argued that these approaches not only harm employees, but lead to weaker companies and worse products. My family’s unlikely homeschooling journey (2022) Many people hold a stereotyped and outdated view of homeschooling, not realizing the explosion of innovative, non-traditional education options available in recent years. My husband and I never planned to homeschool, but we unexpectedly found that our child thrives with this approach. Your Immune System is Not a Muscle (2024) The misleadingly named “Hygiene Hypothesis” is often used to justify the misconception that all microbes are good for us. However, this theory is more accurately reframed as the “Old friends hypothesis”: humans co-evolved with friendly bacteria and some parasites. We did not co-evolve with the crowd infections of mega-cities and 100,000 global flights per day. AI Beyond Elite Institutions Machine learning isn’t just for those at billion dollar companies. These posts highlight unconventional practitioners and offer practical guidance for people in varied domains. Deep Learning: Not Just for Silicon Valley (2017) Our goal at fast.ai is making AI accessible to people outside of elite institutions, who are tackling meaningful problems in low-resource areas. This post introduced some of our earliest international fellows and the diverse range of problems they were working on. I always enjoyed writing about fascinating use cases from our deep learning community. How (and why) to create a good validation set (2017) An all-too-common scenario: a seemingly impressive machine learning model is a complete failure when implemented in production. Advice on one common culprit of this, and how to avoid it. In the early years of fast.ai, I wrote numerous posts with practical advice for machine learning. This article asked the question, “Can A.I. conquer its Excel problem? An Introduction to Deep Learning for Tabular Data (2018) Deep learning is not just for images and text. Companies such as Pinterest and Instacart are also applying it to tabular data, the type of data you might normally put in a spreadsheet. This post caught the attention of a reporter with Fortune, who ended up interviewing me and writing about the topic here. Debunking AI Hype & Holding Tech Accountable The narratives about AI put forth by major tech companies are often misleading about what is necessary, what values matter, and what types of harms can result. Google’s AutoML: Cutting Through the Hype (2018) In a 3-part series, I countered claims that all data scientists need customized, bespoke neural network architectures. While I was nervous about disagreeing with both Google’s CEO and head of AI, my posts led to an invitation to keynote the prestigous ICML AutoML workshop. Seven years later, my critiques have been proved valid, with transfer learning a cornerstone of ML and automated neural network search not commonly used. By 2023, we were supposed to all be using AutoML neural architecture search Five Things That Scare Me About AI (2019) AI ethics is not just a theoretical topic. I was (and still am) alarmed about the harms already being caused to human beings by AI systems irresponsibly applied to healthcare, employment decisions, policing, and more. The Problem with Metrics is a Big Problem for AI (2019) Overemphasizing metrics leads to a variety of real-world harms, including manipulation, gaming, and a myopic focus on the short-term. AI is metric optimization on steroids. I later turned this blog post into an academic paper, together with David Uminsky. Two disturbing case studies I keep returning to are how computerized algorithms have been used to cut healthcare and to fire teachers Deep learning gets the glory, deep fact checking gets ignored (2025) A microbiologist discovered hundreds of errors in a paper that used AI to classify enzymes. This is a case study of how challenging it can be to evaluate AI claims outside our area of expertise, as well as of the misaligned incentives that reward flashy results, but not diligent fact-checking. This has been by far my most popular post on LinkedIn. Immunology & Science Decoding T cells with AI (2024) T cells are a crucial component of the adaptive immune system. Accurately pedicting what they can bind to would impact a range of treatments. Numerous algorithms have been developed for this question, but the problem is far from solved. The surface of a T cell. I’ve enjoyed exploring how AI is being applied to immunology Scientists Just Connected the Dots Between Viruses and… Everything (2025) For a long time, catching frequent viruses was considered both inevitable and harmless. But it turns out that common, seemingly-mild viruses have disturbing long-term health impacts. A thread about viruses Thanks for joining me on this walk through the past! Also, you can subscribe to be notified of new blog posts by submitting your email below: I look forward to reading your responses. Create a free GitHub account to comment below.

17th Nov 2025 1 votes
Scientists Just Connected the Dots Between Viruses and… Everything

Most people catch many viruses in their lives– for example, over 90% of adults have Epstein-Barr virus, and adults catch the flu about once every 5 years. For a long time, catching frequent viruses was considered both inevitable and harmless. But it turns out that common, seemingly-mild viruses have disturbing long-term health impacts. Common respiratory viruses increase the risk of heart attacks and strokes. Viruses are linked to dementia and Alzheimer’s Disease. They can re-awaken cancer cells in patients whose cancer was previously in remission. Persistent infections accelerate aging and undermine longevity. Viruses can be the trigger that kicks off life-long autoimmune diseases. New studies come out each week confirming that viruses can harm the health of your heart, blood vessels, brain, nervous system, and gut. Please pause and let this sink in. If we were to truly internalize this information, there would be massive shifts in the practice of medicine, scientific research, and public policy. Pathogens accelerate aging in many ways. Proal and VanElzakker, 2025 What you can avoid (infections) may be just as important as what you seek out (exercise, healthy foods). This news seems depressing. It’s too late, viruses are everywhere, everyone has already caught them– what can be done? There actually is a lot we can do. First of all, developing new anti-viral therapies and treatments should be a top priority. Second, regardless of what infections you’ve already had, preventing or reducing future infections will have a positive impact. There is exciting work happening towards both of these goals, including AI-assisted drug design, patient-led biomedical research, initiatives to improve indoor air quality and new technologies for cleaning the air. How can viruses cause all these bad outcomes when some people who catch them are fine? Human health is complicated. Disease development involves a complex interplay of factors: infections, underlying genetics, environment, the microbiome, and more. Let’s return to the example of Epstein-Barr Virus (EBV). EBV has been strongly linked to Multiple Sclerosis, prolonged fatigue, and 6 different types of cancer. Given that almost everyone has had EBV, even though “only” a percentage of people develop these lasting impacts, this is a major cause of suffering. Viruses tilt the probabilities against you A new world and a powerful idea You may wonder why so many of these health issues are on the rise, when viruses are nothing new. Our world has changed drastically in recent decades compared to most of human evolution. We live in a hyperconnected age of global mega-cities and record numbers of large international flights now. We spend our time indoors in crowded, poorly ventilated buildings. These factors have allowed viruses to travel faster and farther than ever before. The misleadingly named “Hygiene Hypothesis” is often used to justify the misconception that all microbes are good for us. However, this theory is more accurately reframed as the “Old friends hypothesis”: humans co-evolved with friendly bacteria and some parasites. Viruses are not our friends, but rather enemies. We did not co-evolve with these crowd infections of mass travel, mega-cities, and indoor confines. Not all infections are the same! Modern crowd infections are causing huge harm. Figure from Rook, 2014 The idea that viruses are contributing so much to human suffering and long-term disease is powerful. It will transform how we approach medicine, health, and aging, if we let it. This revelation is one of the key reasons that I decided to make a mid-life career pivot, stepping back from fulfiling work in AI to return to graduate school in Microbiology-Immunology, a journey I have been chronicling here on my blog. I hope to spend the next few decades applying my machine learning skills to problems at the intersection of infections, multi-omic data sets, the microbiome, and chronic disease. Below, I will share some of what has captured my attention and upended my old views on disease and medicine. Viruses have many ways to wreak havoc Viruses have evolved to evade, outmatch, commandeer, and otherwise hurt our immune systems. Here is an incomplete and overlapping list of ways that viruses can harm us: 1. Persistence Some viruses quietly stick around for years or decades after our initial illness. They may re-awaken later to cause more problems, or they may spawn surprising issues that we don’t recognize as part of our initial infection. When they persist in our cells, viruses can impact gene expression, hijacking processes our cells need to gain nutrition and energy. Dr. Amy Proal, a researcher in this area, says that treating persistent infections will be necessary to combat aging and extend healthspan. 2. Autoimmunity During an infection, sometimes our immune cells get confused into attacking our own tissue that may “look” similar to the virus (this process is known as molecular mimicry). Once it has mistakenly learned to attack self-tissue, the immune system may continue to do so, even after the virus has been defeated. This is just one of several ways by which viruses can trigger autoimmune diseases such as Lupus, Multiple Sclerosis, Rheumatoid Arthritis, or Type 1 Diabetes. A confused antibody decides to attack a pathogen, as well as the similar-looking myelin covering of the nerves, causing Guillain-Barré syndrome (Comic from Creative Med Doses) 3. Microbiome changes You might expect a stomach bug like norovirus to change the gut microbiome for the worse. Surprisingly, respiratory viruses such as Influenza, RSV, and Covid all harm the gut microbiome too. This is bad news, since the gut microbiome helps to regulate the immune system and produces neurotransmitters for our brain. 4. Immune Dysregulation There are a bunch of ways that the immune system can malfunction (including the ones listed above). Measles can cause immune amnesia, where the immune system forgets previous infections it had learned to fight, leading people to catch the exact same diseases again. There is growing evidence that covid has a negative impact on the immune system as well. 5. Reactivation of other pathogens Infection with a new virus can wake up old infections that were sleeping quietly in your cells. It is unfair, but sometimes viruses will gang up on you, re-activating other viruses (or bacteria) that weren’t bothering you before. 6. Cardiac damage Chickenpox/Shingles, Influenza, and Covid all raise the risk of heart attacks and strokes. Viruses have many ways of harming our cardiac systems: inflammation, damage to the blood vessles, increased blood clots, and damage to the heart. From a meta-analysis of 48 studies about respiratory viruses triggering heart attacks & strokes (Nguyen, et al, 2025) 7. Cancer Cancer involves a failure of the immune system to kill cells that have gone rogue and turned over to the dark side. In 2008, it was estimated that viral infections contribute to 15-20% of human cancer cases. Additional research further linking viruses and cancer has come out since then, so the percentage may be higher now. Both flu and covid infections can reawaken “sleeping” cancer cells that had previously been in remission or cause cancer to spread. An article from MD Anderson on 8 Viruses that Cause Cancer 8. Cumulative impacts You might hope that you could catch a virus, get it over, and be done with it. Unfortunately, that is often not the case. A young college student was fine after having covid twice, but then struggled to walk short distances after her 3rd infection. A Colorado newspaper columnist was skiing, biking, mountain climbing, and running half-marathons up until his 5th covid infection. At this point, he developed pain, fatigue, and migraines that prevent him from doing the activities he loves most. These are not just isolated anecdotes, research confirms the cumulative dangers of repeat infections. In children, a second covid infection is more likely to cause Long Covid than the first infection. Whatever your previous history, reducing risk of future infections is a worthwhile goal. The above mechanisms are not exclusive. For example, some microbiome changes can make it easier for pathogens to pass from the gut into the bloodstream and provoke an autoimmune reaction (a process I talked about in this 5-minute video) The Paradigm Shift For most viruses, people focus on just a few weeks of initial symptoms. This is the wrong way to think about infections. Viral meningitis or EBV increases your risk of Alzheimer’s or dementia, 5-15 years later. Chicken pox (varicella zoster virus) can reactivate decades afterwards as shingles, which itself then leads to increased risk of stroke for at least the following year. We need to radically change how we think about viruses. There is much we still don’t know about the immune system. Early during the covid pandemic, many experts made definitive statements about the risks of covid, assuming that those who didn’t die in the first few weeks must be completely fine. However, perturbations from infections that initially seem minor can have far-reaching, long-lasting, and time-delayed impacts. There is a ton that is still unknown. Trying to figure out how viruses hijack cell processes, alter microbiomes, and dysregulate the immune system are complex questions. Researching these areas with curiosity, determination, and an open mind will reveal a lot. Reasons for Hope It can be gloomy to think about all the damage viruses can cause. The good news is that we don’t have to resign ourselves to these outcomes. Facing the disturbing reality that many viruses are worse than we thought is just the first step towards coming up with creative new solutions. There are some bright, curious, and determined people focused on these problems, although we need even more hands and brains to get involved. The breadth and depth of the harms caused by viruses can focus biomedical research in new directions. Most viruses do not have effective anti-viral treatments. This creates a huge need. Scientific inquiry can fail catastrophically when those closest to the problem are not included. Patient-led research gives me hope, because it is centered on the expertise of those closest to the problem. I am also optimistic about the use of AI for discovering new drugs and designing immune therapies. On the prevention side, reducing how frequently people get sick will have a big impact. Different viruses spread in different ways. In recent years, we have learned that many infections are airborne. Healthy indoor air is a human right, like access to clean drinking water. The UN recently held a high-level event focused on the right to clean air. There are many measures we can take to reduce transmission of airborne diseases, such as improved ventilation, air purification, and far-UVC technologies. Parliament houses, venues for elites, and barns for pigs have already received these air quality upgrades. We need children in schools, employees in workplaces, and patients in hospitals to get the same protections. Hopefully, we are on the cusp of a clean air revolution, with more people and organizations recognizing that healthy indoor air is essential. N95 masks offer an immediate way to significantly reduce how often you get sick. Thankfully, the N95s available currently are more comfortable and more effective than the surgical or cloth masks that many of us wore back in 2020. On the brink of an indoor air quality revolution Conclusion Viruses can harm our cardiac health and cognition, and increase our chances of cancer. If this revelation is fully realized, it will change how the field of medicine operates, priorities in research funding, and public policy on everything from indoor air quality standards to paid sick leave and school attendance. I believe we are on the threshold of what could be a drastic shift in better understanding, preventing, and treating viruses, thus unlocking longer and healthier lives. Related posts you may also be interested in: 5 Devious Tricks Pathogens Use Against Us Viruses are weirder, worse, & more preventable than you realise Viruses: The Silent Triggers of Autoimmune and Neurodegenerative Diseases Your Immune System is Not a Muscle If you enjoy my posts, please subscribe to be notified of new posts via email: I look forward to reading your responses. Create a free GitHub account to comment below.

6th Oct 2025 1 votes
Deep learning gets the glory, deep fact checking gets ignored

Deep learning is glamorous and highly rewarded. If you train and evaluate a Transformer (a state-of-the-art language model) on a dataset of 22 million enzymes and then use it to predict the function of 450 unknown enzymes, you can publish your results in Nature Communications (a very well-regarded publication). Your paper will be viewed 22,000 times and will be in the top 5% of all research outputs scored by Altmetric (a rating of how much attention online articles receive). However, if you do the painstaking work of combing through someone else’s published work, and discovering that they are riddled with serious errors, including hundreds of incorrect predictions, you can post a pre-print to bioRxiv that will not receive even a fraction of the citations or views of the original. In fact, this is exactly what happened in the case of these two papers: Functional annotation of enzyme-encoding genes using deep learning with transformer layers | Nature Communications Limitations of Current Machine-Learning Models in Predicting Enzymatic Functions for Uncharacterized Proteins | bioRxiv A Tale of two Altmetric Scores This pair of papers on enzyme function prediction make for a fascinating case study on the limits of AI in biology and the harms of current publishing incentives. I will walk through some of the details below, although I encourage you to read the papers for yourself. This contrast is a stark reminder of how hard it can be to evaluate the legitimacy of AI results without deep domain expertise. The Problem of Determining Enzyme Function Enzymes are what catalyze reactions, so they are crucial for making things happen in living organisms. Enzyme Commission (EC) numbers provide a hierarchical classification system for thousands of different functions. Given a sequence of amino acids (the building blocks of all proteins, including enzymes), can you predict what the EC number (and thus, the function) is? This seems like a problem that is custom-made for machine learning, with clearly defined inputs and outputs. Moreover, there is a rich dataset available, with over 22 million enzymes and their EC numbers listed in the online database UniProt. An Approach with Transformers (AI model) A research paper used a transformer deep learning model to predict the functions of enzymes with previously unknown functions. It seemed like a good paper! The authors used a reasonable, well-regarded neural network architecture (two transformer encoders, two convolutional layers, and a linear layer) that had been adopted from BERT. They looked at regions with high attention to confirm that these were biologically significant, which suggests that the model had learned underlying meaning and provided interpretability. They used a standard training, validation, and test split on a dataset with millions of entries. The researchers then applied the model to a dataset where no “ground truth” was known to make ~450 novel predictions. For these novel predictions, they randomly selected three to test in vitro and confirmed that the predictions were accurate. A transformer model, shown on the left, was used to predict Enzyme Commission numbers for uncharacterized enzymes in E. coli. Three of these were tested in vitro (Fig 1a and Fig 4 from Kim, et al.) The Errors The Transformer model in the Nature Communications paper made hundreds of “novel” predictions that are almost certainly erroneous. The paper had followed a standard methodology of evaluating performance on a held-out test set, and did quite well on that (although later investigation suggests there may have been data leakage). The results claimed for enzymes where no ground truth is known were full of errors. For instance, the gene E. coli YjhQ was predicted to be a mycothiol synthase, but mycothiol is not synthesized by E. coli at all! The gene yciO, which evolved from the gene TsaC, had already been shown a decade earlier in vivo to not have the same function as TsaC, yet the Nature Communications paper concluded it did have the same function. Of the 450 “novel” results given in the paper, 135 of these results were not novel at all; they were already listed in the online database UniProt. Another 148 showed unreasonably high levels of repetition, with the same very specific enzyme functions reappearing up to 12 times for genes of E. coli, which biologically implausible. Most of the “novel” results from the transformer paper were either not novel, unusually repetitious, or incorrect paralogs (Fig 5 from de Crecy, et al.) The Microbiology Detective How did these errors come to light? After the model had been trained, validated, and evaluated on a dataset involving millions of entries, it was used to make ~450 novel predictions, and three of these were tested in vitro. It just so happens that one of the enzymes selected for in vitro testing, yciO, had already been studied extensively over a decade earlier by Dr. de Crécy-Lagard. When Dr. de Crécy-Lagard read that deep learning had predicted that yciO had the same function of another gene, TsaC, she knew from her long years in the lab that this was incorrect. Her previous research had shown that the TsaC gene is essential in E. coli even if yciO is present in the same genome and even when yciO gene is overexpressed. Moreover, the yciO activity reported by Kim et al. is more than four orders of magnitude (i.e. 10,000 times) weaker than that of TsaC. All this suggests that yciO does NOT serve the same key function as TsaC. Two enzymes with a common evolutionary ancestor, but different functions (Fig 7 from de Crecy, et al.) YciO and TsaC do have structural similarities, and YciO evolved from an ancestor of TsaC. Decades of research on protein and enzyme evolution have shown that new functions often evolve via duplication of an existing gene, followed by diversification of its function. This poses a common pitfall in determining enzyme function, because the genes will have many similarities with the ones they duplicated and then diversified from. Thus, looking at structural similarities is only one type of evidence for considering enzyme function. It is also crucial to look at other types of evidence, such as neighborhood context of the genes, substrate docking, gene co-occurrence in metabolic pathways, and other features of the enzymes. It is important to look at multiple types of evidence when classifying enzyme function (Fig 2 from de Crecy, et al.) Hundreds of Likely Erroneous Results Spotting this one error inspired de Crécy-Lagard and her co-authors to take a closer look at all of the enzymes found to have novel results in the Kim, et al, paper. They found that 135 of these results were already listed in the online database used to build the training set and thus not actually novel. An additional 148 of the results contained a very high level of repetition, with the same highly specific functions reappearing up to 12 times. Biases, data imbalance, lack of relevant features, architectural limitations, or poor uncertainty calibration can all lead models to “force” the most common labels from the training data. Other examples were proven wrong via biological context or a literature search. For instance, the gene YjhQ was predicted to be a mycothiol synthase but mycothiol is not synthesized by E. coli. YrhB was predicted to synthesize a particular compound, which was already predicted to be synthesized by the enzyme QueD. A form of E. coli with a QueD mutant was unable to synthesize the compound, showing that this is not in fact the function of YrhB. Rethinking Enzyme Classification and “True Unknowns” Identifying enzyme function actually consists of two quite different problems which are commonly conflated: propagating known function labels to enzymes in the same functional family discovering truly unknown functions The authors of the second paper observe, “By design, supervised ML-models cannot be used to predict the function of true unknowns.” While machine learning can be useful for propagating known functions to additional enzymes, there are many types of errors that can occur: including failing to propagate labels when they should, propagating labels when they should not, curation mistakes, and experimental mistakes. Unfortunately, erroneous functions are being entered into key online databases such as UniProt, and this incorrect data may be further propagated if it is used to train prediction models. This is a problem that increases over time. Need for Domain Expertise It is not news that AI work will be more highly rewarded and supported than work that closely inspects the underlying data and integrates deep domain knowledge. The aptly titled “Everyone Wants to do the Model Work, not the Data Work” paper involving dozens of machine learning practitioners working on high-stakes AI projects and found that inadequate-application domain expertise was one of a few key causes of catastrophic failures. Sources of cascading failures in machine learning systems (Fig 1 from Sambasivan, et al.) These papers also serve as a reminder of how challenging (or even impossible) it can be to evaluate AI claims in work outside our own area of expertise. I am not a domain expert in the enzyme functions of E. coli. And for most deep learning papers I read, domain experts have not gone through the results with a fine-tooth comb inspecting the quality of the output. How many other seemingly-impressive papers would not stand up to scrutiny? The work of checking hundreds of enzyme predictions is less glamorous than the work of building the AI model that generated them, yet it is even more important. How can we better incentivize this type of error-checking research? At a time when funding is being slashed, I believe we should be doing the opposite and investing even more into a range of scientific and biomedical research, from a variety of angles. And we need to push back on an incentive system that is disproportionately focused on flashy AI solutions at the expense of quality results. Related Reading: The problem with metrics is a big problem for AI Gaps and Risks of AI in the Life Sciences “AI will cure cancer” misunderstands both AI and medicine You can subscribe to be notified of new blog posts by submitting your email below: I look forward to reading your responses. Create a free GitHub account to comment below.

3rd Jun 2025 1 votes
What AI can tell us about microscope slides

The lavender images below show breast tissue. There are many questions doctors could want to answer using these images: They could want to know whether there are tumors present or not. If there is a tumor, doctors would want to classify its stage, make predictions about how likely the patient is to respond to treatment, and to detect whether the tumor has spread from another organ. All of these are questions which people are now tackling with machine learning. They fall within the area of computational pathology, often abbreviated CPath. In the past year, two CPath AI models were released which achieved state-of-the-art results. Here I will discuss an introduction to this field, what these models do, and what some key challenges are going forward. Breast tissue images from the BACH: Grand challenge on breast cancer histology CPath foundation models There is a powerful idea about how to make more accurate CPath models. Rather than train a model on a single type of tissue and a single task (e.g. identifying cancer in breast tissue), train a model on images of tissue from many different organs (breasts, lymph nodes, lungs, prostate, heart,…) and on multiple different tasks (recognizing cancer, determining the stage and subtype of the cancer, segmenting cells, and predicting treatment outcomes). Patterns learned from one dataset or one task are likely to generalize to others. Such models are known as CPath foundation models. In general, a foundation model is a machine learning model which is trained on a sufficiently diverse large dataset which can then be adapted for a range of downstream tasks. This idea is commonly used in the area of language models such as Chat-GPT and Claude.ai. Language foundation models are trained on many types of language tasks and intended to generalize across different corpuses of text (e.g. wikipedia, reddit posts, academic papers, online conversations, news articles, and more). ImageNet models trained to recognize a huge variety of different pictures often serve as foundation models for images. The success of foundation models within the areas of language and more general images is a key reason why we might expect pathology foundation models to be useful too. Tissues are groups of cells with similar structure and function. Different types of tissue within the human body include nervous, muscle, connective, and epithelial tissue. Image: Wikimedia Two notable CPath foundation models were released in 2024: Prov-GigaPath and UNI. Both models achieved state-of-the-art performance on dozens of pathology tasks (although they were not directly compared to one another). Another relevant paper (from Kaiko.ai) studied the impact of dataset size and model size on CPath model performance. Learning the Vocab Medicine is full of jargon and specialized vocabulary. Pathology refers to the study of disease. It is a broad field, and can include everything from dissecting dead bodies to analyzing blood samples. One key focus of computational pathology is analyzing and interpreting whole slide images (WSIs) and in some cases combined with accompanying meta-data about a patient. Whole slide images refers to the complete microscope slide, although in many cases the region of interest (such as particular cancerous or inflamed cells) may be much smaller, just occupying a subset of the slide. Machine learning (ML) is a subfield of Artificial intelligence (AI) which involves learning from past data, and is increasingly being used with great success in pathology. The focus of most computational pathology ML models is on images of tissue, on microscope slides. That is what we will focus on in this post as well. So Many Tasks! There are many different benchmarks that CPath models can be tested on. These involve numerous datasets: related to different areas of the body, with different sizes, and with different purposes. They also involve a variety of tasks, including binary classification, image segmentation, and outcome prediction. Prov-GigaPath attained state-of-the-art performance on 25 out of the 26 tasks it was evaluated on and UNI attained state-of-the-art performance on 34 different tasks. Here I will give examples of just 3 of these tasks. Task: prostate cancer cell grading In the 1960s, the pathologist Dr. Donald Gleason came up with a grading scale for rating cells as they progressed from normal to prostate cancer. The Gleason Grading system is still widely used and is considered a powerful predictor of how prostate cancer patients will fare. A major medical image conference (MICCAI) held a competition in 2022 for researchers to create algorithms to determine the Gleason grades when given images of prostate tissue. Examples of UNI predictions of Gleason grades for a section of prostate tissue. Figure 3b from the UNI paper The prostate tissue is shown in pink, and segments have been colored in blocks based on where they fall on the Gleason scale. Task: identifying early signs of rejection after a heart transplant Rejection is the main cause of mortality in patients who have received a heart transplant. Since the early stages of rejection can be asymptomatic, it is standard for patients to receive frequent biopsies for 1-2 years following a transplant. These are known as endomyocardial biopsies (EMB), since they remove a small sample of tissue from the inner lining (endo) of the heart (cardial) muscle (myo). Accurately interpreting the results of these biopsies is a key question. Underestimating the chance of rejection could lead to dangerous delays in treatment, but overestimating could lead to alarm and unnecessary follow-ups or treatment. Assessment of the sampled tissue by experienced pathologists has higher variability than many other tasks, such as cancer diagnosis. Deep learning is being used to tackle this task, in models such as Cardiac Rejection Assessment Neural Estimator (CRANE) and the CPath foundation model UNI. Each row shows a different sample of cardiac tissue, with a different medical issue. On the far left are the whole slide images, then zoomed in at higher resolution on a key Region of Interest (ROI). On the far right is a heat map for the most zoomed in area showing which features the algorithm has identified as significant. Figure 3 from the CRANE paper. Task: Genetic Mutations in Cancer For several common genetic mutations in tumors, there are specific drugs known to target those mutations. This has a direct application for clinical treatment. Since genetic mutations can change the form and function of cells, it is reasonable to expect that this information could be deduced from images of the cancer cells. Deep learning models have been built to identify genetic mutations from tissue slides. The benefits of using a computational approach are that it can be scaled as an increasing number of relevant genetic mutations and molecular biomarkers are being discovered. Task-specific models have been built for this, and this is one of the tasks that foundation models can be tested on. Different types of cancer listed along the y-axis and 20 common genetic mutations listed on the y-axis. Figure 1D from Kather, 2020. We need more data One key challenge in the area of CPath foundation models is gathering enough training data. The Cancer Genome Atlas (TCGA) was an ambitious project launched in 2006 by the National Cancer Institute in the USA. Over a 12 year period, samples were collected from over 11,000 patients of 33 different cancer types, and all this data was made publicly available. While this is a rich dataset and a useful resource, all 3 papers we’ve looked at concluded that TCGA is not large enough for effective foundation models. In addition to limited data size, TCGA also has limited diversity, consisting mostly of slides from the primary site of cancer, but not metastasized cancers or different types of tissues. Researchers at Kaiko.ai tested the impact of scaling both the size of their model and the size of the training dataset. While they found limited need to scale model size beyond a certain point, they found that larger datasets continued to lead to increased performance. They concluded that TCGA was likely not large enough and shared their plans to build a larger training set, and are now partnering with cancer centers across Europe to create a dataset for their model. The researchers behind two other CPath foundation models reached the same conclusion about data set size, and gathered massive datasets to train their models. This required partnering with healthcare centers. Prov-GigaPath, a model created by Microsoft Research and Providence Genomics involved data from 30,000 patients across 28 cancer centers (which are part of Providence Healthcare company). UNI, a cPath model created by a team at Harvard, MIT, and the Broad Institute, involved the creation of the Mass-100K: a dataset with over 100K whole slide images across 20 tissue types collected from Mass General Hospital, Brigham & Women’s Hospital, and Genotype-Tissue Expression (GTEx) consortium. These partnerships and curation of training datasets are currently a crucial component of building CPath foundation models. Curating datasets carefully poses many challenges as well. Combining data from different sources, which often use different protocols for how slides are sampled and prepared, can introduce significant biases. Different scales CPath foundation models face the difficulty of capturing both local patterns (that show up in a small tile within a slide) and global patterns across the whole slide. Many tiny tiles are found within a slide. Some models, such as the Hierarchical Image Pyramid Transformer (from several of the same authors as UNI), use hierarchical approaches to deal with these multiple scales. Hierarchical Structure of Whole-Slide Images, Figure 1 from Chen, et al, 2020 Other models, such as Prov-GigaPath, treat the tiles as tokens, encoding both the tiles and the slide as a whole as model inputs. Prov-GigaPath uses both a slide encoder and a tile encoder to take into account these two different scales. Treating slides as tokens, Figure 1a from the Prov-GigaPath paper In pathology clinics, diagnosis and treatment decisions are often made at the patient level, whereas CPath models are often highly focused on regions of interest. Accommodating the multiple relevant scales (small tiles, whole slides, and patient-level) for pathology is a consideration that CPath models need to balance. Going Forward It is still early in the world of CPath and there are many growth opportunities, including the continued need for large and diverse datasets, ways to further optimize model training, tasks which have previously received less focus, and the difficulties of integrating models into clinical work. As the authors of the kaiko.ai paper wrote, “We are still at the very beginning of developing a truly foundational pathology foundation model.” It is a hopeful sign that these models achieve state-of-the-art results on dozens of benchmarks, but it still remains to be seen when and how they will be used in clinical settings. Related Reading: The Most Common and Useful Neural Nets Using AI to Discover New Antibiotics AI and Immunology You can subscribe to be notified of new blog posts by submitting your email below: I look forward to reading your responses. Create a free GitHub account to comment below.

15th Jan 2025 1 votes

More in AI

AI and Existential Dread

There's a lot of polarising discourse right now about the threat AI poses to humanity. Some think it's a farce and others think we face extinction. Here are my thoughts.

23 hours ago 1 votes
The AI-as-Normal-Technology view of loss-of-control incidents

A middle ground between the cybersecurity and AI safety communities

yesterday 2 votes
AI Skills and Job Market, Q3 2026

An overview of the current state of the engineering market and the AI skills that are in demand

3 days ago 1 votes
Why Is AI Bad at Writing?

And Why It’ll Likely Stay That Way—Em Dashes or Not!

a week ago 2 votes
Wendell Berry and the Promise of the Deep Life

Wendell Berry died last week at 92 at his home in Port Royal, Kentucky, where he farmed his land using traditional techniques and wrote with ... Read more The post Wendell Berry and the Promise of the Deep Life appeared first on Cal Newport.

a week ago 1 votes
📚 BoredReading

You seem to be enjoying this.

Join free to unlock everything.

Create free account

Already have an account? Sign in