More from Rachel Thomas, PhD
As a former mathematician, I was used to nobody reading what I wrote. So when I first began blogging in 2015, I never expected that several of my blog posts would go viral or to have multiple journalists contact me (including from NPR, Wired, and Fortune), make the front page of Hacker News (over 10 times), receive conference keynote invitations, and be interviewed on podcasts. I do not consider myself a “natural” writer. In college, I tried to avoid classes that required essays, because writing was a struggle for me. It wasn’t until I was 30 that I set out to practice writing more. I share tips I use for blogging here, which include being willing to put a lot of time into a single post, incorporating high quality information, and having a clear idea of my intended audience. I have selected some of my most popular and impactful posts below. Several of these were originally posted on Medium or fast.ai (the two sites where my writing used to live). They are grouped into clusters based on theme. I hope you might enjoy reading these if you haven’t seen them before! Challenging Conventional Wisdom Questioning widely-held assumptions about tech culture, education, and health has been the basis for several popular posts. If you think women in tech is just a pipeline problem, you haven’t been paying attention (2015) In 2015, I felt burnt out and disillusioned by my experiences working in tech. I was frustrated with how much the popular conversation was still focused on “the pipeline problem”: training young girls to code while ignoring all the adult women being driven out of the tech industry by mistreatment. I spent 9 months researching and writing this post. It went viral and remains my most popular essay. This, together with my other posts, led to me being interviewed and quoted in Wired several times regarding diversity in tech (as well as other AI topics). My first post Trends to Avoid When Founding a Startup (2018) The dominant narrative for Bay Area tech startups is to try to raise venture capital, achieve exponential hypergrowth, and hire lots of computer science PhDs. I argued that these approaches not only harm employees, but lead to weaker companies and worse products. My family’s unlikely homeschooling journey (2022) Many people hold a stereotyped and outdated view of homeschooling, not realizing the explosion of innovative, non-traditional education options available in recent years. My husband and I never planned to homeschool, but we unexpectedly found that our child thrives with this approach. Your Immune System is Not a Muscle (2024) The misleadingly named “Hygiene Hypothesis” is often used to justify the misconception that all microbes are good for us. However, this theory is more accurately reframed as the “Old friends hypothesis”: humans co-evolved with friendly bacteria and some parasites. We did not co-evolve with the crowd infections of mega-cities and 100,000 global flights per day. AI Beyond Elite Institutions Machine learning isn’t just for those at billion dollar companies. These posts highlight unconventional practitioners and offer practical guidance for people in varied domains. Deep Learning: Not Just for Silicon Valley (2017) Our goal at fast.ai is making AI accessible to people outside of elite institutions, who are tackling meaningful problems in low-resource areas. This post introduced some of our earliest international fellows and the diverse range of problems they were working on. I always enjoyed writing about fascinating use cases from our deep learning community. How (and why) to create a good validation set (2017) An all-too-common scenario: a seemingly impressive machine learning model is a complete failure when implemented in production. Advice on one common culprit of this, and how to avoid it. In the early years of fast.ai, I wrote numerous posts with practical advice for machine learning. This article asked the question, “Can A.I. conquer its Excel problem? An Introduction to Deep Learning for Tabular Data (2018) Deep learning is not just for images and text. Companies such as Pinterest and Instacart are also applying it to tabular data, the type of data you might normally put in a spreadsheet. This post caught the attention of a reporter with Fortune, who ended up interviewing me and writing about the topic here. Debunking AI Hype & Holding Tech Accountable The narratives about AI put forth by major tech companies are often misleading about what is necessary, what values matter, and what types of harms can result. Google’s AutoML: Cutting Through the Hype (2018) In a 3-part series, I countered claims that all data scientists need customized, bespoke neural network architectures. While I was nervous about disagreeing with both Google’s CEO and head of AI, my posts led to an invitation to keynote the prestigous ICML AutoML workshop. Seven years later, my critiques have been proved valid, with transfer learning a cornerstone of ML and automated neural network search not commonly used. By 2023, we were supposed to all be using AutoML neural architecture search Five Things That Scare Me About AI (2019) AI ethics is not just a theoretical topic. I was (and still am) alarmed about the harms already being caused to human beings by AI systems irresponsibly applied to healthcare, employment decisions, policing, and more. The Problem with Metrics is a Big Problem for AI (2019) Overemphasizing metrics leads to a variety of real-world harms, including manipulation, gaming, and a myopic focus on the short-term. AI is metric optimization on steroids. I later turned this blog post into an academic paper, together with David Uminsky. Two disturbing case studies I keep returning to are how computerized algorithms have been used to cut healthcare and to fire teachers Deep learning gets the glory, deep fact checking gets ignored (2025) A microbiologist discovered hundreds of errors in a paper that used AI to classify enzymes. This is a case study of how challenging it can be to evaluate AI claims outside our area of expertise, as well as of the misaligned incentives that reward flashy results, but not diligent fact-checking. This has been by far my most popular post on LinkedIn. Immunology & Science Decoding T cells with AI (2024) T cells are a crucial component of the adaptive immune system. Accurately pedicting what they can bind to would impact a range of treatments. Numerous algorithms have been developed for this question, but the problem is far from solved. The surface of a T cell. I’ve enjoyed exploring how AI is being applied to immunology Scientists Just Connected the Dots Between Viruses and… Everything (2025) For a long time, catching frequent viruses was considered both inevitable and harmless. But it turns out that common, seemingly-mild viruses have disturbing long-term health impacts. A thread about viruses Thanks for joining me on this walk through the past! Also, you can subscribe to be notified of new blog posts by submitting your email below: I look forward to reading your responses. Create a free GitHub account to comment below.
Most people catch many viruses in their lives– for example, over 90% of adults have Epstein-Barr virus, and adults catch the flu about once every 5 years. For a long time, catching frequent viruses was considered both inevitable and harmless. But it turns out that common, seemingly-mild viruses have disturbing long-term health impacts. Common respiratory viruses increase the risk of heart attacks and strokes. Viruses are linked to dementia and Alzheimer’s Disease. They can re-awaken cancer cells in patients whose cancer was previously in remission. Persistent infections accelerate aging and undermine longevity. Viruses can be the trigger that kicks off life-long autoimmune diseases. New studies come out each week confirming that viruses can harm the health of your heart, blood vessels, brain, nervous system, and gut. Please pause and let this sink in. If we were to truly internalize this information, there would be massive shifts in the practice of medicine, scientific research, and public policy. Pathogens accelerate aging in many ways. Proal and VanElzakker, 2025 What you can avoid (infections) may be just as important as what you seek out (exercise, healthy foods). This news seems depressing. It’s too late, viruses are everywhere, everyone has already caught them– what can be done? There actually is a lot we can do. First of all, developing new anti-viral therapies and treatments should be a top priority. Second, regardless of what infections you’ve already had, preventing or reducing future infections will have a positive impact. There is exciting work happening towards both of these goals, including AI-assisted drug design, patient-led biomedical research, initiatives to improve indoor air quality and new technologies for cleaning the air. How can viruses cause all these bad outcomes when some people who catch them are fine? Human health is complicated. Disease development involves a complex interplay of factors: infections, underlying genetics, environment, the microbiome, and more. Let’s return to the example of Epstein-Barr Virus (EBV). EBV has been strongly linked to Multiple Sclerosis, prolonged fatigue, and 6 different types of cancer. Given that almost everyone has had EBV, even though “only” a percentage of people develop these lasting impacts, this is a major cause of suffering. Viruses tilt the probabilities against you A new world and a powerful idea You may wonder why so many of these health issues are on the rise, when viruses are nothing new. Our world has changed drastically in recent decades compared to most of human evolution. We live in a hyperconnected age of global mega-cities and record numbers of large international flights now. We spend our time indoors in crowded, poorly ventilated buildings. These factors have allowed viruses to travel faster and farther than ever before. The misleadingly named “Hygiene Hypothesis” is often used to justify the misconception that all microbes are good for us. However, this theory is more accurately reframed as the “Old friends hypothesis”: humans co-evolved with friendly bacteria and some parasites. Viruses are not our friends, but rather enemies. We did not co-evolve with these crowd infections of mass travel, mega-cities, and indoor confines. Not all infections are the same! Modern crowd infections are causing huge harm. Figure from Rook, 2014 The idea that viruses are contributing so much to human suffering and long-term disease is powerful. It will transform how we approach medicine, health, and aging, if we let it. This revelation is one of the key reasons that I decided to make a mid-life career pivot, stepping back from fulfiling work in AI to return to graduate school in Microbiology-Immunology, a journey I have been chronicling here on my blog. I hope to spend the next few decades applying my machine learning skills to problems at the intersection of infections, multi-omic data sets, the microbiome, and chronic disease. Below, I will share some of what has captured my attention and upended my old views on disease and medicine. Viruses have many ways to wreak havoc Viruses have evolved to evade, outmatch, commandeer, and otherwise hurt our immune systems. Here is an incomplete and overlapping list of ways that viruses can harm us: 1. Persistence Some viruses quietly stick around for years or decades after our initial illness. They may re-awaken later to cause more problems, or they may spawn surprising issues that we don’t recognize as part of our initial infection. When they persist in our cells, viruses can impact gene expression, hijacking processes our cells need to gain nutrition and energy. Dr. Amy Proal, a researcher in this area, says that treating persistent infections will be necessary to combat aging and extend healthspan. 2. Autoimmunity During an infection, sometimes our immune cells get confused into attacking our own tissue that may “look” similar to the virus (this process is known as molecular mimicry). Once it has mistakenly learned to attack self-tissue, the immune system may continue to do so, even after the virus has been defeated. This is just one of several ways by which viruses can trigger autoimmune diseases such as Lupus, Multiple Sclerosis, Rheumatoid Arthritis, or Type 1 Diabetes. A confused antibody decides to attack a pathogen, as well as the similar-looking myelin covering of the nerves, causing Guillain-Barré syndrome (Comic from Creative Med Doses) 3. Microbiome changes You might expect a stomach bug like norovirus to change the gut microbiome for the worse. Surprisingly, respiratory viruses such as Influenza, RSV, and Covid all harm the gut microbiome too. This is bad news, since the gut microbiome helps to regulate the immune system and produces neurotransmitters for our brain. 4. Immune Dysregulation There are a bunch of ways that the immune system can malfunction (including the ones listed above). Measles can cause immune amnesia, where the immune system forgets previous infections it had learned to fight, leading people to catch the exact same diseases again. There is growing evidence that covid has a negative impact on the immune system as well. 5. Reactivation of other pathogens Infection with a new virus can wake up old infections that were sleeping quietly in your cells. It is unfair, but sometimes viruses will gang up on you, re-activating other viruses (or bacteria) that weren’t bothering you before. 6. Cardiac damage Chickenpox/Shingles, Influenza, and Covid all raise the risk of heart attacks and strokes. Viruses have many ways of harming our cardiac systems: inflammation, damage to the blood vessles, increased blood clots, and damage to the heart. From a meta-analysis of 48 studies about respiratory viruses triggering heart attacks & strokes (Nguyen, et al, 2025) 7. Cancer Cancer involves a failure of the immune system to kill cells that have gone rogue and turned over to the dark side. In 2008, it was estimated that viral infections contribute to 15-20% of human cancer cases. Additional research further linking viruses and cancer has come out since then, so the percentage may be higher now. Both flu and covid infections can reawaken “sleeping” cancer cells that had previously been in remission or cause cancer to spread. An article from MD Anderson on 8 Viruses that Cause Cancer 8. Cumulative impacts You might hope that you could catch a virus, get it over, and be done with it. Unfortunately, that is often not the case. A young college student was fine after having covid twice, but then struggled to walk short distances after her 3rd infection. A Colorado newspaper columnist was skiing, biking, mountain climbing, and running half-marathons up until his 5th covid infection. At this point, he developed pain, fatigue, and migraines that prevent him from doing the activities he loves most. These are not just isolated anecdotes, research confirms the cumulative dangers of repeat infections. In children, a second covid infection is more likely to cause Long Covid than the first infection. Whatever your previous history, reducing risk of future infections is a worthwhile goal. The above mechanisms are not exclusive. For example, some microbiome changes can make it easier for pathogens to pass from the gut into the bloodstream and provoke an autoimmune reaction (a process I talked about in this 5-minute video) The Paradigm Shift For most viruses, people focus on just a few weeks of initial symptoms. This is the wrong way to think about infections. Viral meningitis or EBV increases your risk of Alzheimer’s or dementia, 5-15 years later. Chicken pox (varicella zoster virus) can reactivate decades afterwards as shingles, which itself then leads to increased risk of stroke for at least the following year. We need to radically change how we think about viruses. There is much we still don’t know about the immune system. Early during the covid pandemic, many experts made definitive statements about the risks of covid, assuming that those who didn’t die in the first few weeks must be completely fine. However, perturbations from infections that initially seem minor can have far-reaching, long-lasting, and time-delayed impacts. There is a ton that is still unknown. Trying to figure out how viruses hijack cell processes, alter microbiomes, and dysregulate the immune system are complex questions. Researching these areas with curiosity, determination, and an open mind will reveal a lot. Reasons for Hope It can be gloomy to think about all the damage viruses can cause. The good news is that we don’t have to resign ourselves to these outcomes. Facing the disturbing reality that many viruses are worse than we thought is just the first step towards coming up with creative new solutions. There are some bright, curious, and determined people focused on these problems, although we need even more hands and brains to get involved. The breadth and depth of the harms caused by viruses can focus biomedical research in new directions. Most viruses do not have effective anti-viral treatments. This creates a huge need. Scientific inquiry can fail catastrophically when those closest to the problem are not included. Patient-led research gives me hope, because it is centered on the expertise of those closest to the problem. I am also optimistic about the use of AI for discovering new drugs and designing immune therapies. On the prevention side, reducing how frequently people get sick will have a big impact. Different viruses spread in different ways. In recent years, we have learned that many infections are airborne. Healthy indoor air is a human right, like access to clean drinking water. The UN recently held a high-level event focused on the right to clean air. There are many measures we can take to reduce transmission of airborne diseases, such as improved ventilation, air purification, and far-UVC technologies. Parliament houses, venues for elites, and barns for pigs have already received these air quality upgrades. We need children in schools, employees in workplaces, and patients in hospitals to get the same protections. Hopefully, we are on the cusp of a clean air revolution, with more people and organizations recognizing that healthy indoor air is essential. N95 masks offer an immediate way to significantly reduce how often you get sick. Thankfully, the N95s available currently are more comfortable and more effective than the surgical or cloth masks that many of us wore back in 2020. On the brink of an indoor air quality revolution Conclusion Viruses can harm our cardiac health and cognition, and increase our chances of cancer. If this revelation is fully realized, it will change how the field of medicine operates, priorities in research funding, and public policy on everything from indoor air quality standards to paid sick leave and school attendance. I believe we are on the threshold of what could be a drastic shift in better understanding, preventing, and treating viruses, thus unlocking longer and healthier lives. Related posts you may also be interested in: 5 Devious Tricks Pathogens Use Against Us Viruses are weirder, worse, & more preventable than you realise Viruses: The Silent Triggers of Autoimmune and Neurodegenerative Diseases Your Immune System is Not a Muscle If you enjoy my posts, please subscribe to be notified of new posts via email: I look forward to reading your responses. Create a free GitHub account to comment below.
DNA sequencing hasn’t lived up to the hype Twenty to thirty years ago, politicians, scientific leaders, journalists, and even Nobel laureates predicted that sequencing the human genome would revolutionize how we treat disease. And while the advances in DNA sequencing that have occurred since then have improved recognition and treatment for some cancers and rare diseases, on the whole the field has not lived up to earlier hype. Time Magazine covers from 1994 and 1999 about genetics In an article titled “Why sequencing the human genome failed to produce big breakthroughs in disease”, a biology professor highlights that most common diseases are not caused by a single gene. In fact, common diseases are often linked to hundreds of gene variants, and even collectively, these variants still account for only a small fraction of disease variance. Here, I want to focus on two other key limitations of DNA sequencing, and how they are now being addressed with new approaches. What DNA can’t tell us First, DNA can’t answer many questions about how cells and organisms work in practice. A neuron in the brain has the same DNA as a liver cell, yet the two have completely different functions. This is because different segments of DNA are turned off or on in different cells. To understand how cells are actually working, you need to know about proteins and RNA (RNA is the intermediary which translates DNA into protein). Proteins are what build the structure of cells, catalyze chemical reactions within the cell, and allow communication between cells. Healthy and unhealthy cells in the same organism will usually have the same DNA. For instance, if some regions of the intestines are experiencing an IBD (irritable bowel disease) flare and others aren’t, they would all have the same DNA, yet likely different RNA and protein levels. A second big problem is that many key sequencing techniques destroy spatial information. You essentially may have to put tissue or cells into a blender in order to get rich information about DNA or RNA sequences. While this data is informative, it turns out that locations of cells within a piece of tissue, and locations of regions within a cell, are also very important! Again, considering the case in which some regions of the intestine are inflamed due to IBD, yet others aren’t, mixing them all together in a blender will lose or distort useful information. Look at the difference between crypts in a healthy segment of the colon (on the left) compared to inflamed crypts (on the right). Differences between a healthy bowel (on the left) and an inflamed bowel (on the right). Source: mypathologyreport.ca Recently, we have seen a rise in breakthroughs that allow us to obtain data about location. Spatial techniques are a necessary and exciting step beyond DNA sequencing. The power of spatial information showed up as a major theme at a conference I attended last year, and spatial techniques have been recognized by Nature Methods as “Method of the year” twice in the last 5 years. A Few Major Areas of Innovation What is genetic sequencing anyway? There are a number of different types of sequencing that have been invented in the last 30 years. Walking through a brief history will illustrate what these technologies are, and what they can and cannot do. To make it concrete, let’s look at the example of how they have been applied to cancer treatment. Sanger Sequencing: This is an older technology dating back to the 1970s, and which was the main way of sequencing DNA up until 2005. Sanger sequencing was one of the methods used in the mid-90s to identify the genes BRCA1 and BRCA2 as key genetic risk factors for breast cancer. The process of discovering BRCA1/BRCA2 involved scientists slowly zeroing in on their chromosomal locations over a period of years. Several other cancer genes were discovered during this time period as well. While Sanger sequencing is effective on smaller amounts of DNA, it can be quite slow to deal with larger volumes. It took over 10 years to sequence the first copy of the human genome using Sanger Sequencing. It is still used today as a simple and reliable way to test for known mutations (such as BRCA1/BRCA2) or for smaller tasks. High-Throughput Sequencing: New technology released in the mid-2000s allowed DNA to be cut into lots of short pieces and for millions of pieces to be sequenced in parallel at once. This approach, called high-throughput sequencing, was significantly faster than Sanger sequencing. High-throughput sequencing has many applications, including to cancer treatment, by making it cheaper and faster to sequence DNA to identify particular mutations which can influence treatment decisions. Method of the Year | Nature Methods Long-read sequencing: High-throughput sequencing has the advantage of high accuracy, but the downside of short sequence lengths. In the 2010s, technologies were released with the opposite set of strengths and weaknesses. Long-read sequencing provides the advantage of long sequence lengths, although the downside of lower accuracy. To compare, high-throughput sequencing uses DNA strands that are a few hundred base pairs long, whereas long-read sequencing uses DNA strands that are tens of thousands base pairs long. Both technologies have different strengths and are widely used today. Single-Cell Sequencing: High-throughput sequencing involves sequencing the DNA of many cells at once, but sometimes it is useful to sequence individual cells. In a tumor, different cells can have different mutations. It is possible that a small subset of the mutations may drive metastasis (the spread of cancer to other areas) or resistance to treatment. Identifying these driver mutations can guide treatment decisions, since particular driver mutations can predict the effectiveness of various drugs. Key mutations may be drowned out in the average if you sequence the entire tumor. This is one reason why it is useful to be able to sequence single cells, and not just obtain the average of many cells. Single-cell sequencing was selected as Nature Methods method of the year in 2013. A Lego interpretation of bulk RNA-seq; single-cell RNA-seq; spatial transcriptomics; and the original organ. Source: Bo Xia, @BoXia7 Multi-Omics: The study of DNA is genomics. DNA alone gives us an incomplete picture of an organism. Epigenomics can provide information about which regions of DNA are active or silenced. To understand how different cells function, as well as cells in different states of disease or health, you also need to know about their RNA and proteins. This data is contained in the field of transcriptomics (transcripts are strands of RNA transcribed from DNA) and proteomics (the proteins in a cell). Metabolomics looks at small molecules (such as sugars, amino acids, and vitamins) within the body and exposomics includes all sorts of environmental exposures. Collectively, these fields are known as -omics or multi-omics. It is valuable to combine multiple types of -omics together for richer sources of information, since each has different strengths, limitations, and insights to offer. Multi-Omics is an exciting area that draws on lots of data, with applications to cancer, infectious disease and immunology. Spatial: Spatial information lets us see all the variation within a section of the body– such as a segment of the intestines, the liver, or a cancerous tumor. This variation can often be significant for understanding disease and treatment prognosis. To better understand why spatial techniques are useful, let us dive into some background about cancer. Tumors aren’t just lumps of bad cells Cancer is defined as excess cell division. I used to think that tumors were just clumps of “bad cells”, where “good” and “bad” were binary states. This is incorrect. Tumors are not uniform, and within the category of “bad” there is a great deal of variation and heterogeneity. Different cells within a tumor may have different mutations from one another. And our immune systems sometimes build complex defense structures within tumors in attempts to more effectively fight them. How close a cancerous cell is to one of these immune structures impacts how likely the body is to destroy the cancerous cell. Notice all the variation within this tumor! Immunologist Dr. Angela Ferguson, who studies head and neck cancers, describes tumors as having a “physical landscape”. She has shown how the organization and structure within a tumor can predict and guide treatment outcomes. Her work found that cancer progression is “landscape-dependent”, where landscape refers to the locations of immune cells and structures within a tumor. It is not enough to study cancer cells in isolation. We need to understand their layout. Methods that effectively put tumor cells into a blender in order to sequence them, disrupting their spatial information, are insufficient on their own. Mapping that spatial information can hold the keys for more effective treatment. Single cell sequencing approaches allow a greater number of genes to be measured, whereas spatial approaches measure fewer genes but also provide location information. Combining these two approaches can prove powerful. Programming Libraries Applied to Spatial -Omics We are living at a time when multi-omics, spatial information, and user-friendly programming libraries are converging for easier exploration and discovery. For instance, below is an image I created using the common Python programming libraries pandas and matplotlib of data the NIH has shared about a rare liver disease. The image shows a slice of liver tissue, with gene expression overlaid in a color scale ranging from purple (low expression) to yellow (highest expression). Using Matplotlib to display a cross-section of liver with gene expression The NIH dataset contain images of slices of liver and expression of many different genes, from both healthy patients and those with a rare liver disease. https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE240429. The plot shows expression of albumin, a protein which helps transport other molecules around the body, overlaid on top of the liver, illustrating which regions produce more or less of this protein. Data science tools are invaluable for transforming and plotting such data. In addition to being able to use standard Python libraries (such as Pandas and Matplotlib) for visualing this data, there are many specialist libraries as well. The Scverse (Sc = single cell) includes a number of Python libraries focused on single cell analysis. The SC Verse includes libaries for single cell sequencing analysis There are also popular R libraries for single cell analysis and multi-omics, such as Harmony, Seurat, and mixOmics. With all of these tools and technologies, it is an exciting time to be working at the intersection of data science and microbiology. The causes of most diseases are complex and multi-factorial, although new approaches of spatial multi-omics provide unprecedented types of useful information. Related Reading: What AI can tell us about microscope slides Gaps and Risks of AI in the Life Sciences AI and Immunology You can subscribe to be notified of new blog posts by submitting your email below: I look forward to reading your responses. Create a free GitHub account to comment below.
The lavender images below show breast tissue. There are many questions doctors could want to answer using these images: They could want to know whether there are tumors present or not. If there is a tumor, doctors would want to classify its stage, make predictions about how likely the patient is to respond to treatment, and to detect whether the tumor has spread from another organ. All of these are questions which people are now tackling with machine learning. They fall within the area of computational pathology, often abbreviated CPath. In the past year, two CPath AI models were released which achieved state-of-the-art results. Here I will discuss an introduction to this field, what these models do, and what some key challenges are going forward. Breast tissue images from the BACH: Grand challenge on breast cancer histology CPath foundation models There is a powerful idea about how to make more accurate CPath models. Rather than train a model on a single type of tissue and a single task (e.g. identifying cancer in breast tissue), train a model on images of tissue from many different organs (breasts, lymph nodes, lungs, prostate, heart,…) and on multiple different tasks (recognizing cancer, determining the stage and subtype of the cancer, segmenting cells, and predicting treatment outcomes). Patterns learned from one dataset or one task are likely to generalize to others. Such models are known as CPath foundation models. In general, a foundation model is a machine learning model which is trained on a sufficiently diverse large dataset which can then be adapted for a range of downstream tasks. This idea is commonly used in the area of language models such as Chat-GPT and Claude.ai. Language foundation models are trained on many types of language tasks and intended to generalize across different corpuses of text (e.g. wikipedia, reddit posts, academic papers, online conversations, news articles, and more). ImageNet models trained to recognize a huge variety of different pictures often serve as foundation models for images. The success of foundation models within the areas of language and more general images is a key reason why we might expect pathology foundation models to be useful too. Tissues are groups of cells with similar structure and function. Different types of tissue within the human body include nervous, muscle, connective, and epithelial tissue. Image: Wikimedia Two notable CPath foundation models were released in 2024: Prov-GigaPath and UNI. Both models achieved state-of-the-art performance on dozens of pathology tasks (although they were not directly compared to one another). Another relevant paper (from Kaiko.ai) studied the impact of dataset size and model size on CPath model performance. Learning the Vocab Medicine is full of jargon and specialized vocabulary. Pathology refers to the study of disease. It is a broad field, and can include everything from dissecting dead bodies to analyzing blood samples. One key focus of computational pathology is analyzing and interpreting whole slide images (WSIs) and in some cases combined with accompanying meta-data about a patient. Whole slide images refers to the complete microscope slide, although in many cases the region of interest (such as particular cancerous or inflamed cells) may be much smaller, just occupying a subset of the slide. Machine learning (ML) is a subfield of Artificial intelligence (AI) which involves learning from past data, and is increasingly being used with great success in pathology. The focus of most computational pathology ML models is on images of tissue, on microscope slides. That is what we will focus on in this post as well. So Many Tasks! There are many different benchmarks that CPath models can be tested on. These involve numerous datasets: related to different areas of the body, with different sizes, and with different purposes. They also involve a variety of tasks, including binary classification, image segmentation, and outcome prediction. Prov-GigaPath attained state-of-the-art performance on 25 out of the 26 tasks it was evaluated on and UNI attained state-of-the-art performance on 34 different tasks. Here I will give examples of just 3 of these tasks. Task: prostate cancer cell grading In the 1960s, the pathologist Dr. Donald Gleason came up with a grading scale for rating cells as they progressed from normal to prostate cancer. The Gleason Grading system is still widely used and is considered a powerful predictor of how prostate cancer patients will fare. A major medical image conference (MICCAI) held a competition in 2022 for researchers to create algorithms to determine the Gleason grades when given images of prostate tissue. Examples of UNI predictions of Gleason grades for a section of prostate tissue. Figure 3b from the UNI paper The prostate tissue is shown in pink, and segments have been colored in blocks based on where they fall on the Gleason scale. Task: identifying early signs of rejection after a heart transplant Rejection is the main cause of mortality in patients who have received a heart transplant. Since the early stages of rejection can be asymptomatic, it is standard for patients to receive frequent biopsies for 1-2 years following a transplant. These are known as endomyocardial biopsies (EMB), since they remove a small sample of tissue from the inner lining (endo) of the heart (cardial) muscle (myo). Accurately interpreting the results of these biopsies is a key question. Underestimating the chance of rejection could lead to dangerous delays in treatment, but overestimating could lead to alarm and unnecessary follow-ups or treatment. Assessment of the sampled tissue by experienced pathologists has higher variability than many other tasks, such as cancer diagnosis. Deep learning is being used to tackle this task, in models such as Cardiac Rejection Assessment Neural Estimator (CRANE) and the CPath foundation model UNI. Each row shows a different sample of cardiac tissue, with a different medical issue. On the far left are the whole slide images, then zoomed in at higher resolution on a key Region of Interest (ROI). On the far right is a heat map for the most zoomed in area showing which features the algorithm has identified as significant. Figure 3 from the CRANE paper. Task: Genetic Mutations in Cancer For several common genetic mutations in tumors, there are specific drugs known to target those mutations. This has a direct application for clinical treatment. Since genetic mutations can change the form and function of cells, it is reasonable to expect that this information could be deduced from images of the cancer cells. Deep learning models have been built to identify genetic mutations from tissue slides. The benefits of using a computational approach are that it can be scaled as an increasing number of relevant genetic mutations and molecular biomarkers are being discovered. Task-specific models have been built for this, and this is one of the tasks that foundation models can be tested on. Different types of cancer listed along the y-axis and 20 common genetic mutations listed on the y-axis. Figure 1D from Kather, 2020. We need more data One key challenge in the area of CPath foundation models is gathering enough training data. The Cancer Genome Atlas (TCGA) was an ambitious project launched in 2006 by the National Cancer Institute in the USA. Over a 12 year period, samples were collected from over 11,000 patients of 33 different cancer types, and all this data was made publicly available. While this is a rich dataset and a useful resource, all 3 papers we’ve looked at concluded that TCGA is not large enough for effective foundation models. In addition to limited data size, TCGA also has limited diversity, consisting mostly of slides from the primary site of cancer, but not metastasized cancers or different types of tissues. Researchers at Kaiko.ai tested the impact of scaling both the size of their model and the size of the training dataset. While they found limited need to scale model size beyond a certain point, they found that larger datasets continued to lead to increased performance. They concluded that TCGA was likely not large enough and shared their plans to build a larger training set, and are now partnering with cancer centers across Europe to create a dataset for their model. The researchers behind two other CPath foundation models reached the same conclusion about data set size, and gathered massive datasets to train their models. This required partnering with healthcare centers. Prov-GigaPath, a model created by Microsoft Research and Providence Genomics involved data from 30,000 patients across 28 cancer centers (which are part of Providence Healthcare company). UNI, a cPath model created by a team at Harvard, MIT, and the Broad Institute, involved the creation of the Mass-100K: a dataset with over 100K whole slide images across 20 tissue types collected from Mass General Hospital, Brigham & Women’s Hospital, and Genotype-Tissue Expression (GTEx) consortium. These partnerships and curation of training datasets are currently a crucial component of building CPath foundation models. Curating datasets carefully poses many challenges as well. Combining data from different sources, which often use different protocols for how slides are sampled and prepared, can introduce significant biases. Different scales CPath foundation models face the difficulty of capturing both local patterns (that show up in a small tile within a slide) and global patterns across the whole slide. Many tiny tiles are found within a slide. Some models, such as the Hierarchical Image Pyramid Transformer (from several of the same authors as UNI), use hierarchical approaches to deal with these multiple scales. Hierarchical Structure of Whole-Slide Images, Figure 1 from Chen, et al, 2020 Other models, such as Prov-GigaPath, treat the tiles as tokens, encoding both the tiles and the slide as a whole as model inputs. Prov-GigaPath uses both a slide encoder and a tile encoder to take into account these two different scales. Treating slides as tokens, Figure 1a from the Prov-GigaPath paper In pathology clinics, diagnosis and treatment decisions are often made at the patient level, whereas CPath models are often highly focused on regions of interest. Accommodating the multiple relevant scales (small tiles, whole slides, and patient-level) for pathology is a consideration that CPath models need to balance. Going Forward It is still early in the world of CPath and there are many growth opportunities, including the continued need for large and diverse datasets, ways to further optimize model training, tasks which have previously received less focus, and the difficulties of integrating models into clinical work. As the authors of the kaiko.ai paper wrote, “We are still at the very beginning of developing a truly foundational pathology foundation model.” It is a hopeful sign that these models achieve state-of-the-art results on dozens of benchmarks, but it still remains to be seen when and how they will be used in clinical settings. Related Reading: The Most Common and Useful Neural Nets Using AI to Discover New Antibiotics AI and Immunology You can subscribe to be notified of new blog posts by submitting your email below: I look forward to reading your responses. Create a free GitHub account to comment below.
More in AI
Inside the harnesses that let AI carry a job across hours, days, and conversations.
how agents turn ruthless on paper first
I talk to a lot of old people, those who were born in the 20th century, and if I ask them what the word “unalive” means, they usually have no idea what I’m talking about, except for some of them who have kids or who study Internet culture. This will, of course, probably seem very weird to most people who were born in or grew up in the 21st century. Just to recap for the olds: “unalive” is the word you use to represent concepts like dying, or death, or killing or being killed, on digital platforms where saying those words accurately will cause the algorithm to punish or censor you. Or, maybe, where the perception is that using those words will result in being censored by the algorithm, and no one is actually willing to find out what happens if you use the forbidden words. This sort of attack on people’s expression started on platforms like TikTok, where nearly all content is distributed through an algorithmic feed, but has since become ubiquitous in nearly all digital media that we see. In fact, these tics are now so prevalent that it’s routine to hear people using this kind of language in everyday life, even though there’s not yet an algorithm to appease in the physical world. I’ve heard people say, out loud, “he unalived himself”, in reference to someone dying by suicide. And all of this has become even more visible in recent days as online conversation has turned to discussion of the horrific lack of accountability around the tragic rape case at Cornell University. Across the Internet, people are routinely referring to the central crime in the case as r*pe or “grape” or even using the 🍇 emoji, without a second thought for what it means that the very word can’t be said online anymore. Or, at least, the assumption is that it can’t be said. To be clear, I am very much in favor of people using content warnings or sensitivity markers for content, and fine with people using abbreviations like “SA” for references to disturbing or triggering topics like sexual assault; we should provide people with as much context and control as possible when choosing what information they want to consume and when. I also know that sometimes, people use lesser terms for stressful subjects like death or assault to create a bit of ironic distance from painful or upsetting topics. But most of the different variations of wording and emojis are coming from trying to appease the platforms, and there’s a heavy cost for those who are worried about being mindful: If someone is using a tool to filter out content, it will no longer be effective because everyone is using misspellings and euphemisms and imagery to get around the algorithm. The spread of censored and mangled syntax is happening because people believe, or have experienced, that platforms will silence them for accurately describing the world in plain language. This shit is terrible, and it has to stop. You Were Not Born With These Constraints One of the things that’s most concerning to me is that an entire generation has grown up not realizing how extreme it is that their very language is being chosen for them by platforms run by people who hate that generation’s ability to express itself, and who hate the things it has to say. From their youngest days, this generation grew up watching people make stupid faces at them for YouTube thumbnails and never had a chance to reflect on the fact that those creators didn’t want to be humiliating themselves by making those expressions — the demands of the algorithms of Big Tech forced them to do that. The rituals of feeding the algorithm are so built into people’s everyday habits that they’re invisible to people who weren’t alive before today’s platforms took over. Every parent of my cohort remembers the first time they heard their toddler finish doing something cute in their living room, and then turn around and say, “please like and subscribe!” afterwards. It’s a ghastly, sickening feeling to confront the fact that our little kids were being brainwashed into thinking that every adorable thing they did should be followed by a prompt to provide data to Google. Over on Instagram, where people originally signed up thinking they were going to see someone’s vacation pictures, or shots of their cousin’s kids, you’re now stuck watching people beg for everyone to reply with cultish phrases in the comments, which will then earn them an obviously AI-generated response in return, all in service of “showing activity” to the algorithm, like it’s an angry god that needs a sacrifice. They’re just not sure exactly what the angry god wants. Your free speech was taken away from you, and the people who did it are the same ones who spent years pretending to care about “free expression”. They contrived examples of lack of free speech on college campuses while squashing protests, and cried crocodile tears about “cancel culture” while getting people fired for political criticism. Now they have no problem with billionaires deciding exactly what words everyone is allowed to say. Larry Ellison is not content with his family owning all of the movies and TV shows — his family has to control what words people are allowed to speak on TikTok, too. Elon Musk isn’t content to merely generate and distribute child sexual abuse material for profit — he wants to silence the messages of the few decent people who are foolish enough to remain on Twitter/X, too. (That’s why I wrote you a guide on how to get your organization off of that cursed platform.) Now that an entire generation has grown up using these Orwellian euphemisms, and all of the Big AI products are trained on the Internet that was created under this regime, do you think today’s AI tools even know that the real, uncensored world exists? If you can’t say “genocide” on any of the major platforms, yet those are the ones all of the Big AI tools used as their training data... well, then the AI tools sure aren’t very likely to know much about genocide, are they? Fuck the Algorithm Our creativity can be constrained by the language we use — our imaginations are limited by what we can think to say. If we’re trained to limit the words we speak just by habit, and those limits are put in place by people whose social, political, cultural and moral goals are the opposite of what we value, then our work is unalive before it is even born. The answer to this is simple: say what you mean. This will take, to some degree, courage. It may even take, I hesitate to say, some sacrifice. When I suggest this course of action to people, they inevitably say, “But it will cost me audience!” or “But what if I lose followers!” or “What if they demonetize me!” Okay, what if they do? What if they do. Are you willing to push on this? To make a point about it? To move to platforms where you can actually say what you mean? Or to remember that you already have platforms where you can say what you mean? On an email newsletter or podcast you can say whatever the hell you want and nobody can stop you. On my blog right here, I can even curse in a headline and it won’t affect anything about how my site operates. (And a reminder: Substack is not an email newsletter, and a Spotify show is not a podcast — they’ll be unaliving your distribution any day now.) If you are a 20th century relic like me, it is incumbent upon you to remind the generation that grew up inside the algorithm that another world is possible, and that we know this because we lived it. We were able to style a MySpace page in any way that we wanted; the code for LiveJournal was entirely open source so there was no part of the algorithm that was unknowable. A blog like the one you’re reading right now could be made by anyone, and put up for pennies, and nobody could stop it from being read by millions of people. (And that last one? It’s still possible.) If you are from this century, forget all that rambling bullshit about ancient history: all that matters is you getting what you deserve, because you’ve been fucked over by the same billionaires who’ve poisoned your planet and infested your world with slop. The best artists around you are invisible to you, and the most important statements by activists that you care about are being silenced. It’s not a conspiracy, it’s a system working as designed. And the proof is as obvious as the fact that the angriest activists you know can’t even talk about systemic abuses or state violence without having to put it in algorithmically approved speech or censoring their captions like they’re going to be read by 5-year-olds. It should make you furious. It is time to kill “unalive”. The response is simple: For every message you put out, start by saying what you mean. Don’t work backwards from what the algorithm wants or what a platform permits. Build a presence on every platform you can, even the ones where you have fewer followers or where you’re harder to find. Tell your audience that your speech being free matters more than corporate convenience. Keep saying it, and keep it positive: independence gets them better art, better information, and more connected communities. Build alliances with other artists, activists and people who share your values, and let them know you’re going to start sharing your work uncensored. Start releasing your work uncensored and see where the platforms push back. (You may be surprised: sometimes you were censoring yourself in anticipation of limits that weren’t even there.) If a platform does try to limit your reach or expression, make a LOT OF NOISE about it. Tell the press, rally your alliance, and spread the word on your other platforms, using the moment to build audience and raise support there. Get others to amplify the parts of your work that don’t violate platform policy, so the controversy drives people to the rest. Find the others pushing back on algorithmic control of expression, and raise and praise their work when they do the same. If we keep accepting the words that are forced upon us by TikTok and Meta and Google and all the rest, while platforms like Twitter/X allow the most hateful and harmful content in the world to be distributed completely unfettered, we’ll only see authoritarianism rise, and the harms against the vulnerable accelerate. But what breaks my heart almost as much is that we’ll see so many brilliant artists and activists and thinkers whose genius will be muted or silenced by mindless, heartless algorithms that capriciously decide who gets to say exactly what words, in what ways. I get angry every time I think about it. The tech tycoons get ever more brazen in what they’re willing to say publicly, boasting about how they’re going to cause the end of the world, or calling for ethnic cleansing, all while putting tighter and tighter reins on the speech and expression of ordinary people. It’s time for “unalive” to die.
More than a year ago I wrote a few posts here that recommended people not to load custom tools into their context (or MCP servers) but to just use more scripts. Most importantly I wrote that Code Is All You Need and I wrote about that MCP needs code. With Pi 1.0 we now added MCP support via Codemode which in some ways is a long time coming, but then also maybe somewhat surprising to some. So I want to share some updated thoughts on this blog on what this all means. What Are Tools When a harness like Pi provides tools for an LLM to call, it does so by supplying some tool definitions which then translate into some token structure on the server side. Whether a model is encouraged to call a tool is the result of the reinforcement learning process. Something I wrote about before if you want to learn more. One of the reasons we strongly lean towards CLI and bash is because it allows easy composition of calls, and because the model also learns how the file system works when it’s trained. So when it invokes a tool like echo foo > /tmp/test.txt the model also learns that after that tool call, there is now a file called test.txt in /tmp. However bash has one fundamental limitation which is that it can only compose programs that run. And there are some things, which are not programs, but native tools to the LLM and they sort of have to be. The most obvious example here is read or view_image. If a multimodal model needs to read an image, it cannot use cat for that because the harness needs to inject the actual image payload into the protocol of the LLM. Another quite vivid example are sub agents. In order to spawn and orchestrate sub agents, it’s tricky to avoid tools that are provided by the harness. While in theory the agent could provide a CLI tool that talks to the outer harness via environment variables and Unix sockets, it’s a rather crude process. It however has another issue, and that is where the code runs. Brains vs Hands To better understand that, it’s important to think a bit more about where all the bits and pieces run. There really usually are two different systems involved. The first is the brain, the harness: it runs on one machine. It’s trusted. The second is often the same machine, but it’s really where the tools are executing: the hands. In Pi we now call this the execution environment, but you can think of it as the target of all the operations. Crucially what is important for us, is that there is a dividing line between the harness brain and the target environment that runs bash and executes the tools. And splitting this in half has some really important consequences. For a start it means that they are running on different file systems and they have different levels of trust. If you for instance use a sandboxing solution like Gondolin your bash stuff will be sandboxed just fine, but the harness itself will not be. Orchestrating The Harness Which brings us to what Codemode really does: it’s a way for the LLM to express and orchestrate complex operations on the harness side, but not the execution environment side. Codemode runs in the harness, in its own sandbox. In case of Pi it’s running in QuickJS within a WASM runtime with intentional limitations: no network, no file system, no timers, limited RAM. The only way is to call more tools. You could also imagine that Codemode could run Scheme or some other language as well. If you are not familiar with Codemode, it’s basically just a way to issue tool calls from within some language, in our case JavaScript. That allows you to compose those calls without necessarily going through the LLM’s context. Credit for naming goes to our friends at Cloudflare who coined it. For instance if you issue a bash call as a regular tool call in the LLM, then we only throw the trailing 2000 lines into the context and if the agent wants more, it needs to look at the overflow file itself. If however the agent issues that invocation via Codemode, then the Codemode side gets larger outputs sent structurally. Most importantly, because Codemode is JavaScript the agent can express concurrent operations and basic workflows. A common way in which you see agents now use this, is to first probe at 5-10 items from some tool response to see what it looks like, and to then write a Codemode script that processes the next n items. Codemode also allows you to throw state into the transcript! That means that one Codemode invocation can stash away data, that the next call in the session can load again. And remember: this is on the harness host, not the sandbox. In case of Pi, Codemode also allows you to issue calls that naturally do not make any sense in Pi’s traditional interface. For instance if you want to generate images with an image model or you want to classify some text with a one shot classifier model, those Pi APIs are exposed via Codemode, but not via regular tools where they would just waste context. What It Looks Like So now that we talked a bunch about it, it’s probably worth being a bit more explicit about it. Let’s walk ourselves through some invocations of Codemode of recent Pi sessions of mine. Note that none of this code is human written. It’s from real sessions of Pi, just re-indented for your viewing pleasure. The agent starts using Codemode automatically either because it’s a task where the model already naturally picks up that tool, or because a user asked it to. Note that Codemode is by default only enabled in Pi when MCP is enabled, but you can turn it on with "defaultTools": ["+codemode"] in the settings. Just ask Pi to enable it for you. Generating Images Let’s start simple with image generation. Image generation is a feature that Pi supports in the AI SDK core, but it’s not a tool that the agent can use. In the past the only way to use image models has been to write a bespoke extension or to have the agent run node itself and use the internal image APIs. However because we expose quite a few of the internal model APIs within Codemode, it means that the agent can use it: const [painter] = await models.getAvailableOfType("image"); const result = await models.generateImages(painter, { input: [{ type: "text", text: "A cute little puppy sitting on a grassy " + "lawn, soft natural light, photorealistic" }], }); if (result.stopReason !== "stop") return result.errorMessage; for (const block of result.output) { if (block.type === "image") image(block); else text(block.text); } Note that the call to image() sends the image back as image content to the LLM. On the harness side it feeds it directly into both the agent, as well as onto disk as a temporary artifact in case the agent wants to be able to pass that image back to bash. Classifying Things Similar things apply to classifier models such as Jev. They also do not fit well into the workflows of an agent through the typical tools. But rather than making a bespoke tool available, Codemode just allows the agent to reach into the AI SDK and invoke those directly. Here you can see how Jev is used to mass process GitHub issues for a quick sentiment analysis: const jev = await models.getModelOfType("classifier", "typesafe", "jev-latest"); const r = await tools.bash({ command: "gh issue list --state open --limit 100 " + "--json number,title,body,comments", }); const issues = JSON.parse(r.output); const results = await Promise.all(issues.map(async (issue) => { const res = await models.classify(jev, { state: { title: issue.title, body: (issue.body || "").slice(0, 4000), comments: issue.comments.slice(-5).map(c => c.body.slice(0, 800)), }, questions: { sentiment: { type: "choice", instructions: "What is the overall sentiment of the author towards pi?", criteria: { positive: "Appreciative, happy, constructive praise", neutral: "Matter-of-fact report or request without emotion", negative: "Frustrated, annoyed, upset, or angry", }, }, frustration: { type: "score", instructions: "How frustrated is the reporter?", criteria: ["not at all", "mildly", "clearly frustrated", "very angry"], }, kind: { type: "choice", instructions: "What kind of issue is this?", criteria: { bug: "Bug report or regression", feature: "Feature request or enhancement", question: "Question or support request", other: "Docs, discussion, meta, spam", }, }, }, }); if (res.stopReason !== "stop") { return { n: issue.number, title: issue.title, error: res.errorMessage }; } return { n: issue.number, title: issue.title, ...res.answers }; })); store("sentiment_results", results); return results .filter(r => !r.error) .sort((a, b) => b.frustration.score - a.frustration.score) .slice(0, 12) .map(r => `#${r.n} ${r.frustration.score.toFixed(2)} [${r.kind.choice}] ${r.title}`); Note how in that above example we also call store() which dumps the result of that execution into the session transcript. A future invocation of Codemode can thus read back that result if it wants to. The Promise.all here is fine, because Pi limits the total number of concurrent tool executions itself to four and maintains a queue for the rest. A more adventurous example is to use Jev to drive a game engine for debugging purposes: Codemode with Jev for Game Debugging Here it knows about my tankctl command and it built itself quickly a minimal harness around it to drive a game loop to assist a user with debugging a problem. Note how it built a 30 step loop in which each step goes back to both the game engine to get a text dump of what’s going on, and then to Jev to determine what to do next: const jev = await models.getModelOfType("classifier", "typesafe", "jev-latest"); const tank = async (cmd) => (await tools.bash({ command: `tools/tankctl "${cmd}"` })).output; await tank("start --map assets/maps/night_arena.map"); const questions = { action: { type: "choice", instructions: "You control the tank '@' in a top-down tank game. " + "Choose the best next action.", criteria: { attack: "an enemy has line of sight to you and you can fire at it", approach: "no enemy has line of sight; drive toward the nearest enemy", dodge: "an enemy shot is heading at you and will hit soon", powerup: "a powerup is close and no enemy threatens you", }, }, }; function commandFor(choice, st) { const p = st.player; const enemy = st.enemies.filter(e => !e.dead) .sort((a, b) => (b.los - a.los) || (a.dist - b.dist))[0]; if (choice === "attack" && enemy) { return `fire_at tank ${enemy.id}; frames 30 until clear,damage,kill`; } if (choice === "dodge") { // move perpendicular to the closest incoming shot const s = st.projectiles.filter(s => !s.yours) .sort((a, b) => a.eta - b.eta)[0]; const dir = s && Math.abs(s.vel[0]) > Math.abs(s.vel[1]) ? (p.pos[1] > s.pos[1] ? "+down" : "+up") : (p.pos[0] > (s ? s.pos[0] : 0) ? "+right" : "+left"); return `input ${dir}; frames 20 until damage; input stop`; } const powerup = st.powerups.filter(u => u.available) .sort((a, b) => a.dist - b.dist)[0]; if (choice === "powerup" && powerup) { return `goto ${powerup.pos[0]} ${powerup.pos[1]} 180`; } return enemy ? `goto ${enemy.pos[0]} ${enemy.pos[1]} 90` : null; } const log = []; for (let step = 0; step < 30; step++) { const st = JSON.parse(await tank("state")); if (st.state !== "playing") break; const threats = st.projectiles .filter(s => !s.yours && s.miss_dist < 1.5 && s.eta < 1.5) .map(s => `incoming shot dist ${s.dist} eta ${s.eta}s`) .join("\n") || "no incoming shots"; const r = await models.classify(jev, { state: { map: await tank("view 8"), threats, hp: st.player.hp }, questions, }); if (r.stopReason !== "stop") { log.push(`#${step} classifier error: ${r.errorMessage}`); break; } const choice = r.answers.action.choice; const cmd = commandFor(choice, st); if (!cmd) break; log.push(`#${step} hp=${st.player.hp} ${choice} -> ${await tank(cmd)}`); } return log.join("\n"); Calling MCP Servers Lastly, Codemode obviously is great for calling MCP servers. And because we do not actually expose any of the MCP tools to the LLM, the agent first uses provided APIs to issue a tool search within Codemode to discover what it might be able to do with the connected servers. This form of progressive discovery makes the whole MCP business work well enough for a lot of use cases today. Here for instance you can see the agent reach for the Sentry MCP straight away, even without discovering the tools, presumably because it has learned during the RL process already about what the Sentry MCP looks like. But it learns from what we inject into the system prompt, that the Sentry server is available to begin with. It’s not completely guessing here. const orgs = await tools.mcp__sentry__find_organizations({}); const { organizations } = orgs.structuredContent; const results = await Promise.allSettled(organizations.map(org => tools.mcp__sentry__find_projects({ organizationSlug: org.slug, regionUrl: org.regionUrl, }) )); return organizations.map((org, i) => { const r = results[i]; if (r.status !== "fulfilled") return { org: org.slug, error: String(r.reason) }; if (r.value.isError) return { org: org.slug, error: r.value.content }; return { org: org.slug, projects: r.value.structuredContent.projects.map(p => p.slug), }; }); Modern MCP Is A Fight I really don’t want to talk too much about MCP here, but MCP is in fact a protocol that greatly benefits from Codemode. The problem in parts is that MCP in practice often targets harnesses that do not (yet?) use Codemode. But the tide is shifting. In the meantime, a temporary crutch has been to do what Cloudflare did, and do Codemode within the MCP server. But now we have Codemode in Codemode which is pretty bad. It means double JSON escaping, easy for smaller models to get confused by and the inner code cannot call the outer tools. So if you for instance use the Cloudflare MCP servers in Pi, the agent needs to write JavaScript and funnel it through more JavaScript. This is really not optimal, but it’s also understandable that this is happening: const accRes = await tools.mcp__cloudflare__execute({ code: `async () => { const r = await cloudflare.request({ method: "GET", path: "/accounts" }); return r.result.map(a => ({ id: a.id, name: a.name })); }`, }); const accounts = JSON.parse(accRes.content.map(c => c.text).join("")); const out = []; for (const account of accounts) { const r = await tools.mcp__cloudflare__execute({ account_id: account.id, code: `async () => { const r = await cloudflare.request({ method: "GET", path: \`/accounts/\${accountId}/workers/scripts\`, }); return r.result.map(s => ({ id: s.id, modified: s.modified_on })); }`, }); out.push({ account: account.name, workers: r.content.map(c => c.text).join("") }); } return out; MCP Desires So to end things off: how well does Codemode work with MCP today? Well … not amazingly well. That’s because MCP servers are not really targeting harnesses that use Codemode yet (though at this point I think most harnesses support it). For this to work well some recommendations: Structured content: Codemode wants calls to return some nicely formatted JSON. So that needs to come back from the server, and many don’t do that yet. The outputSchema system in MCP is great for that. Consistent results: an interesting failure case is when an MCP server does not return consistent data. For instance because it tries to token optimize things depending on how many items are in the result set. This can cause an initial probe with 5 items to succeed, but then fail when the server returns the maximum batch size. Large binary data: today MCP does not yet support large binary data so quite a few use cases that are really interesting do not work well at all yet. You end up with all kinds of weird workarounds such as pre-signed URLs to allow file uploads then to happen through non MCP channels. Composable tool search: the MCP server might know better than the MCP client which tool is appropriate for a task. But there is no good mechanism today that allows a harness to fan out tool searches across multiple MCP servers. It’s all emergent behavior and it does not scale well to multiple active servers. Future of Codemode So where does this leave us? Is this a reversal of what I wrote a year ago where I encouraged CLIs? I don’t think so. In fact, the MCP ecosystem from my perspective picked up on exactly what we pointed out a year ago works: code. But Codemode goes beyond MCP in that it can act as a capable mechanism within the harness to express more freedom for the agent. There are however also some things that we still need to figure out. For one, durability with Codemode is trickier. We might have to adopt some ideas from durable workflow engines here to snapshot invocations. Or maybe, something like Starlark is a better composition language than JavaScript given its deterministic nature. Images, binary data and just the inability of this pattern to work with smaller models is also something that needs to be fleshed out. So it’s for sure not a perfect solution yet, but it’s quite a useful pattern that I expect us to leverage more.
I spent years becoming the engineer people reached for. Now they still reach for me, but for different reasons.