<?xml version="1.0" encoding="UTF-8"?>
<rss  xmlns:atom="http://www.w3.org/2005/Atom" 
      xmlns:media="http://search.yahoo.com/mrss/" 
      xmlns:content="http://purl.org/rss/1.0/modules/content/" 
      xmlns:dc="http://purl.org/dc/elements/1.1/" 
      version="2.0">
<channel>
<title>Rachel Thomas, PhD</title>
<link>https://rachel.fast.ai/</link>
<atom:link href="https://rachel.fast.ai/index.xml" rel="self" type="application/rss+xml"/>
<description>AI, science, education, &amp; ethics</description>
<image>
<url>https://rachel.fast.ai/images/thomas-card.jpg</url>
<title>Rachel Thomas, PhD</title>
<link>https://rachel.fast.ai/</link>
</image>
<generator>quarto-1.9.38</generator>
<lastBuildDate>Mon, 17 Nov 2025 14:00:00 GMT</lastBuildDate>
<item>
  <title>A Decade of Writing Stuff That People, To My Surprise, Actually Read</title>
  <dc:creator>Rachel Thomas</dc:creator>
  <link>https://rachel.fast.ai/posts/2025-11-18-ten-years/</link>
  <description><![CDATA[ 





<p>As a former mathematician, I was used to nobody reading what I wrote. So when I first began blogging in 2015, I never expected that several of my blog posts would go viral or to have multiple journalists contact me (including from NPR, Wired, and Fortune), make the front page of Hacker News (over 10 times), receive conference keynote invitations, and be interviewed on podcasts.</p>
<p>I do not consider myself a “natural” writer. In college, I tried to avoid classes that required essays, because writing was a struggle for me. It wasn’t until I was 30 that I set out to practice writing more. I share <a href="https://rachel.fast.ai/posts/2019-05-13-blogging-advice/index.html">tips I use for blogging here</a>, which include being willing to put <strong>a lot of time</strong> into a single post, incorporating high quality information, and having a clear idea of my intended audience.</p>
<p>I have selected some of my most popular and impactful posts below. Several of these were originally posted on <a href="https://medium.com/@racheltho">Medium</a> or <a href="https://www.fast.ai/">fast.ai</a> (the two sites where my writing used to live). They are grouped into clusters based on theme. I hope you might enjoy reading these if you haven’t seen them before!</p>
<section id="challenging-conventional-wisdom" class="level2">
<h2 class="anchored" data-anchor-id="challenging-conventional-wisdom">Challenging Conventional Wisdom</h2>
<p>Questioning widely-held assumptions about tech culture, education, and health has been the basis for several popular posts.</p>
<section id="if-you-think-women-in-tech-is-just-a-pipeline-problem-you-havent-been-paying-attention-2015" class="level4">
<h4 class="anchored" data-anchor-id="if-you-think-women-in-tech-is-just-a-pipeline-problem-you-havent-been-paying-attention-2015"><a href="https://rachel.fast.ai/posts/2015-07-27-not-pipeline/">If you think women in tech is just a pipeline problem, you haven’t been paying attention</a> (2015)</h4>
<p>In 2015, I felt burnt out and disillusioned by my experiences working in tech. I was frustrated with how much the popular conversation was still focused on “the pipeline problem”: training young girls to code while ignoring all the adult women being driven out of the tech industry by mistreatment. I spent 9 months researching and writing this post. It went viral and remains my most popular essay. This, together with my other posts, led to me being interviewed and quoted in Wired <a href="https://web.archive.org/web/20170331174059/https://www.wired.com/2017/03/hey-tech-giants-action-diversity-not-just-reports/">several</a> <a href="https://web.archive.org/web/20170725111700/https://www.wired.com/story/ap-computer-science-2017/">times</a> regarding <a href="https://web.archive.org/web/20250727061412/https://www.wired.com/story/artificial-intelligence-researchers-gender-imbalance/">diversity in tech</a> (as well as <a href="https://web.archive.org/web/20251001035853/https://www.wired.com/story/the-apple-card-didnt-see-genderand-thats-the-problem/">other</a> <a href="https://web.archive.org/web/20250717150737/https://www.wired.com/story/ai-and-enormous-data-could-make-tech-giants-harder-to-topple/">AI topics</a>).</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-11-18-ten-years/pipeline-long.jpg" class="img-fluid figure-img" style="width:100.0%"></p>
<figcaption>My first post</figcaption>
</figure>
</div>
</section>
<section id="trends-to-avoid-when-founding-a-startup-2018" class="level4">
<h4 class="anchored" data-anchor-id="trends-to-avoid-when-founding-a-startup-2018"><a href="http://www.fast.ai/2018/01/08/startups/">Trends to Avoid When Founding a Startup</a> (2018)</h4>
<p>The dominant narrative for Bay Area tech startups is to try to raise venture capital, achieve exponential hypergrowth, and hire lots of computer science PhDs. I argued that these approaches not only harm employees, but lead to weaker companies and worse products.</p>
</section>
<section id="my-familys-unlikely-homeschooling-journey-2022" class="level4">
<h4 class="anchored" data-anchor-id="my-familys-unlikely-homeschooling-journey-2022"><a href="https://rachel.fast.ai/posts/2022-09-06-homeschooling/">My family’s unlikely homeschooling journey</a> (2022)</h4>
<p>Many people hold a stereotyped and outdated view of homeschooling, not realizing the explosion of innovative, non-traditional education options available in recent years. My husband and I never planned to homeschool, but we unexpectedly found that our child thrives with this approach.</p>
</section>
<section id="your-immune-system-is-not-a-muscle-2024" class="level4">
<h4 class="anchored" data-anchor-id="your-immune-system-is-not-a-muscle-2024"><a href="https://rachel.fast.ai/posts/2024-08-13-crowds-vs-friends/">Your Immune System is Not a Muscle</a> (2024)</h4>
<p>The misleadingly named “Hygiene Hypothesis” is often used to justify the misconception that all microbes are good for us. However, this theory is more accurately reframed as the “Old friends hypothesis”: humans co-evolved with friendly bacteria and some parasites. We did not co-evolve with the crowd infections of mega-cities and 100,000 global flights per day.</p>
</section>
</section>
<section id="ai-beyond-elite-institutions" class="level2">
<h2 class="anchored" data-anchor-id="ai-beyond-elite-institutions">AI Beyond Elite Institutions</h2>
<p>Machine learning isn’t just for those at billion dollar companies. These posts highlight unconventional practitioners and offer practical guidance for people in varied domains.</p>
<section id="deep-learning-not-just-for-silicon-valley-2017" class="level4">
<h4 class="anchored" data-anchor-id="deep-learning-not-just-for-silicon-valley-2017"><a href="https://rachel.fast.ai/posts/2017-02-27-not-just-silicon-valley/index.html">Deep Learning: Not Just for Silicon Valley</a> (2017)</h4>
<p>Our goal at fast.ai is making AI accessible to people outside of elite institutions, who are tackling meaningful problems in low-resource areas. This post introduced some of our earliest international fellows and the diverse range of problems they were working on. I always enjoyed writing about fascinating use cases from our deep learning community.</p>
</section>
<section id="how-and-why-to-create-a-good-validation-set-2017" class="level4">
<h4 class="anchored" data-anchor-id="how-and-why-to-create-a-good-validation-set-2017"><a href="https://rachel.fast.ai/posts/2017-11-13-validation-sets/">How (and why) to create a good validation set</a> (2017)</h4>
<p>An all-too-common scenario: a seemingly impressive machine learning model is a complete failure when implemented in production. Advice on one common culprit of this, and how to avoid it. In the early years of fast.ai, I wrote numerous posts with practical advice for machine learning.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-11-18-ten-years/fortune2.jpg" class="img-fluid figure-img" style="width:60.0%"></p>
<figcaption>This article asked the question, “Can A.I. conquer its Excel problem?</figcaption>
</figure>
</div>
</section>
<section id="an-introduction-to-deep-learning-for-tabular-data-2018" class="level4">
<h4 class="anchored" data-anchor-id="an-introduction-to-deep-learning-for-tabular-data-2018"><a href="https://rachel.fast.ai/posts/2018-04-29-categorical-embeddings/">An Introduction to Deep Learning for Tabular Data</a> (2018)</h4>
<p>Deep learning is not just for images and text. Companies such as Pinterest and Instacart are also applying it to tabular data, the type of data you might normally put in a spreadsheet. This post caught the attention of a reporter with Fortune, who ended up interviewing me and <a href="https://archive.is/uoDHp">writing about the topic here</a>.</p>
</section>
</section>
<section id="debunking-ai-hype-holding-tech-accountable" class="level2">
<h2 class="anchored" data-anchor-id="debunking-ai-hype-holding-tech-accountable">Debunking AI Hype &amp; Holding Tech Accountable</h2>
<p>The narratives about AI put forth by major tech companies are often misleading about what is necessary, what values matter, and what types of harms can result.</p>
<section id="googles-automl-cutting-through-the-hype-2018" class="level4">
<h4 class="anchored" data-anchor-id="googles-automl-cutting-through-the-hype-2018"><a href="https://rachel.fast.ai/posts/2018-07-23-automl3/">Google’s AutoML: Cutting Through the Hype</a> (2018)</h4>
<p>In a 3-part series, I countered claims that all data scientists need customized, bespoke neural network architectures. While I was nervous about disagreeing with both Google’s CEO and head of AI, my posts led to an <a href="https://slideslive.com/38917533/lessons-learned-from-helping-200000-nonml-experts-use-ml">invitation to keynote</a> the prestigous ICML AutoML workshop. Seven years later, my critiques have been proved valid, with transfer learning a cornerstone of ML and automated neural network search not commonly used.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-11-18-ten-years/sundar_pichai.jpg" class="img-fluid figure-img" style="width:60.0%"></p>
<figcaption>By 2023, we were supposed to all be using AutoML neural architecture search</figcaption>
</figure>
</div>
</section>
<section id="five-things-that-scare-me-about-ai-2019" class="level4">
<h4 class="anchored" data-anchor-id="five-things-that-scare-me-about-ai-2019"><a href="https://rachel.fast.ai/posts/2019-01-29-five-scary-things/">Five Things That Scare Me About AI</a> (2019)</h4>
<p>AI ethics is not just a theoretical topic. I was (and still am) alarmed about the harms already being caused to human beings by AI systems irresponsibly applied to healthcare, employment decisions, policing, and more.</p>
</section>
<section id="the-problem-with-metrics-is-a-big-problem-for-ai-2019" class="level4">
<h4 class="anchored" data-anchor-id="the-problem-with-metrics-is-a-big-problem-for-ai-2019"><a href="https://rachel.fast.ai/posts/2019-09-24-metrics/">The Problem with Metrics is a Big Problem for AI</a> (2019)</h4>
<p>Overemphasizing metrics leads to a variety of real-world harms, including manipulation, gaming, and a myopic focus on the short-term. AI is metric optimization on steroids. I later turned this blog post into an <a href="https://www.cell.com/patterns/fulltext/S2666-3899(22)00056-3">academic paper</a>, together with David Uminsky.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-11-18-ten-years/no-accountability.jpg" class="img-fluid figure-img" style="width:65.0%"></p>
<figcaption>Two disturbing case studies I keep returning to are how computerized algorithms have been used to cut healthcare and to fire teachers</figcaption>
</figure>
</div>
</section>
<section id="deep-learning-gets-the-glory-deep-fact-checking-gets-ignored-2025" class="level4">
<h4 class="anchored" data-anchor-id="deep-learning-gets-the-glory-deep-fact-checking-gets-ignored-2025"><a href="https://rachel.fast.ai/posts/2025-06-04-enzyme-ml-fails/">Deep learning gets the glory, deep fact checking gets ignored</a> (2025)</h4>
<p>A microbiologist discovered hundreds of errors in a paper that used AI to classify enzymes. This is a case study of how challenging it can be to evaluate AI claims outside our area of expertise, as well as of the misaligned incentives that reward flashy results, but not diligent fact-checking. This has been by far my most popular <a href="https://www.linkedin.com/posts/rachel-thomas-942a7923_deep-learning-is-glamorous-and-highly-rewarded-activity-7335778935252611073-mndy">post on LinkedIn</a>.</p>
</section>
</section>
<section id="immunology-science" class="level2">
<h2 class="anchored" data-anchor-id="immunology-science">Immunology &amp; Science</h2>
<section id="decoding-t-cells-with-ai-2024" class="level4">
<h4 class="anchored" data-anchor-id="decoding-t-cells-with-ai-2024"><a href="https://rachel.fast.ai/posts/2024-07-09-t-cells/">Decoding T cells with AI</a> (2024)</h4>
<p>T cells are a crucial component of the adaptive immune system. Accurately pedicting what they can bind to would impact a range of treatments. Numerous algorithms have been developed for this question, but the problem is far from solved.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-11-18-ten-years/tcell-niaid-fingernail.jpg" class="img-fluid figure-img" style="width:50.0%"></p>
<figcaption>The surface of a T cell. I’ve enjoyed exploring how AI is being applied to immunology</figcaption>
</figure>
</div>
</section>
<section id="scientists-just-connected-the-dots-between-viruses-and-everything-2025" class="level4">
<h4 class="anchored" data-anchor-id="scientists-just-connected-the-dots-between-viruses-and-everything-2025"><a href="https://rachel.fast.ai/posts/2025-10-07-rethinking-viruses/">Scientists Just Connected the Dots Between Viruses and… Everything</a> (2025)</h4>
<p>For a long time, catching frequent viruses was considered both inevitable and harmless. But it turns out that common, seemingly-mild viruses have disturbing long-term health impacts.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><a href="https://x.com/math_rachel/status/1975660202952417518"><img src="https://rachel.fast.ai/posts/2025-11-18-ten-years/virus-tweet2.jpg" class="img-fluid figure-img" style="width:65.0%"></a></p>
<figcaption>A thread about viruses</figcaption>
</figure>
</div>
<p>Thanks for joining me on this walk through the past! Also, you can subscribe to be notified of new blog posts by submitting your email below:</p>
<script type="text/javascript" src="https://campaigns.zoho.com.au/js/zc.iframe.js"></script>
<iframe scrolling="no" frameborder="0" id="iframewin" width="100%" height="100%" src="https://zcmp-pd.maillist-manage.com.au/ua/Optin?od=11d0c075b7ed49&amp;zx=11a17553b1&amp;tD=156971d471c7b09&amp;sD=156971d471c7cb3">
</iframe>


</section>
</section>

<p><br><br><i>I look forward to reading your responses. Create a free GitHub account to comment below.</i></p> ]]></description>
  <category>machine learning</category>
  <category>advice</category>
  <guid>https://rachel.fast.ai/posts/2025-11-18-ten-years/</guid>
  <pubDate>Mon, 17 Nov 2025 14:00:00 GMT</pubDate>
  <media:content url="https://rachel.fast.ai/posts/2025-11-18-ten-years/collage2.jpg" medium="image" type="image/jpeg"/>
</item>
<item>
  <title>Scientists Just Connected the Dots Between Viruses and… Everything</title>
  <dc:creator>Rachel Thomas</dc:creator>
  <link>https://rachel.fast.ai/posts/2025-10-07-rethinking-viruses/</link>
  <description><![CDATA[ 





<p>Most people catch many viruses in their lives– for example, <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC8752571/">over 90% of adults have Epstein-Barr virus</a>, and adults <a href="https://www.bbc.com/news/health-31698038">catch the flu</a> about once every 5 years. For a long time, catching frequent viruses was considered both inevitable and harmless. But it turns out that common, seemingly-mild viruses have disturbing long-term health impacts.</p>
<p>Common respiratory viruses increase the risk of <a href="https://academic.oup.com/cardiovascres/article/121/9/1330/8164343">heart attacks and strokes</a>. Viruses are linked to <a href="https://www.cell.com/neuron/fulltext/S0896-6273%2822%2901147-3">dementia and Alzheimer’s Disease</a>. They can <a href="https://www.nature.com/articles/s41586-025-09332-0">re-awaken cancer cells</a> in patients whose cancer was previously in remission. Persistent infections <a href="https://www.youtube.com/watch?v=6EvdTlVTazo">accelerate aging</a> and undermine longevity. <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC10051805/">Viruses can be the trigger</a> that kicks off life-long autoimmune diseases. New studies come out each week confirming that viruses can harm the health of your <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC8231753/">heart</a>, <a href="https://academic.oup.com/eurheartj/advance-article/doi/10.1093/eurheartj/ehaf430/8236450?login=false">blood vessels</a>, <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC8231753/">brain</a>, <a href="https://www.cell.com/neuron/fulltext/S0896-6273%2822%2901147-3">nervous system</a>, and <a href="https://gut.bmj.com/content/70/4/698">gut</a>.</p>
<p>Please pause and let this sink in. If we were to truly internalize this information, there would be massive shifts in the practice of medicine, scientific research, and public policy.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-10-07-rethinking-viruses/accelerate-aging.jpg" class="img-fluid figure-img" style="width:50.0%"></p>
<figcaption>Pathogens accelerate aging in many ways. Proal and VanElzakker, 2025</figcaption>
</figure>
</div>
<p>What you can avoid (infections) may be just as important as what you seek out (exercise, healthy foods). This news seems depressing. It’s too late, viruses are everywhere, everyone has already caught them– what can be done?</p>
<p>There actually is a lot we can do. First of all, developing new anti-viral therapies and treatments should be a top priority. Second, regardless of what infections you’ve already had, preventing or reducing future infections will have a positive impact. There is exciting work happening towards both of these goals, including AI-assisted drug design, patient-led biomedical research, initiatives to <a href="https://theconversation.com/why-investment-in-clean-indoor-air-is-vital-preparation-for-the-pandemics-and-climate-emergencies-to-come-265743">improve indoor air quality</a> and <a href="https://www.nature.com/articles/s41598-022-08462-z">new technologies</a> for cleaning the air.</p>
<p>How can viruses cause all these bad outcomes when some people who catch them are fine? Human health is complicated. Disease development involves a complex interplay of factors: infections, underlying genetics, environment, the microbiome, and more. Let’s return to the example of Epstein-Barr Virus (EBV). EBV has been strongly linked to <a href="https://www.science.org/doi/10.1126/science.abj8222">Multiple Sclerosis</a>, <a href="https://pubmed.ncbi.nlm.nih.gov/19564299/">prolonged fatigue</a>, and <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC8752571/">6 different types of cancer</a>. Given that almost everyone has had EBV, even though “only” a percentage of people develop these lasting impacts, this is a major cause of suffering.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-10-07-rethinking-viruses/disease-venn-diagram.jpg" class="img-fluid figure-img" style="width:60.0%"></p>
<figcaption>Viruses tilt the probabilities against you</figcaption>
</figure>
</div>
<section id="a-new-world-and-a-powerful-idea" class="level2">
<h2 class="anchored" data-anchor-id="a-new-world-and-a-powerful-idea">A new world and a powerful idea</h2>
<p>You may wonder why so many of these health issues are on the rise, when viruses are nothing new. Our world has changed drastically in recent decades compared to most of human evolution. We live in a hyperconnected age of global mega-cities and record numbers of large international flights now. We spend our time indoors in crowded, poorly ventilated buildings. These factors have allowed viruses to travel faster and farther than ever before.</p>
<p>The misleadingly named “Hygiene Hypothesis” is often used to justify the misconception that all microbes are good for us. However, this theory is more accurately reframed as the <a href="https://pubmed.ncbi.nlm.nih.gov/24401109/">“Old friends hypothesis”</a>: humans co-evolved with friendly bacteria and some parasites. Viruses are not our friends, but rather enemies. <a href="https://rachel.fast.ai/posts/2024-08-13-crowds-vs-friends/">We did not co-evolve</a> with these crowd infections of mass travel, mega-cities, and indoor confines.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-10-07-rethinking-viruses/old-friends.jpg" class="img-fluid figure-img" style="width:65.0%"></p>
<figcaption>Not all infections are the same! Modern crowd infections are causing huge harm. Figure from Rook, 2014</figcaption>
</figure>
</div>
<p>The idea that viruses are contributing so much to human suffering and long-term disease is powerful. It will transform how we approach medicine, health, and aging, if we let it. This revelation is one of the key reasons that I decided to make a mid-life career pivot, stepping back from fulfiling work in AI to <a href="https://rachel.fast.ai/posts/2023-02-07-school-immunology/">return to graduate school</a> in Microbiology-Immunology, a journey I have been chronicling here on my blog.</p>
<p>I hope to spend the next few decades applying my machine learning skills to problems at the intersection of infections, multi-omic data sets, the microbiome, and chronic disease. Below, I will share some of what has captured my attention and upended my old views on disease and medicine.</p>
</section>
<section id="viruses-have-many-ways-to-wreak-havoc" class="level2">
<h2 class="anchored" data-anchor-id="viruses-have-many-ways-to-wreak-havoc">Viruses have many ways to wreak havoc</h2>
<p>Viruses have evolved to evade, outmatch, commandeer, and otherwise hurt our immune systems. Here is an incomplete and overlapping list of ways that viruses can harm us:</p>
<section id="persistence" class="level3">
<h3 class="anchored" data-anchor-id="persistence">1. Persistence</h3>
<p>Some viruses quietly stick around for years or decades after our initial illness. They may re-awaken later to cause more problems, or they may spawn surprising issues that we don’t recognize as part of our initial infection.</p>
<p>When they persist in our cells, viruses can impact gene expression, hijacking processes our cells need to gain nutrition and energy. <a href="https://www.youtube.com/watch?v=6EvdTlVTazo">Dr.&nbsp;Amy Proal</a>, a researcher in this area, says that treating <a href="https://www.sciencedirect.com/science/article/pii/S1568163725002119">persistent infections</a> will be necessary to combat aging and extend healthspan.</p>
</section>
<section id="autoimmunity" class="level3">
<h3 class="anchored" data-anchor-id="autoimmunity">2. Autoimmunity</h3>
<p>During an infection, sometimes our immune cells get confused into attacking our own tissue that may “look” similar to the virus (this process is known as <a href="https://creativemeddoses.com/topics-list/molecular-mimicry-in-guillain-barre-syndrome/">molecular mimicry</a>). Once it has mistakenly learned to attack self-tissue, the immune system may continue to do so, even after the virus has been defeated. This is just one of several ways by which <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC10051805/">viruses can trigger</a> autoimmune diseases such as Lupus, Multiple Sclerosis, Rheumatoid Arthritis, or Type 1 Diabetes.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-10-07-rethinking-viruses/molecular-mimicry.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>A confused antibody decides to attack a pathogen, as well as the similar-looking myelin covering of the nerves, causing Guillain-Barré syndrome (Comic from Creative Med Doses)</figcaption>
</figure>
</div>
</section>
<section id="microbiome-changes" class="level3">
<h3 class="anchored" data-anchor-id="microbiome-changes">3. Microbiome changes</h3>
<p>You might expect a stomach bug like norovirus to <a href="https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0048224">change the gut microbiome</a> for the worse. Surprisingly, respiratory viruses such as <a href="https://link.springer.com/article/10.1186/s40168-017-0386-z?getft_integrator=asm">Influenza</a>, <a href="https://www.frontiersin.org/journals/immunology/articles/10.3389/fimmu.2018.00182/full">RSV</a>, and <a href="https://www.nature.com/articles/s41575-022-00698-4">Covid</a> all harm the gut microbiome too. This is bad news, since the gut microbiome helps to regulate the immune system and produces neurotransmitters for our brain.</p>
</section>
<section id="immune-dysregulation" class="level3">
<h3 class="anchored" data-anchor-id="immune-dysregulation">4. Immune Dysregulation</h3>
<p>There are a bunch of ways that the immune system can malfunction (including the ones listed above). Measles can cause <a href="https://www.science.org/doi/10.1126/science.aay6485">immune amnesia</a>, where the immune system forgets previous infections it had learned to fight, leading people to catch the exact same diseases again. There is growing evidence that <a href="https://www.bmj.com/content/390/bmj.r1733">covid has a negative impact</a> on the immune system as well.</p>
</section>
<section id="reactivation-of-other-pathogens" class="level3">
<h3 class="anchored" data-anchor-id="reactivation-of-other-pathogens">5. Reactivation of other pathogens</h3>
<p>Infection with a new virus can wake up old infections that were sleeping quietly in your cells. It is unfair, but sometimes viruses will gang up on you, <a href="https://www.jci.org/articles/view/163669">re-activating other viruses</a> (or bacteria) that weren’t bothering you before.</p>
</section>
<section id="cardiac-damage" class="level3">
<h3 class="anchored" data-anchor-id="cardiac-damage">6. Cardiac damage</h3>
<p><a href="https://theconversation.com/chickenpox-and-shingles-virus-lying-dormant-in-your-neurons-can-reactivate-and-increase-your-risk-of-stroke-new-research-identified-a-potential-culprit-194627">Chickenpox/Shingles</a>, Influenza, and <a href="https://theconversation.com/long-covid-viruses-and-zombie-cells-new-research-looks-for-links-to-chronic-fatigue-and-brain-fog-261108">Covid</a> all raise the <a href="https://academic.oup.com/cardiovascres/article/121/9/1330/8164343">risk of heart attacks and strokes</a>. Viruses have <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC8078120/">many ways of harming</a> our cardiac systems: inflammation, damage to the blood vessles, increased blood clots, and damage to the heart.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-10-07-rethinking-viruses/virus-cardiac-events.jpg" class="img-fluid figure-img" style="width:65.0%"></p>
<figcaption>From a meta-analysis of 48 studies about respiratory viruses triggering heart attacks &amp; strokes (Nguyen, et al, 2025)</figcaption>
</figure>
</div>
</section>
<section id="cancer" class="level3">
<h3 class="anchored" data-anchor-id="cancer">7. Cancer</h3>
<p>Cancer involves a failure of the immune system to kill cells that have gone rogue and turned over to the dark side. In 2008, it was estimated that <a href="https://www.mdanderson.org/cancerwise/8-viruses-that-cause-cancer.h00-159774867.html">viral infections</a> contribute to <a href="https://www.sciencedirect.com/science/article/pii/S0925443907002384">15-20% of human cancer</a> cases. Additional research further linking viruses and cancer has come out since then, so the percentage may be higher now. Both flu and covid infections can <a href="https://www.nature.com/articles/d41586-025-02420-1">reawaken “sleeping” cancer cells</a> that had previously been in remission or cause cancer to spread.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-10-07-rethinking-viruses/8-viruses-cancer.jpg" class="img-fluid figure-img" style="width:65.0%"></p>
<figcaption>An article from MD Anderson on 8 Viruses that Cause Cancer</figcaption>
</figure>
</div>
</section>
<section id="cumulative-impacts" class="level3">
<h3 class="anchored" data-anchor-id="cumulative-impacts">8. Cumulative impacts</h3>
<p>You might hope that you could catch a virus, get it over, and be done with it. Unfortunately, that is often not the case. A young college student was fine after having covid twice, but then struggled to walk short distances <a href="https://newrepublic.com/article/177849/biden-democrats-covid-pandemic-2024">after her 3rd infection</a>. A <a href="https://www.aspendailynews.com/opinion/marolt-if-i-haven-t-seemed-like-myself-lately/article_514db18c-11c1-11ef-a507-3fb6e81cde90.html">Colorado newspaper columnist</a> was skiing, biking, mountain climbing, and running half-marathons up until his 5th covid infection. At this point, he developed pain, fatigue, and migraines that prevent him from doing the activities he loves most. These are not just isolated anecdotes, <a href="https://www.nature.com/articles/s41591-022-02051-3">research confirms</a> the <a href="https://www.cidrap.umn.edu/covid-19/covid-reinfection-may-raise-risk-persistent-symptoms-35">cumulative dangers</a> of repeat infections. In children, a <a href="https://www.thelancet.com/journals/laninf/article/PIIS1473-3099(25)00476-1/fulltext">second covid infection</a> is more likely to cause Long Covid than the first infection. Whatever your previous history, reducing risk of future infections is a worthwhile goal.</p>
<p>The above mechanisms are not exclusive. For example, some microbiome changes can make it easier for pathogens to pass from the gut into the bloodstream and provoke an autoimmune reaction (a process I talked about in this <a href="https://www.youtube.com/embed/6BYocBxOfsU?si=iGUjSMqG_f95N4zr">5-minute video</a>)</p>
</section>
</section>
<section id="the-paradigm-shift" class="level2">
<h2 class="anchored" data-anchor-id="the-paradigm-shift">The Paradigm Shift</h2>
<p>For most viruses, people focus on just a few weeks of initial symptoms. This is the wrong way to think about infections. Viral meningitis or EBV increases your risk of Alzheimer’s or dementia, <a href="https://www.cell.com/neuron/fulltext/S0896-6273%2822%2901147-3">5-15 years later</a>. Chicken pox (varicella zoster virus) can reactivate decades afterwards as shingles, which itself then leads to <a href="https://theconversation.com/chickenpox-and-shingles-virus-lying-dormant-in-your-neurons-can-reactivate-and-increase-your-risk-of-stroke-new-research-identified-a-potential-culprit-194627">increased risk of stroke</a> for at least the following year. We need to radically change how we think about viruses.</p>
<p>There is much we still don’t know about the immune system. Early during the covid pandemic, many experts made definitive statements about the risks of covid, assuming that those who didn’t die in the first few weeks must be completely fine. However, perturbations from infections that initially seem minor can have far-reaching, long-lasting, and time-delayed impacts. There is a ton that is still unknown.</p>
<p>Trying to figure out how viruses hijack cell processes, alter microbiomes, and dysregulate the immune system are complex questions. Researching these areas with curiosity, determination, and an open mind will reveal a lot.</p>
</section>
<section id="reasons-for-hope" class="level2">
<h2 class="anchored" data-anchor-id="reasons-for-hope">Reasons for Hope</h2>
<p>It can be gloomy to think about all the damage viruses can cause. The good news is that we don’t have to resign ourselves to these outcomes. Facing the disturbing reality that many viruses are worse than we thought is just the first step towards coming up with creative new solutions. There are some bright, curious, and determined people focused on these problems, although we need even more hands and brains to get involved.</p>
<p>The breadth and depth of the harms caused by viruses can focus biomedical research in new directions. Most viruses do not have effective anti-viral treatments. This creates a huge need. Scientific inquiry can <a href="https://research.google/pubs/everyone-wants-to-do-the-model-work-not-the-data-work-data-cascades-in-high-stakes-ai/">fail catastrophically</a> when those closest to the problem are not included. Patient-led research gives me hope, because it is centered on the <a href="https://rachel.fast.ai/posts/2024-02-20-ai-medicine/index.html">expertise of those closest to the problem</a>. I am also optimistic about the use of AI for <a href="https://rachel.fast.ai/posts/2024-04-03-ai-antibiotics/index.html">discovering new drugs</a> and <a href="https://rachel.fast.ai/posts/2024-08-01-ai-t-cells/index.html">designing immune therapies</a>.</p>
<p>On the prevention side, reducing how frequently people get sick will have a big impact. Different viruses spread in different ways. In recent years, we have learned that many <a href="https://www.theatlantic.com/health/archive/2021/09/coronavirus-pandemic-ventilation-rethinking-air/620000/">infections are airborne</a>. Healthy indoor air is a human right, <a href="https://theconversation.com/why-investment-in-clean-indoor-air-is-vital-preparation-for-the-pandemics-and-climate-emergencies-to-come-265743">like access to clean drinking water</a>. The UN recently held a high-level event focused on the <a href="https://www.science.org.au/news-and-events/news-and-media-releases/australia-to-lead-first-ever-united-nations-indoor-air-quality-global-pledge">right to clean air</a>. There are many measures we can take to reduce transmission of airborne diseases, such as improved ventilation, air purification, and far-UVC technologies. <a href="https://x.com/_CatintheHat/status/1760626083534414110">Parliament houses</a>, <a href="https://slate.com/technology/2023/01/davos-covid-precaution-uv-lights-air-filters.html">venues for elites</a>, and <a href="https://x.com/LongDesertTrain/status/1494126572625924098">barns for pigs</a> have already received these air quality upgrades. We need children in schools, employees in workplaces, and patients in hospitals to get the same protections. Hopefully, we are on the cusp of a <a href="https://www.airclub.org/">clean air revolution</a>, with more people and organizations recognizing that healthy indoor air is essential.</p>
<p><a href="https://www.iqair.com/newsroom/air-pollution-masks-what-works-what-doesn-t">N95 masks offer</a> an immediate way to significantly reduce how often you get sick. Thankfully, the N95s available currently are more comfortable and <a href="https://theconversation.com/what-is-the-best-mask-for-covid-19-a-mechanical-engineer-explains-the-science-after-2-years-of-testing-masks-in-his-lab-175481">more effective</a> than the surgical or cloth masks that many of us wore back in 2020.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-10-07-rethinking-viruses/indoor-air.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>On the brink of an indoor air quality revolution</figcaption>
</figure>
</div>
</section>
<section id="conclusion" class="level2">
<h2 class="anchored" data-anchor-id="conclusion">Conclusion</h2>
<p>Viruses can harm our cardiac health and cognition, and increase our chances of cancer. If this revelation is fully realized, it will change how the field of medicine operates, priorities in research funding, and public policy on everything from indoor air quality standards to paid sick leave and school attendance. I believe we are on the threshold of what could be a drastic shift in better understanding, preventing, and treating viruses, thus unlocking longer and healthier lives.</p>
<p>Related posts you may also be interested in:</p>
<ul>
<li><a href="https://rachel.fast.ai/posts/2024-11-12-devious-tricks/index.html">5 Devious Tricks Pathogens Use Against Us</a></li>
<li><a href="https://rachel.fast.ai/posts/2023-03-07-viruses1/index.html">Viruses are weirder, worse, &amp; more preventable than you realise</a></li>
<li><a href="https://rachel.fast.ai/posts/2023-03-22-viruses2/index.html">Viruses: The Silent Triggers of Autoimmune and Neurodegenerative Diseases</a></li>
<li><a href="https://rachel.fast.ai/posts/2024-08-13-crowds-vs-friends/index.html">Your Immune System is Not a Muscle</a></li>
</ul>
<p>If you enjoy my posts, please subscribe to be notified of new posts via email:</p>
<script type="text/javascript" src="https://campaigns.zoho.com.au/js/zc.iframe.js"></script>
<iframe scrolling="no" frameborder="0" id="iframewin" width="100%" height="100%" src="https://zcmp-pd.maillist-manage.com.au/ua/Optin?od=11d0c075b7ed49&amp;zx=11a17553b1&amp;tD=156971d471c7b09&amp;sD=156971d471c7cb3">
</iframe>


</section>

<p><br><br><i>I look forward to reading your responses. Create a free GitHub account to comment below.</i></p> ]]></description>
  <category>science</category>
  <guid>https://rachel.fast.ai/posts/2025-10-07-rethinking-viruses/</guid>
  <pubDate>Mon, 06 Oct 2025 14:00:00 GMT</pubDate>
  <media:content url="https://rachel.fast.ai/posts/2025-10-07-rethinking-viruses/accelerate-aging-rectangle-2.jpg" medium="image" type="image/jpeg"/>
</item>
<item>
  <title>AI &amp; medicine: promise &amp; peril</title>
  <dc:creator>Rachel Thomas</dc:creator>
  <link>https://rachel.fast.ai/posts/2025-06-17-ai-medicine/</link>
  <description><![CDATA[ 





<p>I recently gave a talk tying together many of the topics I care about most: the potential of AI in immunology, racial and gender bias in medicine, harms of AI in society, the systemic dismissal of patient experience by doctors, and the promise of participatory, patient-led approaches. I hope you’ll watch my 29 minute talk in full here:</p>
<center>
<iframe width="560" height="315" src="https://www.youtube.com/embed/Stg2UfsGALY?si=zn0fa0Kz3IDB9rqy" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="">
</iframe>
</center>
<p>Below is an edited transcript of the talk, with timestamps if you want to jump to specific sections in the video.</p>
<section id="introduction" class="level2">
<h2 class="anchored" data-anchor-id="introduction">Introduction</h2>
<p><a href="https://www.youtube.com/watch?v=Stg2UfsGALY&amp;t=0s">[00:01]</a>: I’ll be speaking about the opportunities and the risk of applying AI in medicine] and some of some of what’s currently happening in the field. As a brief bit of background about me, I earned my PhD in mathematics at Duke University and worked in the tech industry as a software developer and data scientist. Then in 2016, together with Jeremy Howard, I founded Fast AI. Our goal with Fast AI was to make deep learning more accessible and easier to use for people from a variety of different domain expertise areas. Our work was featured in <a href="https://www.economist.com/business/2018/10/27/new-schemes-teach-the-masses-to-build-ai">The Economist</a>, <a href="https://www.forbes.com/sites/mariyayao/2017/04/10/why-we-need-to-democratize-ai-machine-learning-education/">Forbes</a>, <a href="https://www.technologyreview.com/2018/08/10/141098/small-team-of-ai-coders-beats-googles-code/">MIT Tech Review</a>, and elsewhere. And we’ve had millions of students take our courses and use our software library.</p>
<p>I became increasingly focused on the harms that machine learning was causing as well and I wrote <a href="https://www.bostonreview.net/product/redesigning-ai-spring-2021/">three</a> <a href="https://www.amazon.com/Deep-Learning-Coders-fastai-PyTorch/dp/1492045527">book</a> <a href="https://www.oreilly.com/library/view/97-things-about/9781492072652/">chapters</a> on data ethics as well as becoming the founding director of the Center for Applied Data Ethics at the University of San Francisco. More recently, I’ve become really fascinated with applications in microbiology and immunology. I went <a href="https://rachel.fast.ai/posts/2023-02-07-school-immunology/">back to school</a> and earned a master’s in the area because I believe domain expertise is so important and I’m now working on a PhD related to autoimmunity.</p>
</section>
<section id="promise-in-medicine" class="level2">
<h2 class="anchored" data-anchor-id="promise-in-medicine">Promise in Medicine</h2>
<p><a href="https://www.youtube.com/watch?v=Stg2UfsGALY&amp;t=94s">[01:35]</a>: AI is being applied to a lot of areas in medicine which hold a lot of opportunity and seem quite promising. One area is for <a href="https://rachel.fast.ai/posts/2024-04-03-ai-antibiotics/">drug discovery and drug design</a>. The space of potential drugs out there is huge and so it can be very helpful to have machine learning to help iterate through and filter down which candidates may be the most promising for further study. A second big area where AI is being applied in medicine is to reading medical images, such as detecting Parkinson’s disease or strokes from images of the back of the eye.</p>
<p>A third area is taking a sequence of amino acids. Amino acids are the building blocks of proteins and so here on the left, this is hemoglobin written as its sequence of amino acids, and on the right is its 3D structure. So these two things are the same (they are both hemoglobin), but the 3D structure has has more information. The <a href="https://www.youtube.com/watch?v=pB0RxG1NdtA&amp;list=PLtmWHNX-gukLirebdPH8lla41SS78kjLD&amp;index=5">computer program AlphaFold</a>, which is based on deep learning, made international headlines by smashing all previous records on how accurately it could predict structures. The developers of AlphaFold were awarded a Nobel Prize in Chemistry last year for their work. Protein folding has many, many applications. One that particularly interests me is to try to figure out <a href="https://rachel.fast.ai/posts/2024-07-09-t-cells/">what a T cell will bind to</a>, which has applications to cancer vaccines.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-06-17-ai-medicine/hemoglobin.gif" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>AlphaFold transforms sequences of amino acids (on the left) to 3D protein structures (on the right)</figcaption>
</figure>
</div>
</section>
<section id="what-is-machine-learning" class="level2">
<h2 class="anchored" data-anchor-id="what-is-machine-learning">What is Machine Learning?</h2>
<p><a href="https://www.youtube.com/watch?v=Stg2UfsGALY&amp;t=257s">[04:18]</a>: Let me step back for a moment. I have shared some areas that are quite promising to think about of AI and medicine. But what what is AI? What is machine learning? AI is a pretty pretty wide umbrella, and I’m going to focus on deep learning, which is a subfield of machine learning. Let’s go all the way back to the 1950s, when IBM researcher Arthur Samuel coined the term machine learning.</p>
<p>A rules-based approach, which is not using machine learning, involves handcrafting relevant rules. And if you didn’t want to use machine learning, you could teach a computer to play checkers by programming the rules of checkers into the computer. Arthur Samuel had a different idea, which is that he didn’t want to <strong>teach a computer to play checkers</strong>, he wanted to <strong>teach a computer to <em>learn</em> to play checkers</strong> from previous data. <a href="https://ieeexplore.ieee.org/document/5392560">Arthur wrote</a> that instead of specifying in “minute and exact detail” what a computer can do, that he would rather have it learn from data. To apply this to this problem I mentioned earlier of protein folding, rules-based approaches often had scores about hydrogen bonds or different biochemical properties that are that are well known, whereas approaches such as AlphaFold rely on previous protein structures.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-06-17-ai-medicine/samuel.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>Arthur Samuel coined the term “machine learning” in 1959. Image: IBM Watson Media</figcaption>
</figure>
</div>
<p>Machine learning algorithms LEARN to match inputs to outputs. And so for this reason, it’s very important to think about what your inputs are and what your outputs are. This is crucial. A lot depends on the input data. I want you all to keep this in mind when people are trying to sell you something because there’s a lot of snake oil out there, as well as often people may overhype discoveries, so you need to really think about the nature and quality of the input data and if what’s being claimed is even plausible.</p>
<p>Another key moment in the history of machine learning and of deep learning in particular comes from the late 1980s when Yann LeCun and a team developed a deep learning algorithm for the US Postal Service. Deep learning refers to multi-layered neural networks, a family of algorithms, and it involves a lot of matrix math. They used this algorithm to identify handwritten zip codes because the US Postal Service wanted a computer program that could sort letters. It is very useful to automatically sort letters by zip code. However, people have messy handwriting and tere are a bunch of different ways you can write each digit. Yann LeCun used multi-layered neural networks or deep learning to do this. And now when we hear about breakthroughs in AI in the news, typically they’re using deep learning.</p>
</section>
<section id="why-now" class="level2">
<h2 class="anchored" data-anchor-id="why-now">Why now?</h2>
<p><a href="https://www.youtube.com/watch?v=Stg2UfsGALY&amp;t=453s">[07:35]</a>: And so even though this dates back decades, it’s really only since 2012 that we’ve reached this popular moment in deep learning where it’s being applied much more broadly. It’s achieving unprecedented success on a number a number of tasks. And there are three key components of why this is happening now:</p>
<ol type="1">
<li>One is the existence of large enough data sets to train these algorithms. ImageNet is a key example of a data set of images that was curated for a competition.</li>
<li>Compute power, in the form of GPUs. GPU stands for graphic graphics processing unit and it’s what’s used to render video games. The video gaming industry really pushed the computer development along that has been useful for deep learning as well because it’s the same same types of calculations.</li>
<li>Algorithmic advances have made it more feasible to to train and use deep learning. Pictured below is the algorithm behind AlexNet, which won the 2012 ImageNet competition.</li>
</ol>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-06-17-ai-medicine/why-now.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>Three key ingredients to the deep learning revolution: GPUs, large datasets, and algorithmic advances</figcaption>
</figure>
</div>
<p>I was living in San Francisco in 2012 and saw that this seemed like a very promising and exciting area. However, there were a lot of barriers to entry at the time. You had to have done your PhD with a small handful of advisors and a lot of the practical knowledge was not being written down. People would publish theoretical math papers based on it, but it was hard to hard to learn practical details to implement it for yourself if you were outside a few elite institutions.</p>
</section>
<section id="fast.ai" class="level2">
<h2 class="anchored" data-anchor-id="fast.ai">fast.ai</h2>
<p><a href="https://youtu.be/Stg2UfsGALY?si=Ey1SkG04yyksU5Mb&amp;t=552">[09:33]</a>: So in 2016, Jeremy Howard and I started Fast AI to try to address this problem. We really wanted to get experts from all sorts of domains able to use this technology and to focus on the practical side of how to implement this in code. Most people that took our course, I’ll sometimes refer to them as students, but they were working professionals doing this in their part-time, wanting to apply deep learning to the field they were already working in or to pivot their careers. Many went on to get jobs at top top tech companies, to start their own startups, or to work in research. There are a wide range of applications and a wide range of different fields where they’ve applied deep learning.</p>
<p>Our focus was quite unusual in that we were interested in people with limited resources, which is most of the world. The focus of big tech companies is often on experiments that cost $50 million worth of computer time, which almost nobody can do. We were interested in what you can do if you have very little compute power and limited resources and are working in unusual domains that are outside the focus of the the tech industry. This can actually be an advantage. <a href="https://www.fast.ai/posts/2018-04-30-dawnbench-fastai.html">Resource constraints can drive creativity</a>.</p>
<p>In fact, in 2018 a team of Fast AI students beat teams from Google and Intel that had access to way more resources than we did. And the reason behind this is that our constraints forced more creative solutions. I think there’s often a misconception of if you’re trying to do something that’s more accessible or reaching people from non-traditional backgrounds that it must be watered down or that it’s not state-of-the-art. This win of Fast AI over Google and Intel in a competition hosted by Stanford was covered in the <a href="https://www.technologyreview.com/2018/08/10/141098/small-team-of-ai-coders-beats-googles-code/">MIT Tech Review</a> and <a href="https://www.theverge.com/2018/5/7/17316010/fast-ai-speed-test-stanford-dawnbench-google-intel">The Verge</a>. I’m really proud of our students for that. And I love the idea of getting people from different domains using deep learning because you know about problems that nobody else knows about. And that’s that’s really important.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-06-17-ai-medicine/dawnbench.jpg" class="img-fluid figure-img" style="width:75.0%"></p>
<figcaption>News coverage of fast.ai’s win</figcaption>
</figure>
</div>
<p>A question I have been asked many times over the years is, “isn’t this dangerous to make AI accessible to more people?” or “Don’t you worry about it falling into the wrong hands?” And while it’s not implicit in these questions, there’s often an assumption that as long as only billion-dollar elite corporations have access to this, we’ll all be safe, that these are the more ethical stewards of this technology. And I really disagree with that. In fact, I think it’s quite dangerous to have a very homogeneous and elite group in the tech industry creating creating things that impact the entire world. Having greater diversity and more backgrounds is important and very powerful.</p>
</section>
<section id="ml-centralizes-power" class="level2">
<h2 class="anchored" data-anchor-id="ml-centralizes-power">ML centralizes power</h2>
<p><a href="https://www.youtube.com/watch?v=Stg2UfsGALY&amp;t=816s">[13:39]</a>: That said, there is a lot that can and is going wrong, and people that are being harmed. And in many cases these harms can be caused by people from from elite companies or institutions. I spent many years studying and writing about the harms, looking at different case studies, and seeing patterns of key risks in machine learning. A lot of harms fall under this property that machine learning often has the effect of centralizing power. And it does this in several ways.</p>
<p>One is that it can be used at massive scale and cheaply. And that’s often the appeal; it can be used for for cost cutting endeavors, but those often involve a centralization of power. It’s often implemented with no system for recourse and no way to identify or correct mistakes. It’s adding another node of complexity to systems which can be used to evade responsibility. I’ll give examples of all of these in a moment. As systems become more complex, there’s another finger to point to and say, “well, it’s the algorithm’s fault”, which then if nobody feels responsible that doesn’t lead to to systems with good outcomes. It can create feedback loops where it causes an outcome and then amplifies it. I’ll speak about that in more detail in a moment. Sometimes people will say, “yes, machine learning is biased, but so are humans”. But there are published cases where machine learning amplifies biases from humans making them even larger.</p>
</section>
<section id="real-world-harms" class="level2">
<h2 class="anchored" data-anchor-id="real-world-harms">Real world harms</h2>
<p><a href="https://www.youtube.com/watch?v=Stg2UfsGALY&amp;t=934s">[15:34]</a>: A case study that I return to often comes from <a href="https://www.theverge.com/2018/3/21/17144260/healthcare-medicaid-algorithm-arkansas-cerebral-palsy">an algorithm that’s used in many US states</a> to determine home healthcare benefits. And when it was implemented in one state, there was a bug in the code. This software error incorrectly cut care for people with cerebral palsy. And so people like this woman, <a href="https://www.youtube.com/watch?v=Stg2UfsGALY&amp;t=964s">Tammy Dobs, who’s featured here</a>, had drastic and devastating cuts in her care that really impacted her quality of life to suddenly lose this care. And she she was given no explanation. She was like, “Why has this happened?” And they’re just like, “Well, that’s what that’s what the algorithm said.” Although she didn’t just need an explanatio, what she needed was actionable recourse. She needed a way to get this reverted and to receive her care again.</p>
<p>Eventually there was a lengthy court case and they discover the software bug. The patients got their care reinstated. But that is that is not an ideal process at all that they had to go all this time without care. This is actually a common issue: systems are implemented with no way to identify and address mistakes. I also want to point out that it was patients with cerebral palsy who noticed the air first. And that the people who are directly impacted noticed the error and they had no way to surface that information to anyone that would listen. More recently, there has been investigations into how <a href="https://www.statnews.com/2023/03/13/medicare-advantage-plans-denial-artificial-intelligence/">Medicare Advantage is using AI algorithms</a> to cut needed care.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-06-17-ai-medicine/denied-by-ai.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>STAT news coverage of how Medicare Advantage uses AI to deny needed care</figcaption>
</figure>
</div>
<p>This is not unique to the US. I live in Australia and we have had a scandal with Robodebt. This software was implemented by the government to determine if people had been overpaid welfare benefits and then automatically issue them debts that they had to pay back. And the way Robodebt was implemented was later found to be unlawful and used incorrect calculations. It was also set up where the burden of proof was almost impossibly high of what people needed to try to appeal. This ruined lives, including <a href="https://www.theguardian.com/australia-news/2023/feb/20/platitudes-and-false-words-mother-of-robodebt-victim-who-took-own-life-tells-inquiry-of-government-stonewalling">leading to suicides</a>. Robodebt happened at a massive scale. The Australian government went from issuing 20,000 debts per year to 20,000 debts per week. That’s a <a href="https://www.theguardian.com/australia-news/2020/may/29/robodebt-government-to-repay-470000-unlawful-centrelink-debts-worth-721m">50x scale up</a>. There are also a lot of examples from across the EU as well. Human Rights Watch has written a thorough report on this, <a href="https://www.hrw.org/news/2021/11/10/how-eus-flawed-artificial-intelligence-regulation-endangers-social-safety-net">How the EU’s Flawed Artificial Intelligence Regulation Endangers the Social Safety Net</a>.</p>
</section>
<section id="feedback-loops" class="level2">
<h2 class="anchored" data-anchor-id="feedback-loops">Feedback Loops</h2>
<p><a href="https://www.youtube.com/watch?v=Stg2UfsGALY&amp;t=1168s">[19:28]</a>: Feedback loops are another mechanism of centralization of power. Police departments use software to recommend how many police officers they should send to different neighborhoods of you know, we think there’s going to be more crime in X neighborhood and less crime in Y neighborhood, so let’s send more police to X. <strong>However, algorithms don’t just make predictions, they help determine outcomes.</strong> And if you send more police to a neighborhood, they are likely to make more arrests in that neighborhood than a neighborhood where there are no police. That data gets fed back into the algorithm. And so the algorithm might say, “Wow, you’re making a lot of arrests in neighborhood X. Send even more police there.” And you start <a href="https://proceedings.mlr.press/v81/ensign18a.html">creating outcomes and and amplifying them into runaway loops</a>.</p>
<p>This same dynamic of feedback loops shows up in recommendation systems. YouTube recommends videos for you that will often start autoplaying. The YouTube algorithm learned that showing people conspiracy theories often gets them to <a href="https://www.wired.com/story/the-toxic-potential-of-youtubes-feedback-loop/">watch even more conspiracy theories</a>. “Wow, you liked this conspiracy theory. Here’s some more!” Because if people get hooked on conspiracy theories, they want to watch more. And this is true of many recommendation systems across different social media platforms and media platforms.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-06-17-ai-medicine/feedback-loop.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>Algorithms don’t just make predictions; they determine and amplify outcomes</figcaption>
</figure>
</div>
<p>To recap, we’ve seen this centralizing effect. Machine learning centralizes power because it can:</p>
<ul>
<li>be used at massive scale, cheaply</li>
<li>be implemented with no system for recourse &amp; no way to identify mistakes</li>
<li>be used to evade responsibility</li>
<li>create feedback loops</li>
<li>amplify, not just encode, bias</li>
</ul>
<p>The risks of machine learning mirror existing shortcomings in medical system. Disempowerment is already a problem in medicine, and I really worry about how AI could make it worse. And so stepping stepping away from the machine learning for a moment, I want to talk about how how this problem shows up in medicine.</p>
</section>
<section id="failures-of-the-medical-system" class="level2">
<h2 class="anchored" data-anchor-id="failures-of-the-medical-system">Failures of the Medical System</h2>
<p><a href="https://www.youtube.com/watch?v=Stg2UfsGALY&amp;t=1321s">[22:01]</a>: Dr.&nbsp;Ilene Ruhoy is a neurologist. She’s an MD PhD and she developed a 7 cm brain tumor. She was experiencing a variety of disturbing neurological symptoms. She knew something was wrong, but when she went to the doctor, and she went to multiple doctors, <a href="https://www.washingtonpost.com/wellness/interactive/2022/women-pain-gender-bias-doctors/">she recounts</a>, “I was told I knew too much, that I was working too hard, that I was stressed out, that I was anxious.” They were dismissing her symptoms. And this went on for months. Eventually, eventually she did get the MRI, although her tumor was likely much larger at that point than if it had been diagnosed earlier, and got got rushed into emergency surgery. However, this is not not a one-off case.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-06-17-ai-medicine/bias-studies.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>Gender and racial bias in medicine is well-documented across a range of diseases</figcaption>
</figure>
</div>
<p>There are studies on this. Almost <a href="https://www.bbc.com/future/article/20180523-how-gender-bias-affects-your-healthcare">1 in 3 patients with brain tumors</a> had to visit doctors at least five times before they received an accurate diagnosis. It can be exciting to hear about AI that can can read MRIs accurately, but that’s not going to help patients whose doctors won’t take their symptoms seriously enough to order an MRI in the first place. This is not unique to brain tumors. Gender and racial bias are common in medicine. This is well documented and impacts everything from cancer diagnosis delays to how much pain medication is given. As I said before, machine learning is all about matching input data to outputs. Keep in mind this <a href="https://www.bostonreview.net/articles/rachel-thomas-medicines-machine-learning-problem/">input data is very much going to be shaped</a> by the patients who are misdiagnosed, the patients who were incorrectly dismissed or didn’t have their symptoms taken seriously.</p>
</section>
<section id="a-deeper-problem-than-bias" class="level2">
<h2 class="anchored" data-anchor-id="a-deeper-problem-than-bias">A Deeper Problem than Bias</h2>
<p><a href="https://www.youtube.com/watch?v=Stg2UfsGALY&amp;t=1446s">[24:06]</a>: I believe that the problem is actually much deeper than bias. And bias is huge, but fundamentally, we’re talking about the relationship between <a href="https://academic.oup.com/jmp/article-abstract/49/1/58/7329090?redirectedFrom=fulltext">patient expertise and medical authority</a> here. Knowledge is often seen as a one directional flow from doctor to patient, without recognition of patient knowledge. The forms that patient expertise take can really vary. Patient expertise can be an umbrella term to include everything from the lived experience of knowing what is going on in your body to also including the patients that are out there reading lots of medical papers and keeping up on the literature. There are patients that are even publishing medical medical papers. And I’ll give a few examples of that in a moment. A lot of this comes back to the question of whose expertise was or wasn’t included in the design and construction of these these systems. And this is an issue in machine learning as well, even for non-medical applications, about whose expertise is not included. Machine learning systems can <a href="https://research.google/pubs/everyone-wants-to-do-the-model-work-not-the-data-work-data-cascades-in-high-stakes-ai/">fail catastrophically</a> when expertise from those closest to the problem is not included.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-06-17-ai-medicine/whose-expertise-2.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>Whose expertise wasn’t included?</figcaption>
</figure>
</div>
</section>
<section id="a-path-forward" class="level2">
<h2 class="anchored" data-anchor-id="a-path-forward">A Path Forward</h2>
<p><a href="https://www.youtube.com/watch?v=Stg2UfsGALY&amp;t=1537s">[25:37]</a>: I want to end with some examples that give me hope. One of the best events I’ve attended in the past several years was the <a href="https://participatoryml.github.io/">Participatory Approaches to Machine Learning workshop from ICML 2020</a>. It is all online, so you can access the talks and papers. The organizers of this workshop were interested in seeking more cooperative, more democratic, more participatory approaches for how machine learning systems can can be designed and operated. We really need to consider how all stakeholders, including the people who will be most impacted, can and should be involved in in the creation and running of ML systems.</p>
<p>Pratyusha Kalluri, who’s an AI researcher, wrote an excellent article titled, “<a href="https://www.nature.com/articles/d41586-020-02003-2">Don’t ask if AI is good or fair. Ask how it shifts power</a>.” Who is going to have more or less power after after particular AI systems are implemented? That’s what really poses poses a risk for exploitation.</p>
<p>On the medical side, something that really gives me hope is to see efforts of patient-led research. Here is a neat paper that came out just six months ago, where the author of the paper has long COVID and she also has a doctorate in pharmacology. She surveyed <a href="https://www.medrxiv.org/content/10.1101/2024.11.27.24317656v1.full-text">4,000 patients about over 150 different treatments</a> that they’ve tried, and together with her co-authors did clustering of different subtypes of symptoms. She brought a deep knowledge that comes from being an informed and active patient that is plugged into patient communities. I love to see this type of research. There’s also a <a href="https://patientresearchcovid19.com/publication/">Patient-Led Research Collaborative</a> that is publishing innovative work. The way the risks of machine learning mirror existing shortcomings in medicine and the potential negative negative synergy of how those could combine scares me. I look towards ideas of participatory approaches to machine learning and to patient-led research as sources of hope of how we might be able to to do this in more constructive ways.</p>
<p>So with that, if you have any thoughts or questions, I hope you’ll join me in the comment section below. Thank you.</p>
</section>
<section id="related-reading" class="level2">
<h2 class="anchored" data-anchor-id="related-reading">Related Reading:</h2>
<ul>
<li><a href="https://rachel.fast.ai/posts/2023-05-16-ai-centralizes-power/">AI and Power: The Ethical Challenges of Automation, Centralization, and Scale</a> - a deeper dive into how AI centralizes power</li>
<li><a href="https://rachel.fast.ai/posts/2024-09-10-gaps-risks-science/">Gaps and Risks of AI in the Life Sciences</a></li>
<li><a href="https://rachel.fast.ai/posts/2024-11-20-ai-immunology/">AI and Immunology</a></li>
</ul>
<p>You can subscribe to be notified of new blog posts by submitting your email below:</p>
<script type="text/javascript" src="https://campaigns.zoho.com.au/js/zc.iframe.js"></script>
<iframe scrolling="no" frameborder="0" id="iframewin" width="100%" height="100%" src="https://zcmp-pd.maillist-manage.com.au/ua/Optin?od=11d0c075b7ed49&amp;zx=11a17553b1&amp;tD=156971d471c7b09&amp;sD=156971d471c7cb3">
</iframe>


</section>

<p><br><br><i>I look forward to reading your responses. Create a free GitHub account to comment below.</i></p> ]]></description>
  <category>ethics</category>
  <category>machine learning</category>
  <category>science</category>
  <guid>https://rachel.fast.ai/posts/2025-06-17-ai-medicine/</guid>
  <pubDate>Mon, 16 Jun 2025 14:00:00 GMT</pubDate>
  <media:content url="https://rachel.fast.ai/posts/2025-06-17-ai-medicine/denied-ai-thumbnail.jpg" medium="image" type="image/jpeg"/>
</item>
<item>
  <title>Deep learning gets the glory, deep fact checking gets ignored</title>
  <dc:creator>Rachel Thomas</dc:creator>
  <link>https://rachel.fast.ai/posts/2025-06-04-enzyme-ml-fails/</link>
  <description><![CDATA[ 





<p>Deep learning is glamorous and highly rewarded. If you train and evaluate a Transformer (a state-of-the-art language model) on a dataset of 22 million enzymes and then use it to predict the function of 450 unknown enzymes, you can publish your results in Nature Communications (a very well-regarded publication). Your paper will be viewed 22,000 times and will be in the top 5% of all research outputs scored by Altmetric (a rating of how much attention online articles receive).</p>
<p>However, if you do the painstaking work of combing through someone else’s published work, and discovering that they are riddled with serious errors, including hundreds of incorrect predictions, you can post a pre-print to bioRxiv that will not receive even a fraction of the citations or views of the original. In fact, this is exactly what happened in the case of these two papers:</p>
<ul>
<li><a href="https://www.nature.com/articles/s41467-023-43216-z">Functional annotation of enzyme-encoding genes using deep learning with transformer layers | Nature Communications</a></li>
<li><a href="https://www.biorxiv.org/content/10.1101/2024.07.01.601547v2.full">Limitations of Current Machine-Learning Models in Predicting Enzymatic Functions for Uncharacterized Proteins | bioRxiv</a></li>
</ul>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-06-04-enzyme-ml-fails/altmetric.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>A Tale of two Altmetric Scores</figcaption>
</figure>
</div>
<p>This pair of papers on enzyme function prediction make for a fascinating case study on the limits of AI in biology and the harms of current publishing incentives. I will walk through some of the details below, although I encourage you to read the papers for yourself. This contrast is a stark reminder of how hard it can be to evaluate the legitimacy of AI results without deep domain expertise.</p>
<section id="the-problem-of-determining-enzyme-function" class="level2">
<h2 class="anchored" data-anchor-id="the-problem-of-determining-enzyme-function">The Problem of Determining Enzyme Function</h2>
<p>Enzymes are what catalyze reactions, so they are crucial for making things happen in living organisms. Enzyme Commission (EC) numbers provide a hierarchical classification system for thousands of different functions. Given a sequence of amino acids (the building blocks of all proteins, including enzymes), can you predict what the EC number (and thus, the function) is? This seems like a problem that is custom-made for machine learning, with clearly defined inputs and outputs. Moreover, there is a rich dataset available, with over 22 million enzymes and their EC numbers listed in the online database UniProt.</p>
</section>
<section id="an-approach-with-transformers-ai-model" class="level2">
<h2 class="anchored" data-anchor-id="an-approach-with-transformers-ai-model">An Approach with Transformers (AI model)</h2>
<p>A research paper used a transformer deep learning model to predict the functions of enzymes with previously unknown functions. It seemed like a good paper! The authors used a reasonable, well-regarded neural network architecture (two transformer encoders, two convolutional layers, and a linear layer) that had been adopted from BERT. They looked at regions with high attention to confirm that these were biologically significant, which suggests that the model had learned underlying meaning and provided interpretability. They used a standard training, validation, and test split on a dataset with millions of entries. The researchers then applied the model to a dataset where no “ground truth” was known to make ~450 novel predictions. For these novel predictions, they randomly selected three to test <em>in vitro</em> and confirmed that the predictions were accurate.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-06-04-enzyme-ml-fails/kim-fig1a-4.jpg" class="img-fluid figure-img" style="width:75.0%"></p>
<figcaption>A transformer model, shown on the left, was used to predict Enzyme Commission numbers for uncharacterized enzymes in E. coli. Three of these were tested in vitro (Fig 1a and Fig 4 from Kim, et al.)</figcaption>
</figure>
</div>
</section>
<section id="the-errors" class="level2">
<h2 class="anchored" data-anchor-id="the-errors">The Errors</h2>
<p>The Transformer model in the Nature Communications paper made hundreds of “novel” predictions that are almost certainly erroneous. The paper had followed a standard methodology of evaluating performance on a held-out test set, and did quite well on that (although later investigation suggests there may have been <a href="https://www.kaggle.com/code/alexisbcook/data-leakage">data leakage</a>). The results claimed for enzymes where no ground truth is known were full of errors.</p>
<p>For instance, the gene <em>E. coli</em> YjhQ was predicted to be a mycothiol synthase, but mycothiol is not synthesized by <em>E. coli</em> at all! The gene yciO, which evolved from the gene TsaC, had already been shown a decade earlier <em>in vivo</em> to not have the same function as TsaC, yet the Nature Communications paper concluded it did have the same function.</p>
<p>Of the 450 “novel” results given in the paper, 135 of these results were not novel at all; they were already listed in the online database UniProt. Another 148 showed unreasonably high levels of repetition, with the same very specific enzyme functions reappearing up to 12 times for genes of <em>E. coli</em>, which biologically implausible.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-06-04-enzyme-ml-fails/de-crecy-fig5.jpg" class="img-fluid figure-img" style="width:75.0%"></p>
<figcaption>Most of the “novel” results from the transformer paper were either not novel, unusually repetitious, or incorrect paralogs (Fig 5 from de Crecy, et al.)</figcaption>
</figure>
</div>
</section>
<section id="the-microbiology-detective" class="level2">
<h2 class="anchored" data-anchor-id="the-microbiology-detective">The Microbiology Detective</h2>
<p>How did these errors come to light? After the model had been trained, validated, and evaluated on a dataset involving millions of entries, it was used to make ~450 novel predictions, and three of these were tested in vitro. It just so happens that one of the enzymes selected for in vitro testing, yciO, had already been studied extensively over a decade earlier by Dr.&nbsp;de Crécy-Lagard. When Dr.&nbsp;de Crécy-Lagard read that deep learning had predicted that yciO had the same function of another gene, TsaC, she knew from her long years in the lab that this was incorrect. Her previous research had shown that the TsaC gene is essential in <em>E. coli</em> even if yciO is present in the same genome and even when yciO gene is overexpressed. Moreover, the yciO activity reported by Kim et al.&nbsp;is more than four orders of magnitude (i.e.&nbsp;10,000 times) weaker than that of TsaC. All this suggests that yciO does NOT serve the same key function as TsaC.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-06-04-enzyme-ml-fails/de-crecy-fig7.jpg" class="img-fluid figure-img" style="width:80.0%"></p>
<figcaption>Two enzymes with a common evolutionary ancestor, but different functions (Fig 7 from de Crecy, et al.)</figcaption>
</figure>
</div>
<p>YciO and TsaC do have structural similarities, and YciO evolved from an ancestor of TsaC. Decades of research on protein and enzyme evolution have shown that new functions often evolve via duplication of an existing gene, followed by diversification of its function. This poses a common pitfall in determining enzyme function, because the genes will have many similarities with the ones they duplicated and then diversified from.</p>
<p>Thus, looking at structural similarities is only one type of evidence for considering enzyme function. It is also crucial to look at other types of evidence, such as neighborhood context of the genes, substrate docking, gene co-occurrence in metabolic pathways, and other features of the enzymes.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-06-04-enzyme-ml-fails/de-crecy-fig2.jpg" class="img-fluid figure-img" style="width:80.0%"></p>
<figcaption>It is important to look at multiple types of evidence when classifying enzyme function (Fig 2 from de Crecy, et al.)</figcaption>
</figure>
</div>
</section>
<section id="hundreds-of-likely-erroneous-results" class="level2">
<h2 class="anchored" data-anchor-id="hundreds-of-likely-erroneous-results">Hundreds of Likely Erroneous Results</h2>
<p>Spotting this one error inspired de Crécy-Lagard and her co-authors to take a closer look at all of the enzymes found to have novel results in the Kim, et al, paper. They found that 135 of these results were already listed in the online database used to build the training set and thus not actually novel. An additional 148 of the results contained a very high level of repetition, with the same highly specific functions reappearing up to 12 times. Biases, data imbalance, lack of relevant features, architectural limitations, or poor uncertainty calibration can all lead models to “force” the most common labels from the training data.</p>
<p>Other examples were proven wrong via biological context or a literature search. For instance, the gene YjhQ was predicted to be a mycothiol synthase but mycothiol is not synthesized by <em>E. coli</em>. YrhB was predicted to synthesize a particular compound, which was already predicted to be synthesized by the enzyme QueD. A form of <em>E. coli</em> with a QueD mutant was unable to synthesize the compound, showing that this is not in fact the function of YrhB.</p>
</section>
<section id="rethinking-enzyme-classification-and-true-unknowns" class="level2">
<h2 class="anchored" data-anchor-id="rethinking-enzyme-classification-and-true-unknowns">Rethinking Enzyme Classification and “True Unknowns”</h2>
<p>Identifying enzyme function actually consists of two quite different problems which are commonly conflated:</p>
<ul>
<li>propagating known function labels to enzymes in the same functional family</li>
<li>discovering truly unknown functions</li>
</ul>
<p>The authors of the second paper observe, “By design, supervised ML-models cannot be used to predict the function of true unknowns.” While machine learning can be useful for propagating known functions to additional enzymes, there are many types of errors that can occur: including failing to propagate labels when they should, propagating labels when they should not, curation mistakes, and experimental mistakes. Unfortunately, erroneous functions are being entered into key online databases such as UniProt, and this incorrect data may be further propagated if it is used to train prediction models. This is a problem that increases over time.</p>
</section>
<section id="need-for-domain-expertise" class="level2">
<h2 class="anchored" data-anchor-id="need-for-domain-expertise">Need for Domain Expertise</h2>
<p>It is not news that AI work will be more highly rewarded and supported than work that closely inspects the underlying data and integrates deep domain knowledge. The aptly titled <a href="https://research.google/pubs/everyone-wants-to-do-the-model-work-not-the-data-work-data-cascades-in-high-stakes-ai/">“Everyone Wants to do the Model Work, not the Data Work”</a> paper involving dozens of machine learning practitioners working on high-stakes AI projects and found that inadequate-application domain expertise was one of a few key causes of catastrophic failures.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-06-04-enzyme-ml-fails/sambasivan-square.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>Sources of cascading failures in machine learning systems (Fig 1 from Sambasivan, et al.)</figcaption>
</figure>
</div>
<p>These papers also serve as a reminder of how challenging (or even impossible) it can be to evaluate AI claims in work outside our own area of expertise. I am not a domain expert in the enzyme functions of <em>E. coli</em>. And for most deep learning papers I read, domain experts have not gone through the results with a fine-tooth comb inspecting the quality of the output. How many other seemingly-impressive papers would not stand up to scrutiny? The work of checking hundreds of enzyme predictions is less glamorous than the work of building the AI model that generated them, yet it is even more important. How can we better incentivize this type of error-checking research?</p>
<p>At a time when funding is being slashed, I believe we should be doing the opposite and investing even more into a range of scientific and biomedical research, from a variety of angles. And we need to push back on an incentive system that is disproportionately focused on flashy AI solutions at the expense of quality results.</p>
</section>
<section id="related-reading" class="level2">
<h2 class="anchored" data-anchor-id="related-reading">Related Reading:</h2>
<ul>
<li><a href="https://rachel.fast.ai/posts/2019-09-24-metrics/">The problem with metrics is a big problem for AI</a></li>
<li><a href="https://rachel.fast.ai/posts/2024-09-10-gaps-risks-science/">Gaps and Risks of AI in the Life Sciences</a></li>
<li><a href="https://rachel.fast.ai/posts/2024-02-20-ai-medicine/">“AI will cure cancer” misunderstands both AI and medicine</a></li>
</ul>
<p><br></p>
<p>You can subscribe to be notified of new blog posts by submitting your email below:</p>
<script type="text/javascript" src="https://campaigns.zoho.com.au/js/zc.iframe.js"></script>
<iframe scrolling="no" frameborder="0" id="iframewin" width="100%" height="100%" src="https://zcmp-pd.maillist-manage.com.au/ua/Optin?od=11d0c075b7ed49&amp;zx=11a17553b1&amp;tD=156971d471c7b09&amp;sD=156971d471c7cb3">
</iframe>


</section>

<p><br><br><i>I look forward to reading your responses. Create a free GitHub account to comment below.</i></p> ]]></description>
  <category>machine learning</category>
  <category>science</category>
  <guid>https://rachel.fast.ai/posts/2025-06-04-enzyme-ml-fails/</guid>
  <pubDate>Tue, 03 Jun 2025 14:00:00 GMT</pubDate>
  <media:content url="https://rachel.fast.ai/posts/2025-06-04-enzyme-ml-fails/tsac-ycio.jpeg" medium="image" type="image/jpeg"/>
</item>
<item>
  <title>Will gene sequencing finally prove its worth?</title>
  <dc:creator>Rachel Thomas</dc:creator>
  <link>https://rachel.fast.ai/posts/2025-05-13-spatial-omics/</link>
  <description><![CDATA[ 





<section id="dna-sequencing-hasnt-lived-up-to-the-hype" class="level2">
<h2 class="anchored" data-anchor-id="dna-sequencing-hasnt-lived-up-to-the-hype">DNA sequencing hasn’t lived up to the hype</h2>
<p>Twenty to thirty years ago, <a href="https://theconversation.com/why-sequencing-the-human-genome-failed-to-produce-big-breakthroughs-in-disease-130568">politicians, scientific leaders, journalists, and even Nobel laureates</a> predicted that sequencing the human genome would revolutionize how we treat disease. And while the advances in DNA sequencing that have occurred since then have improved recognition and treatment for some cancers and rare diseases, on the whole the field has not lived up to earlier hype.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-05-13-spatial-omics/time-94-99.jpg" class="img-fluid figure-img" style="width:60.0%"></p>
<figcaption>Time Magazine covers from 1994 and 1999 about genetics</figcaption>
</figure>
</div>
<p>In an article titled <a href="https://theconversation.com/why-sequencing-the-human-genome-failed-to-produce-big-breakthroughs-in-disease-130568">“Why sequencing the human genome failed to produce big breakthroughs in disease”</a>, a biology professor highlights that most common diseases are not caused by a single gene. In fact, common diseases are often linked to <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC6391102/">hundreds of gene variants</a>, and even collectively, these variants still account for only a small fraction of disease variance. Here, I want to focus on two other key limitations of DNA sequencing, and how they are now being addressed with new approaches.</p>
</section>
<section id="what-dna-cant-tell-us" class="level2">
<h2 class="anchored" data-anchor-id="what-dna-cant-tell-us">What DNA can’t tell us</h2>
<p>First, DNA can’t answer many questions about how cells and organisms work in practice. A neuron in the brain has the same DNA as a liver cell, yet the two have completely different functions. This is because different segments of DNA are turned off or on in different cells. To understand how cells are actually working, you need to know about proteins and RNA (RNA is the intermediary which translates DNA into protein). Proteins are what build the structure of cells, catalyze chemical reactions within the cell, and allow communication between cells. Healthy and unhealthy cells in the same organism will usually have the same DNA. For instance, if some regions of the intestines are experiencing an IBD (irritable bowel disease) flare and others aren’t, they would all have the same DNA, yet likely different RNA and protein levels.</p>
<p>A second big problem is that many key sequencing techniques destroy spatial information. You essentially may have to put tissue or cells into a blender in order to get rich information about DNA or RNA sequences. While this data is informative, it turns out that locations of cells within a piece of tissue, and locations of regions within a cell, are also very important! Again, considering the case in which some regions of the <a href="https://www.mypathologyreport.ca/pathology-dictionary/cryptitis/">intestine are inflamed due to IBD</a>, yet others aren’t, mixing them all together in a blender will lose or distort useful information. Look at the difference between crypts in a healthy segment of the colon (on the left) compared to inflamed crypts (on the right).</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-05-13-spatial-omics/cryptitis.jpg" class="img-fluid figure-img" style="width:80.0%"></p>
<figcaption>Differences between a healthy bowel (on the left) and an inflamed bowel (on the right). Source: mypathologyreport.ca</figcaption>
</figure>
</div>
<p>Recently, we have seen a rise in breakthroughs that allow us to obtain data about location. Spatial techniques are a necessary and exciting step beyond DNA sequencing. The power of spatial information showed up as a major theme at a <a href="https://www.multiomics2024.org/">conference I attended last year</a>, and spatial techniques have been recognized by Nature Methods as “<a href="https://www.nature.com/collections/dbifijbacd">Method of the year</a>” <a href="https://www.nature.com/articles/s41592-020-01033-y">twice</a> in the last 5 years.</p>
</section>
<section id="a-few-major-areas-of-innovation" class="level2">
<h2 class="anchored" data-anchor-id="a-few-major-areas-of-innovation">A Few Major Areas of Innovation</h2>
<p>What is genetic sequencing anyway? There are a number of different types of sequencing that have been invented in the last 30 years. Walking through a brief history will illustrate what these technologies are, and what they can and cannot do. To make it concrete, let’s look at the example of how they have been applied to cancer treatment.</p>
<p><strong>Sanger Sequencing</strong>: This is an older technology dating back to the 1970s, and which was the main way of sequencing DNA up until 2005. Sanger sequencing was one of the methods used in the mid-90s to identify the genes BRCA1 and BRCA2 as key genetic risk factors for breast cancer. The process of discovering BRCA1/BRCA2 involved scientists slowly zeroing in on their chromosomal locations over a period of years. Several other cancer genes were discovered during this time period as well. While Sanger sequencing is effective on smaller amounts of DNA, it can be quite slow to deal with larger volumes. It took over 10 years to sequence the first copy of the human genome using Sanger Sequencing. It is still used today as a simple and reliable way to test for known mutations (such as BRCA1/BRCA2) or for smaller tasks.</p>
<p><strong>High-Throughput Sequencing</strong>: New technology released in the mid-2000s allowed DNA to be cut into lots of short pieces and for millions of pieces to be sequenced in parallel at once. This approach, called high-throughput sequencing, was significantly faster than Sanger sequencing. High-throughput sequencing has many applications, including to cancer treatment, by making it cheaper and faster to sequence DNA to identify particular mutations which can influence treatment decisions. Method of the Year | Nature Methods</p>
<p><strong>Long-read sequencing</strong>: High-throughput sequencing has the advantage of high accuracy, but the downside of short sequence lengths. In the 2010s, technologies were released with the opposite set of strengths and weaknesses. Long-read sequencing provides the advantage of long sequence lengths, although the downside of lower accuracy. To compare, high-throughput sequencing uses DNA strands that are a few hundred base pairs long, whereas long-read sequencing uses DNA strands that are tens of thousands base pairs long. Both technologies have different strengths and are widely used today.</p>
<p><strong>Single-Cell Sequencing</strong>: High-throughput sequencing involves sequencing the DNA of many cells at once, but sometimes it is useful to sequence individual cells. In a tumor, different cells can have different mutations. It is possible that a small subset of the mutations may drive metastasis (the spread of cancer to other areas) or resistance to treatment. Identifying these driver mutations can guide treatment decisions, since particular driver mutations can predict the effectiveness of various drugs. Key mutations may be drowned out in the average if you sequence the entire tumor. This is one reason why it is useful to be able to sequence single cells, and not just obtain the average of many cells. Single-cell sequencing was selected as Nature Methods method of the year in 2013.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-05-13-spatial-omics/legos.jpg" class="img-fluid figure-img" style="width:80.0%"></p>
<figcaption>A Lego interpretation of bulk RNA-seq; single-cell RNA-seq; spatial transcriptomics; and the original organ. Source: Bo Xia, <span class="citation" data-cites="BoXia7">@BoXia7</span></figcaption>
</figure>
</div>
<p><strong>Multi-Omics</strong>: The study of DNA is genomics. DNA alone gives us an incomplete picture of an organism. Epigenomics can provide information about which regions of DNA are active or silenced. To understand how different cells function, as well as cells in different states of disease or health, you also need to know about their RNA and proteins. This data is contained in the field of transcriptomics (transcripts are strands of RNA transcribed from DNA) and proteomics (the proteins in a cell). Metabolomics looks at small molecules (such as sugars, amino acids, and vitamins) within the body and exposomics includes all sorts of environmental exposures. Collectively, these fields are known as -omics or multi-omics. It is valuable to combine multiple types of -omics together for richer sources of information, since each has different strengths, limitations, and insights to offer. Multi-Omics is an exciting area that draws on lots of data, with applications to cancer, infectious disease and immunology.</p>
<p><strong>Spatial</strong>: Spatial information lets us see all the variation within a section of the body– such as a segment of the intestines, the liver, or a cancerous tumor. This variation can often be significant for understanding disease and treatment prognosis. To better understand why spatial techniques are useful, let us dive into some background about cancer.</p>
</section>
<section id="tumors-arent-just-lumps-of-bad-cells" class="level2">
<h2 class="anchored" data-anchor-id="tumors-arent-just-lumps-of-bad-cells">Tumors aren’t just lumps of bad cells</h2>
<p>Cancer is defined as excess cell division. I used to think that tumors were just clumps of “bad cells”, where “good” and “bad” were binary states. This is incorrect. Tumors are not uniform, and within the category of “bad” there is a great deal of variation and heterogeneity. Different cells within a tumor may have different mutations from one another. And our immune systems sometimes build <a href="https://www.nature.com/articles/s41568-024-00728-0">complex defense structures</a> within tumors in attempts to more effectively fight them. How close a cancerous cell is to one of these immune structures impacts how likely the body is to destroy the cancerous cell.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-05-13-spatial-omics/tls.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>Notice all the variation within this tumor!</figcaption>
</figure>
</div>
<p>Immunologist <a href="https://www.multiomics2024.org/angela-ferguson">Dr.&nbsp;Angela Ferguson</a>, who studies head and neck cancers, describes tumors as having a “physical landscape”. She has shown how the organization and structure within a <a href="https://www.biorxiv.org/content/10.1101/2024.04.18.590189v1.full">tumor can predict and guide treatment outcomes</a>. Her work found that cancer progression is “landscape-dependent”, where landscape refers to the locations of immune cells and structures within a tumor. It is not enough to study cancer cells in isolation. We need to understand their layout. Methods that effectively put tumor cells into a blender in order to sequence them, disrupting their spatial information, are insufficient on their own. Mapping that spatial information can hold the keys for more effective treatment.</p>
<p>Single cell sequencing approaches allow a greater number of genes to be measured, whereas spatial approaches measure fewer genes but also provide location information. Combining these two approaches can prove powerful.</p>
</section>
<section id="programming-libraries-applied-to-spatial--omics" class="level2">
<h2 class="anchored" data-anchor-id="programming-libraries-applied-to-spatial--omics">Programming Libraries Applied to Spatial -Omics</h2>
<p>We are living at a time when multi-omics, spatial information, and user-friendly programming libraries are converging for easier exploration and discovery. For instance, below is an image I created using the common Python programming libraries pandas and matplotlib of data the NIH has shared about a rare liver disease. The image shows a slice of liver tissue, with gene expression overlaid in a color scale ranging from purple (low expression) to yellow (highest expression).</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-05-13-spatial-omics/liver-alb.jpg" class="img-fluid figure-img" style="width:80.0%"></p>
<figcaption>Using Matplotlib to display a cross-section of liver with gene expression</figcaption>
</figure>
</div>
<p>The NIH dataset contain images of slices of liver and expression of many different genes, from both healthy patients and those with a rare liver disease. https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE240429. The plot shows expression of albumin, a protein which helps transport other molecules around the body, overlaid on top of the liver, illustrating which regions produce more or less of this protein. Data science tools are invaluable for transforming and plotting such data.</p>
<p>In addition to being able to use standard Python libraries (such as Pandas and Matplotlib) for visualing this data, there are many specialist libraries as well. The <a href="https://scverse.org/">Scverse</a> (Sc = single cell) includes a number of Python libraries focused on single cell analysis.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-05-13-spatial-omics/sc-verse.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>The SC Verse includes libaries for single cell sequencing analysis</figcaption>
</figure>
</div>
<p>There are also popular R libraries for single cell analysis and multi-omics, such as <a href="https://portals.broadinstitute.org/harmony/index.html">Harmony</a>, <a href="https://satijalab.org/seurat/">Seurat</a>, and <a href="https://mixomics.org/">mixOmics</a>.</p>
<p>With all of these tools and technologies, it is an exciting time to be working at the intersection of data science and microbiology. The causes of most diseases are complex and multi-factorial, although new approaches of spatial multi-omics provide unprecedented types of useful information.</p>
</section>
<section id="related-reading" class="level2">
<h2 class="anchored" data-anchor-id="related-reading">Related Reading:</h2>
<ul>
<li><a href="https://rachel.fast.ai/posts/2025-01-16-cpath-foundation/">What AI can tell us about microscope slides</a></li>
<li><a href="https://rachel.fast.ai/posts/2024-09-10-gaps-risks-science/">Gaps and Risks of AI in the Life Sciences</a></li>
<li><a href="https://rachel.fast.ai/posts/2024-11-20-ai-immunology/">AI and Immunology</a></li>
</ul>
<p>You can subscribe to be notified of new blog posts by submitting your email below:</p>
<script type="text/javascript" src="https://campaigns.zoho.com.au/js/zc.iframe.js"></script>
<iframe scrolling="no" frameborder="0" id="iframewin" width="100%" height="100%" src="https://zcmp-pd.maillist-manage.com.au/ua/Optin?od=11d0c075b7ed49&amp;zx=11a17553b1&amp;tD=156971d471c7b09&amp;sD=156971d471c7cb3">
</iframe>


</section>

<p><br><br><i>I look forward to reading your responses. Create a free GitHub account to comment below.</i></p> ]]></description>
  <category>science</category>
  <guid>https://rachel.fast.ai/posts/2025-05-13-spatial-omics/</guid>
  <pubDate>Mon, 12 May 2025 14:00:00 GMT</pubDate>
  <media:content url="https://rachel.fast.ai/posts/2025-05-13-spatial-omics/time-94-99.jpg" medium="image" type="image/jpeg"/>
</item>
<item>
  <title>The Missing Medical Data Holding Back AI</title>
  <dc:creator>Rachel Thomas</dc:creator>
  <link>https://rachel.fast.ai/posts/2025-01-24-missing-data/</link>
  <description><![CDATA[ 





<p>Amidst the barrage of executive orders and announcements Trump made during his first week in office are two which highlight common misunderstandings and obfuscations about the relationship between <a href="https://rachel.fast.ai/posts/2023-05-16-ai-centralizes-power/">data, AI, and power</a>. These announcements contradict each other:</p>
<ul>
<li>He announced a <a href="https://apnews.com/article/trump-health-communications-cdc-hhs-fda-1eeca64c1ccc324b31b779a86d3999a4">freeze on communications</a> from scientific and health agencies, stopping them from sharing health data with the public, including infectious disease data about the emerging H5N1 bird flu outbreak. The CDC was unable to send out its <a href="https://www.medpagetoday.com/infectiousdisease/generalinfectiousdisease/113905">weekly MMWR report</a> for the first time in 60 years. Without data, it can be nearly impossible to understand, much less respond, to growing threats.</li>
<li>He announced a <a href="https://apnews.com/article/trump-ai-openai-oracle-softbank-son-altman-ellison-be261f8a8ee07a0623d4170397348c41">large investment in AI</a>, including promises to use AI to analyze health records, help doctors care for patients, and treat cancer.</li>
</ul>
<p>Hiding health-related data from the public does not improve the lives of patients, particularly since those with cancer are vulnerable to infection. As I wrote in a <a href="https://rachel.fast.ai/posts/2024-02-20-ai-medicine/">previous post</a>, claims that AI will cure cancer are often used as superficial marketing, ignoring the ways that AI could further disempower patients (whose knowledge and experiences are already disregarded by the medical system).</p>
<p>The orders for scientific agencies to pause communications are a particularly vivid and drastic example of the <a href="https://rachel.fast.ai/posts/2021-10-12-medicine-political/">political nature of data</a>. However, choices about what data is collected and who it is made available to have been shaped by social influences, financial factors, and power disparities long before this past week.</p>
<section id="missing-data-sets" class="level2">
<h2 class="anchored" data-anchor-id="missing-data-sets">Missing Data Sets</h2>
<p>A decade ago, Mimi Onuoha coined the concept of <a href="https://github.com/MimiOnuoha/missing-datasets">Missing Data Sets</a>. She uses this term specifically to refer to “<em>blank spots that exist in spaces that are otherwise data-saturated</em>.” So it’s not just that this data does not exist, but that it ought to exist.</p>
<p>Originally writing in 2016, Onuoha gives the example of how “<em>traditionally there has been little history of standardized and rigorous data collected about police brutality… The group who would make the most sense to monitor this issue—the law enforcement agents who create the data set in the first place—have no incentive to actually gather such data, which could prove incriminating</em>.”</p>
<p>Missing Datasets are not just a USA-based phenomenon. In early 2024, the <a href="https://theconversation.com/the-uk-government-aims-to-stop-publishing-stats-on-homeless-peoples-deaths-heres-why-thats-a-problem-223879">UK government announced</a> that it would stop collecting and publishing statistics about homeless people’s deaths, even as rising evictions and a housing shortage were making the issue more urgent than ever. Public outcry led the government to reconsider.</p>
<p>As another example, <a href="https://thesicktimes.org/2024/02/06/how-we-track-covid-19-with-wastewater-and-other-data-sources/">wastewater data provides</a> a reliable indicator of covid rates, even in the absence of testing. The health department in my home state (Queensland, Australia) used to monitor covid wastewater data, but <a href="https://www.data.qld.gov.au/dataset/queensland-wastewater-surveillance-for-sars-cov-2">stopped doing so in 2022</a>, even as several other states in Australia have continued to collect and share such data. This was a political decision. The lack of this data for Queensland from late 2022 onwards registers as a blank space compared to <a href="https://www.cdc.gov/nwss/rv/COVID19-nationaltrend.html">other countries</a> and <a href="https://www.health.wa.gov.au/articles/a_e/coronavirus/covid19-wastewater-surveillance">other</a> <a href="https://www.health.nsw.gov.au/Infectious/covid-19/Pages/sewage-surveillance.aspx">Australian states</a>, and compared to Queensland’s own record in 2020-2022. There have been multiple covid waves since 2022. This is missing data.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-01-24-missing-data/covid-wastewater.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>This graph from San Jose, California shows how closely wastewater data (in red) tracked with covid positive tests (green) up until most people stopped testing in mid-2022. In Queensland, we do not have access to this data.</figcaption>
</figure>
</div>
<p>Onuoha explains several reasons that <a href="https://github.com/MimiOnuoha/missing-datasets">data may be missing</a>. Those who could collect it lack incentives to do so, or in some cases may even want to hide or obscure it. Another possible reason is that the data in question may be of a type that resists straight-forward quantification or doesn’t fit with our current modes of collection. Her work also reminds us that data won’t solve all problems, and that in some cases missing data can function as protection for disadvantaged groups.</p>
</section>
<section id="decades-of-work-not-instant-magic" class="level2">
<h2 class="anchored" data-anchor-id="decades-of-work-not-instant-magic">Decades of work, not instant magic</h2>
<p>Some people mistakenly believe AI can just be sprinkled on a problem to achieve a magic solution, without understanding the many necessary underlying ingredients. The example of AlphaFold, whose creators won a Nobel Prize in Chemistry for their work, disproves these misconceptions. AlphaFold can predict the 3D structures of proteins with unprecedented accuracy. While clever algorithmic innovations were pivotal to AlphaFold, its success would not have been possible without a large, high-quality dataset that thousands of researchers contributed to for decades.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-01-24-missing-data/insulin-2ways.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>AlphaFold takes a string of amino acids (on the left) and converts it into a 3D structure (on the right)</figcaption>
</figure>
</div>
<p>The construction of the dataset that made AlphaFold possible began a full 50 years earlier. The Protein Data Bank (PDB) was <a href="https://www.nature.com/articles/newbio233223b0">announced in 1971</a> as a database to store information about the 3D structures of proteins. The files were originally shared and distributed using magnetic tape data storage. Initially, just a handful of entries were added to the database each year. By 1980, there were a total of 69 entries present in the database. However, that number <a href="https://www.rcsb.org/stats/growth/growth-released-structures">grew exponentially</a>, reaching nearly 200,000 entries by 2022. Access to the data was updated from magnetic tapes to the world wide web. In the mid-1990s, a <a href="https://onlinelibrary.wiley.com/doi/10.1002/prot.340230303">bi-annual competition</a> was created, encouraging teams to write computer program to predict 3D protein structures. The existence and quality of the database and competition are what led to the record-shattering creation of AlphaFold.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-01-24-missing-data/PDB-per-year.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>The number of proteins stored in PDB grew exponentially with time</figcaption>
</figure>
</div>
<p>AlphaFold wouldn’t have been possible without the rich data of the PDB. Machine learning engineer <a href="https://www.owlposting.com/p/wet-lab-innovations-will-lead-the">Abhishaike Mahajan writes</a>, “<em>In my opinion, what Alphafold2 pulled off — applying a clever model to a large body of pre-existing data to revolutionize a field — is something that will be extremely hard to replicate. Why? Because we’re almost out of that pre-existing data. If we had enough, sure, throw ML at it and call it a day, just as Alphafold2 did. But we don’t have that luxury anymore</em>.”</p>
<p>Mahajan argues that <a href="https://www.owlposting.com/p/wet-lab-innovations-will-lead-the">“<em>We are running out of data</em>,”</a> and that the next big innovations of AI applied to biology will require new wet-lab techniques to collect larger quantities of different biological data types than currently possible.</p>
</section>
<section id="ai-projects-begin-at-the-wrong-starting-point" class="level2">
<h2 class="anchored" data-anchor-id="ai-projects-begin-at-the-wrong-starting-point">AI projects begin at the wrong starting point</h2>
<p>Many machine learning projects begin with an existing dataset, and ask “what can we do with this data?” However, this ignores all the data that we don’t have and ends up overlooking many important questions. The harder question is, “what are the biggest questions in your area, and what data would be useful to answering them?”</p>
<p>In some cases, technical limitations prevent the collection of more, better, or different data. In the case of PDB and CASP, creating those datasets required large teams working for decades to produce, curate, and organize data.</p>
<p>One thing that is holding back the application of AI to medicine is lack of the <strong>right</strong> data. It is not just that medical data may be scattered or hard to access; in many cases, the most interesting variables aren’t being measured or collected at all. Scores of useful and high impact data are not being gathered. For example, widespread medical biases often lead doctors to incorrectly dismiss the pain or physiological symptoms of patients as “anxiety” or “drug-seeking” behavior. In these cases, doctors may fail to record relevant observations or run tests that could gather valuable data. Promises about what AI can achieve with electronic health records must be tempered with the awareness that the data within is too often biased, incorrect, or missing.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-01-24-missing-data/med-bias-research.jpg" class="img-fluid figure-img" style="width:85.0%"></p>
<figcaption>An AI analysis of electronic health records will be shaped by all the above biases in diagnosis delays, pain mismanagement, and misdiagnoses</figcaption>
</figure>
</div>
<p>As another example, many people have observed that an infection was the trigger for onset of Type 1 Diabetes in themselves or their children, yet they often don’t know which type of virus caused that infection. Determining causality around this requires large longitudinal studies, which are expensive and difficult to run.</p>
<p>AI is great at finding patterns in existing data. However, AI will not be able to solve this problem of missing and erroneous underlying data.</p>
</section>
<section id="asking-the-right-questions" class="level2">
<h2 class="anchored" data-anchor-id="asking-the-right-questions">Asking the right questions</h2>
<p>The most fundamental machine learning question, prior to the questions I mentioned above, is “what type of society do we want to live in?”</p>
<p>When <a href="https://archive.is/OdM2b">school districts implemented a computer program</a> to automatically fire supposedly low-performing teachers, there were numerous issues with the underlying algorithm and data. One key type of data used as input was student test scores, which had often been altered due to cheating scandals by school administrators. The algorithms were also based on the dubious assumption that changes in standardized test scores were a reasonable proxy for teacher quality.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-01-24-missing-data/teacher-fired.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>Washington Post headline about a teacher who was loved by students and parents, but fired by an algorithm</figcaption>
</figure>
</div>
<p>Test scores don’t capture a teacher’s creativity, care for their students, or ingenuity. They don’t capture the challenges their particular students are facing, or how those students may grow and flourish apart from test scores.</p>
<p>This is an example of an algorithm where there was low quality data, the wrong data, and faulty assumptions. However, even if it had been implemented differently, the whole premise would still be a terrible idea. Many of us (myself included) find the idea of a society where we use computerized algorithms to fire teachers to be dystopian. The answer to “How can we best use AI to fire teachers?” is that we shouldn’t.</p>
</section>
<section id="negative-spaces" class="level2">
<h2 class="anchored" data-anchor-id="negative-spaces">Negative spaces</h2>
<p>The decisions around which, where, and how data is collected, and who it is shared with, all exert a strong but often overlooked influence on the shape of scientific research and algorithm creation. As I dive deeper into the intersection of microbiology and machine learning, I am seeking to understand not just the data that I’m shown, but also the factors shaping its collection and the negative spaces of what is missing.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-01-24-missing-data/negative-space.jpg" class="img-fluid figure-img" style="width:45.0%"></p>
<figcaption>Negative Space, image by Boris Thaser, Wikimedia Commons</figcaption>
</figure>
</div>
<p>We are in a disturbing era, in which relevant health and safety data that we previously had access to are actively being cut off. Issues of bias and politics that dismiss the viewpoints of marginalized groups were already bad and seem to be worsening.</p>
<p>While promises of AI improving healthcare may seem appealing, systems that ignore patient expertise risk building on faulty premises and causing more harm than good.</p>
</section>
<section id="further-reading" class="level2">
<h2 class="anchored" data-anchor-id="further-reading">Further Reading</h2>
<ul>
<li><a href="https://dl.acm.org/doi/10.1145/3351095.3372829">Lessons from archives: strategies for collecting sociocultural data in machine learning</a> by Eun Seo Jo and Timnit Gebru, on lessons machine learning practitioners could learn from the library sciences</li>
<li><a href="https://github.com/MimiOnuoha/missing-datasets">On Missing Data Sets</a> by Mimi Onuoha</li>
<li><a href="https://www.researchgate.net/publication/258173934_The_end_of_forgetting_Strategic_agency_beyond_the_panopticon">The end of forgetting: Strategic agency beyond the panopticon</a> by Jonah Bossewitch and Aram Sinnreich, on how information flow relates to power dynamics</li>
<li><a href="https://rachel.fast.ai/posts/2024-02-20-ai-medicine/">“AI will cure cancer” misunderstands both AI and medicine</a> by Rachel Thomas</li>
<li><a href="https://rachel.fast.ai/posts/2021-11-04-data-disasters/">Avoiding Data Disasters</a> by Rachel Thomas</li>
</ul>
<p><br><br></p>
<p>You can subscribe to be notified of new blog posts by submitting your email below:</p>
<script type="text/javascript" src="https://campaigns.zoho.com.au/js/zc.iframe.js"></script>
<iframe scrolling="no" frameborder="0" id="iframewin" width="100%" height="100%" src="https://zcmp-pd.maillist-manage.com.au/ua/Optin?od=11d0c075b7ed49&amp;zx=11a17553b1&amp;tD=156971d471c7b09&amp;sD=156971d471c7cb3">
</iframe>


</section>

<p><br><br><i>I look forward to reading your responses. Create a free GitHub account to comment below.</i></p> ]]></description>
  <category>ethics</category>
  <category>machine learning</category>
  <category>science</category>
  <guid>https://rachel.fast.ai/posts/2025-01-24-missing-data/</guid>
  <pubDate>Thu, 23 Jan 2025 14:00:00 GMT</pubDate>
  <media:content url="https://rachel.fast.ai/posts/2025-01-24-missing-data/negative-space-header.jpg" medium="image" type="image/jpeg"/>
</item>
<item>
  <title>What AI can tell us about microscope slides</title>
  <dc:creator>Rachel Thomas</dc:creator>
  <link>https://rachel.fast.ai/posts/2025-01-16-cpath-foundation/</link>
  <description><![CDATA[ 





<p>The lavender images below show <a href="https://pubmed.ncbi.nlm.nih.gov/31226662/">breast tissue</a>. There are many questions doctors could want to answer using these images: They could want to know whether there are tumors present or not. If there is a tumor, doctors would want to classify its stage, make predictions about how likely the patient is to respond to treatment, and to detect whether the tumor has spread from another organ.</p>
<p>All of these are questions which people are now tackling with machine learning. They fall within the area of <a href="https://www.nature.com/articles/s41374-020-00514-0"><strong>computational pathology</strong></a>, often abbreviated <strong>CPath</strong>. In the past year, two CPath AI models were released which achieved state-of-the-art results. Here I will discuss an introduction to this field, what these models do, and what some key challenges are going forward.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-01-16-cpath-foundation/bach.jpg" class="img-fluid figure-img" style="width:90.0%"></p>
<figcaption>Breast tissue images from the BACH: Grand challenge on breast cancer histology</figcaption>
</figure>
</div>
<section id="cpath-foundation-models" class="level2">
<h2 class="anchored" data-anchor-id="cpath-foundation-models">CPath foundation models</h2>
<p>There is a powerful idea about how to make more accurate CPath models. Rather than train a model on a single type of tissue and a single task (e.g.&nbsp;identifying cancer in breast tissue), train a model on images of tissue from many different organs (breasts, lymph nodes, lungs, prostate, heart,…) and on multiple different tasks (recognizing cancer, determining the stage and subtype of the cancer, segmenting cells, and predicting treatment outcomes). Patterns learned from one dataset or one task are likely to generalize to others.</p>
<p>Such models are known as CPath <a href="https://en.wikipedia.org/wiki/Foundation_model"><strong>foundation models</strong></a>. In general, a <em>foundation model</em> is a machine learning model which is trained on a sufficiently diverse large dataset which can then be adapted for a range of downstream tasks. This idea is commonly used in the area of language models such as Chat-GPT and Claude.ai. Language foundation models are trained on many types of language tasks and intended to generalize across different corpuses of text (e.g.&nbsp;wikipedia, reddit posts, academic papers, online conversations, news articles, and more). ImageNet models trained to recognize a huge variety of different pictures often serve as foundation models for images. The success of foundation models within the areas of language and more general images is a key reason why we might expect pathology foundation models to be useful too.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-01-16-cpath-foundation/types-of-tissue.jpg" class="img-fluid figure-img" style="width:60.0%"></p>
<figcaption>Tissues are groups of cells with similar structure and function. Different types of tissue within the human body include nervous, muscle, connective, and epithelial tissue. Image: Wikimedia</figcaption>
</figure>
</div>
<p>Two notable CPath foundation models were released in 2024: <a href="https://www.nature.com/articles/s41586-024-07441-w">Prov-GigaPath</a> and <a href="https://www.nature.com/articles/s41591-024-02857-3">UNI</a>. Both models achieved state-of-the-art performance on dozens of pathology tasks (although they were not directly compared to one another). Another <a href="https://arxiv.org/abs/2404.15217">relevant paper</a> (from Kaiko.ai) studied the impact of dataset size and model size on CPath model performance.</p>
</section>
<section id="learning-the-vocab" class="level2">
<h2 class="anchored" data-anchor-id="learning-the-vocab">Learning the Vocab</h2>
<p>Medicine is full of jargon and specialized vocabulary. Pathology refers to the study of disease. It is a broad field, and can include everything from dissecting dead bodies to analyzing blood samples. One key focus of <em>computational pathology</em> is analyzing and interpreting <em>whole slide images (WSIs)</em> and in some cases combined with accompanying meta-data about a patient. Whole slide images refers to the complete microscope slide, although in many cases the region of interest (such as particular cancerous or inflamed cells) may be much smaller, just occupying a subset of the slide.</p>
<p><em>Machine learning (ML)</em> is a subfield of Artificial intelligence (AI) which involves learning from past data, and is increasingly being used with great success in pathology. The focus of most computational pathology ML models is on images of tissue, on microscope slides. That is what we will focus on in this post as well.</p>
</section>
<section id="so-many-tasks" class="level2">
<h2 class="anchored" data-anchor-id="so-many-tasks">So Many Tasks!</h2>
<p>There are many different benchmarks that CPath models can be tested on. These involve numerous datasets: related to different areas of the body, with different sizes, and with different purposes. They also involve a variety of tasks, including binary classification, image segmentation, and outcome prediction. Prov-GigaPath attained state-of-the-art performance on 25 out of the 26 tasks it was evaluated on and UNI attained state-of-the-art performance on 34 different tasks. Here I will give examples of just 3 of these tasks.</p>
<section id="task-prostate-cancer-cell-grading" class="level3">
<h3 class="anchored" data-anchor-id="task-prostate-cancer-cell-grading">Task: prostate cancer cell grading</h3>
<p>In the 1960s, the pathologist Dr.&nbsp;Donald Gleason came up with a grading scale for rating cells as they progressed from normal to prostate cancer. The Gleason Grading system is still widely used and is considered a powerful predictor of how prostate cancer patients will fare. A major medical image conference (MICCAI) <a href="https://aggc22.grand-challenge.org/AGGC22/">held a competition in 2022</a> for researchers to create algorithms to determine the Gleason grades when given images of prostate tissue.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-01-16-cpath-foundation/gleason.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>Examples of UNI predictions of Gleason grades for a section of prostate tissue. Figure 3b from the UNI paper</figcaption>
</figure>
</div>
<p>The prostate tissue is shown in pink, and segments have been colored in blocks based on where they fall on the Gleason scale.</p>
</section>
<section id="task-identifying-early-signs-of-rejection-after-a-heart-transplant" class="level3">
<h3 class="anchored" data-anchor-id="task-identifying-early-signs-of-rejection-after-a-heart-transplant">Task: identifying early signs of rejection after a heart transplant</h3>
<p>Rejection is the main cause of mortality in patients who have received a heart transplant. Since the early stages of rejection can be asymptomatic, it is standard for patients to receive frequent biopsies for 1-2 years following a transplant. These are known as endomyocardial biopsies (EMB), since they remove a small sample of tissue from the inner lining (endo) of the heart (cardial) muscle (myo). Accurately interpreting the results of these biopsies is a key question. Underestimating the chance of rejection could lead to dangerous delays in treatment, but overestimating could lead to alarm and unnecessary follow-ups or treatment. Assessment of the sampled tissue by experienced pathologists has higher variability than many other tasks, such as cancer diagnosis. Deep learning is being used to tackle this task, in models such as <a href="https://www.nature.com/articles/s41591-022-01709-2">Cardiac Rejection Assessment Neural Estimator (CRANE)</a> and the CPath foundation model UNI.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-01-16-cpath-foundation/cardiac-tissue-zoom.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>Each row shows a different sample of cardiac tissue, with a different medical issue. On the far left are the whole slide images, then zoomed in at higher resolution on a key Region of Interest (ROI). On the far right is a heat map for the most zoomed in area showing which features the algorithm has identified as significant. Figure 3 from the CRANE paper.</figcaption>
</figure>
</div>
</section>
<section id="task-genetic-mutations-in-cancer" class="level3">
<h3 class="anchored" data-anchor-id="task-genetic-mutations-in-cancer">Task: Genetic Mutations in Cancer</h3>
<p>For several common genetic mutations in tumors, there are <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC9012285/">specific drugs known to target those mutations</a>. This has a direct application for clinical treatment. Since genetic mutations can change the form and function of cells, it is reasonable to expect that this information could be deduced from images of the cancer cells. Deep learning models have been built to identify genetic mutations from tissue slides. The benefits of using a computational approach are that it can be scaled as an increasing number of relevant genetic mutations and molecular biomarkers are being discovered. <a href="https://www.nature.com/articles/s43018-020-0087-6">Task-specific models</a> have been built for this, and this is one of the tasks that foundation models can be tested on.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-01-16-cpath-foundation/cancer-mutations.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>Different types of cancer listed along the y-axis and 20 common genetic mutations listed on the y-axis. Figure 1D from Kather, 2020.</figcaption>
</figure>
</div>
</section>
</section>
<section id="we-need-more-data" class="level2">
<h2 class="anchored" data-anchor-id="we-need-more-data">We need more data</h2>
<p>One key challenge in the area of CPath foundation models is gathering enough training data. <a href="https://www.cancer.gov/ccg/research/genome-sequencing/tcga">The Cancer Genome Atlas (TCGA)</a> was an ambitious project launched in 2006 by the National Cancer Institute in the USA. Over a 12 year period, samples were collected from over 11,000 patients of 33 different cancer types, and all this data was made publicly available. While this is a rich dataset and a useful resource, all 3 papers we’ve looked at concluded that TCGA is not large enough for effective foundation models. In addition to limited data size, TCGA also has limited diversity, consisting mostly of slides from the primary site of cancer, but not metastasized cancers or different types of tissues.</p>
<p>Researchers at Kaiko.ai <a href="https://arxiv.org/abs/2404.15217">tested the impact of scaling</a> both the size of their model and the size of the training dataset. While they found limited need to scale model size beyond a certain point, they found that larger datasets continued to lead to increased performance. They concluded that TCGA was likely not large enough and shared their plans to build a larger training set, and are now partnering with cancer centers across Europe to create a dataset for their model.</p>
<p>The researchers behind two other CPath foundation models reached the same conclusion about data set size, and gathered massive datasets to train their models. This required partnering with healthcare centers. <a href="https://www.nature.com/articles/s41586-024-07441-w">Prov-GigaPath</a>, a model created by Microsoft Research and Providence Genomics involved data from 30,000 patients across 28 cancer centers (which are part of Providence Healthcare company). <a href="https://www.nature.com/articles/s41591-024-02857-3">UNI</a>, a cPath model created by a team at Harvard, MIT, and the Broad Institute, involved the creation of the Mass-100K: a dataset with over 100K whole slide images across 20 tissue types collected from Mass General Hospital, Brigham &amp; Women’s Hospital, and Genotype-Tissue Expression (GTEx) consortium.</p>
<p>These partnerships and curation of training datasets are currently a crucial component of building CPath foundation models. Curating datasets carefully poses many challenges as well. Combining data from different sources, which often use different protocols for how slides are sampled and prepared, can introduce significant biases.</p>
</section>
<section id="different-scales" class="level2">
<h2 class="anchored" data-anchor-id="different-scales">Different scales</h2>
<p>CPath foundation models face the difficulty of capturing both local patterns (that show up in a small tile within a slide) and global patterns across the whole slide. Many tiny tiles are found within a slide.</p>
<p>Some models, such as the <a href="https://arxiv.org/abs/2206.02647">Hierarchical Image Pyramid Transformer</a> (from several of the same authors as UNI), use hierarchical approaches to deal with these multiple scales.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-01-16-cpath-foundation/hierarchical.jpg" class="img-fluid figure-img" style="width:60.0%"></p>
<figcaption>Hierarchical Structure of Whole-Slide Images, Figure 1 from Chen, et al, 2020</figcaption>
</figure>
</div>
<p>Other models, such as Prov-GigaPath, treat the tiles as tokens, encoding both the tiles and the slide as a whole as model inputs. Prov-GigaPath uses both a slide encoder and a tile encoder to take into account these two different scales.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2025-01-16-cpath-foundation/slides-prov-gigapath.jpg" class="img-fluid figure-img" style="width:60.0%"></p>
<figcaption>Treating slides as tokens, Figure 1a from the Prov-GigaPath paper</figcaption>
</figure>
</div>
<p>In pathology clinics, diagnosis and treatment decisions are often made at the patient level, whereas CPath models are often highly focused on regions of interest. Accommodating the multiple relevant scales (small tiles, whole slides, and patient-level) for pathology is a consideration that CPath models need to balance.</p>
</section>
<section id="going-forward" class="level2">
<h2 class="anchored" data-anchor-id="going-forward">Going Forward</h2>
<p>It is still early in the world of CPath and there are many growth opportunities, including the continued need for large and diverse datasets, ways to further optimize model training, tasks which have previously received less focus, and the difficulties of integrating models into clinical work. As the <a href="https://arxiv.org/abs/2404.15217">authors of the kaiko.ai paper</a> wrote, “<em>We are still at the very beginning of developing a truly foundational pathology foundation model.</em>” It is a hopeful sign that these models achieve state-of-the-art results on dozens of benchmarks, but it still remains to be seen when and how they will be used in clinical settings.</p>
</section>
<section id="related-reading" class="level2">
<h2 class="anchored" data-anchor-id="related-reading">Related Reading:</h2>
<ul>
<li><a href="https://rachel.fast.ai/posts/2024-07-24-neural-nets/">The Most Common and Useful Neural Nets</a></li>
<li><a href="https://rachel.fast.ai/posts/2024-04-03-ai-antibiotics/">Using AI to Discover New Antibiotics</a></li>
<li><a href="https://rachel.fast.ai/posts/2024-11-20-ai-immunology/">AI and Immunology</a></li>
</ul>
<p>You can subscribe to be notified of new blog posts by submitting your email below:</p>
<script type="text/javascript" src="https://campaigns.zoho.com.au/js/zc.iframe.js"></script>
<iframe scrolling="no" frameborder="0" id="iframewin" width="100%" height="100%" src="https://zcmp-pd.maillist-manage.com.au/ua/Optin?od=11d0c075b7ed49&amp;zx=11a17553b1&amp;tD=156971d471c7b09&amp;sD=156971d471c7cb3">
</iframe>


</section>

<p><br><br><i>I look forward to reading your responses. Create a free GitHub account to comment below.</i></p> ]]></description>
  <category>machine learning</category>
  <category>science</category>
  <guid>https://rachel.fast.ai/posts/2025-01-16-cpath-foundation/</guid>
  <pubDate>Wed, 15 Jan 2025 14:00:00 GMT</pubDate>
  <media:content url="https://rachel.fast.ai/posts/2025-01-16-cpath-foundation/gleason.jpg" medium="image" type="image/jpeg"/>
</item>
<item>
  <title>AI and Immunology</title>
  <dc:creator>Rachel Thomas</dc:creator>
  <link>https://rachel.fast.ai/posts/2024-11-20-ai-immunology/</link>
  <description><![CDATA[ 





<p>I spent close to 20 years focused on mathematics and data science, including cofounding the research lab fast.ai, which focused on the powerful family of AI algorithms known as deep learning. A few years ago, I decided to make a big pivot and <a href="https://rachel.fast.ai/posts/2023-02-07-school-immunology/">return to school</a> for immunology. What motivated my sudden change amidst a successful career? We are now living in <a href="https://www.theatlantic.com/science/archive/2022/04/how-climate-change-impacts-pandemics/629699/">a pandemicene</a>, a period with increasingly likely pandemics. Climate change and habitat destruction are crowding species into closer and closer contact with humans. <a href="https://rachel.fast.ai/posts/2024-08-13-crowds-vs-friends/">Frequent global travel and mega-cities</a> allow unprecedented opportunities for viruses to spread and mutate. Antibiotic resistance is rising rapidly.</p>
<p>At the same time, we have been learning more and more about the <a href="https://rachel.fast.ai/posts/2023-03-07-viruses1/">long-term consequences of infections</a>– seemingly mild infections can contribute to long-term <a href="https://rachel.fast.ai/posts/2023-03-22-viruses2/">autoimmune diseases, neurodegenerative diseases</a>, and cancer. The human immune system is incredibly complex and there is much we don’t know about it. Immunology is a crucial area to study. In this post, I want to gather some of my writing and talks on how AI is being applied to immunology.</p>
<center>
<iframe width="560" height="315" src="https://www.youtube.com/embed/I3q5-cBebKA?si=MFcccYwqKdSTGth0" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="">
</iframe>
<br> <em>My 30 minute presentation on key areas where deep learning is being applied to immunology</em>
</center>
<section id="using-ai-to-predict-what-t-cells-will-bind-to" class="level3">
<h3 class="anchored" data-anchor-id="using-ai-to-predict-what-t-cells-will-bind-to">Using AI to predict what T cells will bind to</h3>
<p>T cells are one of the most important cell types of our immune systems. Figuring out how to predict what a T cell will bind to (meaning what cells it can recognize as bad and coordinate attacks against) would be a vital medical breakthrough.</p>
<ul>
<li><a href="https://rachel.fast.ai/posts/2024-07-09-t-cells/">Decoding T cells with AI</a></li>
<li><a href="https://rachel.fast.ai/posts/2024-07-24-neural-nets/">The Most Common and Useful Neural Nets</a></li>
<li><a href="https://rachel.fast.ai/posts/2024-08-01-ai-t-cells/">AI’s Quest to Predict T Cell Binding– The Holy Grail of Immunology</a></li>
</ul>
</section>
<section id="mapping-immune-cell-communication-networks" class="level3">
<h3 class="anchored" data-anchor-id="mapping-immune-cell-communication-networks">Mapping immune cell communication networks</h3>
<p>Immune cells talk to each other through complex networks. NLP, math, and immunology are useful in studying these networks.</p>
<ul>
<li><a href="https://rachel.fast.ai/posts/2024-01-23-cytokines1/">Applying AI to Immune Cell Networks</a></li>
<li><a href="https://rachel.fast.ai/posts/2024-02-06-cytokines2/">How Immune Cells Communicate</a></li>
</ul>
</section>
<section id="discovering-new-antibiotics" class="level3">
<h3 class="anchored" data-anchor-id="discovering-new-antibiotics">Discovering new antibiotics</h3>
<p>Bacterial resistance to existing antibiotics is an urgent threat. AI is being used to search for new antibiotics.</p>
<ul>
<li><a href="https://rachel.fast.ai/posts/2024-04-03-ai-antibiotics/">Using AI to Discover New Antibiotics</a></li>
</ul>
</section>
<section id="ethical-risks-of-ai-applied-to-immunology" class="level3">
<h3 class="anchored" data-anchor-id="ethical-risks-of-ai-applied-to-immunology">Ethical risks of AI applied to immunology</h3>
<p>The enthusiasm about AI in medicine is failing to grapple with realities of the system. Recognizing gaps in how AI is applied to scientific research can prevent potential shortcomings and risks.</p>
<ul>
<li><a href="https://rachel.fast.ai/posts/2024-09-10-gaps-risks-science/">Gaps and Risks of AI in the Life Sciences</a></li>
<li><a href="https://rachel.fast.ai/posts/2024-02-20-ai-medicine/">“AI will cure cancer” misunderstands both AI and medicine</a></li>
</ul>
<p>You can subscribe to be notified of new blog posts by submitting your email below:</p>
<script type="text/javascript" src="https://campaigns.zoho.com.au/js/zc.iframe.js"></script>
<iframe scrolling="no" frameborder="0" id="iframewin" width="100%" height="100%" src="https://zcmp-pd.maillist-manage.com.au/ua/Optin?od=11d0c075b7ed49&amp;zx=11a17553b1&amp;tD=156971d471c7b09&amp;sD=156971d471c7cb3">
</iframe>


</section>

<p><br><br><i>I look forward to reading your responses. Create a free GitHub account to comment below.</i></p> ]]></description>
  <category>machine learning</category>
  <category>science</category>
  <guid>https://rachel.fast.ai/posts/2024-11-20-ai-immunology/</guid>
  <pubDate>Tue, 19 Nov 2024 14:00:00 GMT</pubDate>
  <media:content url="https://rachel.fast.ai/posts/2024-11-20-ai-immunology/thumbnail.jpg" medium="image" type="image/jpeg"/>
</item>
<item>
  <title>5 Devious Tricks Pathogens Use Against Us</title>
  <dc:creator>Rachel Thomas</dc:creator>
  <link>https://rachel.fast.ai/posts/2024-11-12-devious-tricks/</link>
  <description><![CDATA[ 





<p>“The human immune system is amazing, with an impressive range of techniques to protect against invaders.”</p>
<p>I see this sentiment expressed frequently by everyone from doctors quoted in newspaper articles to online commentators in parenting groups. And I share it– I loved learning about the intricate processes our bodies deploy against invaders when I took graduate immunology.</p>
<p>However, the point that has struck me again and again as I complete my MS in Microbiology-Immunology is that pathogens (virues, bacteria, fungi, and parasites that cause disease) have a fierce and varied range of techniques to overcome our immune defenses, and even to turn our cells’ own defenses against us. Seeing the ingenuous lengths pathogens are capable of illustrates that it is often preferable to avoid getting infected in the first place.</p>
<p>There is hubris in exalting only the human immune system, without also recognizing the capacity of viruses and bacteria to wreak havoc in clever ways. In this post, I will cover 5 different ways that pathogens subvert our defenses.</p>
<section id="commandeering-the-cells-own-machinery" class="level2">
<h2 class="anchored" data-anchor-id="commandeering-the-cells-own-machinery">1. Commandeering the cell’s own machinery</h2>
<p>The year is 2063. An insect-like robot with a swinging gait and clamp-like feet marches along a beam in a space colony, transporting the large cargo suspended behind it.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-11-12-devious-tricks/motor-protein-short.gif" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>A motor protein moving cargo along a microtubule</figcaption>
</figure>
</div>
<p>Ooops– what I meant to say is that a microscopic motor protein in one of our cells moves a bag of proteins along the cytoskeleton, away from or towards the nucleus. This is not a metaphor; we actually have <a href="https://www.youtube.com/watch?v=y-uuk4Pr2i8">motor proteins that are moving complexes</a> within our cells. The above gif is scientifically accurate, not a fantastical invention.</p>
<p>I used to think that cells were just sacks of liquid, with various items such as the mitochondria or nucleus floating haphazardly around inside. I couldn’t have been more wrong. Just as beams, drywall, and supporting columns are necessary to organize and support the architecture of a building, so too do our cells have an intricate organization arranged by various filaments and microtubules.</p>
<p>Tiny motor proteins that look like robots are used to transport items around. But just like the space colony mentioned above, these robots can be commandeered by enemy invaders. Unfortunately, the virus which plays a role in Alzheimer’s Disease is able to use these motor proteins for its own goals.</p>
<p>Alert! Alert! Herpes Simplex Virus-1 (HSV-1) has breached the cell! When HSV-1 invades a human cell, it <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC2132479/">commandeers our motor proteins</a> to transport the viral capsids (packets containing the viral DNA) to the nucleus. The sooner the viral capsids arrive at the nucleus, the sooner the virus can begin replicating, printing copies of itself to invade the neighboring space colonies, err, I mean neighboring cells. HSV-1 has turned our cell’s own machinery against us to improve its speed and efficiency in infection.</p>
<p>HSV-1 is a virus responsible for causing fever blisters on the mouth. It can live latently in neurons, activating or reactivating after decades. A growing body of research is linking HSV-1 with a <a href="https://www.ox.ac.uk/news/2022-08-02-viral-role-alzheimers-disease-discovered">major role in Alzheimer’s Disease</a>. The ability of HSV-1 to use human motor proteins to speed its path to the nucleus for replication is just one of the many devious tricks that different viruses and bacteria have developed.</p>
</section>
<section id="blocking-signals-from-t-cells" class="level2">
<h2 class="anchored" data-anchor-id="blocking-signals-from-t-cells">2. Blocking signals from T cells</h2>
<p>T cells are a keystone of the human immune system. T cells can rearrange their genes, allowing for the potential recognition of billions of different invaders. And yet they are outmatched by a single-celled organism (<em>T. cruzi</em>) carried in the feces of so-called “kissing bugs” (triatomine bugs). A kissing bug bites you in your sleep. In many cases, you may not even notice, and experience no symptoms in the coming weeks. <em>T. cruzi</em> is a species of trypanasome, a tiny one-celled parasite, and causes Chagas disease. They live in the feces of the bugs can enter the bite wound and travel through your bloodstream. They have a preference for making a home inside the muscles of your heart, the lining of your digestive tract, or the nervous system. <em>T. cruzi</em> can hang out patiently for 10-30 years, but are then ready to wreak havoc. Up to one third of patients experience heart problems caused by destruction of cardiac tissues. One in 10 experience neurological and digestive problems. Some die sudden deaths.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-11-12-devious-tricks/chagas-cdc.jpg" class="img-fluid figure-img" style="width:60.0%"></p>
<figcaption>T. cruzi parasites are stained as purple squiggles against the red blobs of human blood cells</figcaption>
</figure>
</div>
<p>Why haven’t your T cells protected you? One of <em>T. cruzi</em>’s devious tricks is that it is able to block the key signal which coordinates the T cell response (a messenger known as IL-2). IL-2 is crucial for activating T cells. With its messaging blocked by these tiny parasites, T cells don’t activate and respond to the threat, and <em>T. cruzi</em> is able to survive in humans, later causing great destruction. While there is medication available that will work if addressed early, remember, most people have no symptoms initially. The best protection against <em>T. cruzi</em> is to avoid it in the first place, which typically requires access to a home that keeps out insects.</p>
</section>
<section id="making-a-home-in-b-cells" class="level2">
<h2 class="anchored" data-anchor-id="making-a-home-in-b-cells">3. Making a home in B cells</h2>
<p>In addition to T cells, the other type of adaptive immune cells that can rearrange their genes are B cells. B cells produce antibodies, small proteins that bind to a target, marking it for destruction by other immune cells. Through a process of hypermutation, B cells are able to create carefully custom-tuned antibodies to target any enemy they want. The flexibility and nearly infinite customizability of this defense are powerful.</p>
<p>Given the powers of B cells, you might expect that invaders would want to stay away from B cells. Yet Epstein-Barr Virus (EBV) and measles virus actively seek out B cells. EBV is able to set up a life-long home in memory B cells, hanging out for decades. While EBV is best known for causing mono, it can also <a href="https://www.nature.com/articles/s41579-022-00770-5">cause multiple sclerosis</a> and cancer. In multiple sclerosis, the immune system attacks the nervous system, damaging neurons and causing plaques in the brain. If the viral antigens present in EBV look similar to neurons, this could confuse the immune system and lead to autoimmune attack.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-11-12-devious-tricks/EBV.jpg" class="img-fluid figure-img" style="width:65.0%"></p>
<figcaption>Epstein-Barr Virus can cause a lot of trouble</figcaption>
</figure>
</div>
<p>Another virus, measles, plays an amnesia trick on memory B cells, causing them to forget their memories of other invaders that the B cells had previously learned to recognize. This <a href="https://www.bbc.com/future/article/20211112-the-people-with-immune-amnesia">immune amnesia</a> is why children who have been infected with measles are significantly more likely to die of other illnesses later, compared to children who have never had measles. The measles vaccine drastically <a href="https://www.bmj.com/content/311/7003/481">reduces the risk of non-measles related deaths</a> by letting the immune system retain the information it has learned.</p>
</section>
<section id="thriving-in-acid" class="level2">
<h2 class="anchored" data-anchor-id="thriving-in-acid">4. Thriving in acid</h2>
<p>T cells and B cells aren’t the only immune cells targeted by particular pathogens. Macrophage literally means “big eater”. These large cells are a useful component of the immune system, because they are ready to eat your enemies. Macrophages contain bags of acid to acidify and ingest invaders. They share the digested remnants of their foes with other immune cells to activate them and coordinate the immune response. Sounds fierce, right? And yet a number of bacteria have developed strategies to survive inside them.</p>
<p>“No! No! Don’t dump me in acid!” cries the bacteria <em>C. burnetii</em> insincerely, “Anything but that!” In fact, that is exactly what <em>C. burnetii</em> (the bacteria that causes Q fever) wants. It thrives in acid. Acidic environments are where it can replicate, making more and more copies of itself. <em>C. burnetii</em> prospers inside the the acidic pouches in macrophages.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-11-12-devious-tricks/Coxiella-burnetii-NIH.jpg" class="img-fluid figure-img" style="width:65.0%"></p>
<figcaption>C. burnetii (shown in green) replicating inside of a vacuole, a location in macrophages where invaders are supposed to be killed</figcaption>
</figure>
</div>
<p><em>C. burnetii</em> is not the only type of bacteria that has figured out strategies for living inside the “Big Eaters” that are meant to protect us. Tuberculosis is a bacteria that spreads when infected people cough, talk, or even just breathe. Approximately <a href="https://www.who.int/teams/global-tuberculosis-programme/tb-reports/global-tuberculosis-report-2021/disease-burden/mortality">1.2 million people die each year of tuberculosis</a>, almost double the number of people who die of HIV. Tuberculosis that is resistant to multiple antibiotics is an urgently growing problem. And unfortunately, our macrophages are unable to eliminate it. When macrophages take a big bite of an invader, this leads to the invader initially being inside a pouch called a phagosome, within the macrophage. Usually, the next step is for the phagosome (containing the invader) to merge with a pouch of acid (called the lysosome). However, tuberculosis is able to prevent that next step from happening. Tuberculosis loves to just kick back and relax within the phagosome. It can remain latent for years or even decades, but can reactivate if the immune system becomes suppressed.</p>
</section>
<section id="when-we-put-the-brakes-on-translation-viruses-cut-the-brake-lines" class="level2">
<h2 class="anchored" data-anchor-id="when-we-put-the-brakes-on-translation-viruses-cut-the-brake-lines">5. When we put the brakes on translation, viruses cut the brake lines</h2>
<p>Pathogens don’t just target immune cells, but also more general processes found in all our cells. A central action is the process by which DNA is transcribed into RNA, which is then translated into proteins. Proteins are the building blocks for all sorts of important materials within our cells, and proteins allow communication throughout the body. The process of producing protein from RNA is known as “translation”. Viruses love to take over our protein translation machinery for their own purposes.</p>
<p>In my molecular cell biology course, the professor recently taught about a method our cells have for halting protein translation when they have been infected by a virus (for my fellow immunology nerds, we were learning about <a href="https://www.frontiersin.org/journals/microbiology/articles/10.3389/fmicb.2021.757238/full">PKR phosphorylation of eIF2α</a>– but you don’t need to know any details of it for this blog post). It is a bit like our cells stepping on the brakes to stop protein translation.</p>
<p>That is encouraging that we can block those pesky viruses from producing proteins! Good job, human cells!</p>
<p>Then, on the following slide, the professor <a href="https://www.frontiersin.org/journals/microbiology/articles/10.3389/fmicb.2021.757238/full">shared a paper</a> that covers 31 different viruses that sabotage this safety mechanism (via 9 categories of actions) so that they can use the cells’ machinery to produce their own viral proteins regardless. Viruses have effectively cut the brake lines, so they can continue barreling along. And so it goes. Each time I learn about an intricate and ingenuous aspect of the human immune system, the following page seems to contain a clever way that pathogens circumvent or subvert it.</p>
</section>
<section id="viewing-pathogens-with-appropriate-respect-gives-us-a-path-forward" class="level2">
<h2 class="anchored" data-anchor-id="viewing-pathogens-with-appropriate-respect-gives-us-a-path-forward">Viewing pathogens with appropriate respect gives us a path forward</h2>
<p>There is hubris in focusing solely on the impressive coordination of the human immune system, without considering the methods that pathogens have developed for counterattacks. This information is not just an academic point; it has practical impact. Unthinking repetition of the soundbite “the immune system is amazing” can lull people into a false sense of confidence that disease prevention is unnecessary. While it can be discouraging to think about the methods bacteria and viruses have for causing harm and the surprising long-term consequences of some infections, the good news is that many infections are preventable.</p>
<p>Mosquito nets, houses that keep out insects, <a href="https://www.theguardian.com/society/2007/jan/19/health.medicineandhealth3">sewage systems, clean drinking water</a>, <a href="https://www.abc.net.au/news/2024-09-24/covid-safety-schools-course-sick-days-teachers-long-covid/104319032">air purifiers</a>, <a href="https://theconversation.com/masks-work-our-comprehensive-review-has-found-229658">N95s in public indoor spaces</a>, and <a href="https://www.science.org/doi/10.1126/science.abg2025">good ventilation</a> are all effective at reducing risks of infection. In particular, research from the past few years has illustrated that <a href="https://www.theguardian.com/society/2007/jan/19/health.medicineandhealth3">cleaning up indoor air quality</a> would likely stop transmission of numerous pathogens.</p>
<p>I still marvel at the intricacy of the human immune system, but alongside that marvel, I keep a commensurate respect for the serious deviousness of many pathogens.</p>
</section>
<section id="related-reading" class="level2">
<h2 class="anchored" data-anchor-id="related-reading">Related Reading:</h2>
<ul>
<li><a href="https://rachel.fast.ai/posts/2024-08-13-crowds-vs-friends/">Our immune systems evolved for a different world (that didn’t involve 100,000 global flights per day)</a></li>
<li><a href="https://rachel.fast.ai/posts/2024-04-25-microbiome-1/">Surprises of the Microbiome</a></li>
<li><a href="https://rachel.fast.ai/posts/2023-03-22-viruses2/">Viruses: The Silent Triggers of Autoimmune and Neurodegenerative Diseases</a></li>
</ul>


</section>

<p><br><br><i>I look forward to reading your responses. Create a free GitHub account to comment below.</i></p> ]]></description>
  <category>science</category>
  <guid>https://rachel.fast.ai/posts/2024-11-12-devious-tricks/</guid>
  <pubDate>Mon, 11 Nov 2024 14:00:00 GMT</pubDate>
  <media:content url="https://rachel.fast.ai/posts/2024-11-12-devious-tricks/Coxiella-burnetii-NIH-thumbnail.jpg" medium="image" type="image/jpeg"/>
</item>
<item>
  <title>In defense of screen time</title>
  <dc:creator>Rachel Thomas</dc:creator>
  <link>https://rachel.fast.ai/posts/2024-10-29-screen-time/</link>
  <description><![CDATA[ 





<p>My daughter is constantly creating– her passions include making art, writing fiction, coding interactive games, and composing music. Yet, I regularly see news articles and media pundits suggesting that my husband and I are doing things all wrong. The reason for these claims? My daughter uses screen-based tools, at least partially, and in some cases, entirely, to pursue the interests I listed in my first sentence. To give more detail on how she pursues her passions:</p>
<ul>
<li><strong>Art</strong>: She creates digital art in <a href="https://www.sketchbook.com/">Sketchbook Pro</a>, both on her own and in lessons. She also takes in-person art courses that involve a mix of acrylic painting, water colors, and sculpture.</li>
<li><strong>Coding</strong>: She loves coding in <a href="https://scratch.mit.edu/">Scratch</a> and <a href="https://p5js.org/">P5.js</a> to design and build interactive games.</li>
<li><strong>Creative writing</strong>: While her handwriting is on target for her age, she prefers typing, as it lets her write faster and express more complex ideas– we taught her touch typing when she was 5. She writes fiction both individually and collaboratively with friends.</li>
<li><strong>Music</strong>: She plays the piano, takes group singing lessons online, and enjoys composing music. There is a composition component to her piano lessons (which happen over zoom) and she composes music online through <a href="https://musiclab.chromeexperiments.com/">Chrome MusicLab</a> and <a href="https://flat.io/">Flat.io</a>.</li>
</ul>
<p>My daughter is not unique. Our family knows many other children with similar creativity levels who are excelling in their passions beyond what is often believed possible for their age groups.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-10-29-screen-time/music-lab.jpg" class="img-fluid figure-img" style="width:60.0%"></p>
<figcaption>Visual from Chrome Music Lab</figcaption>
</figure>
</div>
<p>Parenting approaches are personal and often polarizing, so I normally don’t write on this topic. However, I am concerned by how several (in my view, false) points are being increasingly repeated by politicians and pundits: that screentime is very harmful for children, that it is essential for children’s well-being to attend in-person school every day (<a href="https://apnews.com/article/covid-flu-school-attendance-4845073e737db87f786e1d8c815a48f7">even</a> <a href="https://www.latimes.com/california/story/2023-08-12/got-a-cold-runny-nose-the-sniffles-no-worries-come-to-school-lausd-says">when</a> <a href="https://www.1news.co.nz/2024/09/26/govt-reveals-new-plan-for-getting-kids-to-school-heres-how-it-will-work/">sick</a>), and that it is important for workers to return to the office (even when their jobs can be done remotely).</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-10-29-screen-time/school-attendance.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>There has been a rise in policies requiring children to attend school while sick</figcaption>
</figure>
</div>
<p>The above points are interlocking, since getting workers commuting to in-person offices requires children being at in-person school. And overemphasizing in-person school attendance overlooks a number of children whose needs aren’t being met by in-person school. It also overlooks a number of online and screen-based options opportunities to build skills, express creativity, and form friendships. Here I want to focus on some innovative ways that we can encourage children’s flourishing, whether that is as a supplement to in-person school, or instead of.</p>
<section id="a-false-binary" class="level2">
<h2 class="anchored" data-anchor-id="a-false-binary">A False Binary</h2>
<p>I often hear statements such as, “I’d rather kids played outside than on a computer.” But these choices aren’t binary! You don’t have to pick just one. My daughter plays outside AND on a computer. She plays in-person sports regularly AND has online hobbies. This year we threw her 2 birthday parties: an in-person one for her local friends AND an online one for her long-distance friends.</p>
<p>The term “screen time” is so broad as to be meaningless, clumping together many disparate activities. A child sitting next to a parent playing a math game together on the computer is screen time. Calling your grandparents on skype is screen time. Collaboratively composing music with your best friend in another country is screen time. Yet, when many people refer to screen time, they often seem to be referring to a child passively watching TV on their own (something which my husband and I almost never let our daughter do, and that she doesn’t seem interested in). With any screen-based activity, it is useful to consider:</p>
<ul>
<li>Is it social or solitary?</li>
<li>Does it involve creating content or consuming content?</li>
<li>Is it educational or not?</li>
</ul>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-10-29-screen-time/venn-diagram.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>A Venn Diagram showing how we think about screentime. We avoid the outside (white) region and mostly stick to the intersections.</figcaption>
</figure>
</div>
<p>It is not that some of these options are always good or always bad, but rather that we need to have a nuanced view on different combinations of these factors and how much time you may want to spend on each.</p>
<p>Screens are a crucial tool and outlet for her expressing her creativity, and it saddens me to hear an increasing media narrative vilifying screen time for kids.</p>
<p>Two of my daughter’s best friends are kids she has never met in person (because they live in a different state and in a different country), yet she talks with each of them several times per week. She also has local friends that we get together with regularly in-person. The fact that she has several close long-distance friends isn’t surprising, because my husband and I have many close online / long-distance friendships as well. Our friendships aren’t any less real because they are over screens. Opening up the possibility for long-distance friendships has given us the option of finding additional friends we are particularly compatible with.</p>
</section>
<section id="increased-options" class="level2">
<h2 class="anchored" data-anchor-id="increased-options">Increased options</h2>
<p>Screen-based learning options make homeschooling more feasible for a wider range of families. Online resources (games, classes, and clubs) can offer opportunities that may not be available locally, can fill in gaps parents may not be qualified to cover, and offer an exciting variety of options. For instance, online courses have allowed my daughter to do the following (which are not available locally for us):</p>
<ul>
<li>have tutors/teachers based in 5 different countries, providing diverse experiences</li>
<li>take science courses including: particle physics for kids, biochem for kids (she has been with the same group of kids for 3 years), and a biology class where a small group of students could share images from their digital microscopes over zoom</li>
<li>participate in clubs for creating collaborative narratives, fanart, and fanfiction around two of her favorite series of books</li>
<li>complete mastery-based math games at her own pace, which allowed her and a few friends to end up many years above their grade levels</li>
</ul>
<p>As I wrote about in a previous post, <a href="https://www.fast.ai/posts/2022-09-06-homeschooling.html">we homeschool</a>, and this leaves our daughter with far more time for hobbies and socializing than in-person school gave us.</p>
</section>
<section id="in-person-school-does-not-work-for-everyone" class="level2">
<h2 class="anchored" data-anchor-id="in-person-school-does-not-work-for-everyone">In-person school does not work for everyone</h2>
<p>There are several categories of kids whose needs are often not met in traditional schools and frequently benefit from online options:</p>
<ul>
<li><a href="https://www.nytimes.com/2020/08/10/opinion/coronavirus-school-closures.html">Neurodivergent kids</a>, who may deal with overstimulating, distracting, or painful sensory experiences in traditional schools</li>
<li>Gifted kids, who are not intellectually challenged in traditional schools</li>
<li><a href="https://www.wired.com/story/pandemic-homeschoolers-who-are-not-going-back/">Kids from marginalized cultures</a>, whose cultural heritage may not be reflected at school, or who may experience discrimination</li>
<li>Kids who are being bullied</li>
<li>Kids who have particular passions they would like to hyper-focus on (which are either not included in their school’s curriculum, or which they don’t get adequate time for)</li>
<li>Medically complex kids, including those who may be getting sick so often that they can’t attend school regularly, those for whom an illness could result in hospitalization, or whose health issues make it difficult to attend school in-person</li>
</ul>
<p>It is an injustice that schools are not inclusive of all children. While individuals pulling their children out of school does not address this injustice (and we should continue to work for <a href="https://www.abc.net.au/news/2024-09-24/covid-safety-schools-course-sick-days-teachers-long-covid/104319032">safer</a> and more inclusive schools), many families have found it necessary to leave systems that are harming their kids.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-10-29-screen-time/homeschool-headlines.jpg" class="img-fluid figure-img" style="width:60.0%"></p>
<figcaption>We were not alone in discovering that homeschooling worked better for our child</figcaption>
</figure>
</div>
</section>
<section id="dont-judge-screentime-by-a-few-bad-examples" class="level2">
<h2 class="anchored" data-anchor-id="dont-judge-screentime-by-a-few-bad-examples">Don’t judge screentime by a few bad examples</h2>
<p>I would caution against writing off screen time or online learning based on a low quality or poorly designed experience. Maybe your child’s school went online in 2020 with no notice and no support, and they were stuck in a zoom with 25 other kids (this is not a good set-up for anyone!). We have found that with small group online tutoring, my daughter is able to be more social and talkative than she was in a large class at in-person school, and she masters more material in less time. However, even amongst small online groups, occasionally she’ll try a teacher or club that isn’t a good match for her, and in those cases, we discontinue it.</p>
<p>Since screentime is often equated with a way to occupy kids while parents do something else, there are often unrealistic expectations, particularly for the most rewarding screen-based tools (which tend to involve a learning curve and require support from a trusted adult). In many cases handing your child a new educational app and then leaving them alone can lead to frustration. Sitting with them and working through the app together can be relationship-building, prevent them from getting overly discouraged, and help them build confidence as you figure it out together. With age and experience, kids will likely be able to do more on their own.</p>
</section>
<section id="inclusion-and-accessibility" class="level2">
<h2 class="anchored" data-anchor-id="inclusion-and-accessibility">Inclusion and accessibility</h2>
<p>Countries around the world, including the <a href="https://www.ippr.org/articles/our-greatest-asset">USA</a>, <a href="https://news.sky.com/story/parents-of-children-with-complex-needs-worried-they-could-be-unfairly-fined-under-new-school-absence-rules-13213524">UK</a>, <a href="https://www.educationtoday.com.au/news-detail/Half-of-Australia-6155">Australia</a>, <a href="https://www.rnz.co.nz/news/political/513250/sickness-related-school-absences-to-be-targeted-under-government-plan">New Zealand</a>, <a href="https://english.kyodonews.net/news/2023/10/d8cd090637bd-japan-school-absenteeism-at-record-high-of-nearly-300000-in-fy-2022.html">Japan</a>, <a href="https://www.brusselstimes.com/1053759/number-of-absences-and-suspensions-among-pupils-in-belgium-on-the-rise">Belgium</a>, and <a href="https://www.cbc.ca/news/canada/school-absence-data-1.7156254">Canada</a> are seeing high levels of school absences. Factors for these record levels of absences include <a href="https://calgaryherald.com/opinion/columnists/opinion-we-dont-know-whats-causing-the-tsunami-of-sick-kids-but-wed-better-figure-it-out-fast">more frequent illnesses</a> (such as covid, RSV, and flu) and <a href="https://www.abc.net.au/news/2024-06-16/children-with-long-covid-dismissed-doctors-myth-virus-harmless/103959078">chronic health issues</a>. In mid-2024, many countries (including the USA and Australia) experienced a particularly high wave of the ongoing covid pandemic. Rates of disability are increasing, and people, <a href="https://www.scientificamerican.com/article/long-covid-is-harming-too-many-kids/">including children</a>, are continuing to develop new health conditions after viral infections.</p>
<p>A number of school districts are responding to these waves of illness by placing heightened emphasis on attendance to in-person school, in many cases pressuring or <a href="https://apnews.com/article/covid-flu-school-attendance-4845073e737db87f786e1d8c815a48f7">even</a> <a href="https://www.1news.co.nz/2024/09/26/govt-reveals-new-plan-for-getting-kids-to-school-heres-how-it-will-work/">requiring</a> <a href="https://www.latimes.com/california/story/2023-08-12/got-a-cold-runny-nose-the-sniffles-no-worries-come-to-school-lausd-says">children</a> and <a href="https://archive.md/T2oHI">teachers to attend school while sick</a>. In parts of the <a href="https://news.sky.com/story/parents-of-children-with-complex-needs-worried-they-could-be-unfairly-fined-under-new-school-absence-rules-13213524">UK</a>, <a href="https://www.indystar.com/story/news/education/2024/08/19/new-law-indiana-school-absences-juvenile-court-missing-school-ips-attendance-policy-criminal-charges/74818189007/">USA</a>, and <a href="https://www.rnz.co.nz/news/political/513250/sickness-related-school-absences-to-be-targeted-under-government-plan">New Zealand</a>, policies are being considered (or implemented) to charge parents fines if their children miss school too often. This is counterproductive, as it is hard to focus when unwell, attending school sick increases the likelihood of spreading illness to others, and <a href="https://www.latimes.com/california/story/2022-07-07/working-through-covid-sleep-rest-infection-test-positive">rest is a key way</a> to reduce risk of post-viral illness. None of these punitive attendance policies address root causes, such as <a href="https://www.abc.net.au/news/2024-09-24/covid-safety-schools-course-sick-days-teachers-long-covid/104319032">improving school air quality</a>, making sure parents have adequate paid leave to stay home with sick children, or decoupling school funding from attendance.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-10-29-screen-time/school-attendance.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>Requiring kids to attend school while sick is only going to lead to more sick students and sick teachers</figcaption>
</figure>
</div>
<p>We are also seeing <a href="https://www.abs.gov.au/media-centre/media-releases/55-million-australians-have-disability">rising rates of disability</a> amongst <a href="https://x.com/Mike_Honey_/status/1842530019991769108?t=Cc8jUI2CayxAIBCpoZHiXw&amp;s=19">working age</a> adults. In the USA, there has been an <a href="https://www.ippr.org/articles/our-greatest-asset">increase in people with cognitive disabilities</a>. A <a href="https://www.bbc.com/news/business-68639144">BBC article</a> notes that in the UK sick people are leaving the workforce at record highs. Online access is a crucial part of accessibility and disability rights. For many disabled people, accessibility improved in 2020 with more remote options for work, medical appointments, conferences, and other events, and there is now a counter-reaction in which these options are being removed.</p>
<p>Every child is different, so what works for my family won’t work for everyone. I am not recommending that all families homeschool, or that nobody should attend in-person school. However, we need to consider innovative approaches to learning and communication, and that includes screen-based approaches.</p>
</section>
<section id="hopes-for-a-creative-and-inclusive-society" class="level2">
<h2 class="anchored" data-anchor-id="hopes-for-a-creative-and-inclusive-society">Hopes for a creative and inclusive society</h2>
<p>I want children to discover and develop passions that they enjoy (astronomy, writing, art, music, math, oceanography, chess, there are so many possibilities!) and to be able to connect with others who share their interests. Taking advantage of computer-based and online options opens up a world of possibility. I also want a society in which disabled people are included in events and opportunities, and in which people can rest when they’re sick or preferably avoid illness in the first place. Fewer people commuting to the office is better for the environment. Traditional school is failing a lot of kids. I hope we can consider a range of innovative educational and social approaches, and not exclude an entire category of valuable ways to communicate, create, and learn (e.g.&nbsp;screentime).</p>
</section>
<section id="related-posts" class="level2">
<h2 class="anchored" data-anchor-id="related-posts">Related Posts</h2>
<ul>
<li><a href="https://www.fast.ai/posts/2022-09-06-homeschooling.html">My family’s unlikely homeschooling journey</a>: Prior to 2020, we never expected to homeschool, and now we have committed to it long-term.</li>
<li><a href="https://www.fast.ai/2022/03/15/math-person/">There’s No Such Thing as Not A Math Person</a>: based on a webinar I gave to parents addressing cultural myths about math and how to support your kids in their math education, even if you don’t see yourself as a “math person.”</li>
</ul>


</section>

<p><br><br><i>I look forward to reading your responses. Create a free GitHub account to comment below.</i></p> ]]></description>
  <category>education</category>
  <guid>https://rachel.fast.ai/posts/2024-10-29-screen-time/</guid>
  <pubDate>Mon, 28 Oct 2024 14:00:00 GMT</pubDate>
  <media:content url="https://rachel.fast.ai/posts/2024-10-29-screen-time/music-lab.jpg" medium="image" type="image/jpeg"/>
</item>
<item>
  <title>Gaps and Risks of AI in the Life Sciences</title>
  <dc:creator>Rachel Thomas</dc:creator>
  <link>https://rachel.fast.ai/posts/2024-09-10-gaps-risks-science/</link>
  <description><![CDATA[ 





<p>AI is being used to tackle high-impact scientific problems, including <a href="https://rachel.fast.ai/posts/2024-04-03-ai-antibiotics/">searching for new antibiotics</a> to combat antibiotic-resistance and predicting <a href="https://rachel.fast.ai/posts/2024-08-01-ai-t-cells/">what foreign substances the immune system can recognize</a> as invaders. Amidst the excitement, it is important to keep in mind the gaps and risks of how AI is applied in science. Recognizing these risks opens up new opportunities and can prevent us from being caught off guard by shortcomings.</p>
<section id="missing-data" class="level2">
<h2 class="anchored" data-anchor-id="missing-data">Missing Data</h2>
<p>A recurring problem within data science is that often the data we would be most interested in does not exist or is too difficult to gather. In response, data scientists often instead use imperfect proxies (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9122957/">Thomas and Uminsky 2022</a>), or alter their research questions to make use of the data they have (<a href="https://arxiv.org/abs/1901.02547">Passi and Barocas 2019)</a>). While it is understandable for teams to work with what data is available, this can unduly influence the field to neglect areas where data is unavailable or more difficult to collect.</p>
<p>A research project titled “The library of missing data” (<a href="https://github.com/MimiOnuoha/missing-datasets">Mimi Onouha 2016</a>) inspected examples of datasets that are not collected and of information that we don’t have. What is not measured or recorded can be very revealing, yet it is harder to see an absence. It would be a valuable project to begin with the most important and high-impact open questions in various scientific fields, and then to determine which types of data would need to be collected, what experiments need to be run, and whether AI could prove useful.</p>
<p>While it is understandable that physical and financial constraints shape what types of data can be gathered, it is important to stay aware of how these influence research questions and findings. In some cases, a research area may be ignored due to lack of data. In other cases, results may be biased due to only cetain types of data being included. Missing data can be a type of representation bias, in which not all groups are adequately represented (<a href="https://dl.acm.org/doi/fullHtml/10.1145/3465416.3483305">Suresh and Guttag 2021</a>). For example, databases of <a href="https://rachel.fast.ai/posts/2024-07-09-t-cells/">T cell receptors</a> are highly biased towards particular genetic alleles and certain viral antigens. This data can still help us reach valuable conclusions, but it would be a mistake to overgeneralize those findings to alleles or antigen types that are not well-represented in the underlying data.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-09-10-gaps-risks-science/suresh-fig1.jpg" class="img-fluid figure-img" style="width:75.0%"></p>
<figcaption>Sources of bias at different points throughout the process of creating machine learning models, Figure 1 from Suresh and Guttag, 2021</figcaption>
</figure>
</div>
</section>
<section id="the-physical-world-is-complex-and-volatile" class="level2">
<h2 class="anchored" data-anchor-id="the-physical-world-is-complex-and-volatile">The physical world is complex and volatile</h2>
<p>In other fields, machine learning models have sometimes performed well in testing but failed to work as expected when deployed in the real world (<a href="https://research.google/pubs/everyone-wants-to-do-the-model-work-not-the-data-work-data-cascades-in-high-stakes-ai/">Sambasivan et al.&nbsp;2021</a>). Issues can arise when the context and quality of the training data is not fully understood (<a href="https://arxiv.org/abs/1803.09010">Gebru et al.&nbsp;2021</a>). Biases and errors in datasets will propagate to downstream research. For example, The ImageNet Large Scale Visual Recognition Challenge (ILSVRC) dataset, often called ImageNet-1k, is widely used in image recognition research and has been cited 32,000 times. Over a quarter of the images in the dataset are of wild animals, yet researchers later found that over 12% of these wild animals are labeled incorrectly (<a href="https://ojs.aaai.org/index.php/AAAI/article/view/26682">Luccioni and Rolnick 2023</a>).</p>
<p>ML systems are trained in clearly defined environments, while the physical world often has complex and volatile underlying phenomena, which are not fully accounted for during training. This can create a brittleness to ML-generated solutions (<a href="https://research.google/pubs/everyone-wants-to-do-the-model-work-not-the-data-work-data-cascades-in-high-stakes-ai/">Sambasivan et al.&nbsp;2021</a>).</p>
</section>
<section id="neglecting-mechanistic-understanding" class="level2">
<h2 class="anchored" data-anchor-id="neglecting-mechanistic-understanding">Neglecting Mechanistic Understanding</h2>
<p>Language models are great at interpolating between underlying information, but not at creating things that are completely new (<a href="https://arxiv.org/abs/2305.18654">Dziri et al.&nbsp;2023</a>). Brand new theories will still require human ingenuity. Upon receiving the Nobel Prize, biologist Sydney Brenner, warned that increased data does not necessarily lead to deeper understanding, saying “We’re drowning in an ocean, or a sea of data but we are starved of knowledge.” (<a href="https://www.nobelprize.org/prizes/medicine/2002/brenner/interview/">Sylwan 2002</a>). Geneticist Paul Nurse, another Nobel Prize-winner, echoed these concerns. He wrote about the need for a renewed emphasis on new ideas and theory, alongside data. Nurse proposed that computer scientists need to be embedded alongside biologists to deeply understand the domain problems, “It is through deep familiarity with the biology — not simply a drive to collect more and more data — that important questions will be asked.” (<a href="https://www.nature.com/articles/d41586-021-02480-z">Nurse 2021</a>).</p>
<p>AI is a powerful tool, but it does not replace other ways of knowing about the world, including qualitative research and benchtop quantitative experiments. Machine learning and AI techniques are great at learning patterns in existing data.</p>
</section>
<section id="difficulty-of-cross-disciplinary-work" class="level2">
<h2 class="anchored" data-anchor-id="difficulty-of-cross-disciplinary-work">Difficulty of Cross-Disciplinary Work</h2>
<p>Interdisciplinary work is challenging, as it can be tough to remain on the cutting edge of two separate fields. At one extreme, you can have research based on sound immunology, but using deep learning techniques which are outdated or perhaps were never considered best practice. At the other extreme, you can have research using state-of-the-art deep learning, yet misunderstanding the core mechanisms of immunology or misconstruing the underlying context of datasets.</p>
<p>Meaningful collaboration between those with in-depth immunology knowledge and those with in-depth deep learning knowledge is necessary. Cross-disciplinary work is inherently difficult. Challenges can include differences of communication and work practices between practitioners with different backgrounds; insufficient expertise in one of the fields; and the difficulty of reconciling different approaches and priorities. Moreover, many academic institutions, journals, and other research centers often tend to value specific types of expertise over others. Deep learning and areas of the life sciences such as immunology are each highly complex, jargon-laden, and fast-moving fields. Finding researchers, or even teams, which can keep pace with both is challenging.</p>
<p>Deep learning practitioners often seek to automate systems without understanding what the greatest needs of domain experts are. When this automation leaves gaps, humans are often expected to fill these gaps in ways that are difficult, tedious, or impractical. In contrast, a more fruitful approach may be to take the strengths of both domain experts and deep learning into account from the start to leverage human strengths (<a href="https://slideslive.com/38917533/lessons-learned-from-helping-200000-nonml-experts-use-ml">Thomas 2019</a>).</p>
</section>
<section id="the-double-edged-sword-of-competitions" class="level2">
<h2 class="anchored" data-anchor-id="the-double-edged-sword-of-competitions">The Double-Edged Sword of Competitions</h2>
<p>Carefully structured competitions can play a valuable role in motivating research. Well-defined benchmarks make it clear how various approaches compare to one another. The winner of the 2012 ImageNet competition to identify objects in photographs, an algorithm called AlexNet, led to the mainstream popularization of neural networks and set off an avalanche of increased research into AI. The winner of the 2020 CASP challenge to take amino acid sequences as input and output the most likely 3D protein structure was a computer program called AlphaFold, which has revolutionized protein structure prediction. While these have been huge breakthroughs for their respective fields, over-focusing on competitions can lead to other research areas being neglected. In other cases, the models created for benchmark tasks may be inappropriately applied to other tasks for which they are less suited (<a href="https://www.sciencedirect.com/science/article/pii/S2666389921001847">Paullada et al.&nbsp;2021</a>).</p>
<p>It takes effort to construct well-structured competitions. Data must be gathered and selected, often through benchtop laboratory experiments, while being cognizant of biases and overall representation within the dataset. It is also crucial to have appropriate splits between the training set and a held-out test set. When data leakage occurs, clues from the test set can be found in the training set, artificially inflating the performance of models (<a href="https://www.spiedigitallibrary.org/conference-proceedings-of-spie/11314/1131416/Hazards-of-data-leakage-in-machine-learning--a-study/10.1117/12.2549313.short#_=_">Samala et al.&nbsp;2020</a>).</p>
</section>
<section id="undervaluing-data-work" class="level2">
<h2 class="anchored" data-anchor-id="undervaluing-data-work">Undervaluing data work</h2>
<p>Too often the arduous task of curating quality datasets has been neglected, undervalued, and not given the same currency for career advancement as algorithmic work (<a href="https://dl.acm.org/doi/10.1145/3351095.3372829">Jo and Gebru 2020</a>). A research paper which interviewed over 60 machine learning practitioners was titled “Everyone wants to do the model work, not the data work”, drawing on a representative quote from one of the interviewees. The researchers found that data work was systematically devalued, leading to issues which impacted model performance, sometimes with disastrous consequences (<a href="https://research.google/pubs/everyone-wants-to-do-the-model-work-not-the-data-work-data-cascades-in-high-stakes-ai/">Sambasivan et al.&nbsp;2021</a>)).</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-09-10-gaps-risks-science/sambasivan-square.jpg" class="img-fluid figure-img" style="width:75.0%"></p>
<figcaption>Physical world brittleness, inadequate domain expertise, and other factors can have compounding negative impacts on the process of building AI models. Figure 1 from Sambasivan, et al, 2021</figcaption>
</figure>
</div>
<p>Many areas within science lack consistent benchmark datasets. This can result in models that are not consistently reproducible and makes it tough to compare performance across different papers. Organizers of the largest academic AI conference, Neural Information Processing Systems (NeurIPS), recognized that there were too few incentives for researchers to work on datasets and benchmarks. In 2021, they opened a new track of the conference, on equal standing with the original track, to better support and incentivize research on datasets and benchmarks (<a href="https://blog.neurips.cc/2021/04/07/announcing-the-neurips-2021-datasets-and-benchmarks-track/">Vanschoren and Yeung 2021</a>). This is an encouraging model, and I hope to see more efforts to prioritize data curation.</p>
</section>
<section id="opportunities-for-ai-in-science" class="level2">
<h2 class="anchored" data-anchor-id="opportunities-for-ai-in-science">Opportunities for AI in Science</h2>
<p>The flip side of these risks and gaps is that they offer opportunities for how we can work to improve the application of AI to science and to aim to avoid potential pitfalls. To ensure that we are making the most of AI possibilities in the life sciences, it would be a good idea to:</p>
<ul>
<li>Embed scientists and AI practitioners on closely integrated teams</li>
<li>Survey the biggest open problems in immunology and determine whether AI could help, and if so, what data is needed (as opposed to always just starting with the data we already have)</li>
<li>Reward dataset collection, through conference tracks and journal issues focused on it, similar to the model adopted by NeurIPS</li>
<li>Invest in research that is:
<ul>
<li>Focused on underlying causal mechanisms</li>
<li>Values multiple ways of generating knowledge</li>
<li>Gathering new types of data</li>
</ul></li>
<li>Develop benchmark datasets and structured competitions where relevant</li>
</ul>
<p>By keeping a broader perspective on different types of expertise, different methods of research, and the need for interdisciplinary work, we can prevent a myopic application of AI to science that risks neglecting important questions.</p>
<p><em>Thank you to Jeremy Howard for feedback on earlier drafts of this post.</em></p>
<p>You can subscribe to be notified of new blog posts by submitting your email below:</p>
<script type="text/javascript" src="https://campaigns.zoho.com.au/js/zc.iframe.js"></script>
<iframe scrolling="no" frameborder="0" id="iframewin" width="100%" height="100%" src="https://zcmp-pd.maillist-manage.com.au/ua/Optin?od=11d0c075b7ed49&amp;zx=11a17553b1&amp;tD=156971d471c7b09&amp;sD=156971d471c7cb3">
</iframe>


</section>

<p><br><br><i>I look forward to reading your responses. Create a free GitHub account to comment below.</i></p> ]]></description>
  <category>ethics</category>
  <category>machine learning</category>
  <category>science</category>
  <guid>https://rachel.fast.ai/posts/2024-09-10-gaps-risks-science/</guid>
  <pubDate>Mon, 09 Sep 2024 14:00:00 GMT</pubDate>
  <media:content url="https://rachel.fast.ai/posts/2024-09-10-gaps-risks-science/liquid-drops.jpg" medium="image" type="image/jpeg"/>
</item>
<item>
  <title>Your Immune System is Not a Muscle</title>
  <dc:creator>Rachel Thomas</dc:creator>
  <link>https://rachel.fast.ai/posts/2024-08-13-crowds-vs-friends/</link>
  <description><![CDATA[ 





<section id="that-which-doesnt-kill-you" class="level2">
<h2 class="anchored" data-anchor-id="that-which-doesnt-kill-you">That which doesn’t kill you…</h2>
<p>Some people compare the immune system to a muscle, suggesting that the more you use it, the stronger it gets. Who doesn’t love an analogy? But is this one accurate?</p>
<p>We can see that not all obstacles make you stronger. Destroy the cartilage in your knee, and it may never fully recover, since cartilage doesn’t grow back. Some bacterial infections can permanently scar the lining of the brain and leave <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5570486/">survivors with lower IQs</a>. More and more evidence is linking viruses to a range of diseases including <a href="https://theconversation.com/link-between-epstein-barr-virus-and-multiple-sclerosis-is-a-crucial-discovery-for-people-living-with-ms-175908">multiple sclerosis</a>, <a href="https://theconversation.com/my-work-investigating-the-links-between-viruses-and-alzheimers-disease-was-dismissed-for-years-but-now-the-evidence-is-building-184201">Alzheimer’s disease</a>, <a href="https://pubmed.ncbi.nlm.nih.gov/36669485/">dementia</a>, <a href="https://www.nature.com/articles/s41591-019-0667-0">type 1 diabetes</a>, and certain types of <a href="https://www.mdanderson.org/publications/focused-on-health/7-viruses-that-cause-cancer.h17-1592202.html">cancer</a>. When does illness make you stronger, and when does it cause permanent harm or leave you with chronic health conditions?</p>
<p>Our immune systems are amazing, and amazingly complex. Certain cells, called memory B cells and memory T cells, are able to “remember” invaders that they have seen before. They can rearrange their genes to create billions of possible memories and to respond more quickly to a future infection with the same pathogen. Is this an example of infection making you stronger? It depends. You can only catch the measles virus once, because you will form memory cells for it. But <a href="https://asm.org/articles/2019/may/measles-and-immune-amnesia">measles also destroys your pre-existing memory cells</a>, meaning that you can now re-catch a bunch of other illnesses that you had already built immunity to. Also, a very small percentage of people (unvaccinated babies may have more risk) will seem to fully recover from measles, and then 6 to 15 years later develop <a href="https://www.ncbi.nlm.nih.gov/books/NBK560673/">brain inflammation that often leads to death</a>.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-08-13-crowds-vs-friends/virus-headlines.jpg" class="img-fluid figure-img" style="width:75.0%"></p>
<figcaption>Viruses have been linked to multiple sclerosis, Alzheimer’s, strokes, and other diseases</figcaption>
</figure>
</div>
<p>Chickenpox is another disease that people typically only have once. However, the virus does not fully go away, but rather lies latent in the nervous system and can reactivate as shingles decades later. When the virus reactivates as shingles, it also leads to the <a href="https://theconversation.com/chickenpox-and-shingles-virus-lying-dormant-in-your-neurons-can-reactivate-and-increase-your-risk-of-stroke-new-research-identified-a-potential-culprit-194627">formation of blood clots</a> that raise the risk of stroke for months afterwards. So while you have “immunity” against getting chickenpox again, it has come with <a href="https://rachel.fast.ai/posts/2023-03-07-viruses1/">long-term risk</a>. While chickenpox and measles are examples of viruses, bacteria can also lie latent for a long time after an infection. For instance, the bacteria types that cause epidemic typhus, brucellosis, and tuberculosis can all activate/reactivate long after initial infection.</p>
<p>Another informative example is Dengue virus, a <a href="https://theconversation.com/dengue-why-is-this-sometimes-fatal-disease-increasing-around-the-world-215008">potentially fatal mosquito-borne disease</a> that affects millions of people annually. It has 4 different, but related types. This is a problem, because the memory your immune system forms to a given type will actually harm you if you get <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10062565/">infected with a different type later</a>. As a result, a person’s second dengue infection is more severe than their first. Having no memory of dengue is better than remembering the wrong version!</p>
<p>There is a lot of confusion on germs and illness. It is good to play in the dirt and to be exposed to microbes, but you should also wash your hands after using the toilet and avoid raw sewage. When is hygiene good and when is it bad? Recently, a recurring question in newspaper articles and parents groups is whether it was harmful to children’s immune systems that they stayed home in 2020 and caught fewer illnesses. Is there a correct amount or type of “training” that the immune system needs? To explore these questions, we first need to inspect a widely misunderstood idea, often referred to as the <em>Hygiene Hypothesis</em>.</p>
</section>
<section id="the-hygiene-hypothesis" class="level2">
<h2 class="anchored" data-anchor-id="the-hygiene-hypothesis">The Hygiene Hypothesis</h2>
<p><em>Allergies</em> are a misfiring of the immune system– when it attacks what should be harmless environmental substances, such as pollen or dust. <em>Autoimmunity</em> is a different type of misfiring of the immune system– when it attacks our own cells, whether those are your neurons (multiple sclerosis), your joints (arthritis), your thyroid (Hashimoto’s disease), or your insulin-producing cells (type 1 diabetes). Both allergies and <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9918670/">autoimmune diseases</a> have risen dramatically in recent decades.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-08-13-crowds-vs-friends/miller-autoimmune.jpg" class="img-fluid figure-img" style="width:65.0%"></p>
<figcaption>A drastic rise in autoimmune disease: All the diseases listed on this graph are autoimmune diseases (Miller, 2022).</figcaption>
</figure>
</div>
<p>First proposed in 1989, the Hygiene Hypothesis offers an explanation for this dramatic rise. Certain types of microbes can help <em>modulate</em> our immune systems. By creating a low level of immune activity, they prevent our immune systems from getting bored and confused, which could lead them to attack the wrong target. <strong>Four main categories of pathogens that humans deal with are viruses, bacteria, fungi, and parasites.</strong> The evidence for pathogens that may be beneficial to the immune system is almost entirely for <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4396177/">parasitic worms and friendly (commensal) bacteria</a>. In contrast, many viruses can even trigger the onset of autoimmune diseases or allergies.</p>
<p>The reason for this makes sense: humans co-evolved with parasites and commensal bacteria, going back to when people lived in small hunter-gatherer tribes. For <a href="https://www.google.com.au/books/edition/The_Cambridge_Encyclopedia_of_Hunters_an/5eEASHGLg3MC?hl=en&amp;gbpv=1&amp;dq=The+Cambridge+Encyclopedia+of+Hunters+%26+Gatherers&amp;printsec=frontcover">over 90% of human history</a>, we lived as hunter-gatherers, and were exposed to very different types of microbes than we are now in crowded cities or poorly ventilated office buildings. Researchers have pointed out that the moniker Hygiene Hypothesis is misleading, and have proposed a more accurate alternative. <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4089149/">The “Old Friends” mechanism</a> describes which microbes we co-evolved with for 300,000 years.</p>
</section>
<section id="our-old-friends" class="level2">
<h2 class="anchored" data-anchor-id="our-old-friends">Our old friends</h2>
<p>Our bodies are full of <a href="https://rachel.fast.ai/posts/2024-04-25-microbiome-1/">peaceful (commensal) bacteria</a>, which can help synthesize vitamins we need, regulate dopamine, and even protect us from infection. Most famously our guts, but also the skin, throat, and bladder, all have distinct microbiomes. Disruption due to antibiotic use, Western diets, and cesarean births has changed our microbiomes, contributing to the rise of autoimmunity and allergies. Not all bacteria are the same– the positive microbes transmitted in a vaginal birth are quite different from harmful bacteria like anthrax or tuberculosis.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-08-13-crowds-vs-friends/parasites-square.jpg" class="img-fluid figure-img" style="width:50.0%"></p>
<figcaption>Our “old friends” parasitic worms may seem gross, but crowd infections can be even grosser</figcaption>
</figure>
</div>
<p>Surprisingly, there are a number of studies showing that <a href="https://elifesciences.org/articles/65180">infections with helminths (parasitic worms)</a> can be <strong>beneficial</strong> in treating allergies and autoimmune disease, including for <a href="https://sci-hub.se/https://link.springer.com/chapter/10.1007/7854_2014_361">multiple sclerosis</a>. A proposed mechanism is that parasites provide a low-level of background activation for the immune system, which prevents excess activation. This in turn can prevent inflammatory and autoimmune disorders. However, helminth infections can also be harmful and are not something you should try at home! Researchers are developing therapies based on <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4374592/">proteins derived from helminths</a>.</p>
</section>
<section id="crowd-infections-are-new-in-human-history" class="level2">
<h2 class="anchored" data-anchor-id="crowd-infections-are-new-in-human-history">“Crowd infections” are new in human history</h2>
<p>Our “old friends” are organisms we co-evolved with for &gt; 50,000 years. By evolving together, our bodies learned to take advantage of these organisms. <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4089149/">Old friends can be contrasted with “crowd infections”</a> which only developed much more recently, since 10,000 BCE, after people began to live in more densely populated townships.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-08-13-crowds-vs-friends/busiest-flights-2.jpg" class="img-fluid figure-img" style="width:75.0%"></p>
<figcaption>Our immune systems did not evolve for this!</figcaption>
</figure>
</div>
<p>A crowd infection is one that would not last in sparsely distributed, small hunter-gatherer groups, because it either kills people, or offers enough immunity until the small tribe has been infected and there is nowhere else to spread. For instance, if a small tribe of hunter-gatherers caught a strain of influenza, the virus would die out once they had all caught it in short order, with no more hosts to spread to. The ability to circulate through megacities or crisscross the globe through international travel in our current world offers influenza far more opportunities to infect and time to mutate. It is able to continue infecting and reinfecting, decade after decade. Even viruses that don’t mutate much, such as measles, are able to continue infecting, due to dense populations and an ongoing supply of new people being born.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-08-13-crowds-vs-friends/old-friends.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>Not all infections are the same! Figure from Rook, et al, 2014</figcaption>
</figure>
</div>
<p>Homo Sapiens first evolved some 300,000 years ago, yet crowd infections are believed to have only developed in the last 12,000 years, a small blip in human history. Humans living in dense cities is a relatively recent development. An even more recent development is that of sealed indoor spaces and frequent international air travel. Many crowd infections, such as measles, mumps, chickenpox, colds, and flu, are airborne, spreading when humans talk and breathe in close contact, with poor ventilation. These infections could not widely spread until the last few hundred years of human history.</p>
<p>When I began studying immunology, something that surprised me is how much of the immune system is focused on fighting parasites. There is an entire branch, including several cell types, devoted to this. It seems like such a mismatch to the modern, industrialized world. <em>“Can I have a few more immune cell types focused on viruses or intracellular bacteria?”</em> I thought, <em>“in exchange for some of these parasite-focused cells that I’m not using??”</em> Our “old friends” are quite different from the crowd infections that plague us now– it would be bizarre to assume that research based on one of these categories will apply to the other!</p>
</section>
<section id="old-friends-hypothesis-is-more-descriptive-than-the-hygiene-hypothesis" class="level2">
<h2 class="anchored" data-anchor-id="old-friends-hypothesis-is-more-descriptive-than-the-hygiene-hypothesis">Old Friends Hypothesis is more descriptive than the Hygiene Hypothesis</h2>
<p>Our “old friends” parasitic worms and beneficial microbes are associated with a reduced risk of allergies and autoimmune diseases. No such relationship exists for crowd diseases. In fact, the opposite is true. Crowd diseases contribute to <a href="https://erj.ersjournals.com/content/19/2/341">allergies</a> and <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10051805/">autoimmune diseases</a>. Comparing the immune system to a muscle that gets stronger with use is overly simplistic and, in many cases, inaccurate. There is huge variety in how various pathogens impact us. Being precise in considering different types of microbes and infections will allow us to better understand human health.</p>
</section>
<section id="related-reading" class="level2">
<h2 class="anchored" data-anchor-id="related-reading">Related reading:</h2>
<ul>
<li><a href="https://elifesciences.org/articles/65180">Gross ways to live long: Parasitic worms as an anti-inflammaging therapy?</a></li>
<li><a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4089149/">Microbial ‘old friends’, immunoregulation and socioeconomic status</a></li>
<li><a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4396177/">Unraveling the Hygiene Hypothesis of helminthes and autoimmunity</a></li>
<li><a href="https://rachel.fast.ai/posts/2024-04-25-microbiome-1/">Surprises of the Microbiome</a></li>
<li><a href="https://rachel.fast.ai/posts/2023-03-22-viruses2/">Viruses: The Silent Triggers of Autoimmune and Neurodegenerative Diseases</a></li>
</ul>


</section>

<p><br><br><i>I look forward to reading your responses. Create a free GitHub account to comment below.</i></p> ]]></description>
  <category>science</category>
  <guid>https://rachel.fast.ai/posts/2024-08-13-crowds-vs-friends/</guid>
  <pubDate>Mon, 12 Aug 2024 14:00:00 GMT</pubDate>
  <media:content url="https://rachel.fast.ai/posts/2024-08-13-crowds-vs-friends/flights.jpg" medium="image" type="image/jpeg"/>
</item>
<item>
  <title>AI’s Quest to Predict T Cell Binding– The Holy Grail of Immunology</title>
  <dc:creator>Rachel Thomas</dc:creator>
  <link>https://rachel.fast.ai/posts/2024-08-01-ai-t-cells/</link>
  <description><![CDATA[ 





<p><em>This post is part 3 in a series on using AI to predict T cell binding. Here are <a href="https://rachel.fast.ai/posts/2024-07-09-t-cells/">part 1</a> and <a href="https://rachel.fast.ai/posts/2024-07-24-neural-nets/">part 2</a>.</em></p>
<p>Better understanding what T cells bind to is a crucial question; some researchers even call it <a href="https://www.nature.com/articles/s41577-023-00835-3">“a holy grail of systems immunology”</a>. While AI is already being used towards solutions, there is still much work to be done in this area!</p>
<p>In <a href="https://rachel.fast.ai/posts/2024-07-09-t-cells/">part 1 of this series</a>, we discussed the importance of determining what a T cell, that multi-purpose hero of the adaptive immune system, can bind to. New research papers on this topic are published each month. Almost all of them use neural networks, and in <a href="https://rachel.fast.ai/posts/2024-07-24-neural-nets/">part 2</a>, we discussed different types of neural networks. Now we will see how different types of neural networks are being applied to T cell binding.</p>
<section id="when-t-cell-binding-makes-a-big-difference" class="level2">
<h2 class="anchored" data-anchor-id="when-t-cell-binding-makes-a-big-difference">When T Cell Binding Makes a Big Difference</h2>
<p>T cells play a role in <a href="https://www.youtube.com/watch?v=6BYocBxOfsU&amp;list=PLtmWHNX-gukLirebdPH8lla41SS78kjLD&amp;index=1&amp;t=172s">autoimmune diseases like Type 1 diabetes</a> (where T cells mistakenly destroy insulin-producing cells in the pancreas) and psoriasis (where T cells attack our own skin). T cells are also crucial to a healthy response to cancer. In cancer, something has gone amiss: cells undergo uncontrolled growth. T cells should put a stop to this, but in many cases T cells have been misleadingly soothed into exhaustion.</p>
<p>An increasing number of cancer therapies are based on T cells, and predicting which T cells are the most promising cancer-fighting candidates is crucial. When all is in harmony, our T cells should keep cancer from starting in the first place by quickly identifying and killing any cells which have gone rogue. T cells do this by recognizing small pieces of protein, called peptides, held out to them on special molecules called MHC. If those proteins are the sign of a good cell gone bad, the T cell can destroy the miscreant cell.</p>
<p>We can help the immune system by hanging up WANTED posters with the faces of outlaw proteins on them. This can let the immune system know who to look out for.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-08-01-ai-t-cells/wanted.jpg" class="img-fluid figure-img" style="width:65.0%"></p>
<figcaption>WANTED poster for Cancer/testis antigen family 45 member A1</figcaption>
</figure>
</div>
<p>In practice, this is done by pairing a cancer peptide (called a “neoantigen”) together with a danger signal, and injecting them into a patient in order to sound the alarm and spur T cells into action. We need to identify which rogue cancer proteins are the most promising candidates for such a therapy. Neural networks are being used for this task!</p>
<p>One key question is <a href="https://pubmed.ncbi.nlm.nih.gov/29403468/">which peptides are the best</a> candidates to use in these therapies. A key criteria is that T cells need to be able to recognize the peptide, which means that MHC molecules need to be able to bind to it. Everyone has different T cells and MHC molecules, since these are created by the most diverse genes in the human genome! The answer must be <a href="https://www.nature.com/articles/s42256-020-00260-4">personalized</a>.</p>
</section>
<section id="applying-neural-networks" class="level2">
<h2 class="anchored" data-anchor-id="applying-neural-networks">Applying Neural Networks</h2>
<p>This seems like a question where neural networks could be helpful. In <a href="https://rachel.fast.ai/posts/2024-07-24-neural-nets/">part 2</a>, we discussed several different architectures (sequences of math operations) that we could choose between. One consideration is whether the input to a problem is fixed length or variable length. Are we identifying pictures of dogs, or reading Anna Karenina? The problem of T cell binding can be (and is!) framed both ways, depending on the team and the approach they choose to take. In general, sequences of amino acids can have a huge variety of lengths. However, T cell receptors and the peptides held out by MHC have lengths within a narrow range. As a result, we could use neural net architectures that have developed for fixed length problems. Or we could take advantage of the engineering and large training sets that already went into creating neural networks for variable length proteins. Strong research is being published of both types, and I am curious to see if there will be a consensus 5 years from now on the best approach.</p>
</section>
<section id="to-use-alphafold-or-not" class="level2">
<h2 class="anchored" data-anchor-id="to-use-alphafold-or-not">To Use AlphaFold or Not?</h2>
<p>As discussed in <a href="https://rachel.fast.ai/posts/2024-07-09-t-cells/">part 1</a>, the computer program <a href="https://www.youtube.com/watch?v=pB0RxG1NdtA&amp;list=PLtmWHNX-gukLirebdPH8lla41SS78kjLD&amp;index=2">AlphaFold</a> made headlines in 2020 for its incredible accuracy in taking a 1D sequence of amino acids as input and producing a 3D protein structure as output. It won an international competition for this task by a huge margin, smashing all previous records. Since T cell receptors (TCRs), MHC molecules, and peptides are all proteins, it makes sense that we could modify it for the T cell problem.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-08-01-ai-t-cells/insulin-2ways.jpg" class="img-fluid figure-img" style="width:80.0%"></p>
<figcaption>Two different ways of representing the protein insulin: a string of letters, or its 3D structure</figcaption>
</figure>
</div>
<p>AlphaFold (and similar models, such as RosettaFold and ESMFold) are excellent at what they do, which is to take a sequence of amino acids and predict the 3D structure of a protein. However, it is possible that these models are simultaneously both overkill and underkill for our T cell question. AlphaFold can handle much longer amino acid sequences than are found in the T cell binding problem, yet it doesn’t account for T cells having not just one, but a <a href="https://elifesciences.org/articles/82813">pair of TWO sequences</a> of amino acids (known as the alpha and beta chains). AlphaFold was trained on a massive data set, yet that dataset did not include much in the way of T cell or MHC specific data. Some researchers have chosen to modify AlphaFold, adding additional layers for prediction, or fine-tuning on a more relevant dataset, whereas other teams have chosen to design neural networks that are more directly suited to the task at hand.</p>
<p>Which approach is better? It’s tough to tell. Reading papers within the field can be confusing, as many boast of great results. Since there is an absence of consistent benchmarks and models are often not evaluated for accuracy on new peptides, it can be tough to compare different models.</p>
</section>
<section id="the-need-for-a-grand-challenge" class="level2">
<h2 class="anchored" data-anchor-id="the-need-for-a-grand-challenge">The Need for a Grand Challenge</h2>
<p>You may wonder why we can’t clearly see which neural network for T cell binding is definitively the best. The ImageNet competition produced a clear winner for recognizing objects in images in 2012: AlexNet, a neural net that was head and shoulders better than the other entries. Similarly, AlphaFold dominated the CASP competition in 2020. However, several factors are necessary to create competitions such as ImageNet or CASP, and these factors do not yet exist for the T cell problem.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-08-01-ai-t-cells/lasker-statues2.jpg" class="img-fluid figure-img" style="width:60.0%"></p>
<figcaption>AlphaFold won the 2023 Lasker Award, “America’s Nobel Prize”. The trophies are based on The Winged Victory of Samothrace statue.</figcaption>
</figure>
</div>
<p>You need a large, well-curated, and representative dataset. This can fail in numerous ways. While data is central to any machine learning project, all too often, data work has been undervalued and neglected. The paper <a href="https://research.google/pubs/everyone-wants-to-do-the-model-work-not-the-data-work-data-cascades-in-high-stakes-ai/">“Everyone wants to do the model work, not the data work”</a> drew on interviews with 53 machine learning practitioners across multiple countries. Building fancy math models is rewarded with career advancement, not the arduous labor of collecting and curating the underlying datasets that make those models possible. In many cases, those tasked with collecting data had it piled on top of already demanding jobs, with inadequate communication, training, or compensation. This led to erroneous measurements, miscommunications, and failed projects. To give one example, when machines in a robotic medical AI project were recalibrated, the change was not documented, leading to inconsistent data that was impossible to interpret.</p>
<p>Data work is particularly challenging in the area of T cell receptors (TCRs). Current datasets cover just a <a href="https://elifesciences.org/articles/82813">tiny fraction</a> of all possible TCRs. It is difficult to scale experimental methods, so less than 1 million unique TCR-peptide pairs have been found experimentally. Of the data that exists, the vast majority just contains one of the two amino acid sequences which make up a TCR, even though both sequences are significant.</p>
<p>You also need everyone to agree on what numerical metrics to use in evaluating success, and for those metrics to closely correspond to the value you care about. Great harm can result <a href="https://rachel.fast.ai/posts/2019-09-24-metrics/">when those metrics are a poor proxy</a> for the underlying goal. Providing a meaningful numeric score and declaring a winner require the right setup and a lot of effort behind the scenes.</p>
<p>Some researchers have called for a <a href="https://www.nature.com/articles/s41577-023-00835-3">“grand challenge”</a> of inferring T cell receptor binding, similar to the grand challenge in prediction protein folding which led to AlphaFold. Standardizing datasets and setting clear metrics for comparisons could motivate innovation. This is a high-impact area within immunology and medicine and I hope that it is a continued focus of research!</p>
<p><em>This post is part 3 in a series on using AI to predict T cell binding. Here are <a href="https://rachel.fast.ai/posts/2024-07-09-t-cells/">part 1</a> and <a href="https://rachel.fast.ai/posts/2024-07-24-neural-nets/">part 2</a>.</em></p>
<p><em>Thank you to Jeremy Howard for feedback on an earlier draft of this post.</em></p>
<p>You can subscribe to be notified of new blog posts by submitting your email below:</p>
<script type="text/javascript" src="https://campaigns.zoho.com.au/js/zc.iframe.js"></script>
<iframe scrolling="no" frameborder="0" id="iframewin" width="100%" height="100%" src="https://zcmp-pd.maillist-manage.com.au/ua/Optin?od=11d0c075b7ed49&amp;zx=11a17553b1&amp;tD=156971d471c7b09&amp;sD=156971d471c7cb3">
</iframe>


</section>

<p><br><br><i>I look forward to reading your responses. Create a free GitHub account to comment below.</i></p> ]]></description>
  <category>machine learning</category>
  <category>science</category>
  <guid>https://rachel.fast.ai/posts/2024-08-01-ai-t-cells/</guid>
  <pubDate>Wed, 31 Jul 2024 14:00:00 GMT</pubDate>
  <media:content url="https://rachel.fast.ai/posts/2024-08-01-ai-t-cells/wanted-rectangle2.jpg" medium="image" type="image/jpeg"/>
</item>
<item>
  <title>The Most Common and Useful Neural Nets</title>
  <dc:creator>Rachel Thomas</dc:creator>
  <link>https://rachel.fast.ai/posts/2024-07-24-neural-nets/</link>
  <description><![CDATA[ 





<p><em>This post can stand alone as a friendly introduction to neural nets, no background required. It is part 2 in my series on using AI to figure out what T cells can bind to. Here are <a href="https://rachel.fast.ai/posts/2024-07-09-t-cells/">part 1</a> and <a href="https://rachel.fast.ai/posts/2024-08-01-ai-t-cells/">part 3</a>.</em></p>
<section id="what-are-neural-networks" class="level2">
<h2 class="anchored" data-anchor-id="what-are-neural-networks">What are Neural Networks?</h2>
<p>Neural networks are a type of machine learning algorithm that are currently in the limelight, powering chatbots like ChatGPT and image generators like Midjourney. Neural networks are just math equations, written in computer code. When people hear the terms “neural networks” or “artificial intelligence,” they may picture humanoid robots, but it is more accurate to imagine a scaled-up version of 10th grade math class.</p>
<p>When considering an AI system, two major categories of tasks are:</p>
<ul>
<li><strong>Classifying things</strong>: These are not predictions in the sense of predicting the future, but rather, guessing the answer to a question. Is this a picture of a chihuahua or a blueberry muffin? Which of these potential drugs would be most likely to target antibiotic resistant bacteria?</li>
<li><strong>Generating things</strong>: Create a picture of <a href="https://medium.com/hackernoon/non-artistic-style-transfer-or-how-to-draw-kanye-using-captain-picards-face-c4a50256b814">Kanye West made out of tiny Captain Picard faces</a>. Generate a new molecule that may work as an antibiotic.</li>
</ul>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-07-24-neural-nets/chihuahua-kanye.jpg" class="img-fluid figure-img" style="width:65.0%"></p>
<figcaption>Classification: muffin or chihuahua? (pic from Karen Zack, <span class="citation" data-cites="teenybiscuit">@teenybiscuit</span>, 2016) // Generation: Kanye out of Picard faces (pic from fast.ai intern Brad Kenstler, 2017)</figcaption>
</figure>
</div>
<p>Some problems in medicine and immunology can be framed either as a classification problem OR as a generative problem. For instance, suppose your goal is to produce a <a href="https://rachel.fast.ai/posts/2024-04-03-ai-antibiotics/">new antibiotic to address antimicrobial resistance</a>. Some scientists are trying to classify which compounds, out of hundreds of millions of existing ones, may have the desired properties, whereas others are trying to generate new compounds. Both approaches are useful.</p>
</section>
<section id="language-models" class="level2">
<h2 class="anchored" data-anchor-id="language-models">Language Models</h2>
<p>You may have heard the term “language model” describing chatbots (such as Chat-GPT or Claude) and automated language translation (such as Google Translate and Skype Translator). One approach to developing sophisticated language capabilities is to write a computer program that predicts the next word in a sentence. Predict the next word and then the next, and you can generate strings of text. This is a language model. Both prediction and generation are used to build language models.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-07-24-neural-nets/google-translate.jpg" class="img-fluid figure-img" style="width:65.0%"></p>
<figcaption>Google Translate relies on language models. The gender bias shown here has since been corrected, although bias remains an issue.</figcaption>
</figure>
</div>
<p>Along the way, the model has learned a number of patterns regarding context, vocabulary, and relationships. Fine tuning lets us use these capabilities for specific problems. For instance, a <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10338337/">language model trained on medical chart notes</a> was used to make predictions about which patients were most likely to be readmitted to the hospital. While the model was trained on a prediction task (predicting the next word in a sentence), it can now be used for generation as well when combined with some additional techniques when it is given additional instructions and training from humans.</p>
<p>When language models are trained, they are not always just tasked with predicting the next word. In a sentence, sometimes random words are covered up, and the computer must predict what goes in that spot. It is a fill-in-the-blank assignment.</p>
<p>Through predicting words, the model can learn which other words to focus on. Suppose I told you, “I am so tired that I can barely keep my XXXX open”. You could likely guess that the missing word is “eyes”. To do so, you may have focused on “tired” and “open”. Some of the word positions add little, if anything, to your efforts of deduction. In a similar way, attention-based neural networks learn what to pay attention to, where to focus.</p>
<p>Like the sequence of words in a novel, a sequence of amino acids is its own language, telling the story of a protein. The same techniques and approaches that were developed for natural languages (such as translating between English and French) can also be applied to the biological language of protein creation.</p>
</section>
<section id="popular-architectures" class="level2">
<h2 class="anchored" data-anchor-id="popular-architectures">Popular architectures</h2>
<p>Neural networks are often created from a few popular architectures, which are particular combinations of mathematical operations. Something that can be confusing is that different architectures can often be combined with one another. Also, most of these architectures can be used for both predictive and generative problems. Here are a few particularly popular ones:</p>
<ul>
<li><strong>Multi-Layer Perceptrons (MLPs)</strong>: the oldest example of a neural network was developed in the 1960s! It is the basis of many systems used today (<a href="https://hdl.handle.net/2027/mdp.39015039846566?urlappend=%3Bseq=13">Rosenblatt 1962</a>).</li>
<li><strong>Convolutional Neural Networks (CNNs)</strong>: are like a filter that you scan across a picture, inch by inch, to recognize what is in it. This is useful for programs such as those bird-watcher apps that will tell you what type of bird you spotted, based on the photo you upload (<a href="https://ieeexplore.ieee.org/document/726791">Lecun et al.&nbsp;1998</a>; <a href="https://papers.nips.cc/paper_files/paper/2012/hash/c399862d3b9d6b76c8436e924a68c45b-Abstract.html">Krizhevsky et al.&nbsp;2012</a>).</li>
<li><strong>Recursive Neural Networks (RNNs)</strong>: can take variable-length input (that is, sequences) and store state. Google began using these for translating between languages in 2016 (<a href="https://ieeexplore.ieee.org/abstract/document/6795963">Hochreiter and Schmidhuber 1997</a>; <a href="https://arxiv.org/abs/1609.08144">Wu et al.&nbsp;2016</a>).</li>
<li><strong>Transformers</strong>: offer a more efficient approach for sequences, which can be any length. These are based on <strong>attention</strong>, as explained above. These have now replaced RNNs in many applications, including Google Translate, which lets a user enter a word or phrase, or even entire paragraphs, and get back the response in another language (<a href="https://arxiv.org/abs/1706.03762">Vaswani et al.&nbsp;2017</a>).</li>
<li><strong>Graph Neural Network (GNNs)</strong>: Allow for geometric relationships in the input to be captured. For example, GNNs can help represent the spatial relationships between different amino acids in a sequence. Principles of GNNs are typically combined with other options in the list above (<a href="https://ieeexplore.ieee.org/document/4700287">Scarselli et al.&nbsp;2009</a>; <a href="https://pubmed.ncbi.nlm.nih.gov/19193509/">Micheli 2009</a>).</li>
</ul>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-07-24-neural-nets/convolution2.gif" class="img-fluid figure-img" style="width:65.0%"></p>
<figcaption>How a convolution works, from deeplearning.net, 2018</figcaption>
</figure>
</div>
<p>How do you choose a neural network for your problem? One key consideration is how much the size of your input varies. Suppose you have a collection of photos and are trying to determine which are of cats and which are of dogs. Your photos are likely of similar sizes, and you can make them the same by adding a frame around the smaller ones. Whole families of neural networks have been developed for such problems with fixed-size input.</p>
<p>Now suppose you are developing a neural network that can read books and predict which word will come next. The difference in length between The Hungry Hungry Caterpillar and Les Miserables is huge. For this problem, we need a network that allows variable-size input. RNNs and Transformers allow variable-length input, whereas CNNs take fixed-size input.</p>
</section>
<section id="the-innovation-of-alphafold" class="level2">
<h2 class="anchored" data-anchor-id="the-innovation-of-alphafold">The Innovation of AlphaFold</h2>
<p>Recall from <a href="https://rachel.fast.ai/posts/2024-07-09-t-cells/">Part 1 of this series</a> that the computer program AlphaFold revolutionized the task of taking a 1D sequence of amino acids (shown on the left), and transforming it into a 3D molecule (shown on the right).</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-07-24-neural-nets/insulin-2ways.jpg" class="img-fluid figure-img" style="width:80.0%"></p>
<figcaption>Two different ways of representing the protein insulin: a string of letters, or its 3D structure</figcaption>
</figure>
</div>
<p>The text of a book is 1-dimensional. It could be written on a single long ribbon, stretched out, and read as such. However, best making sense of amino acid sequences requires an additional dimension. We want to know not just the long ribbon of a single sequence, but the evolutionary relationships of related sequences it is similar to.</p>
<p>A crucial innovation of <a href="https://www.nature.com/articles/s41586-021-03819-2">AlphaFold</a> is that it expanded attention to 2 dimensions. The 2 dimensions are the sequence of the amino acids, as well as a block of sequences of similar sequences. These similar sequences provide important evolutionary clues of what our protein of interest may be like. Like distant cousins in a family tree, they can give us insights of what may be some common features. This set of related sequences is known as a Multiple Sequence Alignment (MSA) and is a key component in bioinformatics. AlphaFold applies attention to the MSA, learning which parts of related sequences to focus on in trying to decode a protein structure.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-07-24-neural-nets/alphafold-2reps.jpg" class="img-fluid figure-img" style="width:65.0%"></p>
<figcaption>AlphaFold represents sequences 2 ways. The MSA representation requires 2D attention. Image from Jumper, et al, 2021</figcaption>
</figure>
</div>
<p>Returning to our list of popular architectures, <a href="https://www.youtube.com/watch?v=pB0RxG1NdtA&amp;t=2s">AlphaFold</a> combines aspects of Transformers (in a new form, named the Evoformer, which has 2D attention) and Graph Neural Networks (to represent the spatial relationships between amino acids). Now that we have covered some of the types of neural networks and the structure of AlphaFold, we are ready to combine this with the information about T cells from <a href="https://rachel.fast.ai/posts/2024-07-09-t-cells/">part 1</a>. <strong>Check out <a href="https://rachel.fast.ai/posts/2024-08-01-ai-t-cells/">part 3</a> to see how AI is being used to predict T cell binding.</strong></p>
<p>And for understanding the ethical risks of AI systems, you may be interested in reading my previous posts:</p>
<ul>
<li><a href="https://rachel.fast.ai/posts/2024-02-20-ai-medicine/">“AI will cure cancer” misunderstands both AI and medicine</a></li>
<li><a href="https://rachel.fast.ai/posts/2023-05-16-ai-centralizes-power/">AI and Power: The Ethical Challenges of Automation, Centralization, and Scale</a></li>
<li><a href="https://rachel.fast.ai/posts/2021-08-17-eleven-ethics-videos/">11 Short Videos About AI Ethics</a></li>
</ul>
<p><em>Thank you to Jeremy Howard for feedback on earlier drafts of this post.</em></p>
<p>You can subscribe to be notified of new blog posts by submitting your email below:</p>
<script type="text/javascript" src="https://campaigns.zoho.com.au/js/zc.iframe.js"></script>
<iframe scrolling="no" frameborder="0" id="iframewin" width="100%" height="100%" src="https://zcmp-pd.maillist-manage.com.au/ua/Optin?od=11d0c075b7ed49&amp;zx=11a17553b1&amp;tD=156971d471c7b09&amp;sD=156971d471c7cb3">
</iframe>


</section>

<p><br><br><i>I look forward to reading your responses. Create a free GitHub account to comment below.</i></p> ]]></description>
  <category>machine learning</category>
  <category>science</category>
  <guid>https://rachel.fast.ai/posts/2024-07-24-neural-nets/</guid>
  <pubDate>Tue, 23 Jul 2024 14:00:00 GMT</pubDate>
  <media:content url="https://rachel.fast.ai/posts/2024-07-24-neural-nets/chihuahua-kanye.jpg" medium="image" type="image/jpeg"/>
</item>
<item>
  <title>Decoding T cells with AI</title>
  <dc:creator>Rachel Thomas</dc:creator>
  <link>https://rachel.fast.ai/posts/2024-07-09-t-cells/</link>
  <description><![CDATA[ 





<p><em>I’m studying the intersection of AI and immunology, two notoriously jargon-heavy fields – but it’s important that scientists share their work with a broader audience. This 3-part series is an experiment: an accessible introduction to AI immunology concepts. This post is part 1 in a series on using AI to predict T cell binding. Here are <a href="https://rachel.fast.ai/posts/2024-07-24-neural-nets/">part 2</a> and <a href="https://rachel.fast.ai/posts/2024-08-01-ai-t-cells/">part 3</a>.</em></p>
<p>T cells are one of the most important cell types of our immune systems, assassinating cells that have been infected by viruses or turned cancerous, and sending commands to other immune cells to organize responses against invaders. T cells can help mobilize against a wide range of threats due to their incredible diversity. They are able to rearrange their genes to recognize billions of different types of infected or rogue cells.</p>
<p>Figuring out how to predict what a T cell will bind to (meaning what cells it can recognize as bad and coordinate attacks against) would be a vital medical breakthrough. Some cancer therapies are based on teaching T cells to better recognize tumor cells. Knowing the answer could help us design new drugs. And not knowing can have disastrous consequences. A <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3743463/">therapy was trialed</a> in which patients were injected with T cells that would kill cells displaying a protein that is commonly found in cancer, but is not displayed by healthy cells. Tragically, it turns out that these T cells can also bind to a different and unrelated protein found in heart muscles, and two patients died of cardiac arrest.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-07-09-t-cells/tcell-niaid.jpg" class="img-fluid figure-img" style="width:65.0%"></p>
<figcaption>The surface of a T cell, image from a scanning electron micrograph, credit: NIAID, creative commons license</figcaption>
</figure>
</div>
<p>In this 3-part series, I will explore efforts to use AI to predict what T cells will bind to, and why it matters.</p>
<section id="t-cells-our-multi-purpose-immune-heroes" class="level2">
<h2 class="anchored" data-anchor-id="t-cells-our-multi-purpose-immune-heroes">T cells, our multi-purpose immune heroes</h2>
<p>At their best, T cells destroy viral-infected cells and stop cancer before we ever develop symptoms. When things go awry, T cells can accidentally attack our own cells – for instance Type 1 Diabetes occurs when T cells mistakenly destroy insulin-producing cells in the pancreas and psoriasis is caused by T cells attacking the skin (<a href="https://www.youtube.com/watch?v=6BYocBxOfsU">check out this 5 minute video I made for more information about confused T cells</a>). Or “exhausted” T cells can sit idly by as tumors rapidly grow.</p>
<p>T cell receptors (TCRs) are protein complexes on the outside of a T cell. The way a T cell recognizes a problematic cell is by binding to a small piece of protein, called a peptide. T cells can bind only to peptides that are presented to them on fancy platters– err, I mean on special molecules called MHC molecules, found on the outside of cells. The binding of TCR-peptide-MHC is like a secret handshake. The TCR is very specific about which peptides it can bind to. Most combinations will not work.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-07-09-t-cells/t-cell-markers.jpg" class="img-fluid figure-img" style="width:65.0%"></p>
<figcaption>Does this pair know the same secret handshake? Sketch I made with my daughter’s markers</figcaption>
</figure>
</div>
<p>Since TCRs are made of protein, and they bind to small pieces of protein, it is helpful to consider some related questions about proteins before we tackle TCR binding in more detail.</p>
</section>
<section id="proteins-workhorses-of-the-cell" class="level2">
<h2 class="anchored" data-anchor-id="proteins-workhorses-of-the-cell">Proteins: workhorses of the cell</h2>
<p>Proteins make life as we know it possible; they are <a href="https://www.newscientist.com/article/mg21128251-300-first-life-the-search-for-the-first-replicator/">responsible for the hard work</a> that goes on within our cells. <a href="https://www.nature.com/scitable/topicpage/protein-function-14123348/">Proteins give our cells</a> shape and structure, and help them stay organized. Most biochemical reactions occur thanks to enzymes, which are entirely made of proteins. And hormones are proteins that send messages around the body.</p>
<p>The 3D shape of a protein is crucial to understanding how it functions within the body and what roles it plays. However, the information we often receive about proteins is not 3D. It is a list of letters – a sequence of amino acids, the building blocks of proteins. These lists can twist and fold, creating spirals, sheets, and other interesting sculptures.</p>
<p>For example, insulin is an important protein messenger that lets our cells know when they should absorb more sugar. Insulin consists of this sequence of 110 amino acids (there are 20 distinct amino acids, and each has a letter representing it):</p>
<p><code>MALWMRLLPLLALLALWGPDPAAAFVNQHLCGSHLVEALYLVCGERGFFYTPKTRREAEDLQVGQVELGGGPGAGSLQPLALEGSLQKRGIVEQCCTSICSLYQLENYCN</code></p>
<p>However, to really understand insulin, we need to know what it <a href="https://alphafold.ebi.ac.uk/entry/P01308">looks like as a 3D molecule</a>. Its shape and structure helps determine which other molecules will interact with it.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-07-09-t-cells/insulin-3.gif" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>Insulin lets our cells know to take in sugar from the blood. Animation from the AlphaFold Protein Structure Database.</figcaption>
</figure>
</div>
<p>The insulin receptor (which is on the surface of cells and can bind to insulin, receiving the message and sending off an internal signal within the cell) is also a protein. It is a sequence of over 1,000 amino acids and has a <a href="https://alphafold.ebi.ac.uk/entry/P06213">far more complicated structure</a>.</p>
</section>
<section id="a-new-champion-alphafold" class="level2">
<h2 class="anchored" data-anchor-id="a-new-champion-alphafold">A new champion: AlphaFold</h2>
<p>This question of converting amino acid sequences (on the left) to 3D structures (on the right) has wide-reaching implications within biology and medicine.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-07-09-t-cells/insulin-2ways.jpg" class="img-fluid figure-img" style="width:80.0%"></p>
<figcaption>Two different ways of representing the protein insulin: a string of letters, or its 3D structure</figcaption>
</figure>
</div>
<p>In 1994, a competition was established for researchers to go head-to-head in trying to best answer this question. Leading up to the competition, as scientists discovered new protein structures, using laboratory techniques such as X-ray crystallography or NMR spectroscopy, they kept some of the results secret. Later, these structures would be freely shared with other scientists, but for now, they must be held apart for the challenge. The somewhat dry name, <a href="https://predictioncenter.org/">Critical Assessment of Structure Prediction (CASP)</a>, belies the importance of this task.</p>
<p>In the competition, each team submits their guesses, generated by computer programs they created. The distance between the guesses and the actual structures is measured in Angstroms, with 1 Angstrom being one just ten-billionth of a meter.</p>
<p>CASP did not make international headlines until 2020. This is when an entry, titled <a href="https://www.nature.com/articles/s41586-021-03819-2">AlphaFold</a> smashed all existing records, and all competitors, by a wide margin. AlphaFold was created by DeepMind, an AI company owned by Google.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-07-09-t-cells/headlines-2021.jpg" class="img-fluid figure-img" style="width:65.0%"></p>
<figcaption>AlphaFold was named the innovation of the year by both Science and Nature Methods magazines for 2021</figcaption>
</figure>
</div>
</section>
<section id="challenges-with-t-cell-data" class="level2">
<h2 class="anchored" data-anchor-id="challenges-with-t-cell-data">Challenges with T cell data</h2>
<p>At first glance, it may seem like AlphaFold would give us a solution to our T cell problem. After all, T cell receptors are made of proteins, and they bind to other proteins. Aren’t protein structures a solved problem? Not so fast. The T cell receptor problem poses many different and additional challenges.</p>
<p>There are a few factors that make the question of what a T cell will bind to– known as <a href="https://www.nature.com/articles/s41577-023-00835-3">T cell receptor specificity</a>– distinctive and more complicated than the problem AlphaFold solves. First of all, AlphaFold was created to predict protein structures from a single chain of amino acids. However, TCRs have 2 key chains of amino acids: alpha and beta. Both are important. These alpha and beta chains can be <a href="https://elifesciences.org/articles/82813">mixed and matched</a> willy-nilly.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-07-09-t-cells/hudson-1bc.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>Here, MHC is represented in black &amp; grey and the TCR is represented in light &amp; dark turquoise. The peptide is in purple. Like a puzzle, we are interested in knowing which pieces fit together. Figure 1bc from Hudson, et al, 2023</figcaption>
</figure>
</div>
<p>Another issue is that we have far less data on TCRs than we do on general proteins. The genes that code for the TCR vary greatly across the population. They are some of the most diverse genes in the entire human genome! This means that the (relatively small) dataset we do have is only a fraction of what exists. AlphaFold is based on machine learning, and machine learning models are highly dependent on the quality, quantity, and nature of the data used to train them.</p>
<p><em>This post is part 1 in a series on using AI to predict T cell binding. Here are <a href="https://rachel.fast.ai/posts/2024-07-24-neural-nets/">part 2</a> and <a href="https://rachel.fast.ai/posts/2024-08-01-ai-t-cells/">part 3</a>.</em></p>
<p><em>Thank you to Jeremy Howard for feedback on earlier drafts of this post.</em></p>
<p>You can subscribe to be notified of new blog posts by submitting your email below:</p>
<script type="text/javascript" src="https://campaigns.zoho.com.au/js/zc.iframe.js"></script>
<iframe scrolling="no" frameborder="0" id="iframewin" width="100%" height="100%" src="https://zcmp-pd.maillist-manage.com.au/ua/Optin?od=11d0c075b7ed49&amp;zx=11a17553b1&amp;tD=156971d471c7b09&amp;sD=156971d471c7cb3">
</iframe>


</section>

<p><br><br><i>I look forward to reading your responses. Create a free GitHub account to comment below.</i></p> ]]></description>
  <category>machine learning</category>
  <category>science</category>
  <guid>https://rachel.fast.ai/posts/2024-07-09-t-cells/</guid>
  <pubDate>Mon, 08 Jul 2024 14:00:00 GMT</pubDate>
  <media:content url="https://rachel.fast.ai/posts/2024-07-09-t-cells/tcell-niaid-fingernail.jpg" medium="image" type="image/jpeg"/>
</item>
<item>
  <title>The microbiomes of wildfires, nanoplastics, and roller derby</title>
  <dc:creator>Rachel Thomas</dc:creator>
  <link>https://rachel.fast.ai/posts/2024-05-07-microbiome-2/</link>
  <description><![CDATA[ 





<p>This is part 2 in a series covering surprising facts about the microbiome. Be sure to check out <a href="https://rachel.fast.ai/posts/2024-04-25-microbiome-1/">part 1 here</a>. A microbiome is the community of microorganisms found living together in a given habitat. The microbiome of the human gut can contribute to <a href="https://www.youtube.com/watch?v=6BYocBxOfsU&amp;list=PLtmWHNX-gukLirebdPH8lla41SS78kjLD">development of autoimmune disease</a>. The microbiome of plant roots can include bacteria traveling on <a href="https://www.youtube.com/watch?v=WpgAA1mh7Io">fungal super-highways</a>. Scientists also study the microbiome of nanoplastics floating in the ocean and the tiny microbes throughout our atmosphere that catalyze precipitation. Microbiomes influence food chains, weather patterns, and the human immune system.</p>
<section id="microbes-are-everywhere" class="level2">
<h2 class="anchored" data-anchor-id="microbes-are-everywhere">Microbes are everywhere</h2>
<p>When I lived in California there were several years in which we had widespread wildfire smoke that was harmful to breathe, polluted the air, and created an eerie yellow hue. I knew that smoke contains toxins, and I incorrectly assumed that it wouldn’t be a suitable habitat for anything living. Counter-intuitively, <a href="https://sci-hub.se/https://www.science.org/doi/10.1126/science.abe8116">smoke is a good habitat for some microbes</a>, as the particulate matter helps block UV rays that might otherwise kill them. Water vapor produced by the combustion of biomass can help prevent microbes from drying out. And convective columns and updraft winds can draw in microbes from long distances.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-05-07-microbiome-2/wildfire-microbiome.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>Microbes in wildfire smoke</figcaption>
</figure>
</div>
<p><a href="https://www.thelancet.com/journals/lanplh/article/PIIS2542-5196(23)00046-3/fulltext">One study on the California wildfires</a> found that hospitalizations for coccidioidomycosis, a dangerous fungal infection, increased by 20% in the month following smoke exposure. This is a worrying threat given that the climate crisis has led to increasing frequency and intensity of wildfires. There is even a nascent field of <strong>pyroaerobiology</strong>, studying the aerosolization and transport of <a href="https://esajournals.onlinelibrary.wiley.com/doi/full/10.1002/ecs2.2507">microbes by wildfires</a>.</p>
<p>An estimated 5-13 million tons of plastic enter the oceans each year, including run-off from polluted rivers. Floating “islands” of plastic debris, including microplastics less than 5 mm in size, are found in the oceans. These plastics impact <a href="https://journals.asm.org/doi/10.1128/msystems.01112-20">over 700 different species of aquatic life</a>, ranging from microscopic phytoplankton to large whales. There are microbes that thrive on microplastics, <a href="https://esajournals.onlinelibrary.wiley.com/doi/abs/10.1890/150017">forming the Plastisphere</a>. Unfortunately, this thin film of bacteria and fungi living on microplastics may make them more appealing to some sea animals, such as loggerhead turtles, to eat. Researchers have found that the Atlantic and Pacific oceans have distinct Plastisphere communities, and that these differences may necessitate different remediation approaches.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-05-07-microbiome-2/microplastics-chesapeake.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>Microplastics from the Patapsco River, pictured at the laboratory of Dr.&nbsp;Lance Yonk, Photo by Will Parson/Chesapeake Bay Program</figcaption>
</figure>
</div>
<p>Researchers are also investigating how microbes can be used to break down these dangerous plastics. One interesting research study collated data on <a href="https://journals.asm.org/doi/10.1128/msystems.01112-20">over 16,000 genes involved in plastic degradation</a>. These genes came from 6,000 microbial species, showing that there are numerous possibilities of potential microbes that may be used for bioremediation. Researchers have even found some microbes can be used <a href="https://www.nature.com/articles/d41586-023-01697-4">to help break down PFAS</a>, often called “forever chemicals”, which are polluting aquatic environments. While it is very difficult to break down the fluorine-carbon bonds found in PFAS, microbes can instead first break chlorine-carbon bonds in the chemicals, which then makes it easier to remove the fluorine atoms.</p>
</section>
<section id="there-is-no-clear-good-vs.-bad-for-microbes" class="level2">
<h2 class="anchored" data-anchor-id="there-is-no-clear-good-vs.-bad-for-microbes">There is no clear “good” vs.&nbsp;“bad” for microbes</h2>
<p><em>Pseudomonas syringae</em> is a plant pathogen that is a scourge of agricultural crops. It infects almost all economically valuable species of crops and <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5972017/">threatens global agriculture production</a>.</p>
<p>But actually it’s not all bad – this same species of bacteria is important to weather precipitation cycles. For temperatures above -40 degrees C, ice formation requires a catalyst, called an ice nucleator. Researchers have discovered that <a href="https://www.pnas.org/doi/full/10.1073/pnas.0809816105">most ice nucleators are biological</a>: tiny microbes that are widely dispersed throughout the atmosphere all over the world, even in Antarctica! And it turns out that <a href="https://www.activeremedy.org/wp-content/uploads/2014/10/brent-christner_2012_cloudy-with-a-chance-of-microbes.pdf"><em>Pseudomonas syringae</em> is a common nucleator for catalyzing precipitation</a>. This example highlights the complexity of microbial ecosystems. Efforts to protect plants from this pathogen must take into account the interrelated impact on precipitation.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-05-07-microbiome-2/p-syringae.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>Left and center: Infected tomato plants and an infected leaf (images from Wikimedia). On the right: the life cycle of P. syringae, from the paper “Cloudy with a Chance of Microbes”</figcaption>
</figure>
</div>
<p>Another example of a microbe that can’t be clearly classified as “good” or “bad” is <em>Prevotella Copri</em>, found in the human gut microbiome. An aptly titled paper, “The curious case of Prevotella copri” describes”puzzling discrepancies” in how various research studies have associated this species with a <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10478744/">range of positive and negative impacts on disease</a>.</p>
<p>The authors share that high abundance of P. copri is associated with non-Western diets, rheumatoid arthritis, Ankylosing spondylitis (a type of inflammatory arthritis), Graves’ orbitopathy (autoimmune condition that can damage vision vision), diabetes mellitus, and cancer. This is particularly surprising, because several of the inflammatory and autoimmune conditions listed are more strongly associated with Western diets. Low abundance of P. copri is associated with Western diets, glucose intolerance, chronic hives, Parkinson’s disease, and food allergies.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-05-07-microbiome-2/p-copri.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>From the paper “The curious case of Prevotella copri”</figcaption>
</figure>
</div>
<p>It is unclear whether abundant <em>Prevotella copri</em> in the gut is good or bad. The paper notes that <em>P. copri</em> has a diverse set of strains, which differ from the reference strain, and may partly explain the discrepancies. The ways that gut microbes interact with each other, our immune systems, and our genetics can be quite complex, and these factors may mediate how P. copri influences a person’s health.</p>
</section>
<section id="the-social-microbiome" class="level2">
<h2 class="anchored" data-anchor-id="the-social-microbiome">The Social Microbiome</h2>
<p>We often talk about microbiomes at an individual level: the microbes inside the guts of a person (or insect) or the microbes on their skin or in their mouth. However, researchers also study the <a href="https://www.nature.com/articles/s41559-020-1220-8"><strong>social microbiome</strong></a>: the metacommunity of all the microbiomes of animals in a social community. Members of a community can influence each other’s microbiomes through their social bonds.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-05-07-microbiome-2/roller-derby-fingernail.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>Roller derby can change your skin’s microbiome, Image: Oklarsson, CC BY-SA 3.0, Wikimedia</figcaption>
</figure>
</div>
<p>One study sampled the <a href="https://peerj.com/articles/53/">skin microbiome of roller derby players</a> before and after games. The three teams in the study each came from a different city and state. They found that prior to a game, the microbiomes were more similar within a team, and different across teams. However, after competing in a game, which involves a fair amount of physical contact, all of the players’ skin microbiomes were more similar to one another.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-05-07-microbiome-2/roller-derby-fig1.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>Red, blue, and green represent the microbiomes of teams from San Jose, CA, Eugene, OR, and DC, respectively. There is much more overlap after games. Figure 1 from Meadow, et al, 2013</figcaption>
</figure>
</div>
<p>Much of the research on the social microbiome has been done on nonhuman primates, where grooming, mating, feeding, and other close contact have been shown to impact the microbiome. Social networks and group membership were found to <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4379495/">determine the microbiomes of baboons</a>, and were more predictive than age or gender.</p>
<p>For some animals, microbiomes inside of scent glands are used to communicate group membership. Hyenas typically live in social clans of 40-80 individuals, and these clans are often overlapping in their territories. They secrete fermentative bacteria from their scent glands. This allows for group-specific odors, making it easier for the hyenas to communicate if they are in the same group as one another. <a href="">One study found</a> that the relative abundance of different types of bacteria was distinctive for different hyena clans.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-05-07-microbiome-2/spotted-hyenas.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>A group of hyenas may have a distinctive composition of microbes, giving them an identifying scent. Photo by Charles Peterson, CC BY-NC-ND 2.0</figcaption>
</figure>
</div>
</section>
<section id="more-to-learn" class="level2">
<h2 class="anchored" data-anchor-id="more-to-learn">More to learn</h2>
<p>Understanding the microbiome is a crucial piece of better understanding the health of plants, humans, and other animals, as well as our environment. Microbes synthesize vitamins, catalyze precipitation, spread disease or protect from it, can remediate or exacerbate the impacts of pollution, and influence food chains. Only with relatively recent advances in genomic sequencing have we been able to better study and compare the balance of species composing different microbiomes. This is a fascinating and growing area of research.</p>


</section>

<p><br><br><i>I look forward to reading your responses. Create a free GitHub account to comment below.</i></p> ]]></description>
  <category>science</category>
  <guid>https://rachel.fast.ai/posts/2024-05-07-microbiome-2/</guid>
  <pubDate>Mon, 06 May 2024 14:00:00 GMT</pubDate>
  <media:content url="https://rachel.fast.ai/posts/2024-05-07-microbiome-2/roller-derby-fingernail.jpg" medium="image" type="image/jpeg"/>
</item>
<item>
  <title>Surprises of the Microbiome</title>
  <dc:creator>Rachel Thomas</dc:creator>
  <link>https://rachel.fast.ai/posts/2024-04-25-microbiome-1/</link>
  <description><![CDATA[ 





<p><em>This is part 1 in a series. You can <a href="https://rachel.fast.ai/posts/2024-05-07-microbiome-2/">read part 2 here</a>.</em></p>
<p>Humans <a href="https://microbiomejournal.biomedcentral.com/articles/10.1186/s40168-020-00875-0">co-evolved together</a> with our microbiomes, the bacteria and fungi that colonize our intestines, coat our skin, and inhabit our mouths, throats, and numerous other body locations. Microbiomes are significant because they influence how we process nutrients, fight pathogens, and what diseases we are susceptible to. This is not just true of humans, but plants and insects also have microbiomes that influence their metabolism and immune systems.</p>
<p>A microbiome is the community of microorganisms found living together in a given habitat– this does not just include living on a host. Scientists study the microbiomes of nanoplastics floating in the ocean, covered in bacteria and fungi, and of the tiny microbes throughout the atmosphere that catalyze precipitation. Even wildfire smoke has a microbiome. These environmental microbiomes influence food chains, weather patterns, and the spread of disease. We can’t fully understand the impacts of pollution and climate change without studying them.</p>
<p>Studying the microbiome is crucial for better understanding the health and disease of humans, other animals, plants, and our environment. In this 2-part series, I will cover a grab bag of surprising and fascinating facts about the microbiome.</p>
<section id="when-the-immune-system-attacks-friendly-bacteria" class="level2">
<h2 class="anchored" data-anchor-id="when-the-immune-system-attacks-friendly-bacteria">When the immune system attacks friendly bacteria…</h2>
<p>Commensal bacteria, often referred to as “friendly”, are bacteria that live on human skin or in the gut and help us. The human gut microbiome can help synthesize vitamins we need, regulate dopamine, and protect us from infection. However, so-called friendly bacteria can cause grave problems when they cross out of the gut into the bloodstream. Microbes can’t necessarily be classified as “good” or “bad”, and this impact of what happens when microbes cross from the gut to the blood is an example of this.</p>
<p>I created a short 5 minute video to explain these ideas, geared for a general audience:</p>
<center>
<iframe width="560" height="315" src="https://www.youtube.com/embed/6BYocBxOfsU?si=aSkr2Zmp5HRGwg0Q" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="">
</iframe>
</center>
<p>If “friendly” microbes stray into the bloodstream, our immune system may mount an attack against them by producing antibodies and activating T cells (the two primary approaches of our adaptive immune systems). The small regions on a bacteria recognized by T cells or antibodies are known as <strong>epitopes</strong>. All bacteria have many epitopes, and genetics helps determine which particular epitopes your immune system is able to recognize. Two different people will likely produce T cells recognizing different epitopes for the same bacteria.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-04-25-microbiome-1/delcherico.jpg" class="img-fluid figure-img" style="width:60.0%"></p>
<figcaption>Bacteria can leak between the cells lining the intestines into the bloodstream, where they may trigger an immune reaction (Fig 1, Del Cherico, et al, 2022</figcaption>
</figure>
</div>
<p>If the epitopes our immune system attacks are too similar to regions on our own cells, autoimmune disease can develop. This means that the <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9737253/">friendly bacteria that has leaked into the bloodstream</a> (where it does not belong) “looks” too similar to some of your own cells. For instance, if the bacteria “looks” like your pancreatic cells (which produce insulin), your T cells may begin mistakenly attacking these insulin-producing cells, leading to Type 1 Diabetes.</p>
<p>The microbiome can also play a role in how likely microbes are to leak out of a person’s gut in the first place! Low levels of certain commensal bacteria increases gut permeability. This heightened permeability allowing the microbes to leave the intestines. This process has been <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9737253/">linked to progression of Type 1 Diabetes</a> to more advanced stages. Note that commensal bacteria are just one potential trigger for the development of autoimmune diseases. As I discussed in a <a href="https://rachel.fast.ai/posts/2023-03-22-viruses2/">previous post</a>, viral infections can also trigger Type 1 Diabetes and other autoimmune diseases.</p>
</section>
<section id="bacteria-travel-on-fungal-highways-in-the-plant-microbiome" class="level2">
<h2 class="anchored" data-anchor-id="bacteria-travel-on-fungal-highways-in-the-plant-microbiome">Bacteria travel on fungal highways in the plant microbiome</h2>
<p>Plants also have microbiomes. One component of the plant microbiome are fungi that surround the plant roots and the bacteria that travel along these “fungal highways.” The fungi exude sugars that the bacteria use as an energy source along their journeys. The bacteria are able to mineralize organic phosphate, which is beneficial to the fungus and to the plant. Together, the fungus, plant, and bacteria are in symbiotic relationships.</p>
<p>Watch <a href="https://www.youtube.com/watch?v=AnsYh6511Ic">this short 90 second video</a> to see bacteria traveling along a fungal highway in real-time. You won’t believe how fast they are!</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-04-25-microbiome-1/fungal-highway-3.gif" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>From the Science Friday video on fungal highways</figcaption>
</figure>
</div>
<p>For more on fungal highways, <a href="https://www.youtube.com/watch?v=WpgAA1mh7Io">here is a 3 minute Science Fridays video</a> which discusess how fungal networks optimize traffic to avoid traffic jams. It’s created by Christian Baker, who does a great job of explaining science in a succinct and engaging way. In this video, he interviews math professor Dr.&nbsp;Marcus Roper of UCLA who studies the fluid dynamics of bacterial travel.</p>
<p>Other researchers are looking at the use of fungal highways to allow <a href="https://pubmed.ncbi.nlm.nih.gov/16047804/">pollutant-degrading bacteria to travel</a>, which could help remove pollutants in a broader area as part of bioremediation efforts.</p>
<p>Plant microbiomes have a strong influence on their immune function. They help prevent the colonization and growth of pathogens. Unfortunately, <a href="https://www.nature.com/articles/s41579-023-00900-7">climate change is altering the microbiomes of plants</a>, which can make them more susceptible to disease. There are several mechanisms by which this can happen. Climate change alters the composition of microbes found in soil, which is the first line of defense for plants against pathogens. Plants release chemicals and metabolites into the soil, which are useful for attracting beneficial microbiomes, and climate change can harm plants’ capacity to do this (at least in the same quantities as before). Both plants and microbes may migrate into new areas due to changes in climate, further disrupting previous equilibriums.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-04-25-microbiome-1/climate-change-3ab.jpg" class="img-fluid figure-img" style="width:80.0%"></p>
<figcaption>Some of the ways that climate change can impact plant microbiomes, Fig 3 from Singh, et al, 2023</figcaption>
</figure>
</div>
</section>
<section id="study-of-the-microbiome-is-both-old-and-very-new" class="level2">
<h2 class="anchored" data-anchor-id="study-of-the-microbiome-is-both-old-and-very-new">Study of the microbiome is both old and very new</h2>
<p>My microbiome course began with an in-depth unit covering different methods for sequencing genomes. There have been <a href="">multiple revolutions</a> in gene sequencing technology in recent decades, and new machines for gene sequencing are released regularly. Our growing understanding of the microbiome would not be possible without this technology to sequence numerous bacterial species in a sample, and in many ways the field is very young, with lots of open questions remaining.</p>
<p>Even though microbiome research has only taken off in the last 20 years with the advent of new sequencing technologies, ideas about the microbiome and its role in health have been around for a long time. <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3426293/">Antonie van Leewenhoek used a microscope</a> to compare the microbes in his oral and fecal microbiota in the 1680s.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-04-25-microbiome-1/metchkinoff-pacbio.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>Left: Élie Metchnikoff, Right: a modern PacBio sequencing machine, which costs half a million dollars</figcaption>
</figure>
</div>
<p>Élie Metchnikoff, who won the Nobel Prize in Medicine 1908, is best known as the father of cellular immunity. He discovered that there are immune cells that can ingest pathogens. He named them “eating cells” (phagocytes), and we now refer to these immune cells as “big eaters” (macrophages). In addition to discovering cellular immunity, <a href="https://www.frontiersin.org/articles/10.3389/fpubh.2013.00052/full">Metchnikoff is considered</a> one of the founders of probiotics. He believed consuming sour milk (yogurt) would be beneficial for the microbiome and he laid the conceptual framework for fecal transplantation to modify the microbiome, over 100 years ago. Questions about whether and how we can alter our microbiomes are still being <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5374383/">studied and debated</a> today, as are the roles of genetics, lifestyle, and environment in influencing the microbiome.</p>
</section>
<section id="stay-tuned" class="level2">
<h2 class="anchored" data-anchor-id="stay-tuned">Stay tuned…</h2>
<p>Please subscribe to my blog to be notified of Part 2, which will describe the microbiomes of nanoplastics and wildfire smoke, as well as some puzzling examples of microbes that are both good and bad.</p>
<p>In the meantime, you may be interested in a previous post of mine that discussed the microbiomes of insects, which can help produce both essential nutrients and defensive toxins to use against predators. Read more in <a href="https://rachel.fast.ai/posts/2024-01-09-insects/">4 Things I Learned About Bugs</a>.</p>
<p>Submit your email below to be notified of new posts:</p>
<script type="text/javascript" src="https://campaigns.zoho.com.au/js/zc.iframe.js"></script>
<iframe scrolling="no" frameborder="0" id="iframewin" width="100%" height="100%" src="https://zcmp-pd.maillist-manage.com.au/ua/Optin?od=11d0c075b7ed49&amp;zx=11a17553b1&amp;tD=156971d471c7b09&amp;sD=156971d471c7cb3">
</iframe>


</section>

<p><br><br><i>I look forward to reading your responses. Create a free GitHub account to comment below.</i></p> ]]></description>
  <category>science</category>
  <guid>https://rachel.fast.ai/posts/2024-04-25-microbiome-1/</guid>
  <pubDate>Wed, 24 Apr 2024 14:00:00 GMT</pubDate>
  <media:content url="https://rachel.fast.ai/posts/2024-04-25-microbiome-1/fungal-highway-fingernail.jpg" medium="image" type="image/jpeg"/>
</item>
<item>
  <title>Using AI to Discover New Antibiotics</title>
  <dc:creator>Rachel Thomas</dc:creator>
  <link>https://rachel.fast.ai/posts/2024-04-03-ai-antibiotics/</link>
  <description><![CDATA[ 





<section id="a-growing-threat" class="level2">
<h2 class="anchored" data-anchor-id="a-growing-threat">A Growing Threat</h2>
<p>The WHO estimates that antimicrobial resistance (AMR) contributed to approximately 5 million deaths worldwide in 2019 (<a href="https://www.who.int/news-room/fact-sheets/detail/antimicrobial-resistance">“Antimicrobial Resistance” 2023</a>). In the USA, AMR results annually in 8 million additional days spent in the hospital and over $20 billion in healthcare costs (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4206945/">Bush et al.&nbsp;2011</a>). The problem will only grow in magnitude as antibiotic resistant genes in bacteria spread. AMR is compounded by the fact that development of new antibiotics has stalled over the past few decades. Historically, antibiotics were primarily discovered from natural products generated by soil microbes. This approach led to an initial burst of discoveries in the mid-20th century, but since then, soil screening has produced diminishing returns. Other antibiotics were created as derivatives of known antibiotics, but these have a limited ability to outpace antibiotic resistance (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8463434/">Randall and Davies 2021</a>).</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-04-03-ai-antibiotics/drug-funnel.jpg" class="img-fluid figure-img" style="width:60.0%"></p>
<figcaption>The narrow funnel of drug discovery (Figure from U Michigan Alzheimer’s Disease Center)</figcaption>
</figure>
</div>
<p>Developing new drugs is an expensive, time-consuming, and failure-prone process, with only a small percentage of potential drug candidates making it from discovery through clinical testing to regulatory approval. Antibiotic research has been particularly neglected due to unfavourable financial incentives. New antibiotics are typically only used if standard antibiotics fail, and antibiotics are given for a limited course. Thus, antibiotic research is not seen as profitable and pharmaceutical companies have largely avoided it (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4273861/">Durrant and Amaro 2015</a>). Artificial intelligence holds promise for accelerating research to address this threat, although there are also limitations and risks.</p>
</section>
<section id="possibilities-with-deep-learning" class="level2">
<h2 class="anchored" data-anchor-id="possibilities-with-deep-learning">Possibilities with Deep Learning</h2>
<p>Machine learning uses data to find patterns, as opposed to relying on manually coded rules. Deep learning is a powerful family of machine learning algorithms, responsible for the recent breakthroughs in artificial intelligence. The field of has seen massive improvements in accuracy and speed, and reductions in cost. Key reasons for this progress are the affordability of GPUs (processors used in video gaming which can quickly perform numerical computations in parallel), increases in available data, and research advances in the underlying algorithms (<a href="https://www.nature.com/articles/nature14539">LeCun, Bengio, and Hinton 2015</a>).</p>
<p>Deep learning-assisted approaches to drug discovery offer a potentially faster and cheaper way to identify new antibiotic candidates. It is estimated that there are 10<sup>30</sup>–10<sup>60</sup> potential drug-like chemicals; this scale means it is impossible to explore even a fraction of these (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8429579/">Melo, Maasch, and de la Fuente-Nunez 2021</a>). Algorithms can help determine what compounds to prioritize.</p>
</section>
<section id="protein-folding-approaches-with-ai" class="level2">
<h2 class="anchored" data-anchor-id="protein-folding-approaches-with-ai">Protein Folding Approaches with AI</h2>
<p>AlphaFold has revolutionized protein structure prediction (<a href="https://www.nature.com/articles/s41586-021-03819-2">Jumper et al.&nbsp;2021</a>, my <a href="https://youtu.be/pB0RxG1NdtA">summary here</a>) and there has been interest in using it to aid in antibiotic discovery. Researchers applied AlphaFold to the essential proteins from <em>E. coli</em> (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9446081/">Wong et al.&nbsp;2022</a>). They then used molecular docking programs to predict if these sites would be able to bind with over 300 different existing drug compounds. When they tested the predictions experimentally in the lab, they found that accuracy was relatively low. One limitation of the approach is that AlphaFold does not differentiate between active vs.&nbsp;inactive conformations. The researchers highlighted the need for better scoring functions to accurately rank the binding poses and for benchmarking datasets to evaluate performance (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9446081/">Wong et al.&nbsp;2022</a>).</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-04-03-ai-antibiotics/lyu-2023-1a.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>Comparisons between actual protein structures in turquoise versus AlphaFold predictions in yellow (Figure 1a from Lyu, et al, 2023)</figcaption>
</figure>
</div>
<p>A more recent pre-print (not yet peer-reviewed) offered a potential explanation of AlphaFold’s poor performance on predicting existing drugs (<a href="https://www.biorxiv.org/content/10.1101/2023.12.20.572662v1">Lyu et al.&nbsp;2023</a>, <a href="https://www.nature.com/articles/d41586-024-00130-8">news coverage</a>). The authors proposed that minor differences in structure may cause AlphaFold to miss existing drugs, but that it may be identifying equally promising ones. Starting with 2 known protein structures, the researchers used AlphaFold to virtually screen millions of potential drugs and then synthesized a few hundred of the most promising candidates. For comparison, they also followed a more traditional process to screen and synthesize drug candidates. Even though AlphaFold and the traditional process yielded completely different drug candidates, they had similar success rates of how many of the candidates bound to the proteins of interest as predicted (<a href="https://www.biorxiv.org/content/10.1101/2023.12.20.572662v1">Lyu et al.&nbsp;2023</a>).</p>
</section>
<section id="predicting-antibacterial-activity-with-ai" class="level2">
<h2 class="anchored" data-anchor-id="predicting-antibacterial-activity-with-ai">Predicting Antibacterial Activity with AI</h2>
<p>Other work has not relied on protein structure prediction software like AlphaFold, but instead created deep learning models directly for the underlying problem. Researchers used training sets of thousands of small molecule drugs to build models predicting which will have antibacterial activity (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8349178/">Stokes et al.&nbsp;2020</a>; <a href="https://www.nature.com/articles/s41589-023-01349-8">Liu et al.&nbsp;2023</a>). Message-passing neural networks, a type of deep learning model, allow for iteratively exchanging information between adjacent atoms and bonds. This captures some of the structural information of a protein, without explicitly predicting protein structure. A model was created using a dataset of thousands of repurposed drugs, and then applied to chemical libraries containing over 100 million molecules (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8349178/">Stokes et al.&nbsp;2020</a>). The output was a ranking of the most promising drug candidates based on predicted antibacterial activity, chemical structure, and availability. This approach predicted that Halicin, which is structurally distinct form existing classes of antibiotics, would be an effective drug. When tested empirically, Halicin demonstrated a broad-spectrum impact against bacteria including M. tuberculosis, carbapenem-resistant <em>Enterobacteriaceae</em>, <em>C. difficile</em>, and multi-resistant <em>A. baumannii</em> (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8349178/">Stokes et al.&nbsp;2020</a>).</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-04-03-ai-antibiotics/stokes-2020.jpg" class="img-fluid figure-img" style="width:75.0%"></p>
<figcaption>The process used to discover Aubacin with a message-passing neural network (Figure 1 from Stokes, et al, 2020)</figcaption>
</figure>
</div>
<p>A similar approach was used by the same lab to train a deep learning model to predict antibiotics that would be effective against multi-resistant <em>Acinetobacter baumannii</em> (<a href="https://www.nature.com/articles/s41589-023-01349-8">Liu et al.&nbsp;2023</a>). The WHO has designated drug resistant <em>A. baumannii</em> as one of the highest priorities for urgent development of new antibiotics. The model discovered a new antibiotic drug, Aubacin, which is structurally distinct from existing classes of antibiotics. In a lab experiment, Aubacin was effective at supressing <em>A. baumannii</em> in a mouse wound. It was also effective when tested against 41 different strains of resistant A. baumannii from the CDC Antibiotic Resistant Isolation bank (<a href="https://www.nature.com/articles/s41589-023-01349-8">Liu et al.&nbsp;2023</a>).</p>
</section>
<section id="data-sources-and-tools" class="level2">
<h2 class="anchored" data-anchor-id="data-sources-and-tools">Data Sources and Tools</h2>
<p>AI models are highly dependent on the data used to train them. Training data shapes the accuracy of a model, helps set what possibilities are considered, and determines the context in which it is appropriate to use. One useful data source is the <a href="https://www.broadinstitute.org/drug-repurposing-hub">Drug Repurposing Hub</a> (<a href="https://www.nature.com/articles/nm.4306">Corsello et al.&nbsp;2017</a>), which contains a hand-curated collection of thousands of existing drugs, including ones where safety was established in clinical trials, but that never obtained regulatory approval. Applying existing drugs to new diseases other than that which they were originally developed for is a pragmatic approach that can have a rapid and lower-cost impact. Another useful tool is the <a href="https://zinc15.docking.org/">ZINC15 database</a>, which includes data on small molecules, including their biological activity and chemical properties (<a href="https://pubs.acs.org/doi/full/10.1021/acs.jcim.5b00559">Sterling and Irwin 2015</a>). Both the Drug Repurposing Hub and ZINC 15 databases have been used in neural network approaches (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8349178/">Stokes et al.&nbsp;2020</a>; <a href="https://www.nature.com/articles/s41589-023-01349-8">Liu et al.&nbsp;2023</a>).</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-04-03-ai-antibiotics/drug-repurposing-hub-S2.jpg" class="img-fluid figure-img" style="width:75.0%"></p>
<figcaption>Diverse indications for drugs in the Drug Repurposing Hub, from 2017 launch. It has since grown in size to now include nearly 8,000 compounds (Figure S2 from Corsello, et al, 2017)</figcaption>
</figure>
</div>
<p><a href="https://www.rdkit.org/">RDKit</a> is an open-source library that can be used to compute molecular features. Some researchers use these computed features to augment the data gathered purely from existing structures and ligands, using one of the above databases together with RDKit-computed features as inputs to their models (<a href="https://www.nature.com/articles/s41589-023-01349-8">Liu et al.&nbsp;2023</a>).</p>
</section>
<section id="risks-and-limitations" class="level2">
<h2 class="anchored" data-anchor-id="risks-and-limitations">Risks and Limitations</h2>
<p>The research described above of antibiotics discovered through deep learning-assisted processes is still in preliminary stages, with none of those drugs having yet begun clinical trials, much less reached regulatory approval. In other fields, machine learning models have sometimes performed well in testing but failed to work as expected when deployed in the real world (<a href="https://research.google/pubs/everyone-wants-to-do-the-model-work-not-the-data-work-data-cascades-in-high-stakes-ai/">Sambasivan et al.&nbsp;2021</a>). Issues can arise when the context and quality of the training data is not fully understood (<a href="https://dl.acm.org/doi/10.1145/3458723">Gebru et al.&nbsp;2021</a>). ML systems are trained in clearly defined environments, while the physical world often has complex and volatile underlying phenomena, which are not fully accounted for during training. This can create a brittleness in ML-generated solutions. Moreover, there is a conflicting incentive system in which the creators of ML models are rewarded with prestige, but the arduous work used to gather and curate the data used to train those models is often neglected and taken for granted (<a href="https://research.google/pubs/everyone-wants-to-do-the-model-work-not-the-data-work-data-cascades-in-high-stakes-ai/">Sambasivan et al.&nbsp;2021</a>).</p>
<p>However, given the urgency of the antimicrobial resistance crisis, it is important that we consider a variety of innovative solutions for tackling it. Deep learning can be a powerful tool, and approaches to antibiotic discovery should continue to be pursued.</p>
<p><em>You can subscribe to be notified of new blog posts by submitting your email below:</em></p>
<script type="text/javascript" src="https://campaigns.zoho.com.au/js/zc.iframe.js"></script>
<iframe scrolling="no" frameborder="0" id="iframewin" width="100%" height="100%" src="https://zcmp-pd.maillist-manage.com.au/ua/Optin?od=11d0c075b7ed49&amp;zx=11a17553b1&amp;tD=156971d471c7b09&amp;sD=156971d471c7cb3">
</iframe>


</section>

<p><br><br><i>I look forward to reading your responses. Create a free GitHub account to comment below.</i></p> ]]></description>
  <category>machine learning</category>
  <category>science</category>
  <guid>https://rachel.fast.ai/posts/2024-04-03-ai-antibiotics/</guid>
  <pubDate>Tue, 02 Apr 2024 14:00:00 GMT</pubDate>
  <media:content url="https://rachel.fast.ai/posts/2024-04-03-ai-antibiotics/thumbnail.jpg" medium="image" type="image/jpeg"/>
</item>
<item>
  <title>“AI will cure cancer” misunderstands both AI and medicine</title>
  <dc:creator>Rachel Thomas</dc:creator>
  <link>https://rachel.fast.ai/posts/2024-02-20-ai-medicine/</link>
  <description><![CDATA[ 





<p>AI has made remarkable strides in the medical field, with capabilities including the <a href="https://www.nature.com/articles/s41586-023-06555-x">detection of Parkinson’s disease via retinal images</a>, <a href="https://www.nature.com/articles/s41591-023-02361-0">identification of promising drug candidates</a>, and <a href="https://www.nature.com/articles/s41586-023-06160-y">prediction of hospital readmissions</a>. While these advance are exciting, I’m wary of the practical impact AI will have on patients.</p>
<p>I recently watched a late-night talk show where skeptics and enthusiasts debated AI safety. Despite their conflicting views, there was one thing they could all agree upon. “AI will cure cancer,” one panelist declared, and everyone else confidently echoed their agreement. This collective optimism strikes me as overly idealistic and raises concerns about a failure to grapple with the realities of healthcare.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-02-20-ai-medicine/zhou-ext-fig-6.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>AI models can use retinal images to accurately detect Parkinson’s disease, stroke risk, and other medical issues (Extended Figure 6 from Zhou, et al, 2023)</figcaption>
</figure>
</div>
<p>My reservations about AI in medicine stem from two core issues: First, the medical system often disregards patient perspectives, inherently limiting our comprehension of medical conditions. Second, AI is used to disproportionately benefit the privileged while worsening inequality. In many instances, claims like “AI will cure cancer” are being invoked as little more than superficial marketing slogans. To understand why, it is first necessary to understand how AI is used, and how the medical system operates.</p>
<section id="how-automated-decision-making-is-used" class="level2">
<h2 class="anchored" data-anchor-id="how-automated-decision-making-is-used">How Automated Decision Making is Used</h2>
<p>Automated computer systems, often involving AI, are increasingly being used to make decisions that have a big impact on people’s lives: determining who gets jobs, housing, or healthcare. Disturbing patterns are found across numerous countries and a range of systems: there is typically no way to surface or correct errors, and all too often the goal is to increase corporate and government revenues by denying poor people resources they need to survive.</p>
<p>A woman in France had her food benefits reduced by a computer program and can no longer afford enough to eat. She talked to a program officer, who said <a href="https://www.hrw.org/news/2021/11/10/how-eus-flawed-artificial-intelligence-regulation-endangers-social-safety-net">the cut was due to an error in the computer program</a>, but was unable to change it. She did not have her food benefits reinstated. The system was designed such that the computer is always considered correct, even when humans recognize an error. This woman was not alone; she was one of 60,000 people who went hungry due to these errors.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-02-20-ai-medicine/hrw.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>From the Human Rights Watch report on automated decision making in the EU, https://www.hrw.org/news/2021/11/10/how-eus-flawed-artificial-intelligence-regulation-endangers-social-safety-net</figcaption>
</figure>
</div>
<p>A man in Australia was told that he had been overpaid welfare benefits and that he was now in debt to the government. The debt was an error, based on an intentionally faulty calculation as part of the computational RoboDebt program. However, the man had no way to contest it. Despondent, <a href="https://www.theguardian.com/australia-news/2023/feb/20/platitudes-and-false-words-mother-of-robodebt-victim-who-took-own-life-tells-inquiry-of-government-stonewalling">he died by suicide</a>. This was not a one-off incident. The Australian government was later found to have wrongly created debts for hundreds of thousands of people.They were putting poor people into debt with a flawed calculation system, ruining lives efficiently at scale. The government had <a href="https://www.theguardian.com/australia-news/2020/may/29/robodebt-government-to-repay-470000-unlawful-centrelink-debts-worth-721m">increased the number of poor people it put into debt</a> each week by 50x, compared to before RoboDebt.</p>
<p>A woman in the USA with cerebral palsy needed a health aid to help her get out of bed in the morning, to get her meals, and to complete other basic tasks. Her care was drastically <a href="https://www.theverge.com/2018/3/21/17144260/healthcare-medicaid-algorithm-arkansas-cerebral-palsy">cut due to a computer bug</a>. She was given no explanation and no option for recourse as her quality of life drastically plummeted. Only through a lengthy court case was it finally revealed that many people with cerebral palsy had wrongly lost their care due to a computer error.</p>
</section>
<section id="patterns-in-automated-decision-making" class="level2">
<h2 class="anchored" data-anchor-id="patterns-in-automated-decision-making">Patterns in Automated Decision Making</h2>
<p>These examples always flow in the same direction. <a href="https://digitalrepository.unm.edu/nmlr/vol50/iss3/2/">Professor Alvaro Bedoya</a>, the founding director of the Center on Privacy and Technology at the Georgetown University Law Center, wrote “<em>It is a pattern throughout history that surveillance is used against those considered ‘less than’, against the poor man, the person of color, the immigrant, the heretic. It is used to try to stop marginalized people from achieving power.</em>” The same pattern is found in the role of technology in decision systems.</p>
<p>The goal of many automated decision systems is to increase revenues for governments and private companies. When this is applied to health and medicine, the goal is often achieved by denying poor people food or medical care. People often trust computers to be more accurate than humans, in a bias known as automation bias. A <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3240751/">systematic review of 74 research studies</a> found that automation bias exists across a range of fields, including healthcare, exerting a consistent influence. This bias can make it harder for people to recognize errors in automated decision-making. Moreover, implementing mechanisms to identify and correct errors is often seen as an unnecessary expense.</p>
<p>In all of the cases above, the people most impacted (those losing access to needed food or medical care, or unjustly being thrown into debt) recognized the errors in the system earliest. Yet the systems were built with no mechanism for recognizing mistakes, for allowing the participation of those impacted, nor for providing recourse to those harmed. Unfortunately, this can also be the case in medicine.</p>
</section>
<section id="how-the-medical-system-operates" class="level2">
<h2 class="anchored" data-anchor-id="how-the-medical-system-operates">How the Medical System Operates</h2>
<p>An AI algorithm that reads MRIs more accurately would not have helped neurologist Ilene Ruhoy, MD, PhD, when she developed a 7 cm brain tumor. The key obstacle to her treatment was getting fellow neurologists to believe her symptoms and even order an MRI in the first place. “<em>I was told I knew too much, that I was working too hard, that I was stressed out, that I was anxious</em>,” <a href="https://www.washingtonpost.com/wellness/interactive/2022/women-pain-gender-bias-doctors/">Dr.&nbsp;Ruhoy recounts</a>. Eventually, after her symptoms worsened further, she was able to get an MRI and <a href="https://www.webmd.com/women/features/women-doctors-symptoms-dismissed">urgently sent in for a 7 hour surgery</a>. Because of the delay in her diagnosis, her tumor was so large that it could not be completely removed, which has led to it growing back since her first surgery.</p>
<p>Dr.&nbsp;Ruhoy’s experience is sadly common. While Dr.&nbsp;Ruhoy lives in the USA, a study in the UK found that <a href="https://www.bbc.com/future/article/20180523-how-gender-bias-affects-your-healthcare">almost 1 in 3 patients with brain tumors</a> had to visit doctors at least 5 times before receiving an accurate diagnosis. Again, MRI-reading AI can not help these patients whose doctors won’t order an MRI in the first place. On average, it takes <a href="https://academic.oup.com/rheumap/article/4/1/rkaa006/5758274">Lupus patients 7 years</a> to receive a correct diagnosis, and 1 in 3 are initially misdiagnosed with doctors incorrectly claiming mental health issues are the root of their symptoms. Even healthcare workers are often shocked at how quickly they are dismissed and disbelieved once they become patients. For instance, <a href="https://www.theatlantic.com/health/archive/2021/11/health-care-workers-long-covid-are-being-dismissed/620801/">interviews with a dozen healthcare workers</a> revealed that their colleagues shifted to discarding their expertise as soon as they developed Long Covid.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-02-20-ai-medicine/nothing-wrong-2.jpg" class="img-fluid figure-img" style="width:70.0%"></p>
<figcaption>‘Everybody was telling me there was nothing wrong’, a BBC article by Maya Dusenbery</figcaption>
</figure>
</div>
<p>This disregard of patient experience and patient expertise severely limits medical knowledge. It results in delayed diagnoses, misdiagnoses, missing data, and incorrect data. AI is great at finding patterns in existing data. However, AI will not be able to solve this problem of missing and erroneous underlying data. Furthermore, there is a negative feedback loop around lack of medical data for poorly understood diseases: doctors disbelieve patients and dismiss them as anxious or complaining too much, failing to gather data which could help illuminate the disease.</p>
</section>
<section id="the-wrong-data" class="level2">
<h2 class="anchored" data-anchor-id="the-wrong-data">The Wrong Data</h2>
<p>Even worse, often research problems are reformulated to shoe-horn inadequate data sources in. Analyzing electronic health record data is cheaper than searching for new causal mechanisms. Medical data is often limited by the categories of billing codes, by what doctors choose to note from a patient’s account, and what tests are ordered. The data are inherently incomplete.</p>
<p><a href="https://www.bbc.com/future/article/20180518-the-inequality-in-how-women-are-treated-for-pain">Medical bias</a> is widespread, with studies documenting that doctors give <a href="https://academic.oup.com/painmedicine/article/13/2/150/1935962">less pain medication to Black patients</a> than to white patients for the same conditions. On average, <a href="https://www.bbc.com/future/article/20180523-how-gender-bias-affects-your-healthcare">women have to wait months or years longer</a> than men to get an accurate diagnosis for the same conditions. This impacts the data that is collected and will be used for AI. <a href="https://aclanthology.org/D17-1323/">Multiple research studies</a> have shown that AI not only encodes existing biases, but <a href="https://arxiv.org/abs/1901.09451">can also amplify their magnitude</a>. At heart, these biases often pivot on not believing marginalized people about their experiences: not believing when they say that they are in pain nor how they report their symptoms.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-02-20-ai-medicine/hirsch2.jpg" class="img-fluid figure-img" style="width:75.0%"></p>
<figcaption>From Aubrey Hirsch’s powerful comic, “Medicine’s Women Problem”, https://thenib.com/medicine-s-women-problem/</figcaption>
</figure>
</div>
<p>On a deeper level, ignoring patient expertise limits what hypotheses are devised, and can slow research progress. These issues will propagate into AI, unless researchers seek ways to include meaningful patient participation. These problems are not unique to any one country or any one type of medical system. Patients across the USA, UK, Australia, and Canada (the 4 countries I am most familiar with) are all well-documented to experience these issues.</p>
</section>
<section id="being-honest-about-the-risks-and-opportunities" class="level2">
<h2 class="anchored" data-anchor-id="being-honest-about-the-risks-and-opportunities">Being Honest About the Risks and Opportunities</h2>
<p>While AI holds transformative possibilities for medicine, it is important that we are clear-eyed about both the risks and the opportunities. Many talks on Medical AI give the impression that the only thing holding medicine back is a lack of knowledge. The many other factors that influence medical care, including the systematic disregard for patients’ knowledge of their own experiences, are often ignored in these discussions. Ignoring these realities will lead people to design AI for an idealized medical system that does not exist. There is already a clear pattern in which AI is used to centralize power and harm the marginalized. In medicine, this could lead to patients, who are already disempowered and often disregarded, having even less autonomy or voice.</p>
<p>In my view, some of the most promising areas of research are participatory approaches to machine learning and patient-led medical research. The <a href="https://participatoryml.github.io/">Participatory Approaches to Machine Learning workshop at ICML</a> included a powerful collection of talks and papers on both the need and the opportunities for designing systems with greater participation of those impacted. AI ethics work on topics of contestability (building in ways for participants to contest outputs) and actionable recourse are necessary. For medical research more generally, the <a href="https://patientresearchcovid19.com/">Patient Led Research Collaborative</a> (focused on Long Covid) is an encouraging model. I hope that we can see more efforts within medical AI to center patient expertise.</p>
</section>
<section id="for-further-reading-watching" class="level2">
<h2 class="anchored" data-anchor-id="for-further-reading-watching">For further reading / watching:</h2>
<ul>
<li><a href="https://www.youtube.com/watch?v=vVRWeGlMkGk&amp;list=PLtmWHNX-gukLQlMvtRJ19s7-8MrnRV6h6&amp;index=3">AI, Medicine, and Bias: Diversifying Your Dataset is Not Enough</a> - my keynote talk at the Stanford AI in Medical Imaging symposium</li>
<li><a href="https://www.bostonreview.net/articles/rachel-thomas-medicines-machine-learning-problem/">Medicine’s Machine Learning Problem</a> - my Boston Review article on medicine and AI</li>
<li><a href="https://www.bbc.com/future/article/20180523-how-gender-bias-affects-your-healthcare">‘Everybody was telling me there was nothing wrong’</a> - BBC article by Maya Dusenberry</li>
<li><a href="https://thenib.com/medicine-s-women-problem/">Medicine’s Women Problem</a> - comic by Aubrey Hirsch</li>
<li><a href="https://rachel.fast.ai/posts/2023-05-16-ai-centralizes-power/">AI and Power: The Ethical Challenges of Automation, Centralization, and Scale</a> - a deeper dive into how AI centralizes power</li>
<li><a href="https://www.hrw.org/news/2021/11/10/how-eus-flawed-artificial-intelligence-regulation-endangers-social-safety-net">How the EU’s Flawed Artificial Intelligence Regulation Endangers the Social Safety Net</a> - Human Rights Watch</li>
</ul>
<p><em>Thank you to Jeremy Howard and Krystal South for providing feedback on earlier drafts of this post.</em></p>
<p><br><br></p>
<p>You can subscribe to be notified of new blog posts by submitting your email below:</p>
<script type="text/javascript" src="https://campaigns.zoho.com.au/js/zc.iframe.js"></script>
<iframe scrolling="no" frameborder="0" id="iframewin" width="100%" height="100%" src="https://zcmp-pd.maillist-manage.com.au/ua/Optin?od=11d0c075b7ed49&amp;zx=11a17553b1&amp;tD=156971d471c7b09&amp;sD=156971d471c7cb3">
</iframe>


</section>

<p><br><br><i>I look forward to reading your responses. Create a free GitHub account to comment below.</i></p> ]]></description>
  <category>ethics</category>
  <category>machine learning</category>
  <category>science</category>
  <guid>https://rachel.fast.ai/posts/2024-02-20-ai-medicine/</guid>
  <pubDate>Mon, 19 Feb 2024 14:00:00 GMT</pubDate>
  <media:content url="https://rachel.fast.ai/posts/2024-02-20-ai-medicine/zhou-cover-2.jpg" medium="image" type="image/jpeg"/>
</item>
<item>
  <title>How Immune Cells Communicate</title>
  <dc:creator>Rachel Thomas</dc:creator>
  <link>https://rachel.fast.ai/posts/2024-02-06-cytokines2/</link>
  <description><![CDATA[ 





<p><em>This post is part 2 in a series. Be sure to <a href="https://rachel.fast.ai/posts/2024-01-23-cytokines1/">read part 1 here</a>.</em></p>
<p>A key obstacle hindering medical research for a range of diseases is our lack of understanding of the immune system. <a href="https://www.freethink.com/health/human-immunome">Harvard Professor Wayne Koff described</a> his decades of HIV research, “<em>slowly, over time, we began to see that we understood a lot about HIV at the molecular level, but we didn’t know anything about ourselves. And the reason is that the immune system is incredibly complex</em>.”</p>
<p>Many diseases involve the immune system overreacting, underreacting, or having a mistargeted reaction. Deepening our understanding of the immune system is crucial for better understanding and treating a range of diseases. In <a href="https://rachel.fast.ai/posts/2024-01-23-cytokines1/">part 1 of this series</a>, I shared about the networks by which immune cells communicate and coordinate via the use of protein messengers called <em>cytokines</em>. These networks are complex and not fully mapped. Studying these networks is part of determining how and why the immune system reacts as it does. Here, I will share some more research on immune cell-cytokine networks.</p>
<section id="immune-cells-coordinate-like-a-social-network" class="level2">
<h2 class="anchored" data-anchor-id="immune-cells-coordinate-like-a-social-network">Immune Cells Coordinate Like a Social Network</h2>
<p>Immune cells and the cytokines they use to communicate can be analyzed using the same techniques as are used for social networks. Networks can be represented as matrices, in which each column/row represents a node, and the matrix values represent the presence or strength of connections between them. This type of representation underlies how the network of the internet is represented for search engines such as Google. Matrix math is everywhere! I covered many applications of it in the computational linear algebra course that Jeremy Howard and I designed and taught at University of San Francisco and fast.ai.</p>
<p>The <a href="https://www.nature.com/articles/ni.3693">paper ImmProt</a> sampled immune cells from the bloodstream of humans and used proteomics (the large-scale analysis of proteins) to study immune communication networks. The researchers considered 28 cell types in both steady and activated states, and over 10,000 proteins. ImmProt used several classic math techniques to analyze the network, including principal component analysis, supervised clustering, and Lasso regression analysis.</p>
<p>Cytokines bind to receptors on the surface of the cell that is receiving them. Receptors for some cytokines are found on many different types of immune cells (these receptors are said to be “broadly expressed”), while receptors for other cytokines may be expressed on a very limited number of cell types. The researchers found that there were two types of cytokine communication patterns: there were either many types of cells sending and few types receiving for a given cytokine, or there were few types of cells sending and many receiving. This is a directed, asymmetric information exchange.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-02-06-cytokines2/rieckmann-6h.jpg" class="img-fluid figure-img" style="width:65.0%"></p>
<figcaption>A single type of message (in this case, CCL5/3) can be sent to many different cells (figure 6h from ImmProt)</figcaption>
</figure>
</div>
<p>The authors looked at the number of in-going and out-going communication between various cell types, and determined how much cross-talk there is between different clusters of cell types. In the above diagram, the outer circle lists cell types and the middle circle lists receptor and cytokine gene names. This picture shows that eosinophils and basophils (two cell types that evolved primarily to address parasites) send the cytokine CCL5/3 to a variety of other cells that receive the message using the receptor CCR3. The color of the lines indicates the type of cell sending the cytokine.</p>
<p>While many of the relationships found were well established in existing scientific papers, they also discovered new relationships, and validated two of these new relationships experimentally: one in-going and one out-going. <a href="https://www.nature.com/articles/ni.3693">The authors wrote</a>, “<em>The immune system displays an almost infinite diversity in cell types and activation states, and achieving completeness in both dimensions is particularly challenging for human cells</em>.”</p>
</section>
<section id="ai-does-not-replace-other-ways-of-knowing-about-the-world-benchtop-research-is-still-crucial" class="level2">
<h2 class="anchored" data-anchor-id="ai-does-not-replace-other-ways-of-knowing-about-the-world-benchtop-research-is-still-crucial">AI does not replace other ways of knowing about the world: benchtop research is still crucial</h2>
<p>AI is a powerful tool, but it does not replace other ways of knowing about the world, including <a href="https://rachel.fast.ai/posts/2022-06-01-qualitative/">qualitative research</a> (as I wrote about together with Louisa Bartolo) and benchtop quantitative experiments. Machine learning and AI techniques are great at learning patterns in existing data. However, they are limited by what types of data have been collected so far, and it is still crucial to continue running new laboratory-based experiments. For this reason, the recent <a href="https://www.nature.com/articles/s41586-023-06816-9">Immune Dictionary paper</a> was a valuable contribution to the field. The researchers gathered direct measurements of over &gt;1,400 cytokine-cell type pairings in an in vivo experiment.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-02-06-cytokines2/immune-dictionary-1.jpg" class="img-fluid figure-img" style="width:60.0%"></p>
<figcaption>Steps for creating a dictionary of immune gene expression signatures (Figure 1A from Immune Dictionary)</figcaption>
</figure>
</div>
<p>This study was significant because it considered all major immune cell types and all major cytokines (studies typically look at only 5 immune cell types) and was conducted in vivo (not in a culture). The researchers injected 86 different cytokines into individual mice and measured the responses of 17 types of immune cells in the mouse lymph nodes. They used single-cell RNA sequencing, which is more direct than considering ligand and receptor expression association data. They released their findings as an “Immune Dictionary” and developed <a href="https://www.immune-dictionary.org/app/home">accompanying computer software</a>. This software was then used to identify cytokine networks in tumors after checkpoint blockade therapy (a cancer therapy that helps to reactivate exhausted immune cells).</p>
<p><a href="https://www.nature.com/articles/s41586-023-06816-9">One key finding</a> was that the responses induced by cytokines are highly cell-specific. Rarer types of immune cells expressed a greater number of cytokine types than more common cell types. One rare cell type, typically not even covered in immunology courses, was found to express the highest number of distinct cytokines, influencing nearly every other cell type. This result suggests that rare immune cell types are crucial for cell-to-cell communication.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://rachel.fast.ai/posts/2024-02-06-cytokines2/immune-dictionary-software.jpg" class="img-fluid figure-img" style="width:75.0%"></p>
<figcaption>Software for the dictionary of 17+ immune cell types responding to 86 cytokines in vivo</figcaption>
</figure>
</div>
<p>Another important finding was that every immune cell type could be polarized into multiple states depending on the combination of cytokines that it received. The <a href="https://www.broadinstitute.org/news/new-dictionary-immune-responses-reveals-far-more-complexity-immune-system-previously-thought">two key researchers were surprised</a> at this level of plasticity and complexity in immune cell types, with even the most well-studied cytokines inducing more complex responses than previously expected.</p>
</section>
<section id="an-ongoing-area-of-research" class="level2">
<h2 class="anchored" data-anchor-id="an-ongoing-area-of-research">An Ongoing Area of Research</h2>
<p>This is an exciting area of research to follow. There is still much work to be done in more thoroughly understanding the intricate and complex ways immune cells coordinate with one another to respond to threats. Immune cell-cytokine networks are a great illustration of the power of interdisciplinary work, since NLP, mathematics, and bench top immunology research all provide important insights into the problem. And this is just one of several important problems in immunology where AI is being applied!</p>
<p>You can subscribe to be notified of new blog posts by submitting your email below:</p>
<script type="text/javascript" src="https://campaigns.zoho.com.au/js/zc.iframe.js"></script>
<iframe scrolling="no" frameborder="0" id="iframewin" width="100%" height="100%" src="https://zcmp-pd.maillist-manage.com.au/ua/Optin?od=11d0c075b7ed49&amp;zx=11a17553b1&amp;tD=156971d471c7b09&amp;sD=156971d471c7cb3">
</iframe>


</section>

<p><br><br><i>I look forward to reading your responses. Create a free GitHub account to comment below.</i></p> ]]></description>
  <category>machine learning</category>
  <category>science</category>
  <guid>https://rachel.fast.ai/posts/2024-02-06-cytokines2/</guid>
  <pubDate>Mon, 05 Feb 2024 14:00:00 GMT</pubDate>
  <media:content url="https://rachel.fast.ai/posts/2024-02-06-cytokines2/rieckmann-network.jpg" medium="image" type="image/jpeg"/>
</item>
</channel>
</rss>
