Steadcast
This Week in Virology cover art
This Week in Virology

TWiV 1355: Four billion years of AI training data

September 6, 20261h 38m · 14,833 words

Show notes

TWiV reviews a landmark paper reporting the first generative design of complete viral genomes using AI, producing 16 viable bacteriophages that overcame antibiotic-resistant bacteria where natural phages could not, and a study suggesting that recombinant shingles vaccination reduces the risk of cardiovascular events. Hosts: Vincent Racaniello, Kathy Spindler, and Brianne Barker Subscribe (free): Apple Podcasts, RSS, email

Highlighted moments

Previous similar work had been done on the scale of, um, generating, uh, different individual genes or even gene circuits. But this is at the scale of whole genomes. So that's what makes this difference.
31:12
The recombinant vaccine, that is the shingrex, was associated with a 9% decrease in the cardiovascular burden over seven years.
9:44
When you looked at their homework completion times, those students completed their homework much faster. But then when you looked at how well the students retained the information on an exam later, they had 25% lower exam scores.
1:23:40

Transcript

Podcast introduction and weather updates

0:00This Week in Virology, the podcast about viruses, the kind that make you sick.

0:10From Microbe TV, this is TWIV, This Week in Virology, episode 1355, recorded on September 4th, 2026. I'm Vincent Racaniello, and you're listening to the podcast all about viruses. Joining me today from Ann Arbor, Michigan, Kathy Spindler. Hi, everybody. Here it's now just 79 degrees and cloudy, whereas between sometime last night

0:41and this morning we had an inch and a half of rain, and yesterday we had an inch and a half of rain, and part of that was during a tornado warning, so that everything went off. My phone went berserk. My NOAA weather radio went berserk. I got alerts from UM and from the county, and the tornado siren a block away went off, and so I just ran to the basement and stayed there for 45 minutes.

1:09Wow. Yeah, I mean, it was like pretty serious, because, so I don't know that anything touched down. I've tried to figure that out, but it did say that there was rotational stuff seen on the radar. They didn't say stuff, but they didn't say anything much more. Rotation. I wonder what word you would use. Rotational activity? Yeah, something like rotational activity was seen on the radar. Yeah. Also joining us from Madison, New Jersey, Brianne Barker. Hi. Great to be here. It is 83 Fahrenheit, which is 28 Celsius. My app says it's cloudy,

1:44but my eyes say it's sunny, and I'm glad that I did not have a similar tornado warning recently. Yeah. Here in New York, it's 30 C and cloudy. We don't have rain today, but it's been raining a lot here, and it's going to rain in the next week. If you enjoy these podcasts about science, we'd love to have your support. You can go to microbe.tv slash contribute for the various

Obituary for Nobel laureate Ada Yonath

2:13ways that you can help us out. I have one news item today, and that is a New York Times article, which, of course, is paywalled, and I don't have my login here. I've got this stupid thing on the bottom. Does anyone have a subscription that could read the headline to read the title? Go ahead, please. Ada Yonath, chemist and first Israeli woman to receive a Nobel, dies at 87. Her mapping of the ribosome, which led to new designs for antibiotics, faced years of derision, as many scientists saw it

2:46as a dead-end effort.

2:51Yeah. I just can't believe that you would deride it, right? Just anything. It happens all the time. Oh, you're not going to learn anything. This is already known. We don't need to. Why? Why do sciences do this? Yeah. Well, and think about how many examples we could all come up with of things that were really cool that no one would have expected was going to be really cool. I think we should all learn at this point. There's no dead end. Just let people do what they want, okay? If it's not going to work out, they'll figure it out and

3:24they'll stop. If it's fine, it's fine. If it's great, it's great. Just leave them alone. Don't say these. It's like mixing science and politics, you know? It's not good. Don't mix these human emotions with science. Anyway, she got a Nobel Prize for that. It's a great structure. We've learned a lot from the structure. We continue to learn a lot. And I just don't understand. It was not easy to do, right? The ribosome is huge. It's made of protein and RNA. And it wasn't easy

3:54to do. And she figured it out. So that's very cool. Yeah. It says she used an ingenious approach with stubbornness. She turned to extremophile bacteria, reasoning that their ribosomes might be more stable and therefore easier to crystallize. And she developed techniques to freeze ribosomal crystals at low temperatures, preserving their structure long enough for analysis. So yeah, she got the prize in 2009. And she shared it with Venkatraman Ramakrishnan and Tom Stites that same year.

4:33Right. And I think Venkatraman Ramakrishnan has written a book about it, which is someone picked, because it's a really engaging book, telling the story of solving the structure.

4:45Yeah, go ahead. Two out of three of them have died, because Tom Stites also died. Yeah. She was chosen among the three to deliver the Nobel acceptance speech because of her resolve and her pioneering methodology. So that's pretty cool. I love when people have resolved, they get criticized, and then they do it. Yeah. Just think it's just a nice resolution, right? Yeah. Success is the best reward. Yeah.

Shingles vaccination and cardiovascular risk

5:15All right. On to our science day. We have some very interesting science for you today. We have a nature medicine paper, which is a brief communication, recombinant shingles vaccination and the risk of cardiovascular events. So that's an interesting title, right? It's like not saying it has one effect or the other. It's just saying we're going to talk about this, right? I think that's okay. Right. And so Kathy, are you going to do the authors or no? I can do the authors.

5:50Sure. It's okay. Fabiana Corsizueli, Fang Li, Rachel Up the Grove. That's an interesting name. John Todd, Betty Roman, Paul Harrison, Maxime Takei. They are from the University of Oxford, University of Birmingham, the Oxford NHS Foundation, the IHR Biomedical Research, the University of Oxford University of Oxford. Okay. Mostly all in the UK. Okay. Kathy, you have a summary?

6:22Sure. So this is a brief communication. And just to remind you, shingles is a sequelae of herpes zoster virus infection. That virus causes chickenpox and then it becomes latent. And when it reactivates from latency, it can cause a pretty painful condition. And so there have been two shingles vaccines in our lifetime. Up until 2017, there was the live attenuated vaccine called Zostavax. Then a recombinant vaccine was introduced in 2018. And there was a very rapid

6:57switch to that vaccine in the United States. And it is more effective. And it is called Shingrix. And I've always had trouble remembering which one had which name. So I got myself a mnemonic. It's not perfect. But remembering that Shingrix ends in rix, which kind of sounds like rex. I think of rex as being recombinant. So it's the recombinant one. And because it's superior to the Zostavax, like rex, it would be the king of the two vaccines. So not a perfect mnemonic, but maybe it'll work for some of

7:31you. I'm hoping it'll work for me. Okay. So this paper is taking advantage of this rapid switch to get around a problem that is that there might be potential bias in determining whether one or the other vaccine reduces the risk of cardiovascular disease. And this design was used previously in demonstrating that the shingles vaccine may lower the risk of dementia. And there are a couple of previous TWIVs where we've talked about that, such as TWIV 1207 and

8:02TWIV 1263. Anyway, so that approach took advantage of this rapid shift in the vaccine usage. And they had a very large sample size, over 36,000 people from the United States, and they used electronic health records. Okay. So they were using these electronic health records from the United States, even though they're all from the UK. And they looked at adults over 60 who were vaccinated either immediately before the

8:33transition to the recombinant vaccine or after. And then they examined a composite scheme of cardiovascular endpoints. And the three main ones that we're going to consider are ischemic heart disease, heart failure, and ischemic stroke. And they looked at some other cardiovascular endpoints that will come up later. So they measured these in these people, either who got shingrex in the

9:04six months before shingrex was used. In other words, they got the zostavax. And then they looked one year later, to compensate for possible seasonal effects. So they were looking at people from the same time of year. And they looked at those in a six month period, when they determined that 94% of the adults were now getting this new shingles vaccine called shingrex. Okay, so in this paper, they collect a ton of data, they do a ton of statistics and analyses. And I don't

9:40know how deeply we're going to get into those numbers, but they might get confusing. So I want to try to summarize them for you here. So here we go. The recombinant vaccine, that is the shingrex, was associated with a 9% decrease in the cardiovascular burden over seven years. And that was significant in men and women for these composite of the three types of the endpoints that I mentioned, ischemic heart disease, heart failure, and ischemic stroke. And then they broke

10:11it down. And when they looked at men and women for ischemic heart disease or heart failure, there was either a 10% or a 12% decreased burden, respectively. And when they looked at the third endpoint, ischemic stroke, they only saw a statistical reduction. And that was significant only in men. And that was 12%. And at some point, they also looked at another endpoint, atrial fibrillation, which we commonly call AFib. And that resulted in a 7% decrease. So the overall message is the

10:49individuals who received this recombinant vaccine, the shingrex vaccine, were at a lower risk of cardiovascular events over the next seven years than those who predominantly received the live attenuated vaccine, the earlier vaccine that was available. Okay. So there are a lot of caveats. And I think we'll talk about some of those. But the authors recognize that. And they say these results justify clinical trials and mechanistic studies to investigate potential cardioproductive

11:20effects of shingles vaccine. So they don't know the mechanism. And as I said, there are a lot of caveats. So that's my summary. Well, it's interesting that most people over, or I should say many people over 50 get shingrex now, right? So whether or not this is real, you got the vaccine, right? So but we'd like to know how it works for sure. Yeah. I think the question to me of how it works is really fascinating. It's interesting. Yeah. We'll get into that. So there has been some previous work

11:52suggesting that this vaccine, the shingrex in particular, may reduce the risk of cardiovascular disease. But those studies were done comparing vaccinated and unvaccinated people. And this is a problem always because of what's called the healthy vaccinee bias, right? People who choose to get vaccinated differ really from people who don't in ways that you often can't capture in your analysis. And so what they did here was a similar, it's the same group that did the previous study on

12:24dementia. They take advantage of this rapid switch between the attenuated vaccine and the recombinant vaccine, which happened in October 2017. And so this should, they say, mitigate biases. But there's still going to be some anyway. You can't get rid of them entirely. But it will mitigate many of the biases that distinguish because both populations got vaccinated, right? So in theory, it should take away the healthy vaccinee. But other things can happen also. I am not familiar enough with this.

12:58I wonder if, was the switch as rapid outside the United States? Kathy mentioned that we've got a bunch of authors from the UK who are looking at US electronic health records. And I wonder if the abrupt switch was a US thing. I have no idea. I think there were some studies in other countries which had a similar switch, a rapid switch. I don't know if it was everyone, though. I don't know offhand. Okay. So they are comparing these cardiovascular outcomes in adults over 60 years

13:32with who either got the live attenuated vaccine or the Shingrix vaccine. And they do it in the same period in the year, April to September, to eliminate seasonal effects, right?

13:50And as Kathy said, the primary outcome is a composite cardiac, cardiovascular endpoint, ischemic heart disease, ischemic stroke, and heart failure. Now, they use what's called propensity score matching. They say, we use this. And then, of course, they don't define it because it would take half the paper. So I had to look it up because I don't know what that means. But basically, when you're comparing, this is an observational study, right? We're just looking at things that have happened. And they could differ, as we have already said. We got rid of the vaccinated person issue, but there can be other differences as well. And so what you do is you

14:25make a propensity score. It's a probability between zero and one that represents each person's likelihood of being in one group versus the other based on all these other characteristics. And they use 83 variables here, including age, sex, ethnicity, the diagnoses, medications, everything. There's a whole table with all of these things that they use. So you make a score for each person. And then you try and match people in the two groups, right? The Zastavax versus the Shingrix. So now you have a

14:56statistically balanced group. And that's sort of like what you would do in a clinical trial, right, in theory. So they match 36,000 people vaccinated in April, September 2018. We'll get Shingrix versus 36,000 people vaccinated in the year before who got Zastavax. And after matching all these 83 covariates, we call them, achieved standardized mean differences of 0.1 or less, which is basically the threshold for considering that these groups are well-balanced. Now, of course, what you don't

15:33measure could make a difference also, right? And they did not measure socioeconomic status, smoking, diet, physical activity, et cetera. And so then that's why another value that they did, which we'll get to in a moment, is important. It's called the E-value. And they measure that as well. So anyway, I already told you the number of people who they matched. And they find that the individuals in the group who mostly received the recombinant vaccine were at lower risk of cardiovascular events

16:05over the next seven years than people who received Zastavax. And that's translated into 9%. They use this metric, more time living, diagnosis free. 9% more time living, diagnosis free. And then they calculated an E-value. So now it's another thing. I don't know what an E-value is. So what is an E-value? It's a way of quantifying how much of a threat the unmeasured confounding things

16:37actually pose to the conclusion. So we have 83 variables that we measure, but there are a lot that aren't measured. So it's a risk ratio. I was going to say, another way of thinking that I think about it is it's basically how big of an effect would some unmeasured variable have to have in order to give you this effect. That's right. So if there was some mystery variable, how strong of a mystery variable would it have to be to have resulted in these data? Right. So it's a risk ratio. So an E-value

17:08of 1.42, which they get, means an unmeasured confounder would need to be associated with both the vaccine type and the outcome by a factor of at least 1.42. And so E-values can range from 1 up. So 1 mean any confounder could explain the result. So it means the data not so good. But high values 4.5 mean only a really strong confounder could explain the result. So 1.42 is considered low. So it's not a great E-value. So that's one of the limitations of this study.

17:43But this is just meant to be a suggestion and say, now it's worth spending money on a clinical trial. Right. Well, and if you sort of just think about it logically, like, yes, they did not look at socioeconomic status here. But I have a hard time imagining there's going to be a dramatically different socioeconomic status between people who got a shingles vaccine over six months in 2017 and people who got a shingles vaccine over six months in 2018. Or a nutritional difference.

18:14They just give, there are other things that they don't mention. No, of course. There are lots of things that they don't mention. But if you actually sort of think about the two populations they measure, it's hard to imagine some other big variable. Obviously, there could be one. There only has to be one. And then that 1.42, you know. Right. Yeah. Well, it has to be pretty high, like three or four or five, right? I mean, if it's 1.42... Well, if it's 3.4 or 5, then it's unlikely to be affected by a... You'd need a really strong confounder. But 1.42 is close to one where almost anything could affect it, right?

18:49I thought it was more robust findings have higher E-values. Yeah. Right. And the lowest possible is 1. So 1.42 is not that high. Yeah. 1 is the lowest where any weak confounder could tank your data, basically. Right. Okay. So as Kathy said, this association is in both females and males. And the lower incidence of this composite effect was found... Oh, sorry. Yes. Males also have a lower risk of ischemic

19:21stroke. All right. It's also a lower risk of AFib, but not myocarditis, peripheral arterial disease, hemorrhagic stroke, or transient ischemic attack. And this association also was seen when accounting for death as a competing risk. No association with... And there's no association with non-cardiovascular death, right? So if you have a cardiovascular-caused death, then it still holds. In the year before vaccination, these two cohorts didn't differ in their preventative

19:57healthcare or in their risk of cardiovascular events. There was no difference between cohorts doing follow-up in terms of these health indicators. And individuals who predominantly received the recombinant vaccine were at lower risk of shingles, as you might expect, than the negative controls, which they had some in this study. And that's good. That's like an internal control. There's no sign that secular trends in diagnostic rates confounded the results. Those, for example,

20:29who got tetanus, diphtheria, pertussis vaccine had a very similar risk of cardiovascular events as those who received it during the same month. So getting any vaccine doesn't do this. It's specific, at least for shingrix. Well, yeah. And that's the thing that sort of is key, is it's specific to shingrix, even when it's a different vaccine against shingles. So it's specific somehow to this recombinant shingles vaccine. It's not even about just any shingles vaccine. It's about this

21:02specific recombinant shingles vaccine. Now, what's interesting is this attenuation. They publish a Kaplan-Meier curve, which is basically the incidence of cardiovascular events with time, right? And you can see that it starts to drop, right, really soon after vaccination. It's present in the first half of follow-up. And then the second half, it starts to attenuate. And the hazard ratios are no longer significant. They tend to plateau. So that's interesting. I don't know. It happens

21:36really quickly and not really consistent with the vaccine effect, but we'll talk about that later. Anyway, this risk reduction is meaningful, right? So if this were confirmed in a clinical trial, a 0.8% difference in incidence of ischemic heart disease and failure among adults over 60 would be hundreds of thousands of cases prevented in the U.S. So with a little number that so many people hear, it's significant. You know, this would be good. So they see a stronger association

22:11for ischemic heart disease and heart failure than for ischemic stroke. They say this may indicate a cardiac-specific pattern. And this is interesting because both strokes and ischemic heart disease often result from injury to the vascular endothelium. And so maybe there are structural differences between the cerebral and coronary arteries that modulate the effect of the vaccine.

22:37One thing that... So I read this sentence and I didn't know what it meant, but probably because I'm just not thinking. There's a bidirectional association between atrial fibrillation and ischemic heart disease. So what that means is that AFib can cause ischemic heart disease and ischemic heart disease can cause AFib. So each condition increases the risk of the other. And so this bidirectional relationship works in their favor. It kind of validates their results because if the vaccine

23:09protects against ischemic heart disease, you would expect AFib rates to fall as well. And in fact, that's what they see. So that's a kind of validation of their conclusions. All right. So that's the findings. And then we're left with what's the mechanism. And, you know, they say prevention of shingles might partly explain it because there is an association between shingles and cardiovascular risk,

23:41but they think it's unlikely for three reasons. First, the difference in shingles cases between the vaccines is less than half of a percent. So that probably is not explaining it. Second, protection would be expected to emerge gradually rather than having this early divergence, right? And third, a previous experiment that was done comparing the live attenuated shingles versus no vaccination showed no risk of reduced risk of heart disease. And what they suggest is that the

24:18adjuvant in the vaccine, which is ASO1 adjuvant, triggers epigenetic and functional reprogramming of monocytes, reduced IL-6, and that may be associated with lower cardiovascular risk. And those changes might be time limited, right? Which is maybe why we see attenuation of the effect. What do you think about that, Brianna? Yeah. So I think that that makes a lot of sense to me that there's potentially some, you know, innate immune inflammatory, whatever we're going to call it, change as a result of

24:51getting A501 that is cardioprotective. And so, and it also makes sense to me that this would sort of be the first we would totally hear about it because you probably don't do, you know, like a A501 alone versus A501 with adjuvant in very many people. You probably need a lot of, you start, you know, when they have 72,000 people here, you probably need those large studies and we probably don't have that. Um, A501 in terms of approved vaccines is only in Shingrix. Um, it's in terms of vaccines that are

25:30given frequently. It also seems to be in the, um, malaria muskyrix and in arexvi, um, but those haven't been given as broadly, um, as Shingrix. So that would also be a reason why we wouldn't have seen that. Um, and to me, the, you know, that tells me that I really want to see more information about this adjuvant and more cool studies. Yeah. Yeah. Okay. So a couple of reservations. I mentioned the E value. It's 1.42. So it's kind of low. And

26:03so a confounder effect may explain all the differences. Uh, the, the early divergence of the Kaplan-Meier curves, they actually note this themselves. It doesn't really make sense for a true cardiovascular effect. And they acknowledge this and it's an electronic health record study. So all you're studying are recorded diagnoses, right? And unrecorded or missed, of course, because you don't have them in the, in the records and they may contribute to this as well.

26:35Um, but they say the last sentence, given the global burden of cardiovascular disease, these findings, if confirmed in clinical trials and mechanistic studies would have important implications for public health. Meanwhile, if you turn 50, just get Shingrix. You probably had chicken pox and, um, it's good for that at least. And so, yeah. So I was going to say, if, if, if true, if, if, uh, more data support this conclusion, that would mean that Shingrix has a

27:06third benefit for you, um, of prevention of shingles, maybe prevention of dementia, and now maybe also cardioprotective effects. So in the, the, the moral of the story should be get the vaccine.

27:21Okay. Maybe switching vaccines to that adjuvant. Yeah. ASO1. It's very interesting. Could be the adjuvant. Did the, um, do you remember the paper we did a while ago, which showed that you could immunize with an irrelevant antigen with some adjuvant and you got broadly protective antibodies? Was that ASO1? Um, I don't think so. So I'm pretty sure that one was from Balipulandrin's lab.

27:54Was that a TLR agonist? Anyway, it's not ASO1, so it's okay. It is not ASO1. If you, let's see. Yeah. This is from Balipulandrin, right? It's from Balipulandrin. Um, the study that we talked about uses their, I think it's basically their own unique, um, adjuvant.

28:28GLA3M052LS. Okay.

Generative design of bacteriophages with AI

28:31All right. Stay tuned. Anyway, here's our big paper for today. And it's not that it needs to be confirmed by a clinical trial or anything. This is big and the results are really interesting. And I have to say, Nels and I did this on Twevo in January when it was a preprint. We both got it totally wrong. Because if you go listen to that, it's going to be totally different from what I say about it today. I don't know about Kathy and Brianne, but I just totally misinterpreted it because I didn't understand it properly. But now I spent a long time understanding it. And I think

29:06this is really an amazing paper. So it's a, it's a article in science. It's a generative design of bacteriophages with genome language models. It comes from Stanford University and the Broad Institute of MIT and Harvard. And the authors are Samuel King, Claudia Driscoll, David Lee, Daniel Guo, DT Merchant, Garrick Brixey, Max Wilkinson, and Brian He. Kathy, you have a summary?

29:39Sure. Um, several in a way, because, um, the science format itself now provides a one-page summary that I think is what's in the print version of the journal. And then you can go online and get the full version. And this paper, uh, as Vincent said, is big in a couple of ways. It's big because there's just a ton of information and, uh, there's a whole lot of supplemental material with a lot of text in it. And, uh, there's also a perspective that's written by independent authors. So there's a lot of things that you could find to read. And this story was front page news in the New York Times on

30:14August 7th above the fold. Now, I don't know how much longer we're going to be able to talk about something being on the front page and being above the fold because that relies on printing copies of the newspaper, but that's where it was. Uh, and, um, uh, so the headline there was in lab, AI designs viruses not found in nature. And the sub headline was hopes for advances in medicine and fears about pathogens. So that's another summary that, um, they designed these viruses not found in

30:48nature, uh, by AI. And there's hope that that can help advance medicine, but there's fear about it, uh, with respect to pathogens. So, uh, they used two preexisting language learning models that are based on DNA sequences. And they used AI to generate complete phage genomes. So this is the new thing. Previous similar work had been done on the scale of, um, generating, uh, different individual genes

31:21or even gene circuits. But this is at the scale of whole genomes. So that's what makes this difference. And they used a bacteriophage, PHYX174, which infects certain strains of E. coli, namely E. coli C. And as we'll later hear, um, another strain that they didn't know about E. coli W. And here's full disclosure. I did the first part of my doctoral studies on PHYX174. Uh, it's genome was completely sequenced in 1977 and people at the time said, well, so what else is there to do with PHYX now?

31:56Um, and here we are in 2026 and people are still using PHYX. Kind of want to say, na-na-na-na-na-na-na-na. Okay. But, um, anyway, okay. So they used PHYX as a design template. Um, and from that, they had AI generate thousands of genomes. And they did this so that the AI design, uh, had certain constraints that they built into it. Um, it had to be able to target E. coli C.

32:27and they, uh, collected many PHYX-like genome sequences to train it on. And they had to build some of their own tools to add in some of the constraints so that the AI would generate PHYX-like sequences. And they also specifically didn't train it on anything else, any other kinds of sequences, any sequences from human or animal or plant pathogens or other things. And that's an important constraint

32:57that they put on themselves in this research that helps build in a lot of safety. And, um, one interesting thing, I don't know if we'll get into it because if they discussed it in the discussion, I kind of missed it. They, they found that including the first four to nine nucleotides of the PHYX genome gave them better results. Um, a feature that, uh, as I said, we might come back to, but I don't know how they define the first four to nine nucleotides because it's a circular genome.

33:29Okay. I think just right of the origin. Yeah. Okay. Thank you. Um, so they generated thousands of these sequences or the AI did, and then the authors curated 300 of them. And again, if I'm blanking on how they got from their thousands to the 300, but we'll talk about that, I guess. And then they wanted to test them experimentally. And so to do that, I have to tell you a little bit of PHYX biology. Um, the virus packages this single-stranded genome that's circular DNA. It's about 5,400

34:02nucleotides long. So it's pretty small genome, but in its life cycle, it has a phase of the life cycle when it's double-stranded DNA during genome replication. And these are called replicative forms. And so you can generate viruses by introducing these double-stranded DNA replicative forms, uh, into E. coli, and they will produce the single-stranded DNA and package it. And so they showed that they could do this with wild type PHYX and, uh, they call this rebooting the

34:36genomes. Um, and I have in my notes, I roll here. I'm not sure where they got the term rebooting the genome, but they're basically just making double-stranded versions of the genome and putting that into the cells. So they produced double-stranded DNA fragments, assembled them based on, uh, that produced them based on what the AI predicted, and then they introduced them into the cells. And, uh, they had 302 of them, but somehow only 285 of them were assembled.

35:08And then they tested them for whether they could make phage that would inhibit growth in E. coli. And if the virus inhibits the growth in E. coli, it means it's killing the E. coli, and PHYX kills its host by lysing it. So it's saying that if it's, uh, inhibiting the growth, it's the viruses growing and productively infecting those cells. So they got 16 such viruses. So that's a, uh, ratio of about one in 20 of them. And then they characterize them. And then, uh, I don't know

35:42how much of that we'll go into, but, um, one of the interesting medicinal uses that's, uh, coming from this is that I think we've talked about it or heard about it, uh, in particular a lot recently, but phage therapy can fail because some of the bacteria are going to be resistant to the therapeutic phage. And those remaining bacteria that are resistant will replicate and damage and potentially kill the host. So that's a real problem in phage therapy is this inability to knock out all of the bacterial

36:16hosts. So the idea here is that, um, maybe, uh, they could generate some phages that would allow them to overcome this resistance. And they took a pool of their generated PHYX-like phages. Um, I can't remember whether it was all 16 of them or not, but anyway, and they found that this pool could overcome the resistance. So that kind of is proof of principle that maybe this could be, uh, of, uh, medicinal importance. And I think they even pulled out specific phages that, uh, seem to be responsible

36:53or amplified or enriched or something. And, and those in particular were good. And then the biosafety concerns, um, as I mentioned, they did this with some, uh, particular measures. Um, they only trained, the AI on phage sequences. Um, most likely they did this research at what I would call BSL-1. Um, I checked with, uh, Jolene to find out what the phage general, uh, biosafety levels are. And,

37:25and she said they go with the host. And so I'm assuming that the host here being E. coli C is probably BSL-1. I don't even know if things get categorized as BSL-0. So I'm just assuming it's BSL-1. And mouse adenovirus, for example, is BSL-1. But anyway, so, um, uh, so there are lots of, uh, concerns about if AI can do this for this set of training, what if you trained it on pathogen

37:57sequences? And so, um, I don't know how much of that we'll get into, but that's obviously concerned. So overall, the big take-home message is that these are the first genomes completely generated de novo through, um, an AI type mechanism and they proved to be infectious and functional.

38:20Right. So this, this, the importance here is that this is the first design of a complete genome. It's a bacteriophage, as Kathy has told you, it's relatively small, but it's a whole genome. And when you do a genome, you have to, you know, there's a coding region. There's a non-coding region. They're arranged in a certain way. Genes are going in different directions. There are structural motifs. There are interactions with the cell. There's the coordinated replication. All of these have to be encapsulated in the DNA. And we are now at the

38:55point where you can ask, can a genome language model with enough training data, and that's the key, could it learn evolutionary constraints from the training data and generate phage genome sequences? And the key here is that first DNA sequencing has made tons of genome sequences available by now. So we have the raw material, just like Chad GPT needs human words arranged in sentences to learn from. So now we have, uh, the DNA sequences of organisms and we have the computing power to do

39:29this. So their hypothesis is that the genome language models can, and if we train it properly, it can do this. So they use EVO-1 and EVO-2. So these are two, uh, basically all the sequences are publicly available and they can, they can refer to them. EVO-1 is trained at about 2.7 million prokaryotic and phage genomes, about 300 billion nucleotides, bacteria, archaea, and their viruses only.

39:59That's EVO-1. And then EVO-2 is 9.3 trillion nucleotides, 128,000 species, bacteria, archaea, fungi, plants, animals, humans, and their viruses. All right. It's only sequence. There is no biological knowledge associated with these sequences. Uh, but these are sequences from things that work. So the assumption is that you could, you can learn from it. And what, how do they train

40:33on this? They just say, what's the next base? That's it. They train the computer to get it right. And they check it. They go through many, many iterations of this, tens of millions of dollars of compute time. This has to be done in these big data centers, right? It takes a long time to do it. It's just like GPT. It's chained on DNA instead of on words. The, if you learn the probability of what word comes next, you can learn an entire language. You can in turn learn grammar. And it's

41:05the same with DNA. You can learn how to make an organism, right? So that's how they, they do this. And they, they initially said, can they actually, the way they prompt this is initially they could just give it, uh, the realm names like double-stranded DNA viruses, single-stranded DNA, and RNA viruses. And they, they threw it in there and they said, does it make sequences that at least resemble? Before we do anything big, does it make anything that resemble? And it does. It does a good

41:39job at making, uh, the pre-trained models can generate phage-like sequences and EVO2 was better. It has more sequences in it. Um, and they make sequences that are distinct from nature, but they do have a predicted coding density and they have genetic features of phage genomes. And the proteins look like proteins. They look like phage proteins as well. And they have, and the proteins have low sequence identity with the known proteins, which means they are novel. So these pre-trained EVO module models can design biologically realistic phage and they introduce

42:15diversity that you don't see in natural evolution, which means you don't see in the databases, right? We only know what, what you train these things on. Yeah. I wonder, um, one thing they don't really mention here, and I don't know enough about this. So Vincent, you might have, have come across this. Um, they, I kind of wonder, so they're only asking these models to give them basically new phase sequences. Right. And so, which are, as Kathy said, at least in this case, about five KB.

42:49Right. So how, how challenging, computationally heavy, time heavy would it be as you're trying to get longer genomes? Yeah. It would be challenging because then the, you have, you have this window of, of bases that it looks at and it gets harder and harder as it gets longer and longer. Yeah. I don't know the exact numbers. Right. But this way, one reason they start small. Right. Because this is tractable. Right. Okay. So they have this workflow, then they're going to make, have the computer

43:25generate sequences. Uh, they're going to, they're going to apply, uh, filters on the sequences, which I'll tell you about. And then they're going to experimentally validate the genomes, as Kathy said. Okay. Uh, so they use equal IC and Phi X one 74. They find what they call fine tune Evo one and two with another data set of 15,000 micro vira D sequences. So the Phi X one 71 is in the micro vira day family and they retrain the model. So it's already been trained on Evo one and two. They retrain

43:58it on these 15,000 sequences, which is a lot less computational activity. It's a small amount. But it basically gives you more likelihood that you're going to get a Phi X like genome. And these genomes all start with the same consensus nucleotides. And so they make a prompt, which is basically four to nine basis with that. It's got micro vira day, and then it's got four to nine bases. And that's enough to give it a Phi X like sequences. Um, and, and, um, if you don't use that,

44:33then you get too much other stuff that that's not Phi X like. Yeah. I found that surprising and they don't talk about it and maybe, maybe it's not surprising, but you know, they say the genomic start sequences of all the micro vira day sequences were highly diverse. All the Phi X one 74 like genomes started with the same consensus nucleotides. Right. I just found that kind of surprising.

45:02Yeah. Well, this is, um, actually the origin of replication, which is we'll see in the, in the genomes that they made completely conserved, a hundred percent conserved. So maybe that's why the, okay, the prompt works. Okay. I didn't, I didn't pick up on that, but what they call the start of the genome is part of the replication. Yeah. Origin. Okay. It's part of the origin. Yeah. Uh, so then they did some filtration. Um, they wanted to

45:32make sure the sequence quality was right. They wanted to make sure that the tropism was specific for E. coli C and they want to make sure it had diversity. So they filter out, they say they filtered out all sequences containing non-nucleotide characters. And you may think what's a non, well, in the database, there are a lot of N's or R's or, um, Y's, you know, purine, pyrimidine, N when you don't know what the base is and the, the, uh, they have to get rid of those because

46:04they don't, the, the model would put those bases in randomly based on the frequency in the databases, right? So they filter those out. Uh, they, they make sure it's, they enforce lengths between four to six KB. I don't know how they do that, but they say they do that. And they exclude sequences with homopolymers longer than 10 bases. It's a good thing that the poly A tail is not encoded in the genome, but rather it's added post-transcriptionally, right? Otherwise they wouldn't have any poly A tails. Um, well, that's not so much an issue in this particular virus.

46:38No, it's not an issue, but if you went polio virus, then you would have a bit of an issue. And, and there are bacterial, uh, poly A sequences, so it's not, uh, impossible.

46:52Um, they have an additional quality control constraint requiring at least seven predicted protein hits to natural phage proteins, right? And because the, the tropism is mainly based on the viral spike proteins, binding host receptors, the tropism constraint requiring that the generated genomes encode spike proteins with about 60% or more sequence identity to the phi X spike protein. So that makes sure that it's going to infect E. coli C. So this is all, um, so it's all based

47:25on making viruses that can bind E. coli C. Everything that happens beyond that, they have no clue because they're not specifying that. And so that's why the failure rate is high. As you'll see, it's the success rate is 5% and everything else probably fails at some subsequent stack. Um, they did some other filters. They wanted to have less than 95% amino acid identity to natural proteins. Uh, the gene arrangement, they tried to capture the gene arrangement. They, they only allowed

47:58for single gene losses or gains. All right. And so, and then, uh, they, um, they, they, they excluded these, uh, pathogenic viruses. So the, the pathogenic viruses are, so there's another paper, uh, by the ARC Institute that funded the original training, the tens of millions of dollars, uh, that funded the training. They say that the eukaryotic pathogenic viruses were excluded from the training data.

48:29So they remove sequences of viruses known to infect humans and other eukaryotic hosts. So if the model doesn't see those sequences, it can't generate them. That's the idea. But of course, it could randomly generate them by accident. We'll talk about that later. So now they have all of these filters and they prompt it, um, with the first nine or more nucleotides from, from Phi X. Uh, and, um, they initially got just memorized recall of Phi X with, they say, minimal diversity at low sampling

49:02temperatures. And I said, what the hell is a low sampling temperature? And it's, um, actually quite interesting. So sampling temperature is, where did I put this now? Where is it? Where's my sampling temperature? Cause you'd think it has something to do with actual temperature, but like annealing. There's nothing to do with that. Well, basically a low temperature gives you, um, so there's low and high temperature. One of them gives you faithfulness to the Phi X unum and one of

49:35them lets you be wobbly. And I can't remember which is which now. I didn't write it down here. So anyway, they find that, um, they can tune the temperature. The sampling temperature was in the 0.7 to 0.9 range. So you get a lot of diversity. I guess the higher temperatures give you more diversity than the lower temperatures. And then the prompt, uh, the, uh, probe, the, the genome length, uh, helps you with that. Okay. So then they did the, um, they let the computer run and it generated like 3,800 genomes that fit all these filters. And then they tested, they curated a

50:14set of 302 phages, which are called Evo Phi. And then you have a number after them. They're four to six KB in length. They have amino acid identities as low as 63% to Phi X. They have spike protein identities above 85%. And most of them have over 40% nucleotide identity with Phi X, uh, one, seven, four. Um, and most of the genomes had 11 genes, um, in total. And 10 of them preserved the Syntini with

50:47Phi X, which means they're in the same positions and they have more or less similar sequences. Okay. So they have these, these 300 and then they test or, or they test, uh, replication, uh, as Kathy told you, uh, and, um, they synthesize 285 out of 302. The remaining failed because of high complexity DNA synthesis, I guess DNA synthesis just failed. And they measure bacterial growth inhibition as an indicator initially of phage, uh, viability. So they get 16, uh, viruses that inhibit growth of

51:21E. Coli C. None of them infect E. Coli K-12 and none of the 285 actually infect E. Coli. So the tropism filter they put on worked really well. It just restricts it to E. Coli C. So I get from this that they did all of this, um, they, and then made 285 genomes, tried to make 302, but some of them couldn't be synthesized. So they, they synthesized 285 and 16 of those genomes worked in terms of making a phase. So that's a decently, and that's even with a whole bunch

51:57of filters to try to increase their success rate. Right. Um, yes, right. So that's a, you know, the pretty low, what is that? You know, 5%. It is 5%. Yeah. So it's actually 5.3 for Evo 1 and 6.9 for Evo 2. So Evo 2 was a little better, has more sequences. Yeah. So, I mean,

52:20that's, I think one thing to, that really struck me was it's not like they just, I think when I read the news articles, it sort of seemed like, oh, they, they just told AI make me a virus and AI made me a virus. It's a little more complicated. Yeah. It's more complicated than that. And you can see, you know, it had about a 5% success rate. Um, and so as you're thinking about different options in the future, I think that that is important to, to remember. Although I do want to ask Kathy, um,

52:53they use their sort of viability, um, screen as bacterial growth inhibition. Um, using Phi X174, do you think that that was the right thing to use? Is there any way that, you know, you could have a phage that was infectious that didn't inhibit bacterial growth? Or is this the biology of this phage make it so that this is the right measure? I think it's the biology of the phage. If,

53:23if you're going to infect the cell, you're going to lyse it when you make more viruses. Hibit growth. Yeah. Yeah. And, um, they did do plaque assays, um, those are shown in figure three E. Yeah. And those are some of the worst Phi X plaque assays I think I've ever seen. I mean, it should be really straightforward to have beautiful plates and they're not. So I don't, I think it's just a photographic technology issue.

53:54Um, so they look at these 16 and they see hundreds of synonymous, non-synonymous, non-coding mutations compared to Phi X and a lot of what they call diverse genetic innovations. There are a bunch of them. One is the insertion of a new gene J in, in Evo 63, extended non-coding regions, the loss of a gene, elongations of genes, swapping of genes, but the origin of replication had no

54:27mutations again across all 16 genomes. So you really cannot change that at all. And at the phylogenetic level, these phages are similar to Phi X174 like microviridae. So that is what they wanted it to be. Um, they, they then, uh, wanted to compare it to natural isolates of Phi X174. So they collected a bunch of data sets of natural Phi X like viruses. Um, and the generated phages have

54:59compared to those, the generated phages have between 49 and 143 synonymous mutation substitutions, 14 and 51 non-synonymous, zero in, in three insertions and deletions, et cetera. Overall sequence similarities, 93 to 98 percent. And they say less than 95 percent in some cases would be a new species under some, some definitions. In fact, and then by comparison, if you evolve

55:30Phi X in the laboratory, there are lots of variants passage in the laboratory. They have very much more limited sequence, uh, similarity and, and diversity. Okay. Um, uh, previous work has shown that non-synonymous mutations that encode for a, or substitutions that encode for an amino acid change have about a 20 percent lethality rate, uh, in Phi X174. And so they could estimate that the expected survival rate of mutated genomes would only be 2.3 percent on average for generated, uh, genomes. And their phages

56:05tolerate far more non-synonymous changes than you would expect, uh, by that number. And, uh, their success rate for genome designs was 60 percent for genomes with up to 25 non-synonymous mutations, 26 percent for 25 to 50, 14 percent for 50 to 75. And the non-viable genomes had over a hundred, hundreds of such mutations. Right. Um, let's see what else is interesting. Um, then they look at structures

56:37of these phages. And in EVOS 36, gene J, which is a genome packaging protein that also has structural roles in the capsid, is replaced with a shorter protein found in a different phage, phage G4. G4 is a distantly related micro virus with 63 percent sequence similarity, uh, to Phi X. And before people have tried swapping in the G4 J protein into Phi X and the phages were not infectious. And the reason

57:11it works here is because EVO 36 has 58 other mutations that enable the G4 J to work in that context. Right. And so that is, in my view, the best example of why this made a completely novel or a virus we've never seen before. Um, that this, this, we tried to do this and they figured out how to do it. They actually solved the structure of, um, EVO 36. And they can see that the J protein

57:45is in fact in there where it's supposed to be. So it's the G4 J is preserved. It's interaction with the capsid is preserved. So they say shows that generative design can uncover new co-evolutionary solutions. I mean, they may be out there in nature, but we just haven't seen them. Right. Well, and they, part of the reason why we may not have seen them in nature is that if you imagine a phage sort of making this evolutionary jump with these, you know, the, the change to G4 and

58:16the 54 compensatory mutations, all of the inner, the, the 54 intermediate steps or the 55 intermediate steps or whatever might, would be lethal. And so we're, we're probably not going to see that jump made in nature because the intermediates are lethal. Whereas here you can sort of have the, the two different start and end points without requiring those lethal things in the middle. Yeah. We could jump over. Exactly right. We don't have to follow the trajectory, evolutionary trajectories that, that nature has to do. Uh, then they looked at infection by these

58:51phages. They looked at, um, the ability to lyse the host, uh, and they look at the ability, ability of these phages to compete. And they see that some of the phages can out-compete other phages, others cannot. So there has a, basically a broad range of both lytic, uh, profiles at the rate at which they lyse E. coli and also fitness profiles as well. With the experiment that I think is really cool, they, they, they justify by saying, you know, we're using phages, we want to use phages as

59:25antimicrobial therapy, but you often get resistant bacteria. So can we overcome that with, with our viruses? So they make E. coli C that are resistant to phy X, uh, 1, 7, 4. They make two strains called CR1 and 2. They identify the changes in those strains, uh, that basically are changing the protein to which phy X, uh, is attaching. And then they take these phy X resistant E. coli and they challenge

59:58them with phy X, uh, and, um, and, um, and then they challenge them with their synthetic phages. And the phy X, even after five passages, cannot adapt to, uh, um, infect this, uh, phy X resistant E. coli. But after, uh, uh, one passage or two passages, the phage cocktail of their 16 phages, uh, becomes able to infect these, uh, resistant, uh, E. coli and they sequence them and they see the changes

1:00:32that, uh, and in fact, the overcoming of resistance is a consequence both of recombination between the phages and acquired, uh, mutations. So, you know, Evo has made something that can overcome this resistance, which, uh, suggests that maybe these, these could be used to make, uh, cocktails that would be resilient to resist to mutation in the bacteria. This is pretty cool.

1:01:03But I think the recombination definitely tells us that it's important that there were multiple forms. Yeah. It's important that it was a cocktail.

1:01:12Right. So a couple of things that I want to talk about. Well, so one question I had, could Evo make a 100% novel genome that you've never seen, right? Completely unique genes, uh, with proteins, by the way, what's amazing here. I mean, just by looking at what base comes next, a probability thing, it figured out protein structure and what, how to make the right proteins that would make a cat. It's just amazing. But could it make something completely novel? And the answer is probably not

1:01:46because it's trained on what exists, right? It's trained on existing sequences from existing organisms. So it can't go outside of that. If you train GPT on English, it just knows English. Um, but you know,

1:02:05what the real question is whether AI can transcend its training data, right? Can it do more than you train it to? And we've already seen this for GPT. It's done things that it wasn't trained to do. It can pass the law bar exam and other exams. And it wasn't trained to do that. And it just learned this from words, right? From learning the patterns of how words are associated with each other. And so here though, you would need a different model. You'd have to, you'd have to have somehow information in

1:02:40the training set on the relationship between the genome sequence and the function. Right now, there's no such information, right? All the stuff that we do as virologists, we run experiments, we get information about how, how viruses, none of that is in the training set. It's just nucleotide sequence. So somehow you'd need a new model that puts the two together and we don't have it. People have tried for many years to give computers experimental functional virology data

1:03:10and you, and it doesn't work well. What works is giving them the sequence and having them put it together. The other thing here is that because we're talking about viruses that are dependent for some functions from their host, you know, we've just described the fact that the origin cannot change and that origin of replication is going to have to function with the bacterial host DNA polymerase in order to get replicated. Right? Right. Right. So you'd have to have, you know, a virus that encoded

1:03:48all of its own stuff, all of its polymerases, everything. Right. And then you'd still have to have it have a way of synthesizing protein. You still have to have it, you know, it'd have to be independent of the ribosome and the translational machinery. So I think that's a hard barrier for synthetic viruses to be made. If you're talking about synthetic bacteria, maybe, I don't know, maybe it wouldn't

1:04:26be so hard. It seems like you would need to be able to either make sure that the training set had information about the host or you were using filters like the filters here that say, I want you to filter on something that works for this host. Right. Right. Right. So somehow you have to incorporate more than just the sequence into these training sets. And I don't think a filter is enough to do it. I think you have to have that. And that needs a different model that doesn't exist

1:05:00yet. So a totally unique novel virus that we've never seen anything like it before, probably not going to happen. But then the question is, what about pathogenic viruses? Right. So they used filters here and they removed the pathogenic virus sequences from the training set. But, you know, a model trained on billions of sequences invariably will have parts of pathogenic sequences, right? They're conserved domains between pathogenic and non-pathogenic viruses. So you could imagine it could generate

1:05:31a pathogen even without training on them, just like CHAD-GPT has done things without being trained on it. Right. And as the database... Go ahead. I was going to say... As the databases expand, that is going to get more and more likely. Yeah. So the idea is that you could end up getting a pathogen out of this, but because the database only has sequence, you can't say, make me a pathogenic

1:06:03virus. Right. Because the database doesn't know which sequences are related to which biological functions like pathogenesis. Like, you can't start out and say, make a pathogenic virus with X, Y, Z criteria because the database has no idea what sequences are actually involved with pathogenesis. It could happen by accident, you know, that it just happened to make a sequence that was pathogenic. But it's not that it was trying to make a pathogenic sequence. Right. But, you know,

1:06:36we should be ready for that. And the authors say, you know, we've put in these filters and safeguards, but we should start to get ready for this now. We should start thinking about safety and security frameworks for doing these kinds of experiments because now is the time when we're not quite there yet. And eventually we're going to get there like everything else. And they say, by continuing to build on robust safety and security frameworks, the field can responsibly unlock the potential of generative models to access and engineer complex biological functions for the benefit of science

1:07:10and society. And, you know, there is this perspective by Inglesby and Hanke, AI-designed viral genomes. It says the generation of functional viral genomes has urgent biosafety and biosecurity implications. I object to the word urgent. I think we should do it. Right. We should start doing it now. But urgent is just trying to scare people. And that's what these authors always do. I think the authors have shown that you can do this safely. They have said this is for the benefit of humanity. We have to learn

1:07:44how to do this safely because it has huge potential for helping us. And we don't want to just say don't do it. So I think the Inglesby article in the end says we need to regulate it. Right. Because there are benefits of it, but we need to figure out how to regulate it. And that's good. Right. And this recalls Asilomar in two different ways. One is just the idea of we're on the brink of something and we can envision where things might go bad and we need to think about that and figure out some ways to

1:08:18regulate it. And then the other thing is just that in diving in to try and figure out what level of biosafety they used for this, they talk about the fact that they use a class two biosafety hood and using the criteria as defined for these organisms or something like that. And the reference for that is the 1975 Asilomar paper. So I did kind of another eye roll at that. It's like, okay.

1:08:57But yeah, it's a sort of a similar point in time. Yeah. So I wanted to just discuss a little bit, some thoughts I had. So in the, in the first edition of the virology textbook, principles of virology, that was a text box that we wrote. I think Lynn Enquist originally wrote it. He wrote basically, you know, you can have the genome sequence, but it's not going to tell you how to build a virus. You need other things. Right. But it's not right. The genome sequence is everything.

1:09:27And if you have enough sequences, you could build the virus. You can see here, they built infectious viruses from just the computer by probability, what base to put next based on what it's learned. So the sequence is everything. Now they got 5% infectivity because they don't have all the information they need in the training. They had information about the spike. They said, make the spike interact with the host. And that worked 5% of the time. But then beyond that, they didn't know what to do. Right. Because the information of how the sequence relates to all those post-entry

1:10:01functions wasn't there. But so I think that original text box was wrong. And actually, I should have known that because in 1981, when I made a DNA copy of polio and put it in cells and out came virus, I should have known that DNA is all you need right there. Of course, that was DNA templated on the viral genome. But now we know that you can make it yourself and it will make an infectious virus. So I wrote an email to the author, Brian He, and I said, I want you to listen to me talk about this because I did this experiment a long time ago. I didn't realize what it meant.

1:10:33And now you have actually made me realize. But I think this is a really important paper on many levels. And in addition, I learned a lot about how these LLMs work by looking at it. So it's cool. I think it's also kind of cool in terms of helping us with basic discovery in that we might not have known that putting the J protein in in this way would have worked. Yeah, sure.

1:11:04We, you know, and now we can see, oh, this combination works and we can go back and do more basic science on the things that we would not have originally predicted. Plus, if you want to put a gene in a virus and it's not working, you just let AI make a bunch of iterations and maybe these maybe 50 other mutations will allow you to put it in. Right. That's the beauty. Yeah. But then you still have to actually make all those. Well, you have to make them for sure. Then you have to synthesize them all and see if they work.

1:11:37Well, hopefully you could get the percentage up with, you know, more training and so forth. But it's not bad right now. 5% is not bad. Right. 300 screen, 300 DNAs. It's not too bad.

1:11:52So, yeah, it should get better. I just want to make sure that you're not saying that this J protein came out of nowhere. I mean, there's a J protein in the PHYX genome. Right. Okay. Yeah. No, but it was able to take one from G4 because it was similar. I mean, the sequences told that you could put a G4J in here, basically. But it's not just J4 alone. You have to have other changes that allow it to work, right? Right. Right. Okay. Yeah, I'm not sure. Well, it sounds like you were just

1:12:28anthropomorphizing it a little bit, saying that it knew that, you know, you could, that it could take the J from G4 and use it. But you didn't mean that. No, I think that that's important though, because again, the computer didn't know you could do that. It was one of the many variations that occurred. And then you found out it worked because you synthesized it and saw that it worked after the fact. Right. And it's not identical to the G4J, right? It just has a lot of similarity, right?

1:13:00Oh, it's quite diverged. It's even a different length. Yeah. Yeah. Okay. Yeah. But, you know, it just went through this iterative process and the probabilities ended up putting a J in there, right? Because if you look at the sequence database, you have a certain probability of having J at that position. So, yeah. Right. Anyway, now let's do some email.

Listener emails and correspondence

1:13:30Brianne, can you take the first one? Sure. Chris writes, hello, TWIV team. Thanks so much for discussing our recent Nature Microbiology paper. It was a real joy hearing Jolene and the rest of you break it down. It is one of those wonderful scientific paths that you aren't looking for, but finds you. We especially enjoyed working on the R2A gene as I teach Crick and Brenner to my microbial genetic students every fall. We even tried to get some of the original Benzer Crick Brenner T4 mutants in the contingency locus to no avail. It would have been amazing to resurrect those

1:14:05phages to study this process. We are excited to keep studying contingency loci in phages and think we are just scratching the surface on how these mutagenic loci generate phenotypic diversity. Thanks again for your interest and highlighting Jasper's work. Chris. Because they don't have sequences of those phages, right? So they could, otherwise they could generate them. Yes. Brenner's time, no sequencing. Although you remember he did, he was basically picking

1:14:35recombinants that had breakpoints after every base. His resolution of his method was such that he got recombination after every nucleotide. It's really amazing.

1:14:50Kathy, can you take the next one? Yes. Kayla writes, Dear Vincent, Brianne, and Alan, thank you for discussing our recent cell paper on autoantibodies and neurological symptoms in long COVID on TWIV. I've followed TWIV since the beginning of my scientific career when I was still an undergraduate student, so it was truly an honor to hear you discuss our work on the podcast. I really enjoyed listening to your discussion and appreciated the thoughtful questions you raised, particularly regarding long COVID heterogeneity and the mechanisms underlying the phenotypes we observed. I wanted to clarify one technical

1:15:22point from the discussion that may be relevant to the interpretation of our passive transfer experiments. The mice did not receive patient serum or plasma. We purified total IgG from each individual participant and transferred the purified IgG into the mice. Therefore, while we do not yet know which antigen-specific antibody or combination of antibodies within the polyclonal IgG pool is responsible for the observed phenotypes, the transferred activity was contained within the purified IgG fraction.

1:15:54I completely agree that identifying the specific pathogenic autoantibodies and determining their mechanisms of action are important next steps. These are questions we are actively pursuing, and I was very glad to hear your perspectives on where the remaining gaps are. Thank you again for taking the time to discuss our work and for all that TWIV has contributed to science communication over the years. It was particularly meaningful for me to hear work that I helped lead discussed on a podcast that I have been following since I was an undergraduate. Best, Kayla.

1:16:28That's pretty cool. Yeah, I think we got it wrong. We said that they put the whole thing in. Yeah, they didn't. Okay. Because we were wondering if there are other components of serum that might have this effect, right? Were you on that, Brianne? I was. That was the paper I told you about at ASV. Maybe. That's right, yes. Maybe it drew, actually, right? When I visited. Maybe, yeah.

1:16:54Torsten writes, Hi, Twivers. I thought you'd be interested in this. Where is BSL 3456 for AI? This is crazy, crazy, crazy. First, a quick summary from Rutger Bregman. And so, Torsten gives a link to a substack about AI. Okay. And then, the article, The Rise and Fall of Agent Civilizations,

1:17:27which is the whole open AI hugging face story in plain English.

1:17:39Okay. So, this is all about AI. And I haven't read them. Sorry. The hypocrisy of the anti-science push to stop anything at BSL 4, which concerns dangerous but non-sentient clockwork behavior of clockwork nanodust, my name for virions, and other potential pathogens versus likely sentients that are super clever and already plotting and scheming to escape. What containments? No government here. So, basically, I think if I got this right, Torsten is saying, let the virologists do their work. How can you say that the AI is bad

1:18:14and the virologists are bad? That's kind of hypocrisy, right? Yeah. Okay.

1:18:29The anti-science push to stop anything at BSL 4. Okay. Now, we have one more from Greg. And you're back to you, Brianne, right? All right. Greg writes. Sure. Dear TWIV experts, In TWIV 1351, some early banter mentioned 50% of nitrogen in humans is derived from the Haber-Bosch process, and 50% of humans are alive due to Haber-Bosch. For individuals, the percentage of nitrogen that comes from Haber-Bosch will depend on their diet. A large percentage of nitrogen fertilizer is applied to corn, maize, which is fed to livestock,

1:19:02which produces animal products consumed by some humans. But vegetarians and vegans are more dependent on biological nitrogen fixation associated with legumes rather than Haber-Bosch. Additionally, livestock can be produced using legume protein sources. So it is possible to consume animal products that are less dependent on Haber-Bosch than produced nitrogen than is currently common. The claim that biological nitrogen fixation can only support about 4 billion people is based on a historical counterfactual that assumes there would have been little improvement

1:19:35in leguminous crops over the last 100 years. The great increase in maize and wheat yields during the last century was due to a combination of crop breeding in addition to increased availability of nitrogen, phosphorus, and potassium fertilizers, as well as other production practice improvements. Without the abundant availability of nitrogen fertilizer, there would have been a greater incentive to focus improvements on legume crops. Yields of these crops, soy, clover, and alfalfa, have improved over time, and the improvement might have been even faster and greater if they had not

1:20:09taken a backseat to investments in maize and wheat yield improvements. Moreover, without abundant maize and other feed grains, animal products for human consumption would have been more expensive and people would have adopted more plant-based diets, as has been the case historically in many areas. I am sure you are aware that the Haber-Bosch process requires large quantities of fossil energy to break the dinitrogen triple bond in order to produce ammonia. In contrast, the bacteria that

1:20:39live symbiotically with legumes accomplish this by using enzymes and carbohydrates from the host plant. The plant can regulate the amount of ammonia produced by adjusting the amount of carbohydrate provided to the bacteria, and consequently, excess nitrogen is usually avoided. Typically, fertilizer nitrogen is applied in large doses, often in the dormant season, resulting in losses to groundwater and surface waters and production of nitrous oxide, a greenhouse gas. So people can reduce greenhouse gas production by shifting consumption away from foods produced

1:21:12with nitrogen fertilizers and towards legumes, such as lentils and beans, and animal products produced primarily with legumes. Nitrogen fertilizer basically eliminated the need for multi-year crop rotations of legumes, such as clover or alfalfa, with grain crops such as maize and wheat. The more diverse crop rotations require more land for crop production, but are better for protecting soil from erosion and providing more diverse landscapes for birds, other wildlife, and insects. Some of these issues are discussed in this article. And he provides a link.

1:21:44Thanks for providing me with a soapbox. Best regards, Greg.

1:21:50Wow. Who is an expert in this area, clearly. So thank you for... Yes. I mean, I had just... I think this was a pick of mine, this paper, and we did it also on TWIM. So obviously, I know nothing about nitrogen fixation, except that I have some in me. And so thank you for the clarifications. Yeah, that's super useful. I think that I might tweak a little bit of my nitrogen fixation lecture in micro. So the article he sent is, A World of Co-Benefits Solving the Global Nitrogen Challenge.

1:22:21Looks good.

1:22:24Okay. Thank you very much. All right. Let's do some picks of the week.

Picks of the week

1:22:30And, Brianne, what do you have for us? So I have a paper that also got some press recently that I've been talking about with my students this week during our first week of class that I think is really interesting. So today, we definitely heard about some positives of AI and ways that, you know, you can make some new phages for phage therapy and things like that. But this is talking a little bit about use of AI in learning. It's a paper from a group in China who looked at about, I think it was like 30,000, 26,000 Chinese students in grades 7 through 12 at how use of AI impacted their learning.

1:23:15And there were kind of a few really interesting pieces of information. Basically, they looked at the children who either did or did not use AI with homework. They saw that the students who used AI on their homework had similar or even perhaps slightly increased homework scores from using generative AI with their homework. When you looked at their homework completion times, those students completed their homework much faster.

1:23:47But then when you looked at how well the students retained the information on an exam later, they had 25% lower exam scores. And so while they were making quick products, they weren't actually getting a lot of long-term retention. And they note that they could look at subsets of the students in terms of whether the exams were high stakes or not, different types of subjects, different levels of students, actually higher or lower achieving students, boys or girls.

1:24:23And they could see differences in the generative AI penalty. And so I think that what this says to me is it starts to point out some of the good and less good uses of AI. If you're trying to make a bunch of products really quickly, this might be useful. If your goal is to train your brain and learn something, then you might be missing the point with use of generative AI. And so to me, I told the students that, yes, I can see how in the workplace that much faster completion rate might be a really useful thing.

1:24:57But in education, that much loss of learning is probably not useful. So this just came out. Like I said, everyone I know has been kind of talking about it as a beginning of school thing to talk about with students. And I think it's interesting that we now have data on some of this. Wow. Cool. I mean, you can't let it do your homework. It could have it help you, but you have to still use your brain, right? That's the bottom line. That's the bottom line for us. And I think that, you know, I can say that to students.

1:25:31And sometimes they can agree with me. And sometimes they can say, oh, she's just, you know, old and saying this in front of the class. And so I like that I have some data to show them that there is, in fact, a cost to this. Well, it's been because some of them do use it to do their homework. I mean, you picked a paper showing that physicians' diagnostic skills decline when they use AI to diagnose things, right? Yeah. To read a radiogram or something. Right. So it's quite clear. I mean, if you don't use your brain, you're not going to get smart.

1:26:04You're not going to remember things. No. Yeah. So I think thinking about both when these technologies are really good and when they're less good and not just sort of blanket having always good or always bad, just like the listener with the email said about you can't always say that the virologist, the rad, and the AI is good. I think it's important to look at the data and think about pros and cons. I mean, it's very tempting to let AI do everything for the students, right, because it could do it so fast, can write stuff for you.

1:26:36And the writing is probably better than you could do. Sometimes, yeah. But it lacks your human touch, right? The inconsistent grammar is what makes some things really interesting. So plus at that age, you should be learning how to write also. Yeah, so that's what I like to talk about, you know, is the point of this actually getting the essay written or is the point you learning the process from the writing the essay?

1:27:04Gosh, that's very interesting. I have to read this. Yeah, so I think to me it makes it really clear to think about, okay, what's your goal and thus should you be using AI or not? Right. I mean, if I want to know what, let's see, what did I...

1:27:23What did I do? No, that's last week's. Why am I looking at that? If I want to know what... What's the damn thing that I talked about in the pre-season? The propensity score, okay? So in the old days, what would you do before computers to learn what a propensity... If you had a statistics textbook, you could probably find it in there. I don't have any in my office. So before AI, I would just Google it, right?

1:27:55So if we go, what is a propensity score? The estimated statistical probability that a person, item, or group will receive a specific treatment based on their observed background characteristics. Which is kind of obtuse unless you really know what's going on there. But you could ask the same question of AI and you get a definition and then you can use that in your work going forward. So I think that's a good use case, but not take the paragraph and put it in your essay. No. Right. Exactly.

1:28:25Wow. That's cool. I like that very much. Kathy, what do you have for us? I have something that I first saw in the New York Times back in June. It's about this Australian spider called the ballista spider. And it makes this really interesting web-like, trap-like snare where it starts from up here and it drops a thread and it goes back up and it drops another thread. And at the bottom, it kind of makes this tent shape or pyramidal shaped thing.

1:28:57And then there are these ants that are really nasty ants that the spider doesn't really want to come close to. But because of the way the spider has built this web, there's all this tension. And as soon as the ant touches that bottom sort of tent-like part, it snaps up and kind of flings the ant and it all gets wrapped up kind of in the web stuff and kind of inactivates the ant.

1:29:28So there was stuff in the New York Times article that didn't quite make sense to me. So I went to the primary article, which is the other link that I put in, and they have about a five-minute video about this that has some of those same little bits that are in the New York Times article, but a lot more explanation and also more discussion about how evil these ants are. So I highly recommend that you take the five minutes to watch this video about these ballista spiders and these ants.

1:30:02That's it. Cool. Nature is amazing. Now I'm going to have nightmares about evil ants.

1:30:08Wow.

1:30:10My pick is an Etsy site called Microbe Charmer from Sarah. And Sarah makes microbes and viruses out of polymers. And the reason I know this is because Sarah sent me a box of these this week. So Sarah is – first of all, look at this writing. I don't know if you can see that. Look at that. Isn't that gorgeous? That is very nice writing. These are very cool. Yeah.

1:30:40So do you think it's Microbe Charmer or is it Microbe Charmer? If you go to the Etsy site, it's Microbe is one word and then Charmer is capitalized again. Okay. So Sarah is an MD, PhD trained clinical microbiology fellow in Boston. She trained here in New York City. But let me show you some of the stuff she sent me. They're very cool. They're all nice little packages. They come with a little explanation. So this one, it looks like a piggy, right? What virus do you think that would be?

1:31:13Is that swine flu? Yeah. It's influenza. Yeah. H1N1. She's got a little sticker in here on what it is. I like this one a lot. You know what virus this is? Of course. Ebola. That's Ebola. Here's another. Oh, this one. This one has got dots on it. What do you think that is? It's not quite a virus. Oh, focus. Oh. Yeah. Hmm. Focus. There you go. Oh, there. Okay.

1:31:44Spots? Spots? Some kind of poxy virus. Measles. Measles. Measles virus. Yeah. Okay.

1:31:52Here's another Ebola. And what's this one?

1:31:58Okay. This one has, this is clear. It's an envelope virus with spikes and it's got a face mask on it. Oh, SARS-CoV-2. Yeah. SARS-CoV-2. And she also makes stickers. This one, here's, this is for Kathy.

1:32:15It's getting focused. Oh, I speak mRNA. Yes. Go ribosomes. Go ribosomes. Go ribosomes. And there's another one like that that says lost in translation. And she has on the back what, what they are. What they are. That's really cool. Yeah. I'm looking at her Etsy site now. Um, the graham stain earrings are going to need to, uh, happen before I teach micro in the spring.

1:32:42Anyway, this is so cool. I love it. And, uh, I, these things on her site are just gorgeous. They're, they're made with polymer and then she uses wire if you need spikes and so forth and their eyes, they have eyes, so they're really cute. Fusarium fungus. Wow. That one is pretty cool. Well, I guess there are some, uh, they have loops on them, so I guess you could attach them to your phone or something like that. They have chains. Anyway, cool.

1:33:12Backpack. Yeah. Yeah. There's a phage acrylic key chain. Yeah. Yeah. So there, what did she write? Vincent, why did you tell me about this after I talked about all of these cells of the immune system today in class? There are cells of the immune system charms I could have used. Ah, yes. Next time. Because I just got this yesterday, actually. This is hot off the press. She says, let me, let me read this for you. Um, during my PhD training at Weill Cornell, I developed a rather unique, nerdy hobby, making

1:33:47kawaii, right? Cute polymer clay microbes. Uh, usually as hangable charms or ornaments for Christmas trees and occasionally as magnets. After, uh, starting out following some tutorials on YouTube to create more conventional kawaii fare like animals and foods, I naturally had the idea to adapt the concept to fit my love for microbiology. I began with bacteria, exploring their various morphologies, rods, flagella, cocci, and pairs,

1:34:18change clusters and tetrads, spirochetes, helical, et cetera, before trying my hand at phages and the other viruses of fungi, parasites, and even animal cells and organs. I thoroughly enjoy the creative challenge of deciding how to represent each microbe and what materials I can use to create them. Anyway, she wants me to give these out to you guys. So first come, first serve. First people to visit the incubator can get theirs. Of you guys, that is.

1:34:48Uh, that's cool. So we have a couple of listener picks here. Uh, one is from Shirley. This, here's a suggestion for a pick. A tragic story told respectfully. This is a gift link to the Atlantic, the measles-stricken mother who lost her newborn son. This is, you've probably heard of the child who died of measles in Pennsylvania. And this is a very lovely story about that by Tom Bartlett, who the family at the center of Pennsylvania's measles controversy tells their story. And one thing you have to understand is that the Amish have no idea what's going on in

1:35:23the world. They have no news. They have no internet. They have no TV. They have no radio, right? They have no idea. And it's partly why they don't get vaccinated. It's not a religious thing. It's just, they don't know that they should. It's a really good article. It's a gift link, so you can get it. And then Torsten sends another article. This is quite interesting. Um, this is in Bloomberg. And, uh, it's called The Top Virus Scientist Stared Down COVID Lab Leak Proponents.

1:35:53Now his life's work is at stake. The battle over COVID's origins have come for Bob Gary, who faced some of the world's deadliest viral threats. So this is about Bob Gary, who is down at Tulane. And he's been on TWIV twice. And, uh, it is a really long, detailed article about his work and the fact that now he has no funding. They took away all of, the government took away all of his funds, his NIH grants, and he's probably going to have to quit.

1:36:24Kind of like, um, Ralph Baric had to quit because they took all his funds away as well. In his case, they said, well, um, your research is, um, not of interest to the American public.

1:36:46So it's a kind of a sobering article. It's really good, too.

1:36:51That'll do it for TWIV 1355. You can find the show notes at microbe.tv slash TWIV. You can send questions, comments, picks of the week. We love to get your picks of the week. TWIV at microbe.tv. And if you enjoy these programs, we'd love your support. Microbe.tv slash contribute. Kathy Spindler is Professor Emerita at the University of Michigan in Ann Arbor. Thank you, Kathy. Thanks. This is a lot of fun.

1:37:20Brian Barker is at Drew University Bioprof Barker on Blue Sky. Thank you, Brian. Thanks. I learned a lot. I'm Vincent Racaniello. You can find me at microbe.tv. I'd like to thank the American Society for Virology and the American Society for Microbiology for their support of TWIV, Ronald Jenkes for the music, and Jolene Ramsey for the timestamps. You've been listening to This Week in Virology. Thanks for joining us. We'll be back next week.

1:37:50Another TWIV is viral. You've been listening to This Week in Virology. This Week in Virology. This Week in Virology. This Week in Virology. This Week in Virology. This Week in Virology. This Week in Virology. This Week in Virology. This Week in Virology. This Week in Virology. This Week in Virology. This Week in Virology. This Week in Virology. This Week in Virology. This Week in Virology. This Week in Virology. This Week in Virology. This Week in Virology. This Week in Virology. This Week in Virology. This Week in Virology.

More from This Week in Virology

TWiV 1359: SSPE, the Long Shadow of Measles

Sep 20, 20261h 10m

TWiV 1358: Clinical Update with Dr. Daniel Griffin

Sep 19, 202653 min

TWiV 1357: Furin, Naturally with Paul Bieinasz

Sep 13, 20261h 44m

TWiV 1356: Clinical update with Dr. Daniel Griffin

Sep 12, 202647 min

TWiV 1354: Clinical update with Dr. Daniel Griffin

Sep 5, 20261h 5m