The broadcast is now starting. All attendees are in listen-only mode.
Hi, everyone. My name's Jeremy Wilkinson. I'm PacBio's Global Segment Lead for Microbial Genomics, and I'll be the host and moderator of today's webinar. I want to welcome you all to this webinar, which is a rebroadcast from a webinar in September of last year, with the topic focused on advances in unlocking human microbiome clarity with PacBio HiFi sequencing. Today, we have two human-focused microbiome talks from experts in the field, and then myself is available here for live Q&A at the end. Just a couple of housekeeping things. We have a lot of great material to cover today, and the presentations portion of this webinar is pre-recorded and is, again, a rebroadcast of a webinar from September of 2025. It will be followed by a live Q&A session with just me.
Please ask any questions about PacBio HiFi sequencing for microbial applications that I can help you to answer. You're welcome to submit your questions at any point during the webinar by typing them in the area provided on your attendee control panel, and I'll get to them either in the chat or at the end. All attendees are muted. We have also uploaded a few pieces of literature to your control panel, feel free to download those. We are recording this webinar, and we'll make it available here in the next couple of weeks. You'll get an email with a link to the recording. Immediately following this webinar, you'll receive a brief survey, please do take the time to fill that out, as it does help us plan for future webinars. Some more housekeeping on audio.
If you do encounter audio issues when we switch between this live conversation and the pre-recorded talks and then back to the live Q&A, you may have the wrong audio mode selected, please check that audio tab on the control panel to make sure you have selected the correct audio output depending on the device you're using. It also helps, it has in my case in the past, to leave the webinar and then rejoin it in case you can't hear anything. If you do continue to experience audio issues, we apologize for the inconvenience. Like I said, we are recording this webinar, and you'll have access to it in the next couple of weeks. If you are attending on the phone, the audio quality is better if you use headphones or earphones, and it's also better if you're connected to Wi-Fi versus cellular data.
Like I said, this is a one-hour webinar. We have recorded talks and a live Q&A at the end. For the live Q&A, it'll just be me, please ask any questions you have about HiFi microbial genomic sequencing. This is a rebroadcast from a webinar from September 2025. Due to PacBio's new SPRQ-Nx chemistry, now making high-accuracy long-read metagenomics even more affordable than before, we wanted to put this webinar out there again. There's going to be a cost and throughput slide in my section, just note that those have changed for the better with SPRQ-Nx . The agenda is as follows. I'll be giving a brief introduction to HiFi microbiome sequencing. Ashlee Earl will be giving a talk on connecting the dots in recurrent urinary tract microbiomes.
Jake Minich will be giving a talk on culture-independent metapangenomics enabled by long-read sequencing. With that, we can get started. All right. I'm going to give a brief introduction to HiFi sequencing and specifically to microbial genomics and microbiome sequencing. What PacBio HiFi sequencing is, they are sequence reads that are both long and highly accurate, this is more complete and accurate long-read sequencing. With these highly accurate long reads, we can enable a broad range of applications for human, microbial, and plant, and animal on both our Vega benchtop system, which is the one there on the left, and the Revio high-throughput system, which is the one on the right. Of course, with the Revio system, you can do more of these applications at the same time or expand into these applications in the future as your laboratory needs grow.
For microbial genomics, the applications can be metagenomics, microbial isolate assembly, epigenetics, viral sequencing, targeted sequencing, or full-length gene sequencing like 16S or ITS or 18S, for example. Like I said, HiFi reads really combine the best of two worlds. This is high accuracy and long read lengths. You get the accuracy that you would typically get with short-reads, but at 100 times longer length. This is possible because we circularize the DNA in the library prep process using SMRTbell adapters, these are hairpin adapters. We make multiple sequencing observations on the same stretch of DNA from a single molecule as the polymerase sequences around that, goes around that molecule. This washes out any random errors, which are shown as these little red dots in this graph, and allows for highly accurate consensus sequence to be called from those subreads. This results in what we call a HiFi read that is typically an average of 99.9% or greater accuracy or Q30. The consensus sequence generation, this HiFi reads generation actually happens on the instrument, there's no need for additional compute off-board for that HiFi read generation.
If you want to explore our HiFi data and see for yourself, we do have many data sets available. These are either on our GitHub page, pb-metagenomics-tools, or pacb.com/datasets. These are all different types of microbial genomics data sets. HiFi sequencing delivers more comprehensive and higher quality data for microbial genomics, and these are from culture to culture-free based approaches.
There are really four primary ways that folks utilize HiFi sequencing, plus the fifth, which I will talk about next, which is viral sequencing. This is first off microbial whole genome isolate sequencing. This is where you can obtain single contig chromosomes, so complete genomes, plus you get the complete plasmids as well, which we know plasmids are really important for inferring things like resistance. You get consensus accuracies that are greater than Q40 and oftentimes greater than Q50. You are able to also detect restriction modification motifs, so you can get the epigenetic and methylation information. With all of these things, you can do things like identify strains, antimicrobial resistance genes, and mobile vectors that may be mediating spread. The important thing is short-read polishing or hybrid-based assembly is not needed because you have high accuracy long reads.
For all of these applications that I'm talking about here, they are at competitive costs with other technologies, and you'll see that in a couple of slides from now. The next application is full-length gene sequencing for microbiome research. This can be 16S for bacteria, ITS for fungal, 18S for microeukaryotes. You can also sequence full operons, rRNA operons, like 16S, ITS, 23S, for example, for bacteria. What this gives you compared to partial gene short-read based approaches is you get more classified reads. Drastically more reads are being classified, and importantly, those are being classified down to species and strain level. For shotgun metagenomics, the first application is read-based profiling, which is mapping back to reference databases. When you have these high accuracy long reads, this is really good at mapping back to reference databases.
This is because you have on average about eight genes per sequence read, and you're able to get more and richer functional information. About 90% of the reads have at least one gene located within it, which gives you about, on average, four to five functional annotations per HiFi read. You are also able to profile taxonomy with high precision and recall. For shotgun metagenome assembly, you are able to obtain many high-quality metagenome-assembled genomes or MAGs, many of which are actually single contig, complete or near complete MAGs. The accuracy really comes into play here. It's not only long reads, but it's high accuracy, which is different compared to some other long reads, which helps you with resolving these MAGs at the strain level. Many of these MAGs are actually down to the strain level and not collapsed at the species level.
The same thing goes as is with the microbial isolate sequencing. There's no reason to do hybrid-based assembly or short-read polishing because you have high accuracy. Lastly, HiFi sequencing gives you the ability to distinguish one virus within a crowd. There's lots of different ways that you can do viral sequencing with PacBio. You can use an amplicon-based approach. You can use target enrichment with hybrid capture. You can do full-length RNA sequencing, or you can use a metagenomics approach to capture everything that's within there and pull out all of the different viruses. The last slide here is these are all the microbial genomics applications with the two PacBio HiFi systems, the Vega system on the left and the Revio system on the right. These are just the recommended maximum number of samples per SMRT Cell.
For the microbiome sample types here, these are based off the human gut microbiome. This also shows you estimated yield per sample and cost estimates per sample. These cost estimates are based off the cost of library prep plus sequencing, PacBio reagents for the whole genome applications, which are the first three, and then the last two, which are the amplicon-based applications. This also includes the cost of PCR, so PCR library prep and sequencing. What you can see is that the cost here that are highlighted in bold and magenta are very competitive of what it would cost to do with other technologies. There are three different options for these whole genome sequencing applications, the first three. This is because we have different library prep kits that are tailored towards your needs based off scale, automation capabilities, so on and so forth.
We have our HiFi Plex Prep Kit 96. We have our HiFi prep kit 96. We have the Ampli-Fi protocol, which is really useful for ultra-low input sample types like swabs, for example, or water. We have SeqWell LongPlex Kit, which is a PacBio compatible partner, which does the shearing and bar coding upstage up front of the SMRTbell library prep with PacBio. For our 16S, we have standard 16S, which is just a single monomer of 16S or any amplicon in this case. We have Kinnex, which you'll hear about later, which just concatenates multiple of those 16S amplicons together to take full advantage of the HiFi sequencing, the long reads.
When you have Kinnex, it just allows you to do more samples per SMRT Cell and get a higher read depth per sample than you can with the monomer-based approaches. That's it for me. Thanks for your time. Now we're going to move into the talks from our experts. Thanks.
Okay, thank you so much for this opportunity to talk about some of our research focused on recurrent urinary tract infections and how we've been using PacBio Kinnex technology to start to connect the dots. Okay. UTIs, many of us are not strangers to this type of infection. In fact, it's predicted that at least 50% of women at some point in their lifetimes will have a UTI. It turns out about 20%-30% of those will go on to have something called recurrent urinary tract infections. You'll hear me call it rUTI. These are infections that can occur multiple times in a matter of months and can last sometimes for years. Obviously debilitating for the individual, but also it's terrible for society.
Turns out, in the U.S. alone, 15% of our antibiotics go to treating these infections, and we know that they're becoming increasingly difficult to treat due to antimicrobial resistance. It's a big problem. Yet, despite the issue, we know actually very little about the drivers of rUTI. We've known for a long time that the gut has a role. We are known to be our own reservoirs of our uropathogens. There's linkage between the gut and the bladder. It turns out in some recent work with our great collaborators at WashU team led by Scott Hultgren, we believe that the gut plays perhaps an even greater role in potentially interacting the microbes in the gut, interacting with the host immune system, and potentially setting up susceptibility in the bladder.
However, that work aside, we also have a great interest in this organ that you see in the middle, the vagina. It's thought to also be a reservoir for uropathogens, and we have a number of questions focused on whether this organ with its microbes may also have a role in setting up susceptibility to rUTI. Really that's the subject of the work I'm going to be telling you about today. Trying to understand whether the vaginal microbiota, in particular, play a role in susceptibility to recurrent UTI. To get at this question, we have a study that we call the UPS study, the Uropathogen Surveillance study.
We enrolled a total of 70 individuals in this study, roughly split between two groups, controls, women who had one or no UTIs in their lifetime, and a set of rUTIs, individuals who've had at least three UTIs in the past year leading up to enrollment. Each of these women agreed to let us into their lives, answer lots of questions, but also provide lots of samples. Each individual was part of our study for 24 weeks. Every 12 weeks, they'd come in, give us blood, urine, stool, as well as periurethral and vaginal swabs at each of the visits. These women were also given at-home collection kits so that should a time when a UTI occurred, they could provide us with samples at that time from the convenience of their home.
They were asked to provide samples for 12 weeks weekly, so that we could follow them longitudinally over that time. Okay, lots of samples. I think there were 300 or so visits in total, thousands of samples. What I want to focus on today are the vaginal swab samples. This just gives you a visual snapshot of what the sampling looked like. On the left-hand side, we have our control participants. The x-axis is showing over time. We have individuals giving us their three samples. You'll notice on the right are our rUTI participants, lot more sampling happening. As we expected, these individuals were much more likely to have UTIs than controls. In fact, 17 of our 34 rUTI participants had a recurrence, and we captured 23 UTIs and lots and lots of vaginal swabs throughout.
I'll talk a little bit about another journey that these vaginal swabs went down, but the first question we had was, are there differences in the vaginal microbiomes of these individuals, controls versus rUTI, that might give us some sense of what could be different and potential pointers to drivers? Fine question, but a difficult one to answer using tried and true shotgun metagenomic techniques. It turns out when you just sequence the DNA coming from these vaginal swabs en masse, this is what it looks like. You get for each of the eight samples I'm showing you here, on the y-axis is showing the percentage of reads that map to the host, right? These are really dominated by host, which makes it very challenging to look at the microbial communities.
This is where we partnered with PacBio very early in their development of the Kinnex platform, and it's been really quite an important system for us. Just in a quick snapshot, we took all of our vaginal swabs, put it through their Kinnex system, which enables us to take a peek at the full length 16S rRNA gene from our communities. Effectively, what the technology does is it allows you to really take advantage of the throughput, the high throughput of these high volume new PacBio machines like the Revio, while also optimizing for the output and also the accuracy of these instruments. Effectively, the 16S amplicons are generated, they're concatenated, and then run on these machines. Here's just some stats on the number of larger concatenated reads that we got from our full experiment.
Once we've done the informatic, the multiplexing, the number of 16S amplicons of about the expected length of 1.5 kb and really terrific quality. Okay, what can we do with these data? Quite a few things. I'll just start with looking at the species level data, the results of looking at these data sets. Here we've organized the data such that each column here represents an individual and collating all of their data across all time points into one snapshot per individual. On the y-axis is the relative abundance of the various different bacterial species that we could detect using the Kinnex data. I'm sure you can see there's lots of different colors. It seems that there are certain species that really predominate, which we expected to see.
If we look at these profiles through the lens of something called community state types, which is a common method for the vaginal microbiome field, we can see that our individuals represented the five known state types and roughly the abundances we would expect in normal populations. Great. Of course, we wanted to know if these state types mattered in terms of rUTI susceptibility. Here's the data organized in a way where we could start to take a peek at this question. What we found is basically what your eyes are likely already telling you, which is when we organize these data in the control bucket or the rUTI bucket, there really was no difference in the number of individuals who had any of these different community state types. That didn't seem to be particularly explanatory or interesting.
One thing, however, that we know is that these CSTs can be dynamic over time in an individual. Here we have one of our rUTI participants over time, and we can see that they start off with this blue CST1, and they change to this green CST3, and they come back, and we see that the Participant 5 is not unique, that this can happen quite often. We wondered whether maybe the CST type, whether there was any connection between that CST type and whether or not an individual has an active UTI. Joseph Entner, a fantastic computationalist in the group, who wrangled a lot of the data and generated the results I'm going to be showing you today, did this analysis just asking is the individual at higher risk of having a UTI given their community state type?
The answer was largely not really, didn't seem to matter a ton. Although there was the CST1, which is typified by the species called L. crispatus, that seems to be least associated with active UTI. L. crispatus is actually quite interesting, going back decades, there's been a lot of work thinking about this as a probiotic for rUTIs. It's been a bit of a mixed bag in terms of results, but it's interesting that we're seeing this in our data as well. One of the real strengths, I think, of this approach is going even deeper, going beyond the species level and down to something that we call amplicon sequence variants. Instead of a species, it's really getting to the level of a subspecies.
Here, what I'm showing is for each of the species that we can observe across our individuals, Joseph has now just delineated with these tick marks where a particular ASV begins and ends. You see, it gives us a lot more complexity in drilling down in what these organisms really are. We've gone from now over about 300 species across all of our data sets to over 8,000 ASVs. We're really super excited about diving into these data to give us a sense of diversity over time and over space, different habitats in these individuals, since we have the urines, we've got the gut samples, we've got the periurethral swabs.
What I want to focus on right now is that we were excited to take another shot at asking that question of whether there was something special about our rUTIs versus our controls, just in general. That's these data here, and it turns out when you look at it this way, there are in fact, a couple of Lactobacillus crispatus ASVs that seem to typify one group versus the other. This ASV2, and to a bit lesser extent, ASV5, seem to be enriched in our rUTI group as compared to our controls, which tend to have a bit more of this ASV3. That was pretty cool. Of course, we immediately want to go to that next level of asking, well, does ASV2 and ASV5, do they matter? If so, how? It's a little limiting to just take the ASVs at face value.
You got to start to dig in a little deeper. What Joseph did, was he went to public repositories. He grabbed all of the high-quality Lactobacillus crispatus genomes that he could get his hands on, then he assessed whether or not these different references. Each tip along this tree represents a different reference genome that he pulled from public repositories. Now he's annotating on this outside ring, whether or not these references have any of these particular ASVs. It turns out, like many bacterial species, Lactobacillus crispatus has more than one 16S sequence in its genome. You can see that in general, these three sit throughout the tree of life for L. crispatus and often co-occur, including our ASV2 and 5 of special interest.
Okay, that is interesting, but it doesn't really allow us to start to dive in easily into function, right? Because are we looking in this part of the tree or this part of the tree, or this part of the tree? It's not clear at this point. In a tip of the hat to multi-omics, I want to share with you the integration of a different data type that has started to give us some more power to dive in. It turns out for every vaginal swab that we processed for the Kinnex system, we also had a companion vaginal swab that was used to plate or grow bacteria on plates in the laboratory. Really the goal here was to be able to assess these vaginal communities for uropathogens. We know, we strongly suspect uropathogens are present, but they are very low abundance.
We suspected, and I think it is true, we miss them quite a lot, even in these deeply sequenced 16S profiles that we were generating. These plates, what we do is we get everything that grows on these plates. We swipe all of the biomass off the plate, stick it in a tube, and extract the DNA. In this case, we put it through a more standard short-read metagenomic sequencing process to get at these plate sweep metagenomes. When we looked at those, here's data coming from these metagenomes. Each column similarly is coming from a single participant, in this case, in a single time point. They're coming from a single participant sort of aggregated together. What we find is that in fact, yep, we see there are uropathogens, these darker gray species that are shown here.
We also see Lactobacillus crispatus. It turns out Lactobacillus crispatus is a facultative anaerobe, meaning that it can grow in the presence of oxygen and did so on these plates. What was cool about that is it gave us an opportunity to dive a bit deeper. I just won't talk about it much, but we've, over the years, developed techniques that allow us to go from a species level view within complex community samples down to strains. Effectively, let's say in this example, we have this one blue chunk of the bar. We can discern from it in this cartoon example that there are, for instance, three different strains of Lactobacillus crispatus. We can discern relative abundances, and importantly, we can place them on the tree of life.
Then we can, of course, compare what we find in one sample to what we find in other samples. Long story short, what the strain predicted, sort of strains told us was that these rUTI-associated L. crispatus strains are actually coming from this single subclade within the tree. So we're really excited about that. All right, let me just quickly take us through hopefully the take-homes that I'm leaving you with today, including how fantastic the Kinnex system is. It's high throughput, high resolution, and it's really just a wonderful workhorse, particularly when you're dealing with communities that are chock-full of host. Based on these data, we've been able to glean some really interesting ASV-level insights that we think are pointing in a direction between the association between certain Lactobacilli and rUTI risk. A quick shout-out again to integrating data.
I think both the data from the Kinnex were fantastic, then bringing in this plate swipe data, plate sweep data together, powerful, but there are some blind spots. Together really helped us to zero o n this specific subclade that we're very much looking forward to diving into further, so that we can begin to really unpack what is going on at the gene and functional level that can hopefully help us understand why in the world this specific type of crispatus is thriving in this rUTI vaginal ecology. Potentially point us in directions for what might be driving rUTI susceptibility. With that, lots of thanks. I want to first start with just the WashU team, an incredible set of collaborators led by Scott Hultgren. I want to shout out Henry Schreiber and Karen Dotson. Lots of great funding over the years.
I'm very humbled and lucky to work with an incredible team of folks at the Broad Institute. I want to point out Joseph Entner in particular, who drove the vast majority of the analyses that I talked about today. Last but certainly not least, I want to thank PacBio and Jeremy Wilkinson in particular, who's been a superior supporter of all of our work, and a wonderful person to be working with. All right, with that, thank you very much.
Awesome. Well, thank you all for coming to this microbiome session and using HiFi. I'm super excited to share with you some of the results of our recent study. Across the world, around 25% of kids zero to five are undernourished. Most of these are children which are short for their age or stunted. Wasting, which is the more severe phenotype, is a lower incidence. There's a lot of reasons for this. It's not just about getting sufficient nutrition, but it can also be access to clean water, hygiene, both macronutrients, micronutrients, the intersection of infectious disease, including parasites, malaria, educational status. About 10 years ago, there was a really important study by the Gordon Lab at WashU showing that the microbiome actually plays a role. It may even be partially causative of the most severe form of wasting.
To this day, most of the research has been focused on wasting, which is the more severe phenotype, whereas stunting has been less studied. I would say that also for all of the microbiome research within this field, the absolute majority of it has used short-read sequencing, with most of it being short-read 16S sequencing, which is still the dominant methodology used in the field. 16S is great, but it only allows you to look at changes in abundance of different organisms and with a relatively low resolution. Shotgun metagenomics, primarily with short-reads, allows you to get at a little bit more information, such as the gene content. Again, you're really restricted to looking at genes that go up or down in a system. With short-reads, several new approaches have been trying to assemble the genome.
If you assemble the genomes, then you can look at the gene content within a given species or a genera. One of the challenges with short-reads is that while you can get partial genomes, without having a complete genome, it's really difficult to actually compare and do pangenome analyses or comparative genomics of these microbes, because one or two genes missing could be really important. In most of the cases, you could only have 10%, 20%, or even 30% of the genome missing. Long reads provides this really unique opportunity where we may be able to get complete genomes, which then could allow us to do comparative genomics and pangenome analyses in the context, in this case, of childhood undernutrition.
We wanted to apply this to a dataset, but first we really needed to understand which sort of methodologies would be the best in terms of generating high quality and high numbers of complete metagenome-assembled genomes or CMAGs, which I'll refer to. In the beginning, this experiment, which really started about two years ago, our lab had been doing nanopore for probably the past 10 years. We had a PromethION, so I ran all these samples through the PromethION methodology. We had also recently acquired a Revio, so processed through the Revio. Later on we also added the Illumina as a comparison, and specifically their synthetic long-read approach with our collaborators at UCSD. Specifically, we looked at a cohort from Malawi, sub-Sahara Africa over here.
This was from an interventional clinical trial that occurred back in 2015, and there were two villages in Malawi where participants were enrolled, and they, generally speaking, were about a year and a half of age when they started, and then they were tracked for 11 months. You can see that between these two villages, there are differing levels of malnutrition, so differing levels of acute malnutrition, different levels of stunting, different levels of environmental enteric dysfunction, diarrhea is kind of a phenotype of that, but it's a little bit more than that. It's inflammation on the small intestine. We took four children from each village, so the fecal samples of these participants, and looked at a total of six time points, or five time points for one of them, across this entire study design. Then we did deep sequencing on different long reads.
I developed this high molecular weight, high throughput DNA extraction methodology, which I can go into more if you're interested. We're currently writing a protocols paper on this. As I mentioned, we processed these eight participants in six time points, so 47 total samples through these varying technologies. The SMRTbell 3.0 kit using Megaruptor to shear. Then onto the PacBio Revio, a total of 14 runs. Also using their PromethION with, at the time, the best or the highest accuracy base calling, and also with the R10.4.1 flow cells and the Kit 14 chemistry, which has the boosted accuracy. We were using the best approaches available at the time for the technologies, and then we also ran this through a synthetic long-read approach with Illumina.
Our primary benchmarker endpoint we're looking at is the total numbers and the quality of the complete MAGs, in this way defined by completeness, more than 90%, less than 5% contamination, and then a single contig. Immediately, when we compare the top 100 longest contigs from these technologies, you can see that long-read approaches or single molecule approaches, both PacBio and ONT, are clearly an order of magnitude higher than the short-read approaches. From here, it's pretty clear that short-reads is really a thing of the past. Long reads, you get way longer assemblies. It was also clear that the PacBio approach generally, across all the contigs, assembles into longer contigs, and that this was really specific for the assembly method which was used.
When we looked at these three deep sequence samples, and then the subsampling all the way down to 1 GB, this also becomes very clear in that the total number of complete MAGs, regardless of the assembly methodology, was always higher for PacBio than it was for ONT. When you look at the results also vary within the assembly methodologies for each approach. That's also important to keep in mind. In terms of the amount of sequencing data to complete MAGs, the PacBio Revio generated 2.3x more CMAGs than ONT. Of course, when we compare this and we put a dollar amount on this, I basically calculated all of the costs up until this point, and this included DNA extraction, library prep, sequencing costs, and looked at the total amount of data input, then the total amount of MAGs or complete MAGs.
We can come up with a dollar amount, and you can see again that the long-read approach is even almost an order of magnitude cheaper than short-read approaches. Our short-read approach was actually generated zero complete MAGs. I had to go out to another resource where they generated almost 10 TB of data, and you can see that even with that, we're almost generating the same number of complete MAGs. Between PacBio and Nanopore as well, the cost per MAG or complete MAG is much basically cheaper with PacBio. Why is this? Why is PacBio generating more MAGs, more complete MAGs? I think there's a few things. I think, in this case, in terms of sequencing yield, while it wasn't significant, they're pretty similar in yields.
I will say that our recent runs with the Revio, I think our latest or best was like 80 GB or 85 GB. This has improved, especially with SMRT chemistry. The amount of usable reads, like Nanopore produces more reads, but when you look at the quality, this is the accuracy of the reads. It's like two orders of magnitude higher than ONT. I think especially for microbial genomes, this is, in my opinion, really the sweet spot of the technology that I just don't know if it'll ever be matched. This level of quality really helps when you're trying to assemble very similar microbes or even strains or species within a complex community. The read lengths, median and 50 and whatnot, was also larger for PacBio. Of course, ONT, you do get longer reads.
If you're looking for these ultra-long reads, you can get longer. That being said, when you look at the quality of the genomes, we can see that, and these are only in the complete MAG comparison, that both the completeness and contamination had a higher quality for PacBio as compared to ONT. Again, this is comparing best versus best. We also looked at using this ideal approach developed by Mick Watson, that the genomes themselves were also of higher quality. If you have an error introduced into your genome, then when you go to annotate it, the genes themselves could be truncated artificially. If your genes are shorter in your annotated genome as compared to what's in the database, then you would see a difference here. That's essentially what you see in this plot to the right.
Again, that also was significant. Why else might having full genomes be valuable to your dataset? We looked at this in the context of we created this hybrid genome library or a genome atlas of the microbes in this dataset, generating 986 complete CMAGs. We looked at the classification rate of if you were to look at the sequence reads, when you just look at the standard sourmash non-redundant database versus this database plus your custom database. When you use a custom database, you can drastically increase the classification rate of your sequences. Without the custom database, over half of your data is unclassified, meaning you don't have any sort of taxonomic assignment. Whereas when you utilize genomes from your data in your dataset, you can increase this classification rate to 82% for a total of 44%-60%.
We apply this broadly to the dataset, and we found that the overall microbial diversity in children with poor linear growth was higher than it is with improving or good linear growth. This was also associated with breastfeeding status, and these associations were significant both for alpha diversity and also beta diversity, meaning compositional differences in the microbiomes. We wanted to increase the sample sizes, we sequenced around another 40 participants, five time points. Over 200 additional samples utilizing the SeqWell LongPlex methodology for the library prep, which allows you to do a very fast and efficient library prep using transposon-based enzymes. We sequenced this across these pools, across a total of four different Revio runs, and we also included one matched run with Revio versus Vega.
This was handled in collaboration with PacBio, and a special shout-out to both Ian and Primo in the Apps lab, also collaboration with SeqWell, specifically Joe Miller. What this allowed us to do is to apply machine learning because we had a larger sample size, and we could classify the association, or we could classify whether children had an improving or worsening linear growth at 78%. We're hopeful that this will increase with additional samples. What's really interesting is that when you have this strain-level and even species-level resolution, you can really start to discern differences in associations of individual, say, species with a growth outcome, or in this case, a nutrition outcome. Prevotella, there's the most dominant genus in our dataset. I think there's maybe 30 or 40 species. What you can see is that it depends.
Some of the species were associated with improving linear growth, whereas others were associated with poor linear growth. This is really critical to have this level of resolution to be able to distinguish the importance between these various features or species or strains in your dataset. We also then, because we had these genomes, we could apply pangenome approaches and specifically microbial GWAS. This was performed at the genus level for the top 10 abundant genera. What we found, there's a total of 89 genetic associations with linear growth. Most of these were in the Prevotella, followed by Faecalibacterium, and then there were a few associations in Megasphaera and Bifido. Of course, all of these targets need to be validated in future studies. This sort of approach was really not feasible without doing long-read sequencing, and in my opinion, is really exciting.
The other kind of final thing we did was look at genome stability. Because we had a time series of these participants, we could ask the question, well, how does the genomes of a given species change over time? In this case, we would say we would look at the earliest time point, then we'd look at instances where that genome or that species was found in at least two additional time points. So they had to have at least three time points from which it was found. Then we could compare the genome similarity of these latter time points to the initial time point. We looked at this in these eight participants. Again, four had improving linear growth, four had worsening linear growth.
What we found, which was really interesting, is that children with poor linear growth had a higher change or differences in the genome Average Nucleotide Identity over time. This suggests that either the genome species are evolving or they're being replaced by similar organisms. It suggests stability or instability in the microbial genomes as an association or an indicator of gut health or linear growth status. In conclusion, undernutrition is really important globally. It needs more funding. It needs more attention. We still really don't understand what are the exact causes from the microbiome, but we know that stunting is very high in Malawi, environmental enteric dysfunction, also very high. We generated a really unique dataset using varying technologies. We show that long-reads is superior, with PacBio being the most efficient way to generate complete MAGs.
We apply this approach to our long-read dataset and specifically show how across time, strain diversity, stability, and overall diversity was associated with linear growth status. We also performed pangenome analyses and microbial GWAS from these genomes. We also show how genome stability could be a new way of understanding gut health, which is really only feasible using long-reads. We're happy to report that just last week, this finally came out in Cell, and there's a really nice write-up from the PR team at Salk. I'd encourage you, if you're interested, to read more about it. All the data is public. It's always been public. Finally, I've just started my lab at Baylor, and this picture was taken today. This is the day that I got my Vega, so I was pretty stoked.
At any rate, I haven't found a proper hat for the Vega or a name yet, but if you're interested, I'm hiring PhD students, postdocs. What I'm really interested in next steps is can we use this microbiome to develop predictive diagnostics of future health outcomes? I'm really interested to understand more of the mechanisms of the microbiome and looking at other types of features, metabolomics, proteomics. I'm also very interested in applying this to agricultural systems, specifically aquaculture fish farming related systems. How can we improve animal health, animal production with the microbiome? With that, really like to thank Todd Michael at the Salk, my collaborators Mark Manary and Kevin Stevenson at WashU, Mike Tisa from Baylor College of Medicine, Rob Knight and Omar, and then some of his other trainees at UC San Diego, and then various folks at PacBio and SeqWell.
With that, I can take any questions. Thank you.
Okay, now we're going to move into the Q&A portion. Just a reminder that if you do have any questions, please use the Questions tab on the control panel, and I'll try to address any questions, given the amount of time we have. We already have a few questions, so I can get started on those. Again, please ask questions that's related to HiFi sequencing for microbial genomics, and try not to ask any questions based on the scientific projects that were presented. First question is, if we were to do a full bacterial genome mapping with PacBio HiFi sequencing, how can we know if it's true? My understanding of this question is you could do an alignment to the reference database, so like NCBI or RefSeq, using any type of aligner, like BLAST or Minimap, and that's how you can know if it's true.
Another question here is a question on extraction. Can only assume the extraction steps should produce very high molecular weight DNA. Are there any specific extraction protocols that need to/should be done? My experience in extracting high quality long DNA is somewhat painful. My context is extracting DNA from bacterial isolates. There are various different kits that you can use from commercial providers that still work well for extracting microbiome samples for PacBio HiFi sequencing, and they typically will yield anywhere from about 6 kb to about 12 kb in length of the DNA, which is still really suitable for HiFi sequencing. Those kits are the various different QIAGEN microbiome kits, the various Zymo Research microbiome kits, and then the kit that Jake Minich there just presented, and Rob Knight's lab as well has gone with is the MagMAX microbiome kit.
They found that it produces the longest fragments and still gives them the correct profiles. Next question is, We're investigating neuropeptide genes and parasitic genome sequence with PacBio and HiSeq, and assembly using FALCON, Juicer, 3D- DNA. However, the current assembly and annotation appear incomplete, and we have not identified any candidate neuropeptide genes despite their presence in close related species. Could this reflect limitations of existing genomic resources rather than true biological absence? Would additional long-read sequencing transcriptomics or alternative annotation approaches improve detection? We would appreciate any recommendations for relevant workflows or resources. The state-of-the-art assembler to use for PacBio data is hifiasm out of Heng Li's lab, I would encourage you to utilize that.
The other thing is perhaps the coverage isn't just enough with the long-read aspect of this. You may want to increase coverage of the genome through sequencing if it's not high enough. You typically don't need much genome coverage with HiFi. Usually 20x is suitable. I would just look into that. Obviously incorporating transcriptomics will also help for capturing genes in case it's being actively expressed there. Those are things that I would look into. A question on what's MAGs here. MAGs are metagenome-assembled genomes. You sequence a metagenome, you do a metagenome assembly to get the contigs, and then you take those contigs and you bin them into MAGs, and you de-replicate them. MAGs are, again, metagenome-assembled genomes from the metagenome sequence data that's been assembled.
In the context that we typically talk about them is medium quality MAGs are greater than 50% complete. High quality MAGs are greater than 90% complete. The single contig or complete MAGs that Jake was talking about are also greater than 90% complete, less than 5% contamination, and they only have single contig. They're one contig MAGs, whereas high quality MAGs are less than 20 contigs. There's lots of various different levels of MAGs when you talk about them, but it's typically those classifications. Next question came in here just now. Do you have any troubleshooting solutions for sequencing chlamydia in microbiome samples? Most bacteria were sequenced, but not chlamydia. I have seen it. I have seen it specifically in 16S results. Not sure if you're talking about for 16S sequencing or if you're talking about for shotgun metagenomics.
I have seen it in full-length 16S data. Chlamydia has been detected using that. Maybe if you're not seeing it in shotgun metagenomics, it's just because it's at such a low abundance and you just haven't gotten the appropriate depth to capture it. Whereas 16S is slightly more sensitive even with less data, maybe that's the cause for that. How do you deal with host contamination or interference? That's a great question. There are various extraction kits that can help with heavy host type removal, specifically, again, for the shotgun metagenomics. For the amplicon- based, you don't have to worry about it as much, although there is some mispriming not for the target region that can amplify up the host. Particularly in humans, this happens sometimes with 16S because one of the human chromosomes has similar sequences that the primers anneal to.
For shotgun metagenomics, there's kits from both QIAGEN and Zymo that we've seen work well. Like the QIAamp kit from QIAGEN and the HostZERO from Zymo. I would implore you to try those out if you're doing shotgun metagenomics. Okay, thanks. You confirmed 16S using Kinnex. You should be able to detect it with the chlamydia with 16S. If not, again, perhaps the depth's just not high enough, but typically it should be fine. Another thing that you could do to add in nucleotides, which lots of people do in order to get better species and strain resolution, is to go beyond just 16S and incorporate at least the ITS region that sets between the small subunit and the large subunit, so sets between 16S and 23S.
You can incorporate that region, and it would make the fragment size like 2.5 kb, or you can do the whole operon, which is like 4.5 kb, and that will add more nucleotides for better resolution. Another question, does PacBio offer a workflow similar to Oxford Nanopore's host depletion approach where host DNA contamination is removed during sequencing? No, we don't have adaptive sampling. That's what you're talking about. If so, how does this compare to wet lab depletion methods like dimensional enzymatic host DNA removal, term sensitivity, repeatability, clinical and environmental metagenomic samples? Yeah, we don't have adaptive sampling like what Nanopore has. We have seen really good success, particularly in the collaboration with QIAGEN on the QIAamp kit, where we tested it on an oral sample, which typically can have as much as 50% human contamination.
When we sequenced that extracted DNA on PacBio on the Revio, when we mapped back to the human genome, using Minimap, we actually found only, I think it was 0.4% of the reads mapped to humans. We got 99.6% of the reads were actually microbial, which was great success. I think that's a good point. Obviously, there's heavier host sample types than oral, but I think that was a good start. Chlamydia is certainly present, as confirmed by PCR sequencing of 16S on 100 samples did not show chlamydia. Do you have any suggestion on how to improve that? I would also make sure that it's absolutely prevalent in the database that you're using.
Whether you're using SILVA, Greengenes2, GTDB, et cetera, making sure that there's various variants or strains of chlamydia that are present in the database, because that could also be an issue. It also could be based on the primer sequences, I would look into making sure that we do use degenerate primers in our protocol, but I would try to make sure that the primer sequences are actually able to prime to chlamydia because there are some various bacteria examples where it's just still not able to be detected, even though you're able to capture a majority of them. You can look into the primer sequences. All right. We are at time, and we are through all the questions. I appreciate everyone attending. Again, when you leave, please fill out that survey. It only takes a couple of minutes, and it really helps us.
Thanks everyone for attending. Bye