<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://spaam-community.github.io/spaam-community.github.io-dev//feed.xml" rel="self" type="application/atom+xml" /><link href="https://spaam-community.github.io/spaam-community.github.io-dev//" rel="alternate" type="text/html" /><updated>2026-07-21T07:39:41+00:00</updated><id>https://spaam-community.github.io/spaam-community.github.io-dev//feed.xml</id><title type="html">SPAAM Community</title><subtitle>Standards, 
Precautions, and 
Advances in 
Ancient 
Metagenomics
</subtitle><author><name>SPAAM Community</name></author><entry><title type="html">SPAAM Steering Committee Elections 2026</title><link href="https://spaam-community.github.io/spaam-community.github.io-dev//news/2026/07/21/elections/" rel="alternate" type="text/html" title="SPAAM Steering Committee Elections 2026" /><published>2026-07-21T00:00:00+00:00</published><updated>2026-07-21T00:00:00+00:00</updated><id>https://spaam-community.github.io/spaam-community.github.io-dev//news/2026/07/21/elections</id><content type="html" xml:base="https://spaam-community.github.io/spaam-community.github.io-dev//news/2026/07/21/elections/"><![CDATA[<p>This year, SPAAM will hold the first elections for the Steering Committee. Following the <a href="https://www.spaam-community.org/constitution/">constitution</a>, the Steering Committee should consist of 12 members with terms of 2 years tiered in two cohorts overlapping for one year, so in the future every year, 6 new members will be elected into the Steering Committee.</p>

<p>Since we are in the process of introducing the new system, this year we will elect only 3 new members into the Steering Committee. In 2027, the terms of 6 of the current members will end and we will start the regular election scheme.</p>

<p><strong>Deadline: August 15th</strong></p>

<p>Meet the candidates and cast your vote <a href="/assets/media/SPAAM_Steering_Committee_elections2026_standalone.html">here</a>.</p>]]></content><author><name>SPAAM Community</name></author><category term="News" /><summary type="html"><![CDATA[This year, SPAAM will hold the first elections for the Steering Committee. Following the constitution, the Steering Committee should consist of 12 members with terms of 2 years tiered in two cohorts overlapping for one year, so in the future every year, 6 new members will be elected into the Steering Committee.]]></summary></entry><entry><title type="html">Emerging Ancient RNA Virus Research</title><link href="https://spaam-community.github.io/spaam-community.github.io-dev//blog/2026/05/08/blog-RNA-viruses/" rel="alternate" type="text/html" title="Emerging Ancient RNA Virus Research" /><published>2026-05-08T00:00:00+00:00</published><updated>2026-05-08T00:00:00+00:00</updated><id>https://spaam-community.github.io/spaam-community.github.io-dev//blog/2026/05/08/blog-RNA-viruses</id><content type="html" xml:base="https://spaam-community.github.io/spaam-community.github.io-dev//blog/2026/05/08/blog-RNA-viruses/"><![CDATA[<p><em>by <a href="https://benjaminguinet.github.io">Benjamin Guinet</a></em></p>

<p>Helmholtz Institute for One Health (HIOH), Greifswald, Germany</p>

<p><strong>🏆 Winner of the 1st SPAAM blog post competition</strong></p>

<h2 id="introduction">Introduction</h2>
<p>Influenza, COVID-19, dengue, Ebola, and rinderpest are examples of RNA viral diseases notorious for triggering outbreaks that close borders, overwhelm health systems, and destabilize economies and agriculture worldwide [1]. Yet despite their profound societal and ecological impacts, the origins and long-term evolutionary dynamics of many of these viruses remain poorly understood. Most evolutionary inferences rely on contemporary viral diversity and short epidemic time series spanning only decades, a narrow window compared with the millennia over which RNA viruses have circulated, diversified, and repeatedly crossed species barriers throughout human and animal history.
In recent years, paleovirology has emerged as a powerful approach for investigating deep evolutionary history of viruses through genomic “fossils” known as endogenous viral elements (EVEs). These sequences arise from viral integration into the host germ line and, in rare cases, can persist for millions of years when retained by selection, thereby preserving molecular traces of ancient infections [2–7]. However, most EVEs are too ancient to resolve the recent processes that shaped the emergence of human and domesticated animal diseases.
A powerful way to access these periods is the direct analysis of authentic ancient pathogen genomes recovered from host remains. By extending observations into the past, such genomes can for instance recalibrate molecular clocks, uncover extinct lineages, clarify host shifts, and link viral emergence to ecological and demographic change [3, 8–10]. However, this approach has thus far been applied predominantly to bacteria and DNA viruses [11, 12]. This imbalance is evident in community resources such as Ancient MetagenomeDir, where fewer than 10% of cataloged ancient viral genomes correspond to RNA viruses, despite RNA viruses comprising more than half of known mammal-associated viruses today [13, 14]. This disparity likely reflects technical constraints rather than biological reality, as protocols optimized for the recovery and sequencing of highly degraded RNA remain rarely implemented. Consequently, RNA viruses are systematically underrepresented in ancient genomic datasets, creating a major blind spot in reconstructions of pathogen history.
In this context, it appears timely to consider how these limitations might be addressed in order to better integrate RNA viruses into paleovirological research. This mini-review therefore aims to synthesize emerging strategies for the recovery, authentication, and analysis of RNA viruses from archival and ancient materials. By bringing together recent technical advances while acknowledging the challenges that remain, it outlines a practical roadmap to facilitate their inclusion in paleovirology and to extend evolutionary inferences beyond contemporary outbreaks. Ultimately, these approaches may enable a more direct and nuanced reconstruction of the origins, emergence, and long-term dynamics of the RNA viruses that have shaped human and animal disease. 
Ancient RNA as a viable molecular archive
A major barrier to ancient RNA virus research has long been the assumption that RNA is too unstable to persist beyond short historical timescales [15]. This view is grounded in well-known biochemical considerations, including rapid postmortem RNA fragmentation, ubiquitous RNase activity, and the predominantly single-stranded nature of most RNA molecules [16]. Consequently, for buried archaeological remains in particular, these processes were thought to rapidly eliminate RNA, leading to the expectation that ancient specimens would rarely contain recoverable RNA and that ancient virus research should therefore focus primarily on EVEs.
However, greater optimism surrounded pathology and archival specimens, especially formalin-fixed tissues, since formalin fixation inactivates RNases [17, 18] and has long been used for routine tissue preservation since its introduction in the early 1890s [19]. Early studies indeed demonstrated that RNA could persist postmortem when preserved by cold, desiccation, or chemical treatment and be detected in such material, although extensive cross-linking and fragmentation introduced substantial technical challenges that complicated sequencing-based recovery [20].
This momentum enabled the broader application of RNA-based approaches to pathological and archival tissues. Prominent examples include the reconstruction of influenza virus genomes from early twentieth-century specimens (Fig. 1), which demonstrated that the strain responsible for the 1918 Spanish influenza pandemic originated from an avian-derived lineage that recently adapted to humans, clarifying its evolutionary source and early diversification [21]. Another notable example is the sequencing of a 1912 measles virus genome, which refined estimates for the virus’s host jump from cattle to humans to approximately 3,000 years ago, a period when increasing population density likely enabled sustained transmission in a newly susceptible human population [9] (Fig. 1). More recently, authentic viral RNA has been recovered from even older archival specimens stored in alcohol since the eighteenth century (Fig. 1). Although these findings have not yet undergone peer review, researchers reported the recovery of an almost complete Rhinovirus A genome from a human lung specimen. Because this sample predates the widespread adoption of formalin fixation, it demonstrates that centuries-old museum collections preserved using alternative methods can retain recoverable informative viral RNA [22].
These results also raised a broader question: if viral RNA could be recovered from archival material, under what conditions might it persist for much longer periods? Addressing this question requires moving beyond historical specimens toward naturally preserved material in which degradation processes are slowed or arrested.</p>

<h2 id="beyond-the-recovery-of-historical-rna-pathogen-molecules">Beyond the recovery of historical RNA pathogen molecules</h2>
<p>Recent progress in ancient host genomics has demonstrated that nucleic acids can persist far longer than previously assumed under favorable environmental conditions [23]. Consistent with this, ancient microbial DNA has been recovered from specimens spanning from hundreds to millions of years old [11, 24–28].
On the RNA side, recent studies have shown that RNA, previously assumed to be far more fragile, can also persist over millennial timescales when preservation conditions are favorable. For example, RNA has been recovered from a 14,300-year-old Pleistocene canid preserved in permafrost [29], as well as from extinct species such as the thylacine and from permafrost-preserved mammoths dated to approximately 37,000 years ago [30, 31] (Fig. 1).
Building on these insights, the first convincing evidence that RNA viruses can be recovered from ancient material is now beginning to emerge. In plants, RNA preservation appears to be more common than in animals, likely because intact RNA is required for seed viability and germination, facilitating its persistence over extended timescales. Consistent with this, a Chrysovirus was isolated from approximately 1,000-year-old maize samples [32] (Fig. 1). However, in animals, recovery presents a substantially greater challenge, as RNA molecules typically degrade rapidly after death due to enzymatic activity and environmental damage. Nevertheless, recent metatranscriptomic analyses of naturally mummified Adélie penguin remains from Antarctica reported the recovery of authentic RNA virus genomes spanning centuries to millennia, including diverse lineages of Picornavirus and Rotavirus [33] (Fig. 1).</p>

<p><img src="/assets/media/Timeline_RNA_virus2.svg" alt="timeline" width="100%" class="center" /></p>

<p><em><strong>Figure 1.</strong> Ancient host and viral RNA discoveries placed on a timeline with modern RNA virus sequencing activity. The timeline uses a segmented non-linear year axis in which deep time (∼40,000–10,000 BCE), the Holocene (10,000 BCE–0), the historical period (0–1900 CE), the 20th century (1900–2000), and the genomic era (2000–present) are displayed at progressively expanding scales to allow simultaneous visualization of sparse ancient RNA discoveries and dense modern sequencing. Annotated host ancient RNA events (top; circles) and ancient viral RNA events (bottom; squares) were curated manually from the literature. Modern RNA virus data (bars) correspond to complete NCBI RNA virus assemblies, filtered to representative RNA virus families found in mammals, and plotted yearly according to collection date metadata. These assembly counts are intended to reflect relative temporal trends in sequencing activity rather than absolute numbers of RNA virus genomes generated, which are substantially higher in practice.</em></p>

<p>Together, these findings show that, under favorable preservation, ancient RNA virus genomes can persist far longer than previously assumed. The challenge ahead is to move beyond proof-of-concept discoveries toward systematic, reproducible detection and authentication, supported by strategic sampling, optimized laboratory workflows, and robust bioinformatic and evolutionary frameworks to reliably integrate these data into reconstructions of viral evolution.</p>

<h2 id="future-directions-in-the-field">Future directions in the field</h2>
<h3 id="efficiently-targeting-rna-viruses">Efficiently targeting RNA viruses</h3>
<p>As discussed above, pathology collections containing formalin-fixed specimens, as well as in some cases alcohol-preserved samples, can serve as valuable and often preferred sources of material for the discovery of relatively recent RNA viruses. When the objective is to investigate much older infections, naturally mummified remains with preserved internal tissues represent especially promising targets and should be prioritized, where ethical and curatorial considerations permit.
Because tissue tropism determines viral load, careful selection of sampling sites within remains is critical. However, most preserved tissues analyzed in ancient biomolecular studies consist of skin or muscle, which are rarely the primary targets of RNA viruses. In contrast, high viral loads typically occur in organs such as lung, intestine, liver, spleen, and oral or nasal mucosa. For acute infections, viral abundance also varies with disease stage and time of death [34–36], often peaking early and declining rapidly, so individuals dying late during infection may contain little detectable RNA. Consequently, both tissue availability and host disease stage shape recovery success. These dynamics are also frequently age-structured, with juveniles often showing higher viral loads, prolonged shedding, or increased mortality for many RNA viruses [37–39].
Dental calculus may represent a promising alternative substrate for investigating RNA viruses. Although it contains PCR inhibitors, extractions from ancient remains often yield high nucleic acid loads, sometimes exceeding those of other substrates [40]. These mineralized deposits may therefore preserve RNA viruses with oral or respiratory tropism. Supporting this, SARS-CoV-2 RNA has for instance been detected in dental calculus [41].    <br />
Whether ancient viral RNA can survive over longer timescales remains unclear, but calculus may be particularly informative in ruminants and other social species that accumulate substantial deposits over their lifetime [40].</p>

<h3 id="rna-extraction-and-sequencing">RNA extraction and sequencing</h3>
<p>Recovering ancient RNA viruses requires workflows that minimize assumptions about which pathogens are present and that accommodate extensive molecular degradation. A practical approach begins with the extraction of total RNA, followed by metatranscriptomic sequencing in dedicated ancient DNA/RNA facilities, where DNA is enzymatically removed and fragmented RNA is converted into cDNA libraries suitable for short-read sequencing. These libraries enable broad, unbiased screening of viral, microbial, and host RNA.
Once RNA viruses are detected, targeted capture approaches can increase sequencing coverage and enable the recovery of near-complete genomes. Capture panels can be designed using modern viral diversity or draft ancient consensus sequences, allowing high sensitivity while minimizing strong assumptions about the targeted sequences, as hybridization-based methods remain flexible and can enrich even highly divergent molecules.</p>

<h3 id="bioinformatic-methodology-and-challenges">Bioinformatic methodology and challenges</h3>
<h4 id="detectability">Detectability</h4>

<p>Once ancient RNA molecules are sequenced, they must be detected and classified using appropriate analytical tools. Although RNA viruses often show higher short-term substitution rates than DNA viruses or host genomes [42], their long-term evolution is constrained by functional and host-dependent pressures, allowing conserved genes, particularly those encoding essential proteins such as the RNA-dependent RNA polymerase (RdRp), to retain detectable homology across deep timescales. However, over archaeological intervals, sequence divergence, postmortem damage, and fragmentation can reduce similarity between ancient fragments and modern references below the detection limits of standard mapping approaches. This challenge is further compounded by the incomplete representation of RNA virus diversity in public databases [43], which may bias detection when only distant references are available.
Therefore, while traditional nucleotide mapping strategies may be sufficient for relatively recent, century-old specimens, they might become less effective for more ancient samples. In these cases, shifting from nucleotide- to amino acid-based analyses can improve sensitivity, as protein sequences are more robust to synonymous substitutions and better preserve homology across evolutionary distances. Because applying BLASTX to all raw reads is computationally prohibitive, the search space can first be reduced using efficient k-mer classifiers such as Metabuli [44] to identify candidate viral reads. Faster homology search tools such as DIAMOND or MMseqs2 [45, 46] can then be applied to the filtered set against comprehensive databases such as NT. Together, these approaches enable sensitive homology detection while keeping computational demands manageable and complement traditional mapping pipelines used in ancient DNA studies.</p>

<h4 id="authentication-of-ancient-rna-molecules">Authentication of ancient RNA molecules</h4>

<p>Detection is only part of the problem, as authentication remains a central challenge. When sufficient closely related viral genomes provide a measurable temporal signal, molecular dating represents the most robust and reliable approach for validating viral age. Complementary evidence could be obtained from molecular damage patterns, but while such patterns are well characterized for ancient DNA, equivalent criteria for ancient RNA are still lacking. Studies of ancient host RNA suggest that RNA can exhibit characteristic damage signatures, including cytosine deamination influenced by secondary structure [31]. Systematic evaluation of these patterns will therefore be essential for distinguishing authentic ancient signals from modern contamination and technical artifacts, such as errors introduced during cDNA synthesis or spurious alignments of very short sequences [47]. 
RNA evolutionary models</p>

<p>Viral genomes recovered from ancient material provide rare, time-stamped calibration points that extend far beyond the narrow temporal window of modern sampling [9, 48–50]. Yet extracting deep evolutionary signal from these data remains challenging. Molecular dating approaches based solely on recent tip calibrations tend to overestimate short-term substitution rates and consequently compress deeper timescales. Because long-term purifying selection and substitutional saturation progressively erase observable substitutions, apparent evolutionary rates decline with temporal depth, a discrepancy known as the time-dependent rate phenomenon (TDRP) [51]. As a result, divergence times inferred from modern sequences alone are often systematically biased toward artificially recent estimates.
Extending the temporal depth of calibration through the inclusion of historical or ancient genomes can partially mitigate this bias. A clear illustration comes from measles virus, where incorporation of a 1912 genome together with selection-aware Bayesian molecular clock models shifted the estimated divergence between measles virus and rinderpest virus from medieval estimates based only on modern data (mean ∼899 CE) to the first millennium BCE (mean ∼528 BCE), nearly 1,500 years earlier than previously inferred [9]. This revision highlights how deeper calibration and models that explicitly accommodate time-varying evolutionary rates can substantially reshape reconstructions of RNA virus origins.
Today, methodological advances are beginning to address these limitations. Mechanistic molecular clock approaches, such as the “prisoner-of-war” model, explicitly link the apparent slowdown of evolutionary rates to substitutional saturation and functional constraint, thereby producing substantially older and more realistic divergence estimates than conventional clocks [52]. More broadly, continued progress will depend on analytical frameworks that jointly accommodate temporal rate variation, lineage-specific dynamics, and selective constraint while minimizing biases introduced by purifying selection, saturation, and uncertain calibrations [53]. Coupled with authentic ancient genomes, these models will transform isolated historical sequences into robust temporal anchors, enabling more reliable reconstructions of the long-term evolution of RNA viruses.</p>

<h2 id="conclusion">Conclusion</h2>
<p>Ancient RNA virus research is entering a pivotal phase. Proof of concept has now been demonstrated in humans and wildlife, and barriers once considered prohibitive are steadily being overcome. Progress now hinges on strategic sampling of biologically relevant tissues, robust standards for authentication, and innovative computational approaches. Meeting these challenges will allow RNA viruses to be integrated into ancient pathogen research, delivering a more balanced and historically grounded view of viral evolution and turning ancient genomics into a tool for addressing contemporary questions in virology.</p>

<h2 id="acknowledgements">Acknowledgements</h2>
<p>I thank Prof. Sébastien Calvignac-Spencer for insightful and helpful discussions.</p>

<h2 id="references">References</h2>

<p>[1] Rachel E. Baker et al. Infectious disease in an era of global change. Nature Reviews Microbiology, 20(4):193–205, 2022. doi: 10.1038/s41579-021-00639-z. URL https://www.nature.com/articles/s41579-021-00639-z.</p>

<p>[2] Aris Katzourakis and Robert J. Gifford. Endogenous viral elements in animal genomes. PLOS Genetics, 6(11):e1001191, 2010. doi: 10.1371/journal.pgen.1001191. URL https://dx.plos.org/10.1371/journal.pgen.1001191.</p>

<p>[3] Pakorn Aiewsakun and Aris Katzourakis. Endogenous viruses: Connecting recent and ancient viral evolution. Virology, 479–480:26–37, 2015. doi: 10.1016/j.virol.2015.02.011. URL https://linkinghub.elsevier.com/retrieve/pii/S0042682215000549.</p>

<p>[4] Gabriel Luz Wallau. RNA virus EVEs in insect genomes. Current Opinion in Insect Science, 49:42–47, 2022. doi: 10.1016/j.cois.2021.11.005. URL https://doi.org/10.1016/j.cois.2021.11.005.</p>

<p>[5] Yiqiao Li et al. Endogenous viral elements in shrew genomes provide insights into pestivirus ancient history. Molecular Biology and Evolution, 39(10):msac190, 2022. doi: 10.1093/molbev/msac190. URL https://doi.org/10.1093/molbev/msac190.</p>

<p>[6] Benjamin Guinet et al. Endoparasitoid lifestyle promotes endogenization and domestication of dsDNA viruses. eLife, 12:e85993, 2023. doi: 10.7554/eLife.85993. URL https://doi.org/10.7554/eLife.85993.</p>

<p>[7] Edward C. Holmes. The evolution of endogenous viral elements. Cell Host &amp; Microbe, 10(4):368–377, 2011. doi: 10.1016/j.chom.2011.09.002. URL https://doi.org/10.1016/j.chom.2011.09.002.</p>

<p>[8] Arthur Kocher et al. Ten millennia of hepatitis B virus evolution. Science, 374(6564):182–188, 2021. doi: 10.1126/science.abi5658. URL https://www.science.org/doi/10.1126/science.abi5658.</p>

<p>[9] Ariane Düx et al. Measles virus and rinderpest virus divergence dated to the sixth century BCE. Science, 368(6497):1367–1370, 2020. doi: 10.1126/science.aba9411. URL https://www.science.org/doi/10.1126/science.aba9411.</p>

<p>[10] Livia V. Patrono et al. Archival influenza virus genomes from Europe reveal genomic variability during the 1918 pandemic. Nature Communications, 13(1):2314, 2022. doi: 10.1038/s41467-022-29614-9. URL https://www.nature.com/articles/s41467-022-29614-9.</p>

<p>[11] Arthur Kocher, Johannes Krause, and Maria A. Spyrou. Insights into infectious diseases through ancient pathogen genomics. Nature Reviews Microbiology, 2025. doi: 10.1038/s41579-025-01259-7. URL https://www.nature.com/articles/s41579-025-01259-7.</p>

<p>[12] Kelly E. Blevins et al. Ancient DNA insights into diverse pathogens and their hosts. Nature Reviews Genetics, 27(1):96–111, 2026. doi: 10.1038/s41576-025-00912-4. URL https://www.nature.com/articles/s41576-025-00912-4.</p>

<p>[13] James A. Fellows Yates et al. Community-curated and standardised metadata of published ancient metagenomic samples with AncientMetagenomeDir. Scientific Data, 8:31, 2021. doi: 10.1038/s41597-021-00816-y.</p>

<p>[14] Tomoko Mihara et al. Linking virus genomes with host taxonomy. Viruses, 8(3):66, 2016. doi: 10.3390/v8030066. URL https://www.mdpi.com/1999-4915/8/3/66.</p>

<p>[15] Sarah L. Fordyce et al. Long-term RNA persistence in postmortem contexts. Investigative Genetics, 4:7, 2013. doi: 10.1186/2041-2223-4-7. URL https://link.springer.com/article/10.1186/2041-2223-4-7. PMCID: PMC3662605.</p>

<p>[16] Urmi Chheda et al. Factors affecting stability of RNA—temperature, length, concentration, pH, and buffering species. Journal of Pharmaceutical Sciences, 113(2):377–385, 2024. doi: 10.1016/j.xphs.2023.11.023. URL https://www.sciencedirect.com/science/article/pii/S0022354923004987.</p>

<p>[17] Mythily Srinivasan, Daniel Sedmak, and Scott Jewell. Effect of fixatives and tissue processing on the content and integrity of nucleic acids. The American Journal of Pathology, 161(6):1961–1971, 2002. doi: 10.1016/S0002-9440(10)64472-0. URL https://linkinghub.elsevier.com/retrieve/pii/S0002944010644720.</p>

<p>[18] Kelly A. Speer et al. A comparative study of RNA yields from museum specimens, including an optimized protocol for extracting RNA from formalin-fixed specimens. Frontiers in Ecology and Evolution, 10:953131, 2022. doi: 10.3389/fevo.2022.953131. URL https://www.frontiersin.org/articles/10.3389/fevo.2022.953131/full.</p>

<p>[19] H. Puchtler and S. N. Meloan. On the chemistry of formaldehyde fixation and its effects on immunohistochemical reactions. Histochemistry, 82(3):201–204, 1985. doi: 10.1007/BF00501395. URL https://doi.org/10.1007/BF00501395.</p>

<p>[20] Marc R. Friedländer and M. Thomas P. Gilbert. How ancient RNA survives and what we can learn from it. Nature Reviews Molecular Cell Biology, 25:417–418, 2024. doi: 10.1038/s41580-024-00726-y. URL https://doi.org/10.1038/s41580-024-00726-y.</p>

<p>[21] YongLi Xiao et al. High-throughput RNA sequencing of a formalin-fixed, paraffin-embedded autopsy lung tissue sample from the 1918 influenza pandemic. The Journal of Pathology, 229(4):535–545, 2013. doi: 10.1002/path.4145. URL https://pathsocjournals.onlinelibrary.wiley.com/doi/10.1002/path.4145.</p>

<p>[22] Erin E. Barnett et al. Recovery of an 18th century human rhinovirus genome through ancient RNA isolation of human lungs. bioRxiv, 2026. doi: 10.64898/2026.01.29.702071. URL https://www.biorxiv.org/content/10.64898/2026.01.29.702071v1. Preprint.</p>

<p>[23] Love Dalén et al. Deep-time paleogenomics and the limits of DNA survival. Science, 382(6666):48–53, 2023. doi: 10.1126/science.adh7943. URL https://doi.org/10.1126/science.adh7943.</p>

<p>[24] Martin Sikora et al. The spatiotemporal distribution of human pathogens in ancient Eurasia. Nature, 643(8073):1011–1019, 2025. doi: 10.1038/s41586-025-09192-8. URL https://doi.org/10.1038/s41586-025-09192-8.</p>

<p>[25] Christina Warinner et al. Ancient human microbiomes. Journal of Human Evolution, 79:125–136, 2015. doi: 10.1016/j.jhevol.2014.10.016. URL https://doi.org/10.1016/j.jhevol.2014.10.016.</p>

<p>[26] Benjamin Guinet et al. Ancient host-associated microbes obtained from mammoth remains. Cell, 188(23):6606–6619.e24, 2025. doi: 10.1016/j.cell.2025.08.003. URL https://linkinghub.elsevier.com/retrieve/pii/S0092867425009171.</p>

<p>[27] Antonio Fernandez-Guerra et al. Two-million-year-old microbial communities from the Kap København Formation in North Greenland. bioRxiv, 2023. doi: 10.1101/2023.06.10.544454. URL http://biorxiv.org/lookup/doi/10.1101/2023.06.10.544454. Preprint.</p>

<p>[28] Katherine Hearne et al. AncientMetagenomeDir Dating MetaDataset highlights need for standardised radiocarbon reporting in ancient DNA. bioRxiv, 2026. doi: 10.64898/2026.02.06.704039. URL https://doi.org/10.64898/2026.02.06.704039. Preprint posted February 6, 2026.</p>

<p>[29] Oliver Smith et al. Ancient RNA from late Pleistocene permafrost and historical canids shows tissue-specific transcriptome survival. PLOS Biology, 17(7):e3000166, 2019. doi: 10.1371/journal.pbio.3000166. URL https://dx.plos.org/10.1371/journal.pbio.3000166.</p>

<p>[30] Emilio Mármol-Sánchez et al. Historical RNA expression profiles from the extinct Tasmanian tiger. Genome Research, 33(8):1299–1316, 2023. doi: 10.1101/gr.277663.123. URL http://genome.cshlp.org/lookup/doi/10.1101/gr.277663.123.</p>

<p>[31] Emilio Mármol-Sánchez et al. Ancient RNA expression profiles from the extinct woolly mammoth. Cell, 189(1):52–69.e22, 2026. doi: 10.1016/j.cell.2025.10.025. URL https://linkinghub.elsevier.com/retrieve/pii/S0092867425012310.</p>

<p>[32] Mahtab Peyambari et al. A 1,000-year-old RNA virus. Journal of Virology, 93(1):e01188-18, 2019. doi: 10.1128/JVI.01188-18. URL https://journals.asm.org/doi/10.1128/JVI.01188-18.</p>

<p>[33] Tjorven Hinzke et al. RNA virus genomes from centuries- to millennia-old Adélie penguin mummies. bioRxiv, 2025. doi: 10.64898/2025.12.17.693957. URL http://biorxiv.org/lookup/doi/10.64898/2025.12.17.693957. Preprint.</p>

<p>[34] Aytul Buğra et al. Postmortem pathological changes in extrapulmonary organs in SARS-CoV-2 RT-PCR-positive cases: A single-center experience. Irish Journal of Medical Science, 191(1):81–91, 2022. doi: 10.1007/s11845-021-02638-8. URL https://link.springer.com/article/10.1007/s11845-021-02638-8.</p>

<p>[35] William J. Moss and Diane E. Griffin. What’s going on with measles? Journal of Virology, 98(8):e00758-24, 2024. doi: 10.1128/jvi.00758-24. URL https://journals.asm.org/doi/10.1128/jvi.00758-24.</p>

<p>[36] Rotem Ben-Shachar and Katia Koelle. Transmission-clearance trade-offs indicate that dengue virulence evolution depends on epidemiological context. Nature Communications, 9:2355, 2018. doi: 10.1038/s41467-018-04595-w. URL https://www.nature.com/articles/s41467-018-04595-w.</p>

<p>[37] Sinead E. Morris et al. Influenza virus shedding and symptoms: Dynamics and implications from a multiseason household transmission study. PNAS Nexus, 3(9):pgae338, 2024. doi: 10.1093/pnasnexus/pgae338. URL https://academic.oup.com/pnasnexus/article/3/9/pgae338.</p>

<p>[38] Andrea Misin et al. Measles: An overview of a re-emerging disease in children and immunocompromised patients. Microorganisms, 8(2):276, 2020. doi: 10.3390/microorganisms8020276. URL https://www.mdpi.com/2076-2607/8/2/276.</p>

<p>[39] Michelle Wille et al. RNA virome abundance and diversity is associated with host age in a bird species. Virology, 561:98–106, 2021. doi: 10.1016/j.virol.2021.06.007. URL https://www.sciencedirect.com/science/article/pii/S0042682221001379.</p>

<p>[40] Jaelle C. Brealey et al. Dental calculus as a tool to study the evolution of the mammalian oral microbiome. Molecular Biology and Evolution, 37(10):3003–3022, 2020. doi: 10.1093/molbev/msaa135. URL https://academic.oup.com/mbe/article/37/10/3003/5848415.</p>

<p>[41] Anoop Kaur Boparai et al. Dental calculus—an emerging bioresource for past SARS-CoV-2 detection, studying its evolution and relationship with oral microflora. Journal of King Saud University - Science, 35(4):102646, 2023. doi: 10.1016/j.jksus.2023.102646. URL https://jksus.org/dental-calculus-an-emerging-bio-resource-for-past-sars-cov2-detection-studying-its-evolution-and-relationship-with-oral-microflora/.</p>

<p>[42] Edward C. Holmes. The comparative genomics of viral emergence. Proceedings of the National Academy of Sciences of the United States of America, 107:1742–1746, 2010. doi: 10.1073/pnas.0906193106. URL https://pnas.org/doi/full/10.1073/pnas.0906193106.</p>

<p>[43] Guillermo Dominguez-Huerta et al. The RNA virosphere: How big and diverse is it? Environmental Microbiology, 25(1):209–215, 2023. doi: 10.1111/1462-2920.16312. URL https://sfamjournals.onlinelibrary.wiley.com/doi/10.1111/1462-2920.16312.</p>

<p>[44] Jaebeom Kim and Martin Steinegger. Metabuli: Sensitive and specific metagenomic classification via joint analysis of amino acid and DNA. Nature Methods, 21(6):971–973, 2024. doi: 10.1038/s41592-024-02273-y. URL https://www.nature.com/articles/s41592-024-02273-y.</p>

<p>[45] Benjamin Buchfink, Chao Xie, and Daniel H. Huson. Fast and sensitive protein alignment using DIAMOND. Nature Methods, 12(1):59–60, 2015. doi: 10.1038/nmeth.3176. URL https://www.nature.com/articles/nmeth.3176.</p>

<p>[46] Martin Steinegger and Johannes Söding. MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nature Biotechnology, 35(11):1026–1028, 2017. doi: 10.1038/nbt.3988. URL https://www.nature.com/articles/nbt.3988.</p>

<p>[47] Sébastien Calvignac-Spencer and Carles Lalueza-Fox. A trunkload of ancient RNA. Cell, 189(1):1–2, 2026. doi: 10.1016/j.cell.2025.12.004. URL https://linkinghub.elsevier.com/retrieve/pii/S0092867425014175.</p>

<p>[48] Luca Nishimura et al. Detection of ancient viruses and long-term viral evolution. Viruses, 14(6):1336, 2022. doi: 10.3390/v14061336. URL https://www.mdpi.com/1999-4915/14/6/1336.</p>

<p>[49] Axel A. Guzmán-Solís et al. Ancient viral genomes reveal introduction of human pathogenic viruses into Mexico during the transatlantic slave trade. eLife, 10:e68612, 2021. doi: 10.7554/eLife.68612. URL https://elifesciences.org/articles/68612.</p>

<p>[50] Ophélie Lebrasseur, Kuldeep Dilip More, and Ludovic Orlando. Equine herpesvirus 4 infected domestic horses associated with Sintashta spoke-wheeled chariots around 4,000 years ago. Virus Evolution, 10(1):vead087, 2024. doi: 10.1093/ve/vead087. URL https://academic.oup.com/ve/article/10/1/vead087/7523751.</p>

<p>[51] Pakorn Aiewsakun and Aris Katzourakis. Time-dependent rate phenomenon in viruses. Journal of Virology, 90(16):7184–7195, 2016. doi: 10.1128/JVI.00593-16. URL https://journals.asm.org/doi/10.1128/JVI.00593-16.</p>

<p>[52] Mahan Ghafari et al. A mechanistic evolutionary model explains the time-dependent pattern of substitution rates in viruses. Current Biology, 31(21):4689–4696.e5, 2021. doi: 10.1016/j.cub.2021.08.020. URL https://linkinghub.elsevier.com/retrieve/pii/S0960982221011246.</p>

<p>[53] Jonathon C. O. Mifsud et al. Recent advances in the inference of deep viral evolutionary history. Journal of Virology, 99(9):e00292-25, 2025. doi: 10.1128/jvi.00292-25. URL https://journals.asm.org/doi/10.1128/jvi.00292-25.</p>]]></content><author><name>Benjamin Guinet</name></author><category term="Blog" /><category term="spaam," /><category term="blog" /><summary type="html"><![CDATA[by Benjamin Guinet]]></summary></entry><entry><title type="html">SPAAM8</title><link href="https://spaam-community.github.io/spaam-community.github.io-dev//events/news/2026/04/01/spaam8/" rel="alternate" type="text/html" title="SPAAM8" /><published>2026-04-01T00:00:00+00:00</published><updated>2026-04-01T00:00:00+00:00</updated><id>https://spaam-community.github.io/spaam-community.github.io-dev//events/news/2026/04/01/spaam8</id><content type="html" xml:base="https://spaam-community.github.io/spaam-community.github.io-dev//events/news/2026/04/01/spaam8/"><![CDATA[<p><img src="/assets/media/Flyer_SPAAM8.jpg" width="600px" /></p>

<p>We’re happy to announce SPAAM8 – a gathering of experts in ancient metagenomics research!</p>

<p>SPAAM8 will be taking place as a hybrid satellite meeting prior to the ICP 2026 in Stockholm!</p>

<h2 id="when">When?</h2>
<p>22nd June 2026</p>

<h2 id="where">Where?</h2>
<p>🇸🇪 This year, the in-person event will be hosted in Stockholm, in the same venue as ICP 2026.</p>

<p>👩‍💻 In case you can’t make it to Stockholm, there will be a possibility to attend online.</p>

<h2 id="abstract-submission">Abstract submission</h2>
<p>Please, submit your abstract through <a href="https://forms.gle/psw9GS3ZyrLVPbbo7">this form</a>. 
<strong>Deadline:</strong> May 1st, 2026</p>

<h2 id="event-registration">Event Registration</h2>
<p>Registration is now open! Please register through <a href="https://forms.gle/bmbTZxi9VGsXbTJ18">this form:</a> <strong>before June 10th, 2026</strong>.</p>

<p><strong>In person attendance fee: €25</strong>. A payment link will be sent to all registered participants in the coming days.
Includes coffee, snacks and catering throughout the day.</p>

<div id="paypal-container-6JQXPUUDC32HL"></div>
<script src="https://www.paypal.com/sdk/js?client-id=BAA3mcOCmIeVqe9TrSpFq410LaHUj0OVyayUmM5jUGvKDs7Vr3DcHeqPx7p82gkwr6DsUaiud8sh0Nytvg&amp;components=hosted-buttons&amp;disable-funding=venmo&amp;currency=EUR">
</script>

<script>
  paypal.HostedButtons({
    hostedButtonId: "6JQXPUUDC32HL",
  }).render("#paypal-container-6JQXPUUDC32HL")
</script>]]></content><author><name>SPAAM Community</name></author><category term="Events" /><category term="News" /><summary type="html"><![CDATA[We’re happy to announce SPAAM8 – a gathering of experts in ancient metagenomics research! SPAAM8 will be taking place as a hybrid satellite meeting prior to the ICP 2026 in Stockholm! When? 22nd June 2026 Where? 🇸🇪 This year, the in-person event will be hosted in Stockholm, in the same venue as ICP 2026.]]></summary></entry><entry><title type="html">Wrangling Sequencing Data. Part 1: File Formats</title><link href="https://spaam-community.github.io/spaam-community.github.io-dev//blog/2026/02/02/blog-seqdata1/" rel="alternate" type="text/html" title="Wrangling Sequencing Data. Part 1: File Formats" /><published>2026-02-02T00:00:00+00:00</published><updated>2026-02-02T00:00:00+00:00</updated><id>https://spaam-community.github.io/spaam-community.github.io-dev//blog/2026/02/02/blog-seqdata1</id><content type="html" xml:base="https://spaam-community.github.io/spaam-community.github.io-dev//blog/2026/02/02/blog-seqdata1/"><![CDATA[<p><em>by <a href="https://www.oliviasmith.me/home">Olivia Smith</a></em></p>

<p>Welcome to “Wrangling Sequencing Data” - yeehaw! There’s a lot of different ways to tackle this, so this is an attempt at the guide to metagenomics data management that I needed when I first started. There’s a substantial amount of ground to cover so this will come in three parts. In Part 1, I’ll go into <em>what</em> kinds of files you might encounter (will be familiar to more experienced folks). In Part 2, I’ll discuss <em>how</em> to get them (hopefully helpful to most). Finally, in Part 3, I’ll briefly discuss a couple of tools for working with all these files, but I’ll leave most of the deeper investigation on those tools up to you.</p>

<p>How I recommend using this guide: read relevant or interesting-to-you sections as you go. There’s a lot of info here from file types to bash loops to tool recommendations, and I don’t expect all of it to be necessary or helpful for everyone.</p>

<p><strong>Also. There’s a tl;dr below each section header.</strong></p>

<h2 id="part-1-file-formats">Part 1: File Formats</h2>

<p>There are 3 different kinds of files ~metagenomicists~ (you!) commonly come across. These are all text files. It’s important to know the difference between them so that you can work with them appropriately.</p>

<p><strong>tl;dr</strong>
Sequencing Reads: List of named reads with base call quality scores
Alignment files: Map of the ordered reads in the overall sequence
Index files: Complementary file which makes it less computationally intensive to scan through a big file
Consensus sequence: FASTA file with list of named, ‘pretty’ sequences</p>

<h3 id="sequencing-reads">Sequencing Reads</h3>

<p><strong>Sequencing reads (FASTQ)</strong> are the most ‘raw’ version of sequencing data that you will get. This is what the sequencing center will return after a sequencing run. These files usually have extension <code class="language-plaintext highlighter-rouge">.fastq</code> or <code class="language-plaintext highlighter-rouge">.fastq.gz</code> (<code class="language-plaintext highlighter-rouge">.gz</code> indicates that the file was compressed with the software tool <code class="language-plaintext highlighter-rouge">gzip</code>). I like to think of them as just a big alphabet soup of reads.</p>

<p>Reads can either be <strong>single-end</strong> (reading in one direction) or <strong>paired-end</strong> (reads going both backwards and forwards). Paired-end sequencing is more expensive, but it improves the quality of assemblies by allowing you to better place the reads in the right location. When they are paired-end, you will most often see <em>two</em> <code class="language-plaintext highlighter-rouge">.fastq</code> files (forward and backward). <strong>If so, you will want both of these files.</strong></p>

<p>The file will be full of reads shown in this seemingly unintelligible format:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>@ERR3654168.1 M_PC0253:13:C3N9HACXX:8:1114:7573:78875/1
TCTGCACAGATTTCGGTGGTACTCTGAAGGCGGAGCAC
+
JJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJ
</code></pre></div></div>

<p>This table shows a line-by-line breakdown of each entry (read) in the file:</p>

<p><img src="/assets/media/blog-seqdata1-fastq-file.png" alt="fastq-file" width="100%" class="center" /></p>

<p>†Some databases or organizations have specific formats for the headers. <a href="https://en.wikipedia.org/wiki/FASTQ_format">Wikipedia</a> has a good explanation of this.</p>

<p>Line 4 is the kind of weird one (but a really important one!). The scores <em>aren’t just numbers</em> like you’d expect, but instead are ASCII characters. We also see this in alignment files, as noted below. The quality scores follow this pattern (increasing quality from left to right, so <code class="language-plaintext highlighter-rouge">!</code> is lowest quality, <code class="language-plaintext highlighter-rouge">~</code> is highest):</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>!"#$%&amp;'()*+,-./0123456789:;&lt;=&gt;?@ABCDEFGHIJKLMNOPQRSTUVWXYZ[\]^_`abcdefghijklmnopqrstuvwxyz{|}~
</code></pre></div></div>

<p><strong>Quality scores are very important</strong>, because it makes it possible to filter out issues in the sequencing data before you work with the files. Even with ancient sequencing data, you will want to do some form of quality control. Often when working with aDNA, the ends of the reads are trimmed (removed, due to damage often occurring there), a quality score threshold is used, and reads are constrained by length (again, since aDNA is damaged and so doesn’t generally produce the normal 250+ bp reads we see in non-degraded DNA.</p>

<h3 id="alignment-files">Alignment Files</h3>

<p>Once you have your sequence read files, the next step is often to create an <strong>alignment</strong>. Alignment software figures out which reads overlap and what order they should go in to recreate the DNA sequence. I’m not experienced with aligning so I won’t discuss how it works any further –  see <a href="https://www.sciencedirect.com/science/article/pii/S0888754317300551">Chowdhury and Garai (2017)</a> for a review on technical details and various tools. Alignment generally creates either a <strong>SAM</strong> or <strong>BAM</strong> file with extension <code class="language-plaintext highlighter-rouge">.sam</code> or <code class="language-plaintext highlighter-rouge">.bam</code>, respectively. Most often you will use BAM files since these are binary (compressed) versions.</p>

<p>If you were to open a BAM file and look at it in the terminal you would see this (command to do so is the first line):</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>$ samtools view OBP001.TF.bam | less

M_K00233:75:HMVJGBBXX:7:2110:5477:22590 0   	1   	882544  37  	73M 	*   	0   	0   	NNNNNNNNNNATTTGGCCCCTCACCAAAAACATTTTCTGACCCCTACCCCAGACCCCGACCCTNNNNNNNNNN   	!!!!!!!!!!JJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJJ!!!!!!!!!!   	XT:A:U  NM:i:0  X0:i:1  X1:i:0  XM:i:0  XO:i:0  XG:i:0  MD:Z:73 RG:Z:ILLUMINA-OBP001.A0101.TF1.1_S0_L007_R1_001.fastq
</code></pre></div></div>

<p>Just like with the fastq files, this isn’t <em>really</em> meant to be read by you! It’s meant to be read by a computer. But knowing what it should look like is helpful when you need a sanity check. It’s good practice to sanity check files when you first download them, especially if you didn’t generate them. Below is a table from the <a href="https://samtools.github.io/hts-specs/SAMv1.pdf">SAMTools documentation</a> describing each column. It’s unusual that you’ll need to work with all of these, but it’s again useful to point out the quality scores for the mapping (MAPQ) and base calling (QUAL). <strong>Remember: you will use these columns for quality filtering before analysis.</strong></p>

<p><img src="/assets/media/blog-seqdata1-bam-file.png" alt="bam-file" width="100%" class="center" /></p>

<h3 id="index-files">Index Files</h3>

<p>In addition to your alignment files, you’ll need a partner file called an <strong>index file</strong>. This file is important because it allows for faster searching when working with the file in SAMtools or BCFtools.</p>

<p>Most often when you download your BAM file, it won’t come with the index file. You can generate it very easily using:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>$ samtools index OBP001.TF.bam
</code></pre></div></div>

<p>By default, this will create a file with the same name and <code class="language-plaintext highlighter-rouge">.bai</code> added to the end (in this case <code class="language-plaintext highlighter-rouge">OBP001.TF.bam.bai</code>). <strong>Note that it is required that the file has the correct naming convention and is in the same location as the alignment file.</strong> But just let SAMtools do its default thing and you’ll be good!</p>

<p>I’m not going to show you this file, because it <em>really really</em> doesn’t make any sense. But you can check that it’s not empty using the bash command for word count, <code class="language-plaintext highlighter-rouge">wc</code>. The <code class="language-plaintext highlighter-rouge">-l</code> flag will return the number of lines and full file path. If it returns 0 or an empty line, something is wrong.</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>$ wc -l OBP001.TF.bam.bai
13652 /scratch1/07869/osmith/Tutorial/OBP001.TF.bam.bai
</code></pre></div></div>

<h3 id="consensus-sequences">Consensus Sequences</h3>

<p>A <strong>consensus sequence (FASTA)</strong> is the most palatable version of a sequence and usually has the file extension <em>.fasta</em>. It is just the ‘refined’ sequence with headers specifying the sequence name (often a gene name, contig, or chromosome, though it can be anything).</p>

<p>These files look like some version of this (where each &gt; marks a new sequence):</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>&gt;NC_000020.11 Homo sapiens chromosome 20, GRCh38.p14 Primary Assembly
NNNNNNNNNNTGTTCAGTCGGGCAGGGAGTGGGAATAGACAAGACCACAAGCAGCTTGGTGCCTCTGAAAGG
GAGAGGGGTGGAGGGGAGACTAGAGAGGTGGGTAGGAATACTGGATTCCACTGACCACGTGCTGGATGTCAT
GCTTAGCCCTCCTGCTCTGTGCCAGGTTAGGCACCTGGTGTTTTACATATATTATATTACATTCTATTACAG
ACAACTCCATAGCAATCCTTTCCTCTCCATTCCATTTCTCTCCACTCCATCCCATTCCATTCCA

&gt;NC_000021.9 Homo sapiens chromosome 21, GRCh38.p14 Primary Assembly
NNNNNNNNNNNNNNNNNNGATCCACCCGCCTTGGCCTCCTAAAGTGCTGGGATTACAG
GTGTTAGCCACCACGTCCAGCTGTTAATTTTTATTTAATAAGAATGACAGAGTGAGGGCCATCACTGTTA
ATGAAGCCAGTGTTGCTCACAGCCTCCCCTTGGTCACTTTTTGTGACTGAAGGGCATGTGTTCAGGCAAG
ATTGTTGGGTGGCTGTGTTTTGTCTTCTTCCAGCTCGGCCATGGAATAGCC

</code></pre></div></div>
<p>While consensus sequences are common in modern genomics (including modern metagenomics and genetic engineering), they are less often seen in paleogenomics. This is because damaged DNA makes it more challenging to confidently call the sequence.</p>

<p>You can also generate index files for consensus FASTA sequences using
<code class="language-plaintext highlighter-rouge">samtools faidx {FILE}</code>
which makes the file less computationally intensive to work with (see ‘Index Files’ above) and will generate a file with the same name plus the <code class="language-plaintext highlighter-rouge">.faidx</code> extension.</p>

<p>The difference between FASTQ, BAM and FASTA files. Created with Biorender.com
<img src="/assets/media/blog-seqdata1-biorender.png" alt="biorender" width="100%" class="center" /></p>

<hr />

<p>That’s all for Part 1, be sure to come back to read Part 2: Download Tools!</p>]]></content><author><name>Olivia Smith</name></author><category term="Blog" /><category term="spaam," /><category term="blog" /><summary type="html"><![CDATA[by Olivia Smith]]></summary></entry><entry><title type="html">Wrangling Sequencing Data. Part 2: File Formats</title><link href="https://spaam-community.github.io/spaam-community.github.io-dev//blog/2026/02/02/blog-seqdata2/" rel="alternate" type="text/html" title="Wrangling Sequencing Data. Part 2: File Formats" /><published>2026-02-02T00:00:00+00:00</published><updated>2026-02-02T00:00:00+00:00</updated><id>https://spaam-community.github.io/spaam-community.github.io-dev//blog/2026/02/02/blog-seqdata2</id><content type="html" xml:base="https://spaam-community.github.io/spaam-community.github.io-dev//blog/2026/02/02/blog-seqdata2/"><![CDATA[<p><em>by <a href="https://www.oliviasmith.me/home">Olivia Smith</a></em></p>

<p>This is Part 2 of the Wrangling Sequencing Data series and it will focus of Download Tools. To find out about different file formats, make sure you have read <a href="https://www.spaam-community.org/blog/2026/02/02/blog-seqdata1/">Part 1</a>.</p>

<p>One of the things that I often found challenging when I first started was figuring out how to find the best databases and download tools. Here, I’ll discuss a couple options.</p>

<p><strong>tl;dr</strong>
You can download practically any file you’d want using a command line HTTP/FTP package (usually wget for Linux, curl for MacOS). Some databases also have browser interfaces or command line APIs; ENA has enaBrowserTools and NCBI has SRAToolkit. Learn and use bash loops for repetitive tasks!</p>

<h3 id="command-line">Command Line</h3>
<h4 id="wget-linux">wget (Linux)</h4>

<p>wget is a tool that can be used to download any HTTP(S) and FTP(S) from the internet, which is perfect for our uses (you can see more <a href="https://www.gnu.org/software/wget/">wget details</a> on the GNU Project page). It is installed by default on most Linux machines. It seems to also be available for Windows PowerShell with a slightly different syntax, but I have never used it myself. wget is not installed on MacOS (Unix) systems by default, but you can add it using another package manager (e.g. brew install wget to install with <a href="https://brew.sh/">Homebrew</a>) or use curl instead (see below).</p>

<p>To download, simply use <em>wget {URL}</em>, as so:</p>
<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>wget ftp://ftp.sra.ebi.ac.uk/vol1/fastq/ERR365/008/ERR3654168/ERR3654168.fastq.gz
</code></pre></div></div>

<p>If your command works, you should see a few statements like this letting you know that it is accessing the data:</p>
<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>Resolving ftp.sra.ebi.ac.uk (ftp.sra.ebi.ac.uk)... 193.62.193.138
Connecting to ftp.sra.ebi.ac.uk (ftp.sra.ebi.ac.uk)|193.62.193.138|:21... connected.
Logging in as anonymous ... Logged in!
</code></pre></div></div>

<p>And then a progress bar will appear. Once the download is done, it will let you know what the saved file is called - in this case the file is called “ERR3654168.fastq.gz”. By default this will be the name of the file from the source.</p>
<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>2023-04-26 11:48:32 (16.8 MB/s) - ‘ERR3654168.fastq.gz’ saved [60429342]
</code></pre></div></div>

<h4 id="curl-macos-unix">curl (MacOS, Unix)</h4>

<p>On MacOS, you can use <a href="https://curl.se/">curl</a> instead! And the good news is it works the same way, but you’ll want to explicitly tell it where to save the file.</p>

<p>Use curl -o {OUT_FILE} {URL} and you’re good to go.</p>
<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>curl –o ERR3654168.fastq.gz ftp://ftp.sra.ebi.ac.uk/vol1/fastq/ERR365/008/ERR3654168/ERR3654168.fastq.gz
</code></pre></div></div>

<p>With curl you should just get your download details right away (it will show 100 under % Total when it’s done):</p>
<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code> % Total	% Received % Xferd  Average Speed   Time	Time 	Time  Current
                                   Dload   Upload  Total     Spent	Left  Speed
 30 57.6M    30 17.6M	0 	0  2829k  	0  0:00:20  0:00:06  0:00:14 3645k
</code></pre></div></div>

<h3 id="european-nucleotide-archive-ena-toolkit">European Nucleotide Archive (ENA) Toolkit</h3>
<p>The <a href="https://www.ebi.ac.uk/ena/browser/home">European Nucleotide Archive</a> (ENA) is where you will find most of the metagenomic data from publications in our field. If you want to use the browser to download files, I’ve created a walkthrough here: <a href="https://drive.google.com/file/d/1rRBSqLbZLe-Y7A0dEfxm43cIwHjYqS02/view?usp=share_link">Downloading Metagenomic Data from ENA</a></p>

<p>The enaBrowserTools toolkit is great for downloading lots of sequences at one time (without writing a bash loop!). Download and install enaBrowserTools using the instructions <a href="https://github.com/enasequence/enaBrowserTools/blob/master/README.md">here</a>.</p>

<h3 id="national-center-for-biotechnology-information-ncbi-toolkit">National Center for Biotechnology Information (NCBI) Toolkit</h3>
<p>NCBI has built a command line tool for accessing their sequencing database called SRA Toolkit. Download and install SRA Toolkit using the instructions <a href="https://github.com/ncbi/sra-tools/wiki/01.-Downloading-SRA-Toolkit">here</a>. The SRA Toolkit quickstart tutorial for downloading fastq files is <a href="https://github.com/ncbi/sra-tools/wiki/08.-prefetch-and-fasterq-dump">here</a>.</p>

<h3 id="bash-loops">Bash loops</h3>
<p>One last thing – it will be a huuuuuuuge help if you learn how to write simple loops in bash to perform downloads. This will make it so that you can take many links, accession numbers, etc. and just let the terminal do its thing while you drink coffee.</p>

<p>Let’s walk through a simple bash loop (the base and $ are just to show that these are two different commands):</p>
<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>(base) $ var=(2 3 6 8 11)
(base) $ for v in ${var[@]}; do echo $v; done
</code></pre></div></div>

<p>Let’s break this down:</p>
<ol>
  <li>Use for to mark the beginning of the loop</li>
  <li>Create a variable var (can be named anything) and give it a value which is an array of integers. Notation for arrays in bash is variable=(value_1 value_2 … value_n). Note that in bash we use spaces to separate values in an array.</li>
  <li>Loop over every element (number) in that variable using the syntax ${YOUR_VARIABLE[@]} . The [@] says that we want to go through each element instead of specifying just one of them (i.e. ${var[1]} would return the value 2).</li>
  <li>Follow this with ; do and our command; in my example I used echo, which is just a print statement. Here again we see the notation for variables which is $ followed by the variable name. In cases where the variable is surrounded by other characters, you can use ${YOUR_VARIABLE} to ensure that your command is executed properly.
Use ; done to say where the loop should end.</li>
</ol>

<p>Here’s an example of a loop to download a few different sets of BAMs using wget (these are not real, don’t try to actually download them):</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>(base) $ ftps=(ftp://ftp.example01.fastq.gz ftp://ftp.example02.fastq.gz ftp://ftp.example03.fastq.gz)
(base) $ for f in ${ftps[@]}; do wget ${f}; done
</code></pre></div></div>

<p>Two things: Remember to always use the three terms for, do, and done, each separated by ;. Also note that you don’t need to use quotation marks around characters the way you do in many other programming languages.</p>

<p>One final note: I constantly use loops to split up jobs across chromosomes (or sequence chunks, whatever works for you) since parallelizing your work is important for speed. Here I use BCFTools to make pileups from an alignment chromosome by chromosome. Check out <a href="https://samtools.github.io/bcftools/bcftools.html">their documentation</a> if you want to learn more about the command, options, and flags.</p>

<p>Looks like this:</p>
<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>for c in {1..22}; do bcftools mpileup -b -o mypileup.${c}.bam myalignment.bam ${c}; done
</code></pre></div></div>

<p>That’s the end of Part 2. Be sure to check out Part 3 on Analysis Tools!</p>]]></content><author><name>Olivia Smith</name></author><category term="Blog" /><category term="spaam," /><category term="blog" /><summary type="html"><![CDATA[by Olivia Smith]]></summary></entry><entry><title type="html">Wrangling Sequencing Data. Part 3: File Formats</title><link href="https://spaam-community.github.io/spaam-community.github.io-dev//blog/2026/02/02/blog-seqdata3/" rel="alternate" type="text/html" title="Wrangling Sequencing Data. Part 3: File Formats" /><published>2026-02-02T00:00:00+00:00</published><updated>2026-02-02T00:00:00+00:00</updated><id>https://spaam-community.github.io/spaam-community.github.io-dev//blog/2026/02/02/blog-seqdata3</id><content type="html" xml:base="https://spaam-community.github.io/spaam-community.github.io-dev//blog/2026/02/02/blog-seqdata3/"><![CDATA[<p><em>by <a href="https://www.oliviasmith.me/home">Olivia Smith</a></em></p>

<p><strong>Don’t forget to check out Part 1 and 2!</strong></p>

<p><a href="https://www.spaam-community.org/blog/2026/02/02/blog-seqdata1/">Part 1: File formats</a></p>

<p><a href="https://www.spaam-community.org/blog/2026/02/02/blog-seqdata2/">Part 2: Download tools</a></p>

<p><strong>tl;dr</strong>
There’s tons of options for working with sequencing reads and consensus sequences, so pick whatever floats your boat and/or what the nearest-to-you scholar uses. You should learn SAMTools and BCFTools for alignment files. My best advice is <strong>(1) read the documentation and (2) read all the options/flags for a function before you use it.</strong></p>

<h3 id="sequencing-reads-fastp-fastqc-gatk">Sequencing reads: fastp, fastqc, GATK</h3>
<p>The first step in many analyses is processing of the raw sequencing reads in some way. A huge number of tools are out there for this, but I like <a href="https://github.com/OpenGene/fastp">fastp</a> (<a href="https://academic.oup.com/bioinformatics/article/34/17/i884/5093234">Chen, et al., 2018</a>) and many folks use <a href="https://gatk.broadinstitute.org/hc/en-us">GATK</a> from the Broad Institute or <a href="https://bioinf.shenwei.me/seqkit/">SeqKit</a> (<a href="https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0163962">Shen, et al., 2016</a>). <a href="https://www.bioinformatics.babraham.ac.uk/projects/fastqc/">FastQC</a> is another nice tool for QC which creates really readable HTML reports!</p>

<p>A few examples of things you might do in ancient metagenomics with reads:</p>
<ul>
  <li>Filter on base call quality</li>
  <li>Trim adapters used for sequencing and barcoding but that we don’t want for our analysis</li>
  <li>Merge paired-end reads</li>
  <li>Align them!</li>
</ul>

<h3 id="alignment-and-index-files-samtools-bcftools">Alignment and index files: SAMTools, BCFTools</h3>
<p>You need to know how to use <a href="http://www.htslib.org/">SAMTools</a> and <a href="https://samtools.github.io/bcftools/bcftools.html">BCFTools</a>. They have great documentation and are really powerful, so are worth the investment to understand. Also this is the way that you will generate the index file to go with your alignment, which is usually required for any tool that uses alignments.</p>

<p>Some examples of what you can do with SAMTools and BCFTools:</p>
<ul>
  <li><a href="https://samtools.github.io/bcftools/bcftools.html#call">Call</a> variants</li>
  <li>Build a <a href="https://samtools.github.io/bcftools/bcftools.html#consensus">consensus</a> sequence</li>
  <li>Make <a href="https://samtools.github.io/bcftools/bcftools.html#mpileup">pileups</a> to count the base calls at each position or compute genotype likelihoods</li>
  <li>Take a random subset or fraction of reads to test on (an option when using <a href="https://www.htslib.org/doc/samtools-view.html">view</a>)</li>
</ul>

<h3 id="consensus-sequences-blast-geneious-benchling-ape">Consensus sequences: BLAST, Geneious, Benchling, ApE</h3>
<p>You’ve perhaps seen these before, as these are the most accessible tools. They often have straightforward user interfaces via applications or browsers! These are great for things like homology analysis (<a href="https://blast.ncbi.nlm.nih.gov/Blast.cgi">BLAST</a>), or plasmid visualization and genetic engineering (<a href="https://www.geneious.com/">Geneious</a>, <a href="https://www.benchling.com/">Benchling</a>, <a href="https://jorgensen.biology.utah.edu/wayned/ape/">ApE</a>). These tools are super powerful and flexible, so anything you can imagine doing with a consensus sequence is probably possible in one of these.</p>

<h3 id="managing-environments">Managing environments</h3>
<p>As you can tell, there’s a broad range of tools and software used in the field. It is good practice to “manage your computational environments” – you can imagine this as organizing boxes of stuff. Environments are isolated computational setups that keep your software dependencies from conflicting and causing problems; each environment has only what is necessary for that task. <a href="https://conda.io/projects/conda/en/latest/user-guide/install/index.html">Conda</a> is a widely used package/environment software which makes installing lots of the tools I’ve mentioned here easy (I use miniconda). You can do similar environment management <a href="https://docs.python.org/3/library/venv.html">with Python</a>. Also available are container platforms like <a href="https://www.reddit.com/r/docker/comments/keq9el/please_someone_explain_docker_to_me_like_i_am_an/">Docker</a> or <a href="https://docs.sylabs.io/guides/3.5/user-guide/introduction.html">Singularity</a>. These are more common when working in industry settings or with high performance computing environments.</p>

<hr />

<p>Olivia Smith is a PhD candidate at The University of Texas at Austin working under the supervision of Dr. Arbel Harpak. You can find her on Twitter at <a href="https://twitter.com/smitholivias">@SmithOliviaS</a>.</p>]]></content><author><name>Olivia Smith</name></author><category term="Blog" /><category term="spaam," /><category term="blog" /><summary type="html"><![CDATA[by Olivia Smith]]></summary></entry><entry><title type="html">SPAAM Newsletter #8</title><link href="https://spaam-community.github.io/spaam-community.github.io-dev//news/2026/01/07/newsletter/" rel="alternate" type="text/html" title="SPAAM Newsletter #8" /><published>2026-01-07T00:00:00+00:00</published><updated>2026-01-07T00:00:00+00:00</updated><id>https://spaam-community.github.io/spaam-community.github.io-dev//news/2026/01/07/newsletter</id><content type="html" xml:base="https://spaam-community.github.io/spaam-community.github.io-dev//news/2026/01/07/newsletter/"><![CDATA[<p>Hello all!</p>

<p>We are excited to share with you the latest <a href="/assets/media/SPAAM_newsletter_8.html">SPAAM Newsletter</a>!
In it, you’ll find:</p>

<ul>
  <li>
    <p>updates on recent and upcoming <strong>events and workshops</strong></p>
  </li>
  <li>
    <p><strong>funding and research opportunities</strong> worth exploring</p>
  </li>
  <li>
    <p>highlights of the <strong>latest publications</strong> in ancient metagenomics
We hope you enjoy this edition and look forward to staying connected with you!</p>
  </li>
</ul>]]></content><author><name>SPAAM Community</name></author><category term="News" /><summary type="html"><![CDATA[Hello all!]]></summary></entry><entry><title type="html">SPAAM Blog Post Competition</title><link href="https://spaam-community.github.io/spaam-community.github.io-dev//news/2025/11/25/blog-post-competition/" rel="alternate" type="text/html" title="SPAAM Blog Post Competition" /><published>2025-11-25T00:00:00+00:00</published><updated>2025-11-25T00:00:00+00:00</updated><id>https://spaam-community.github.io/spaam-community.github.io-dev//news/2025/11/25/blog-post-competition</id><content type="html" xml:base="https://spaam-community.github.io/spaam-community.github.io-dev//news/2025/11/25/blog-post-competition/"><![CDATA[<p><img src="https://spaam-community.github.io/assets/media/SPAAM_blog_competition.png" alt="SPAAM Blog competition" width="400px" class="center" /></p>

<p>The SPAAM Community is thrilled to announce our first-ever blog post competition!</p>

<p>We invite all members of the SPAAM community to submit an original blog post on a topic of their choice related to ancient DNA; whether it’s a reflection on your research, an exciting methodological advance, an ethical discussion, or a creative take on the field.</p>

<h2 id="-what-you-need-to-know">📝 What you need to know:</h2>

<p>• Theme: Anything connected to ancient DNA</p>

<p>• Length: ~1000–2000 words</p>

<p>• Audience: The broader ancient DNA community and interested readers</p>

<p>• Deadline: 15th February 2026</p>

<p>• Submission: please send your piece to the SPAAM email spaam.community@gmail.com with the title “blog post submission”</p>

<h2 id="-the-winner">🏆 The winner:</h2>

<p>You will not only get a chance of having your post featured on the SPAAM website but also win a mystery book prize!</p>

<h2 id="-publication">✨ Publication:</h2>

<p>All submissions will be considered for publication after an editing and proofreading process.</p>

<p>This is a great opportunity to share your perspective, practice science communication, and engage with the community.</p>

<p>📢 We can’t wait to read your stories and insights. Let your creativity shine!</p>]]></content><author><name>SPAAM Community</name></author><category term="News" /><summary type="html"><![CDATA[]]></summary></entry><entry><title type="html">SPAAM Newsletter #7</title><link href="https://spaam-community.github.io/spaam-community.github.io-dev//news/2025/11/07/newsletter/" rel="alternate" type="text/html" title="SPAAM Newsletter #7" /><published>2025-11-07T00:00:00+00:00</published><updated>2025-11-07T00:00:00+00:00</updated><id>https://spaam-community.github.io/spaam-community.github.io-dev//news/2025/11/07/newsletter</id><content type="html" xml:base="https://spaam-community.github.io/spaam-community.github.io-dev//news/2025/11/07/newsletter/"><![CDATA[<p>Hello all!</p>

<p>We are excited to share with you the latest <a href="/assets/media/SPAAM_newsletter_7.html">SPAAM Newsletter</a>!
In it, you’ll find:</p>

<ul>
  <li>
    <p>updates on recent and upcoming <strong>events and workshops</strong></p>
  </li>
  <li>
    <p><strong>funding and research opportunities</strong> worth exploring</p>
  </li>
</ul>

<p>highlights of the <strong>latest publications</strong> in ancient metagenomics
We hope you enjoy this edition and look forward to staying connected with you!</p>]]></content><author><name>SPAAM Community</name></author><category term="News" /><summary type="html"><![CDATA[Hello all!]]></summary></entry><entry><title type="html">ISBA11 Debriefing</title><link href="https://spaam-community.github.io/spaam-community.github.io-dev//blog/2025/10/23/isba11-debrief/" rel="alternate" type="text/html" title="ISBA11 Debriefing" /><published>2025-10-23T00:00:00+00:00</published><updated>2025-10-23T00:00:00+00:00</updated><id>https://spaam-community.github.io/spaam-community.github.io-dev//blog/2025/10/23/isba11-debrief</id><content type="html" xml:base="https://spaam-community.github.io/spaam-community.github.io-dev//blog/2025/10/23/isba11-debrief/"><![CDATA[<p><em>by <a href="https://www.spaam-community.org/steering_committee/">The Steering Committee and SPAAM7 organizing committee</a></em></p>

<p>At the conclusion of the ISBA conference, members of our ancient metagenomics community gathered for a debriefing session to share our impressions and feedback on the ISBA conference program. Below is a summary of the key points discussed.</p>

<h2 id="keynote-sessions">Keynote Sessions</h2>
<p>The keynote presentations were generally very well received. In particular, <strong>Cristina Valdiosera</strong>’s keynote (Biomolecular Echoes: The Unequal Footprint of Human History; https://media.unito.it/?content=11523) stood out as a clear favourite among attendees.</p>

<h3 id="accessibility-of-keynotes">Accessibility of Keynotes</h3>
<p>Some participants found the first keynote by <strong>Joanna Brück</strong> (Decolonizing European prehistory; https://media.unito.it/?content=11522) challenging to follow, especially those without a background in theoretical archaeology. To promote better interdisciplinary understanding, SPAAM will be launching an initiative – starting with a dedicated channel on Element – to connect archaeologists interested in metagenomics. This will hopefully foster further collaborative activities, such as joint seminars.</p>

<h2 id="methodological-talks">Methodological Talks</h2>
<p>Talks focusing on methodologies were highly appreciated. There was consensus within the group that future conferences should dedicate more discussion time to methodological topics. <strong>Mohamed Sarhan</strong>’s presentation (De-novo assembly and reconstruction
of an ancient <em>Streptococcus pyogenes</em> genome from a pre-Columbian mummy: Insights into the evolution of a human adapted pathogen) on de-novo assembly of ancient metagenomes was highlighted as particularly valuable. <strong>Mohamed Sarhan</strong> also won the ISBA Future Fellows prize for the best talk. Congratulations!</p>

<h2 id="scientific-sessions">Scientific Sessions</h2>
<p>There was strong enthusiasm for the Evolution session held on Thursday, August 28, which featured a strong showing of SPAAM-related talks, such as <strong>Kristen Bos</strong>’ talk on treponematoses, <strong>Mario Apata</strong> on oral health in pre-columbian populations, <strong>Maria Lopopolo</strong> on leprosy in pre-contact Americas, <strong>Meriam Guellil</strong> on human Betaherpesviruses, <strong>Iseult Jackson</strong> on Salmonella enterica, <strong>Ian Light-Maka</strong> on Yersinia pestis in Bronze Age sheep, and <strong>Louis L’Hôte</strong> on Sheeppox virus. <strong>Maria Lopopolo, Charlotte Avanzi and Nicolas Rascovan</strong> received the ISBA Distinguished Article Award for their related paper (https://doi.org/10.1126/science.adu7144). Congratulations!</p>

<h2 id="museomics--ethics">Museomics &amp; Ethics</h2>
<p>Presentations covering museomics and associated ethical considerations were also well liked and generated discussion. The importance of  involving museums in the first stages of thinking about projects was highlighted. There was a discussion on who “owns” the samples and on where the data (extracts, libraries) should be stored. It was pointed out that there is a need for collaborations across archives and labs, to keep track of what was sampled, in order to connect researchers and avoid double sampling (for example: libraries or extracts could be shared and re-used across labs). Cross-disciplines dialogue is needed to establish sorts of ‘protocols’ or ‘lists’ of things that a museum/collection owner should ask when a researcher arrives to ask for samples.</p>

<h2 id="sedadna-and-contamination-minimization">sedaDNA and Contamination Minimization</h2>
<p><strong>Arjen de Groot</strong>’s talk (Towards an archaeological workflow for sedaDNA sample collection: methods and best practices for minimizing surface contamination) on methods for minimizing surface contamination in sedaDNA research was another highlight, offering practical insights for the community.</p>

<h2 id="poster-sessions">Poster Sessions</h2>
<p>Many found it challenging to follow the poster sessions, as the espresso talks were sometimes scheduled on different days from the related poster presentations. Additionally, it was often unclear when presenters would be available at their posters. For future conferences, having a clearer list and schedule for poster sessions as well as presenter availability would be beneficial.</p>

<h2 id="best-spaam-talks">Best SPAAM talks</h2>
<p>Just like at the last ISBA10 in Tartu 2023, we took a vote on the best SPAAM talk presented at ISBA this year (with a total of 13 SPAAM talks between plenary and espresso talks). This year’s favorites were all about leprosy! We congratulate:</p>

<ol>
  <li>
    <p><strong>Maria Lopopolo</strong> for the best plenary talk: Uncovering pre-European contact leprosy in the Americas and its enduring persistence</p>
  </li>
  <li>
    <p><strong>Aida Andrades Valtueña</strong> for the best espresso talk: Mycobacterium leprae genomes from Central Europe provide clues into the past diversity of the leprosy bacterium</p>
  </li>
</ol>

<p>Thank you all for your active participation and valuable feedback!</p>]]></content><author><name>The Steering Committee and SPAAM7 organizing committee</name></author><category term="Blog" /><category term="spaam," /><category term="blog" /><summary type="html"><![CDATA[by The Steering Committee and SPAAM7 organizing committee]]></summary></entry></feed>