Real-time Evolution: Monitoring SARS-CoV-2 Mutations via the PED Algorithm

Real-time Evolution: Monitoring SARS-CoV-2 Mutations via the PED Algorithm
Like

Share this post

Choose a social network to share with, or copy the URL to share elsewhere

This is a representation of how your post may appear on social media. The actual post will vary between social networks

Explore the Research

BioMed Central
BioMed Central BioMed Central

Polymorphic edge detection (PED): two efficient methods of polymorphism detection from next-generation sequencing data - BMC Bioinformatics

Background Accurate detection of polymorphisms with a next generation sequencer data is an important element of current genetic analysis. However, there is still no detection pipeline that is completely reliable. Result We demonstrate two new detection methods of polymorphisms focusing on the Polymorphic Edge (PED). In the matching between two homologous sequences, the first mismatched base to appear is the SNP, or the edge of the structural variation. The first method is based on k-mers from short reads and detects polymorphic edges with k-mers for which there is no match between target and control, making it possible to detect SNPs by direct comparison of short-reads in two datasets (target and control) without a reference genome sequence. The second method is based on bidirectional alignment to detect polymorphic edges, not only SNPs but also insertions, deletions, inversions and translocations. Using these two methods, we succeed in making a high-quality comparison map between rice cultivars showing good match to the theoretical value of introgression, and in detecting specific large deletions across cultivars. Conclusions Using Polymorphic Edge Detection (PED), the k-mer method is able to detect SNPs by direct comparison of short-reads in two datasets without genomic alignment step, and the bidirectional alignment method is able to detect SNPs and structural variations from even single-end short-reads. The PED is an efficient tool to obtain accurate data for both SNPs and structural variations. Availability The PED software is available at: https://github.com/akiomiyao/ped .

The novel coronavirus (Severe Acute Respiratory Syndrome Coronavirus 2, SARS‑CoV‑2), first identified in Wuhan at the end of 2019, spread rapidly worldwide by January 2020. In Japan, infections were initially reported among people who had dined together on traditional houseboats and among cruise ship passengers, eventually developing into a full-scale pandemic.

SARS‑CoV‑2 shares similarities with the earlier SARS‑CoV outbreak; however, the amino acid sequence of the spike (S) protein responsible for receptor binding differs substantially. Although both viruses use ACE2 as the cellular entry receptor, the mode of interaction with ACE2 is different in SARS-CoV-2. This resulted in markedly increased binding affinity and infection efficiency for SARS‑CoV‑2.

The more stable binding to the ACE2 protein enabled efficient early replication in airway epithelial cells, leading to extremely high transmissibility. In addition, the high frequency of asymptomatic infections meant that infected individuals often continued normal activities, facilitating explosive global spread.

Although the case fatality rate of SARS‑CoV‑2 is considered lower than that of SARS‑CoV, its high transmissibility led to a massive number of infections. As a result, many people—particularly the elderly and those with compromised immune systems—lost their lives.

The nucleotide sequence of SARS‑CoV‑2 was released at a very early stage by Chinese researchers. A distinctive feature for research on this pandemic was the rapid and wide public availability of next‑generation sequencing (NGS) data derived directly from patient samples. In particular, the United Kingdom analysed sequencing data from a very large number of patients, which became available for download from the NCBI Sequence Read Archive (SRA).

We analysed these downloaded sequences using our virus-tailored modification of PED program to detect genetic variants. By applying PED’s bidirectional alignment method, we were able to efficiently detect both single nucleotide polymorphisms (SNPs) and insertions/deletions (indels) in various sequenced virus genome samples.

PED includes a function that determines homozygous and heterozygous states based on the frequency of detected variant candidates. However, this framework is not appropriate for SARS‑CoV‑2, which is not a diploid organism. In practice, SARS‑CoV‑2 samples often contained mixed infections of multiple variants, and the observed variant frequencies varied widely among mutations.

To address this, we extended PED by adding a function that outputs read counts for each detected variant, specifically targeting organisms like SARS‑CoV‑2 that do not exhibit homozygous or heterozygous genotypes. This enhancement enabled more appropriate confirmation of viral mutations.

In our study, we downloaded virus genome sequences form ca. 50 individuals per sampling date and detected polymorphisms by  our modified PED. During the early stages of the pandemic. As there were many days with fewer than 50 reported, we analysed all available data for those days.

The figure shows an example of the detection frequency of mutations plotted at monthly intervals from the start of the pandemic. From 2020 until around May 2021, the Alpha variant carrying the N501Y (A23063T) mutation was dominant, after which it was replaced by the Delta variant carrying the L452R (T22916G) mutation. We also obtained extensive information on insertion and deletion mutations, which are a particular strength of PED. Although PED was originally developed primarily for detecting polymorphisms in eukaryotic organisms, it produced results for viral variants that were largely consistent with previously reported findings. When NGS data were available, results could be obtained within just a few minutes.

In this way, we were able to observe the evolution of SARS‑CoV‑2 during the pandemic almost in real time. However, eventually we suspended this monitoring analysis because SRA later limited downloads of the metadata for sample collection dates to approximately 100,000 records, preventing access to more recent data.

The ability to analyse mutations directly from publicly available raw sequencing data—without waiting for expert consensus analyses—was a highly significant development made evident during this pandemic.

https://akiomiyao.github.io/ped/covid19/index.html

Follow the Topic

COVID19
Life Sciences > Biological Sciences > Microbiology > Medical Microbiology > Infectious Diseases > COVID19
Genetics Research
Life Sciences > Health Sciences > Biomedical Research > Genetics Research
Medical Genetics
Life Sciences > Biological Sciences > Genetics and Genomics > Medical Genetics
  • BMC Bioinformatics BMC Bioinformatics

    This is an open access, peer-reviewed journal that considers articles describing novel computational algorithms and software, models and tools, including statistical methods, machine learning and artificial intelligence, as well as systems biology.

Related Collections

With Collections, you can get published faster and increase your visibility.

Extracellular vesicle research

BMC Bioinformatics is welcoming submissions to our Collection on Extracellular vesicles research.

BMC Bioinformatics is welcoming submissions to our Collection on Extracellular vesicles research. Extracellular vesicles (EVs) are are small lipid bilayer-delimited particles released by cells that play crucial roles in intercellular communication and various physiological processes. The study of EVs has gained significant attention due to their potential as biomarkers for disease diagnosis, therapeutic targets and drug delivery systems. Advanced bioinformatics tools are essential for analyzing EV data, identifying EV-associated molecules, and understanding their biological functions.

This Collection welcomes submissions on the development of new computational and/or statistical approaches for the study of extracellular vesicles. We encourage contributions that highlight innovative methods for detecting and characterizing EVs and elucidating the molecular mechanisms underlying EV biogenesis and function.

All manuscripts submitted to this journal, including those submitted to collections and special issues, are assessed in line with our editorial policies and the journal’s peer-review process. Reviewers and editors are required to declare competing interests and can be excluded from the peer review process if a competing interest exists.

Publishing Model: Open Access

Deadline: Sep 30, 2026

Bioinformatics and ecology

BMC Bioinformatics is calling for submissions to our Collection on Bioinformatics and ecology.

Ecology has become a data-heavy science. Genomic technologies and computational methods have changed what ecologists can ask and answer, and bioinformatics tools are now central to making sense of the large, messy datasets these studies produce.

This matters for some of the most pressing environmental problems we face, including climate change, habitat loss, non-indigenous species, and biodiversity crisis. Advances in DNA sequencing, comprehensive reference databases, and computational analysis now let researchers characterise microbial and macrobial communities and assess their role in ecosystem function with a level of detail that was not possible a decade ago. In conservation genomics, the same tools support practical decisions about managing endangered species and restoring habitats.

There is plenty of room for the field to develop further. As machine learning methods mature, they should improve predictions of species distributions and how ecosystems respond to environmental pressure. Better taxonomic profiling and genome assembly will also sharpen our understanding of microbial ecology and what it tells us about ecosystem health.

This Collection brings together research that pairs bioinformatics approaches with ecological questions. We are particularly interested in environmental DNA (eDNA) analysis, metagenomics, and the use of machine learning to understand community structure and species distribution. We welcome contributions that include:

  • Ecosystem monitoring using bioinformatics
  • Metagenomics in microbial ecology
  • Applications of machine learning in ecological studies
  • Environmental DNA as a tool for biodiversity assessment
  • Conservation genomics and species management
  • Metabarcoding approaches for large-scale biodiversity studies
  • Bioinformatic pipelines and tools for ecological genomics

This Collection supports and amplifies research related to SDG6: Clean Water and Sanitation, SDG 13: Climate Action, SDG 14: Life Below Water, and SDG 15: Life on Land.

All manuscripts submitted to this journal, including those submitted to collections and special issues, are assessed in line with our editorial policies and the journal’s peer-review process. Reviewers and editors are required to declare competing interests and can be excluded from the peer review process if a competing interest exists.

Publishing Model: Open Access

Deadline: Apr 30, 2027