☆ 4.7 Article

ARAMIS: From systematic errors of NGS long reads to accurate assemblies

BRIEFINGS IN BIOINFORMATICS (2021)

Journal

BRIEFINGS IN BIOINFORMATICS

Volume 22, Issue 6, Pages -

Publisher

OXFORD UNIV PRESS

DOI: 10.1093/bib/bbab170

Keywords

error correction; next-generation sequencing; homopolymer; long read; genome assembly

Funding

CBMSO (CSIC-UAM)
Community of Madrid within the European Youth Employment Initiative (YEI) [PEJD-2018-PRE/BMD-9388, PEJD-2017-PRE/BMD-4828]
Spanish Ministery of Science and Innovation within the European Youth Employment Initiative (YEI) [PEJ2018-005067-P]
Spanish Ministery of Science and Innovation [PTA2017-14628-I]
Fundacion Ramon Areces
Fundacion Banco de Santander

Ask authors/readers for more resources

Protocol

Community support

Reagent

Community support

Automated Summary New
Abstract

Researchers have developed a NGS long-reads indels correction pipeline called ARAMIS, which combines multiple correction software in one step using accurate short reads to address insertions and deletions errors in long-read sequencing. The study found systematic sequencing errors in PacBio sequences affecting homopolymeric regions, and that the type of indel errors introduced during PacBio sequencing are related to the GC content of the organism.

NGS long-reads sequencing technologies (or third generation) such as Pacific BioSciences (PacBio) have revolutionized the sequencing field over the last decade improving multiple genomic applications like de novo genome assemblies. However, their error rate, mostly involving insertions and deletions (indels), is currently an important concern that requires special attention to be solved. Multiple algorithms are available to fix these sequencing errors using short reads (such as Illumina), although they require long processing times and some errors may persist. Here, we present Accurate long-Reads Assembly correction Method for Indel errorS (ARAMIS), the first NGS long-reads indels correction pipeline that combines several correction software in just one step using accurate short reads. As a proof OF concept, six organisms were selected based on their different GC content, size and genome complexity, and their PacBio-assembled genomes were corrected thoroughly by this pipeline. We found that the presence of systematic sequencing errors in long-reads PacBio sequences affecting homopolymeric regions, and that the type of indel error introduced during PacBio sequencing are related to the GC content of the organism. The lack of knowledge of this fact leads to the existence of numerous published studies where such errors have been found and should be resolved since they may contain incorrect biological information. ARAMIS yields better results with less computational resources needed than other correction tools and gives the possibility of detecting the nature of the found indel errors found and its distribution along the genome. The source code of ARAMIS is available at https://github.com/genomics-ngsCBMSO/ARAMIS.git

ARAMIS: From systematic errors of NGS long reads to accurate assemblies

Journal

BRIEFINGS IN BIOINFORMATICS

Publisher

OXFORD UNIV PRESS

Keywords

Categories

Funding

Ask authors/readers for more resources

Protocol

Reagent

Authors

I am an author on this paper

Reviews

Primary Rating

Secondary Ratings

Novelty

Significance

Scientific rigor

Rate this paper

Recommended

ARAMIS: From systematic errors of NGS long reads to accurate assemblies

Journal

BRIEFINGS IN BIOINFORMATICS

Publisher

OXFORD UNIV PRESS

Keywords

Categories

Funding

Ask authors/readers for more resources

Protocol

Reagent

Authors

I am an author on this paper

Reviews

Primary Rating

Secondary Ratings

Novelty

Significance

Scientific rigor

Rate this paper

Recommended

Export Citation

Share Paper