☆ 4.8 Article

The whole alignment and nothing but the alignment: the problem of spurious alignment flanks

NUCLEIC ACIDS RESEARCH (2008)

Journal

NUCLEIC ACIDS RESEARCH

Volume 36, Issue 18, Pages 5863-5871

Publisher

OXFORD UNIV PRESS

DOI: 10.1093/nar/gkn579

Keywords

Funding

National Library of Medicine
National Institutes of Health

Ask authors/readers for more resources

Protocol

Community support

Reagent

Community support

Abstract

Pairwise sequence alignment is a ubiquitous tool for inferring the evolution and function of DNA, RNA and protein sequences. It is therefore essential to identify alignments arising by chance alone, i.e. spurious alignments. On one hand, if an entire alignment is spurious, statistical techniques for identifying and eliminating it are well known. On the other hand, if only a part of the alignment is spurious, elimination is much more problematic. In practice, even the sizes and frequencies of spurious subalignments remain unknown. This article shows that some common scoring schemes tend to overextend alignments and generate spurious alignment flanks up to hundreds of base pairs/amino acids in length. In the UCSC genome database, e.g. spurious flanks probably comprise >18% of the human-fugu genome alignment. To evaluate the possibility that chance alone generated a particular flank on a particular pairwise alignment, we provide a simple 'overalignment' P-value. The overalignment P-value can identify spurious alignment flanks, thereby eliminating potentially misleading inferences about evolution and function. Moreover, by explicitly demonstrating the tradeoff between over- and under-alignment, our methods guide the rational choice of scoring schemes for various alignment tasks.

The whole alignment and nothing but the alignment: the problem of spurious alignment flanks

Journal

NUCLEIC ACIDS RESEARCH

Publisher

OXFORD UNIV PRESS

Keywords

Categories

Funding

Ask authors/readers for more resources

Protocol

Reagent

Authors

I am an author on this paper

Reviews

Primary Rating

Secondary Ratings

Novelty

Significance

Scientific rigor

Rate this paper

Recommended

The whole alignment and nothing but the alignment: the problem of spurious alignment flanks

Journal

NUCLEIC ACIDS RESEARCH

Publisher

OXFORD UNIV PRESS

Keywords

Categories

Funding

Ask authors/readers for more resources

Protocol

Reagent

Authors

I am an author on this paper

Reviews

Primary Rating

Secondary Ratings

Novelty

Significance

Scientific rigor

Rate this paper

Recommended

Export Citation

Share Paper