4.6 Article

CRF: detection of CRISPR arrays using random forest

期刊

PEERJ
卷 5, 期 -, 页码 -

出版社

PEERJ INC
DOI: 10.7717/peerj.3219

关键词

Repeat detection; Random forest; Machine learning; CRISPR; Data visualization

资金

  1. Committee on Faculty Research (CRF) Program, Miami University, Oxford, Ohio, USA
  2. Department of Biology, Miami University, Oxford, Ohio, USA
  3. Office for the Advancement of Research & Scholarship (OARS), Miami University, Oxford, Ohio, USA

向作者/读者索取更多资源

CRISPRs (clustered regularly interspaced short palindromic repeats) are particular repeat sequences found in wide range of bacteria and archaea genornes. Several tools are available for detecting CRISPR arrays in the genomes of both domains. Here we developed a new web-based CRISPR detection tool named CRF (CRISPR Finder by Random Forest). Different from other CRISPR detection tools, a random forest classifier was used in CRF to filter out invalid CRISPR arrays from all putative candidates and accordingly enhanced detection accuracy. In CRF, particularly, triplet elements that combine both sequence content and structure information were extracted from CRISPR repeats for classifier training. The classifier achieved high accuracy and sensitivity. Moreover, CRF offers a highly interactive web interface for robust data visualization that is not available among other CRISPR detection tools. After detection, the query sequence, CRISPR array architecture, and the sequences and secondary structures of CRISPR repeats and spacers can be visualized for visual examination and validation. CRF is freely available at http://bioinfolab.miamioh.edu/crf/home.php.

作者

我是这篇论文的作者
点击您的名字以认领此论文并将其添加到您的个人资料中。

评论

主要评分

4.6
评分不足

次要评分

新颖性
-
重要性
-
科学严谨性
-
评价这篇论文

推荐

暂无数据
暂无数据