☆ 4.6 Article

Deep learning and support vector machines for transcription start site identification

PEERJ COMPUTER SCIENCE (2023)

期刊

PEERJ COMPUTER SCIENCE

卷 9, 期 -, 页码 -

出版社

PEERJ INC

DOI: 10.7717/peerj-cs.1340

关键词

Transcription start site; Bioinformatics; Machine learning; Deep learning; Support vector machine; Long short-term memory; Convolutional neural network

类别

Computer Science, Artificial Intelligence Computer Science, Information Systems Computer Science, Theory & Methods

向作者/读者索取更多资源

Protocol

社区支持

Reagent

社区支持

智能总结 New
摘要

Recognizing transcription start sites is crucial for gene identification. This article compares the performance of support vector machines (SVMs) and deep learning methods, specifically long short-term memory neural networks (LSTMs), in predicting transcription start sites. The results show that deep learning methods are better suited for this task, especially when working with sequence data. Additionally, a method for generating transcription start site datasets and the importance of balanced data are discussed.

Recognizing transcription start sites is key to gene identification. Several approaches have been employed in related problems such as detecting translation initiation sites or promoters, many of the most recent ones based on machine learning. Deep learning methods have been proven to be exceptionally effective for this task, but their use in transcription start site identification has not yet been explored in depth. Also, the very few existing works do not compare their methods to support vector machines (SVMs), the most established technique in this area of study, nor provide the curated dataset used in the study. The reduced amount of published papers in this specific problem could be explained by this lack of datasets. Given that both support vector machines and deep neural networks have been applied in related problems with remarkable results, we compared their performance in transcription start site predictions, concluding that SVMs are computationally much slower, and deep learning methods, specially long short-term memory neural networks (LSTMs), are best suited to work with sequences than SVMs. For such a purpose, we used the reference human genome GRCh38. Additionally, we studied two different aspects related to data processing: the proper way to generate training examples and the imbalanced nature of the data. Furthermore, the generalization performance of the models studied was also tested using the mouse genome, where the LSTM neural network stood out from the rest of the algorithms. To sum up, this article provides an analysis of the best architecture choices in transcription start site identification, as well as a method to generate transcription start site datasets including negative instances on any species available in Ensembl. We found that deep learning methods are better suited than SVMs to solve this problem, being more efficient and better adapted to long sequences and large amounts of data. We also create a transcription start site (TSS) dataset large enough to be used in deep learning experiments.

Deep learning and support vector machines for transcription start site identification

期刊

PEERJ COMPUTER SCIENCE

出版社

PEERJ INC

关键词

类别

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

Deep learning and support vector machines for transcription start site identification

期刊

PEERJ COMPUTER SCIENCE

出版社

PEERJ INC

关键词

类别

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

导出引文

分享论文