☆ 4.6 Article

mbkmeans: Fast clustering for single cell data using mini-batch k-means

PLOS COMPUTATIONAL BIOLOGY (2021)

期刊

PLOS COMPUTATIONAL BIOLOGY

卷 17, 期 1, 页码 -

出版社

PUBLIC LIBRARY SCIENCE

DOI: 10.1371/journal.pcbi.1008625

关键词

类别

Biochemical Research Methods Mathematical & Computational Biology

资金

National Institutes of Health [R00HG009007]
NIH BRAIN Initiative [U19MH114830]
Zuckerberg Initiative DAF, an advised fund of Silicon Valley Community Foundation [DAF2018-183201, CZF2019-002443]
ENS-CFM Data Science Chair
Programma per Giovani Ricercatori Rita Levi Montalcini - Italian Ministry of Education, University, and Research

向作者/读者索取更多资源

Protocol

社区支持

Reagent

社区支持

智能总结 New
摘要

Single-cell RNA-Sequencing (scRNA-seq) is a widely used technology for measuring gene expression at the single-cell level, with analyses often detecting distinct cell subpopulations through clustering algorithms. The development of the mbkmeans package offers a solution for handling large datasets without requiring full data loading into memory. This package provides efficient computation and performance comparisons with other clustering methods.

Single-cell RNA-Sequencing (scRNA-seq) is the most widely used high-throughput technology to measure genome-wide gene expression at the single-cell level. One of the most common analyses of scRNA-seq data detects distinct subpopulations of cells through the use of unsupervised clustering algorithms. However, recent advances in scRNA-seq technologies result in current datasets ranging from thousands to millions of cells. Popular clustering algorithms, such as k-means, typically require the data to be loaded entirely into memory and therefore can be slow or impossible to run with large datasets. To address this problem, we developed the mbkmeans R/Bioconductor package, an open-source implementation of the mini-batch k-means algorithm. Our package allows for on-disk data representations, such as the common HDF5 file format widely used for single-cell data, that do not require all the data to be loaded into memory at one time. We demonstrate the performance of the mbkmeans package using large datasets, including one with 1.3 million cells. We also highlight and compare the computing performance of mbkmeans against the standard implementation of k-means and other popular single-cell clustering methods. Our software package is available in Bioconductor at . Author summary We developed the mbkmeans package () in Bioconductor, an open-source implementation of the mini-batch k-means algorithm. Our package allows for on-disk data representations, such as the common HDF5 file format widely used for single-cell data, that do not require all the data to be loaded into memory at one time.

mbkmeans: Fast clustering for single cell data using mini-batch k-means

期刊

PLOS COMPUTATIONAL BIOLOGY

出版社

PUBLIC LIBRARY SCIENCE

关键词

类别

资金

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

mbkmeans: Fast clustering for single cell data using mini-batch k-means

期刊

PLOS COMPUTATIONAL BIOLOGY

出版社

PUBLIC LIBRARY SCIENCE

关键词

类别

资金

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

导出引文

分享论文