4.8 Article

AlphaFold Protein Structure Database in 2024: providing structure coverage for over 214 million protein sequences

期刊

NUCLEIC ACIDS RESEARCH
卷 -, 期 -, 页码 -

出版社

OXFORD UNIV PRESS
DOI: 10.1093/nar/gkad1011

关键词

-

向作者/读者索取更多资源

The AlphaFold Database Protein Structure Database (AlphaFold DB) has expanded significantly since its initial release in 2021, now containing over 214 million predicted protein structures. Powered by the AlphaFold2 artificial intelligence (AI) system, the database has integrated its predictions into primary data resources such as PDB, UniProt, Ensembl, InterPro, and MobiDB. This manuscript details the enhancements made to data archiving, including the addition of model organisms, global health proteomes, Swiss-Prot integration, and curated protein datasets. The access mechanisms of AlphaFold DB, from direct file access to advanced queries using Google Cloud Public Datasets, are also discussed, along with improvements and added services since its release, such as enhancements to the Predicted Aligned Error viewer and the 3D viewer customization options.
The AlphaFold Database Protein Structure Database (AlphaFold DB, https://alphafold.ebi.ac.uk) has significantly impacted structural biology by amassing over 214 million predicted protein structures, expanding from the initial 300k structures released in 2021. Enabled by the groundbreaking AlphaFold2 artificial intelligence (AI) system, the predictions archived in AlphaFold DB have been integrated into primary data resources such as PDB, UniProt, Ensembl, InterPro and MobiDB. Our manuscript details subsequent enhancements in data archiving, covering successive releases encompassing model organisms, global health proteomes, Swiss-Prot integration, and a host of curated protein datasets. We detail the data access mechanisms of AlphaFold DB, from direct file access via FTP to advanced queries using Google Cloud Public Datasets and the programmatic access endpoints of the database. We also discuss the improvements and services added since its initial release, including enhancements to the Predicted Aligned Error viewer, customisation options for the 3D viewer, and improvements in the search engine of AlphaFold DB. The AlphaFold Protein Structure Database (AlphaFold DB) is a massive digital library of predicted protein structures, with over 214 million entries, marking a 500-times expansion in size since its initial release in 2021. The structures are predicted using Google DeepMind's AlphaFold 2 artificial intelligence (AI) system. Our new report highlights the latest updates we have made to this database. We have added more data on specific organisms and proteins related to global health and expanded to cover almost the complete UniProt database, a primary data resource of protein sequences. We also made it easier for our users to access the data by directly downloading files or using advanced cloud-based tools. Finally, we have also improved how users view and search through these protein structures, making the user experience smoother and more informative. In short, AlphaFold DB has been growing rapidly and has become more user-friendly and robust to support the broader scientific community. Graphical Abstract

作者

我是这篇论文的作者
点击您的名字以认领此论文并将其添加到您的个人资料中。

评论

主要评分

4.8
评分不足

次要评分

新颖性
-
重要性
-
科学严谨性
-
评价这篇论文

推荐

暂无数据
暂无数据