Preprint
Machine Learning

The PRIDE database at 20 years: 2025 update

Yasset Pérez‐Riverol(European Bioinformatics Institute), Chakradhar Bandla(European Bioinformatics Institute), Deepti J Kundu(European Bioinformatics Institute), Selvakumar Kamatchinathan(European Bioinformatics Institute), Jingwen Bai(European Bioinformatics Institute), Suresh Hewapathirana(European Bioinformatics Institute), Nithu Sara John(European Bioinformatics Institute), Ananth Prakash(European Bioinformatics Institute), Mathias Walzer(European Bioinformatics Institute), Shengbo Wang(European Bioinformatics Institute), Juan Antonio Vizcaíno(European Bioinformatics Institute)
November 4, 2024Nucleic Acids Research1,779 citations

1.8k

Citations

32

Influential Citations

Nucleic Acids Research

Venue

2024

Year

Abstract

The PRoteomics IDEntifications (PRIDE) database (https://www.ebi.ac.uk/pride/) is the world's leading mass spectrometry (MS)-based proteomics data repository and one of the founding members of the ProteomeXchange consortium. This manuscript summarizes the developments in PRIDE resources and related tools for the last three years. The number of submitted datasets to PRIDE Archive (the archival component of PRIDE) has reached on average around 534 datasets per month. This has been possible thanks to continuous improvements in infrastructure such as a new file transfer protocol for very large datasets (Globus), a new data resubmission pipeline and an automatic dataset validation process. Additionally, we will highlight novel activities such as the availability of the PRIDE chatbot (based on the use of open-source Large Language Models), and our work to improve support for MS crosslinking datasets. Furthermore, we will describe how we have increased our efforts to reuse, reanalyze and disseminate high-quality proteomics data into added-value resources such as UniProt, Ensembl and Expression Atlas.

Analysis

Why This Paper Matters

This 2025 update on the PRIDE database is significant because PRIDE is the world's leading mass spectrometry-based proteomics data repository and a founding member of the ProteomeXchange consortium. With an average of 534 datasets submitted per month, the database's continuous growth demands robust infrastructure and innovative tools. The paper highlights practical improvements that directly impact the reproducibility and accessibility of proteomics data, which is critical for the broader AI and bioinformatics community that relies on high-quality, well-annotated datasets for training models and validating hypotheses.

The introduction of a chatbot based on open-source large language models (LLMs) is particularly noteworthy. It represents a concrete application of LLMs in scientific data management, potentially lowering the barrier for researchers to query and interact with the repository. This aligns with the trend of using AI to enhance data discovery and usability in large-scale biological databases.

Technical Contributions

  • Infrastructure Upgrades: Implementation of Globus for efficient transfer of very large datasets, a new data resubmission pipeline, and an automatic dataset validation process. These improvements ensure scalability and data quality.
  • PRIDE Chatbot: Built using open-source LLMs, the chatbot allows users to ask natural language questions about the database, facilitating easier access to metadata and dataset information.
  • Crosslinking Data Support: Enhanced support for MS crosslinking datasets, expanding the types of proteomics data that can be deposited and reused.
  • Data Reuse Pipelines: Increased efforts to reanalyze and disseminate high-quality proteomics data into added-value resources such as UniProt, Ensembl, and Expression Atlas, thereby maximizing the impact of submitted data.

Results

The paper reports that the number of submitted datasets has reached an average of 534 per month, a clear indicator of the repository's growth and the effectiveness of the infrastructure improvements. No specific performance metrics for the chatbot or validation pipeline are provided, but the operational success is evident from the submission volume.

Significance

For the AI field, this work demonstrates how LLMs can be practically deployed in scientific data repositories to improve user interaction. The emphasis on data reuse and integration with major bioinformatics resources creates a richer ecosystem for training machine learning models on proteomics data. The infrastructure improvements also set a benchmark for other large-scale data repositories, highlighting the importance of automated validation and efficient data transfer in maintaining data quality and accessibility.