ReNovo: Retrieval-Based De Novo Mass Spectrometry Peptide Sequencing

ML conference · ICLR 2025 · 2024

ML conference · ICLR 2025 · 2024. Shaorong Chen et al. Proteomics is the large-scale study of proteins. Tandem mass spectrometry, as the only high-throughput…
Date 2024-11-27
Type ML conference
Venue ICLR 2025
Publisher OpenReview
Contribution algorithm
Link https://openreview.net/forum?id=uQnvYP7yX9

Abstract

Proteomics is the large-scale study of proteins. Tandem mass spectrometry, as the only high-throughput technique for protein sequence identification, plays a pivotal role in proteomics research. One of the long-standing challenges in this field is peptide identification, which entails determining the specific peptide (sequence of amino acids) that corresponds to each observed mass spectrum. The conventional approach involves database searching, wherein the observed mass spectrum is scored against a pre-constructed peptide database. However, the reliance on pre-existing databases limits applicability in scenarios where the peptide is absent from existing databases. Such circumstances necessitate de novo peptide sequencing, which derives peptide sequence solely from input mass spectrum, independent of any peptide database. Despite ongoing advancements in de novo peptide sequencing, its performance still has considerable room for improvement, which limits its application in large-scale experiments. In this study, we introduce a novel Retrieval-based De Novo peptide sequencing methodology, termed ReNovo, which draws inspiration from database search methods. Specifically, by constructing a datastore from training data, ReNovo can retrieve information from the datastore during the inference stage to conduct retrieval-based inference, thereby achieving improved performance. This innovative approach enables ReN- ovo to effectively combine the strengths of both methods: utilizing the assistance of the datastore while also being capable of predicting novel peptides that are not present in pre-existing databases. A series of experiments have confirmed that ReNovo outperforms state-of-the-art models across multiple widely-used datasets, incurring only minor storage and time consumption, representing a significant advancement in proteomics. Supplementary materials include the code.

Authors

  1. Shaorong Chen · Westlake University, Zhejiang University
  2. Jun Xia · The Hong Kong University of Science and Technology, The Hong Kong University of Science and Technology (Guangzhou), Westlake University
  3. Jingbo Zhou · Westlake University, Zhejiang University
  4. Lecheng Zhang · Westlake University
  5. Zhangyang Gao · Westlake University
  6. Bozhen Hu · Westlake University
  7. Cheng Tan · Westlake University
  8. Wenjie Du · Westlake University
  9. Stan Z. Li · Westlake University

Methods and tools

  • ReNovo: Retrieval-based sequencing

Data used

  • ProteomeTools (HC-PT (NovoBench)) · no public address
  • Seven-species benchmark (NovoBench split) · no public address

Seen in the charts

Back to the full map

Back to top