Bridging the Gap between Database Search and De Novo Peptide Sequencing with SearchNovo
preprint · bioRxiv · 2024
| Date | 2024-10-19 |
| Type | preprint |
| Venue | bioRxiv |
| Publisher | Cold Spring Harbor Laboratory |
| Contribution | adjacent |
| DOI | 10.1101/2024.10.19.619186 |
| Citations (OpenAlex) | 1 |
Peer-reviewed version: Bridging the Gap between Database Search and De Novo Peptide Sequencing with SearchNovo (2025-01-22, ICLR 2025)
Abstract
AO_SCPLOWBSTRACTC_SCPLOWAccurate protein identification from mass spectrometry (MS) data is fundamental to unraveling the complex roles of proteins in biological systems, with peptide sequencing being a pivotal step in this process. The two main paradigms for peptide sequencing are database search, which matches experimental spectra with peptide sequences from databases, and de novo sequencing, which infers peptide sequences directly from MS without relying on pre-constructed database. Although database search methods are highly accurate, they are limited by their inability to identify novel, modified, or mutated peptides absent from the database. In contrast, de novo sequencing is adept at discovering novel peptides but often struggles with missing peaks issue, further leading to lower precision. We introduce SearchNovo, a novel framework that synergistically integrates the strengths of database search and de novo sequencing to enhance peptide sequencing. SearchNovo employs an efficient search mechanism to retrieve the most similar peptide spectrum match (PSM) from a database for each query spectrum, followed by a fusion module that utilizes the reference peptide sequence to guide the generation of the target sequence. Furthermore, we observed that dissimilar (noisy) reference peptides negatively affect model performance. To mitigate this, we constructed pseudo reference PSMs to minimize their impact. Comprehensive evaluations on multiple datasets reveal that SearchNovo significantly outperforms state-of-the-art models. Also, analysis indicates that many retrieved spectra contain missing peaks absent in the query spectra, and the retrieved reference peptides often share common fragments with the target peptides. These are key elements in the recipe for SearchNovos success. The code for reproducing the results are available in the supplementary materials.
Methods and tools
- SearchNovo: DB-search + de novo fusion
Cites (14)
- NovoBench: Benchmarking Deep Learning-based De Novo Peptide Sequencing Methods in Proteomics (2024) both
- AdaNovo: Adaptive De Novo Peptide Sequencing with Conditional Mutual Information (2024) both
- De novo peptide sequencing with InstaNovo: Accurate, database-free peptide identification for large scale proteomics experiments (2023) both
- De novo mass spectrometry peptide sequencing with a transformer model (2022) both
- Computationally instrument-resolution-independent de novo peptide sequencing for high-resolution devices (2021) both
- Personalized deep learning of individual immunopeptidomes to identify neoantigens for cancer vaccines (2020) both
- De novo sequencing of proteins by mass spectrometry (2020) both
- Peptide Sequencing with Deep Learning (2020) crossref
- De novo peptide sequencing by deep learning (2017) both
- MS2PIP: a tool for MS/MS peak intensity prediction (2013) both
- NovoHMM: A Hidden Markov Model for de Novo Peptide Sequencing (2005) both
- PEAKS: powerful software for peptide de novo sequencing by tandem mass spectrometry (2003) crossref
- Implementation and Uses of Automated de Novo Peptide Sequencing by Tandem Mass Spectrometry (2001) both
- De novo peptide sequencing via tandem mass spectrometry (1999) both
Cited by (2)
- Bidirectional Representations Augmented Autoregressive Biological Sequence Generation (2025) semanticscholar
- Curriculum Learning for Biological Sequence Prediction: The Case of De Novo Peptide Sequencing (2025) semanticscholar