Bridging the Gap between Database Search and De Novo Peptide Sequencing with SearchNovo

preprint · bioRxiv · 2024

preprint · bioRxiv · 2024. Jun Xia et al. AO_SCPLOWBSTRACTC_SCPLOWAccurate protein identification from mass spectrometry (MS) data is fundamental to…
Date 2024-10-19
Type preprint
Venue bioRxiv
Publisher Cold Spring Harbor Laboratory
Contribution adjacent
DOI 10.1101/2024.10.19.619186
Citations (OpenAlex) 1

Abstract

AO_SCPLOWBSTRACTC_SCPLOWAccurate protein identification from mass spectrometry (MS) data is fundamental to unraveling the complex roles of proteins in biological systems, with peptide sequencing being a pivotal step in this process. The two main paradigms for peptide sequencing are database search, which matches experimental spectra with peptide sequences from databases, and de novo sequencing, which infers peptide sequences directly from MS without relying on pre-constructed database. Although database search methods are highly accurate, they are limited by their inability to identify novel, modified, or mutated peptides absent from the database. In contrast, de novo sequencing is adept at discovering novel peptides but often struggles with missing peaks issue, further leading to lower precision. We introduce SearchNovo, a novel framework that synergistically integrates the strengths of database search and de novo sequencing to enhance peptide sequencing. SearchNovo employs an efficient search mechanism to retrieve the most similar peptide spectrum match (PSM) from a database for each query spectrum, followed by a fusion module that utilizes the reference peptide sequence to guide the generation of the target sequence. Furthermore, we observed that dissimilar (noisy) reference peptides negatively affect model performance. To mitigate this, we constructed pseudo reference PSMs to minimize their impact. Comprehensive evaluations on multiple datasets reveal that SearchNovo significantly outperforms state-of-the-art models. Also, analysis indicates that many retrieved spectra contain missing peaks absent in the query spectra, and the retrieved reference peptides often share common fragments with the target peptides. These are key elements in the recipe for SearchNovos success. The code for reproducing the results are available in the supplementary materials.

Authors

  1. Jun Xia · The Hong Kong University of Science and Technology, The Hong Kong University of Science and Technology (Guangzhou), Westlake University
  2. Sizhe Liu · University of Southern California, Westlake University
  3. Jingbo Zhou · Westlake University, Zhejiang University
  4. Shaorong Chen · Westlake University, Zhejiang University
  5. Hongxin Xiang · Hunan University
  6. Zicheng Liu · Westlake University
  7. Yue Liu · National University of Singapore, Westlake University
  8. Stan Z. Li · Westlake University

Methods and tools

Cites (14)

Cited by (2)

Seen in the charts

Back to the full map

Back to top