NovoRank: Refinement for De Novo Peptide Sequencing Based on Spectral Clustering and Deep Learning
peer-reviewed · Journal of Proteome Research · 2025
| Date | 2025-02-07 |
| Type | peer-reviewed |
| Venue | Journal of Proteome Research |
| Publisher | American Chemical Society (ACS) |
| Contribution | post-processor |
| DOI | 10.1021/acs.jproteome.4c00300 |
| Citations (OpenAlex) | 2 |
| Venue 2-year citedness | 3.83 |
Abstract
De novo peptide sequencing is a valuable technique in mass-spectrometry-based proteomics, as it deduces peptide sequences directly from tandem mass spectra without relying on sequence databases. This database-independent method, however, relies solely on imperfect scoring functions that often lead to erroneous peptide identifications. To boost correct identification, we present NovoRank, a postprocessing tool that employs spectral clustering and machine learning to assign more plausible peptide sequences to spectra. Prior to de novo peptide sequencing, spectral clustering is applied to group similar spectra under the assumption that they originated from the same peptide species. NovoRank then employs a deep learning model, incorporating both cluster-derived proteomic features and individual spectrum characteristics, to rerank the candidate peptides produced by de novo peptide sequencing. Our results show that NovoRank significantly enhances the performance of various de novo peptide sequencing tools, increasing both recall and precision by 0.020 to 0.080 at the peptide-spectrum match (PSM) level. Notably, NovoRank achieves a recall as high as 0.830 for Casanovo at the PSM level. The source code of NovoRank is freely available at https://github.com/HanyangBISLab/NovoRank and is licensed under Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International.
Methods and tools
- NovoRank: Spectral clustering refinement
Data deposited
- Data for ‘NovoRank: Refinement for De Novo Peptide Sequencing Based on Spectral Clustering and Deep Learning’ (as deposited) · 10.5281/zenodo.14046459
Data used
- 11 human cell lines - Comparative proteomic analysis of eleven common cell lines reveals ubiquitous but varying expressi (as deposited) · PXD002395
- A mass-tolerant database search identifies a large proportion of unassigned spectra in shotgun proteomics as modified pe (as deposited) · PXD001468
- Proteogenomics of colorectal cancer liver metastases (as deposited) · PXD014222
- ProteomeTools (Parts I-III) · PXD004732, PXD010595, PXD021013
Cites (6)
- Sequence-to-sequence translation from mass spectra to peptides with a transformer model (2024) crossref
- Deep learning-driven fragment ion series classification enables highly precise and sensitive de novo peptide sequencing (2024) crossref
- The impact of noise and missing fragmentation cleavages on de novo peptide identification algorithms (2022) crossref
- pNovo 3: precise de novo peptide sequencing using a learning-to-rank framework (2019) crossref
- pNovo: De novo Peptide Sequencing and Identification Using HCD Spectra (2010) crossref
- PEAKS: powerful software for peptide de novo sequencing by tandem mass spectrometry (2003) crossref