Limitations of de novo sequencing in resolving sequence ambiguity

preprint · bioRxiv · 2025

preprint · bioRxiv · 2025. Sam van Puyenbroeck et al. De novo peptide sequencing enables peptide identification from fragmentation spectra without relying on…
Date 2025-08-19
Type preprint
Venue bioRxiv
Publisher Cold Spring Harbor Laboratory
Contribution benchmark
DOI 10.1101/2025.08.19.671052
Citations (OpenAlex) 1

Abstract

De novo peptide sequencing enables peptide identification from fragmentation spectra without relying on sequence databases. However, incomplete spectra create ambiguity, making unambiguous identification challenging. Recent deep learning advances have produced numerous de novo models that predict sequences and refine peptide-spectrum matches under such conditions. Yet, their relative strengths, weaknesses, and ability to handle spectrum ambiguity remain unclear. Here, we benchmark eight state-of-the-art models on three publicly available proteomics datasets, comparing performance using established metrics and quantifying inter-model agreement. We assess post-processing approaches, including iterative refinement, rescoring, and reranking, for their ability to improve identification accuracy, and perform an error analysis to identify common mispredictions and their causes. Model performance varied, with considerable overlap of correct identifications. Post-processing yielded no or only modest improvements. Most sequencing errors were model-independent and driven by limited fragment ion coverage, a limitation also observed in database searches with large search spaces.

Authors

  1. Sam van Puyenbroeck · Ghent University, VIB
  2. Denis Beslic · Robert Koch Institute
  3. Tomi Suomi · University of Turku and Åbo Akademi University
  4. Tanja Holstein · Ghent University, Robert Koch Institute, VIB
  5. Thilo Muth · Federal Institute for Materials Research and Testing (BAM), Max Planck Institute for Dynamics of Complex Technical Systems, Robert Koch Institute
  6. Laura L. Elo · University of Turku, University of Turku and Åbo Akademi University
  7. Lennart Martens · Ghent University, Infrastructure Nationale de Protéomique (ProFI-FR2048), University of Strasbourg, VIB
  8. Robbin Bouwmeester · Ghent University, VIB
  9. Tim Van Den Bossche · Ghent University, VIB
  10. Tine Claeys · Ghent University, VIB

Methods and tools

  • De novo sequence-ambiguity benchmark: Benchmark across 8 leading DL de novo peptide sequencers on three proteomics datasets, showing large overlap of correct calls between models and that post-processing yields only modest gains. The shared error source is limited fragment-ion coverage, a bottleneck that database search shares as well.

Cites (34)

Cited by (1)

Seen in the charts

Back to the full map

Back to top