GyroNovo: Error-Guided Fragment Imputation with Mass-Aware Attention for De Novo Peptide Sequencing

preprint · arXiv · 2026

preprint · arXiv · 2026. Abdellah El Mekki et al. De novo peptide sequencing from tandem mass spectra is essential for identifying peptides without relying on…
Date 2026-09-24
Type preprint
Venue arXiv
Publisher arXiv
Contribution algorithm
DOI 10.48550/arXiv.2609.30542

Abstract

De novo peptide sequencing from tandem mass spectra is essential for identifying peptides without relying on reference databases. Despite advances in deep learning, accurate sequencing remains challenging because experimental spectra are often sparse, noisy, and incomplete, leaving informative b- and y-ion fragments unobserved. Existing methods attempt to recover this missing evidence via latent-space imputation before autoregressive decoding. However, they typically treat imputation as a fixed reconstruction task, without considering which missing fragments are most relevant to decoder errors. Moreover, existing peak representations do not explicitly model mass differences between peaks, despite their fundamental importance. We introduce GyroNovo, a framework with two main contributions. First, we use decoder errors observed during training to adapt the imputation objective, prioritizing fragments associated with frequent decoding errors. We further use the decoder error distribution to construct easy and hard augmented views of each spectrum, enabling the decoder to learn under varying degrees of spectral corruption and missing-fragment severity. Second, we introduce a mass-aware inductive bias into self-attention by using rotary embeddings to encode pairwise mass differences between spectral peaks. Together, these components align missing-fragment recovery with decoder behavior while explicitly incorporating the mass relationships that underlie peptide fragmentation. At inference time, GyroNovo retains a standard encoder-imputer-decoder architecture and requires neither additional inputs nor auxiliary search procedures. Experiments on NovoBench show gains of about 9 percentage points in peptide-level precision and 7 percentage points in amino-acid-level precision over the state-of-the-art baseline. Code: https://github.com/UBC-NLP/gyronovo.

Authors

  1. Abdellah El Mekki · University of British Columbia
  2. Laks V.S. Lakshmanan · University of British Columbia
  3. Muhammad Abdul-Mageed · Mohamed bin Zayed University of Artificial Intelligence, University of British Columbia

Methods and tools

  • GyroNovo: Attacks missing b- and y-ion fragments on two fronts. Rather than treating imputation as a fixed reconstruction task, it uses the decoder errors seen during training to steer the imputation objective toward the fragments that actually cause mistakes, and to build easy and hard augmented views of each spectrum so the decoder learns under varying spectral corruption. It also gives self-attention a mass-aware inductive bias, using rotary embeddings to encode pairwise mass differences between peaks. Inference needs no extra inputs or search. Reports about 9 points of peptide-level and 7 points of amino-acid-level precision over the previous best on NovoBench.

Seen in the charts

Back to the full map

Back to top