Diffusion

6 methods · 2023–2026

Diffusion: Iterative denoising: start from noise over the residue positions and refine repeatedly, so the precursor-mass constraint and the consistency of the whole sequence can be enforced at every step rather than only at the end.

Iterative denoising: start from noise over the residue positions and refine repeatedly, so the precursor-mass constraint and the consistency of the whole sequence can be enforced at every step rather than only at the end.

The earliest of its 6 methods is InstaNovo+ (2023); 5 more have followed.

Methods 6
Papers describing them 8
Authors 28
Active 2023-08-30 to 2026-07-04
Deep learning 6 of 6
Kinds algorithm (5), adjacent
Acquisition DDA (4), DIA

Methods (6)

Oldest first, by the paper that describes each one.

  • InstaNovo+ (2023): Multinomial diffusion
  • DiffNovo-DIA (2025): Transformer-diffusion model for DIA de novo peptide sequencing: DIA-side companion to DiffNovo. From Shiva Ebrahimi’s PhD thesis (UNT, 2025); a PyPI package exists (diffnovo-dia v0.1.2) but the GitHub repo is currently empty and no standalone paper has been published.
  • Casanovo-DM2 (2025): Diffusion decoder variants (Casanovo-DS / DM1 / DM2) plugged into Casanovo’s spectrum encoder; DM2 + DINOISER loss reported the best amino-acid recall.
  • DiffuNovo (2026): Regressor-guided diffusion
  • Diffusion spectrum foundation model (2026): Self-supervised diffusion encoder for tandem mass spectra, trained by masked spectrum prediction on unlabelled data to give general-purpose spectrum embeddings transferable to downstream proteomics tasks instead of representations inherited from a model optimised only for de novo sequencing. Reaches an R² of 0.784 for precursor m/z at scale against a 0.923 baseline while learning compact, organised latent representations, a trade-off between compactness and predictive accuracy; precursor-charge and fragmentation-type classification stay near the majority-class baselines. Ablations indicate the masking and multi-loss objective drive the latent structure, and that the centroid loss needs modulation such as AdaLN to avoid penalising the learned embeddings.
  • PhysNovo (2026): Discrete-diffusion de novo peptide sequencer that folds in physical mass-constraint terms at inference time, so generated sequences respect the observed precursor mass rather than relying purely on the learned amino-acid prior.

Papers describing them (8)

Authors (28)

Alexander Wong, Amandla Mabona, Andreas Hougaard Laustsen, Anne Ljungars, Chi-en Amy Tai, Erwin M. Schoof, Esperanza Rivera-de-Torre, Jakob Berg Jespersen, Jeroen Van Goey, Jingbo Zhou, Jun Xia, Karim Beguir, Kevin Eloff, Konstantinos Kalogeropoulos, Marcin J. Skwark, Nicolas Lopez Carranza, Oliver Morell, Rachel Catzel, Sam P. B. van Beljouw, Shaorong Chen, Shiva Ebrahimi, Stan J. J. Brouns, Timothy P. Jenkins, Ulrich auf dem Keller, Vicent Mwanda, Wanyu Lin, Wesley Williams, Zeyu An

Seen in the charts

Back to the full map

Back to top