Diffusion
6 methods · 2023–2026
Diffusion: Iterative denoising: start from noise over the residue positions and refine repeatedly, so the precursor-mass constraint and the consistency of the whole sequence can be enforced at every step rather than only at the end.
Iterative denoising: start from noise over the residue positions and refine repeatedly, so the precursor-mass constraint and the consistency of the whole sequence can be enforced at every step rather than only at the end.
The earliest of its 6 methods is InstaNovo+ (2023); 5 more have followed.
| Methods | 6 |
| Papers describing them | 8 |
| Authors | 28 |
| Active | 2023-08-30 to 2026-07-04 |
| Deep learning | 6 of 6 |
| Kinds | algorithm (5), adjacent |
| Acquisition | DDA (4), DIA |
Methods (6)
Oldest first, by the paper that describes each one.
- InstaNovo+ (2023): Multinomial diffusion
- DiffNovo-DIA (2025): Transformer-diffusion model for DIA de novo peptide sequencing: DIA-side companion to DiffNovo. From Shiva Ebrahimi’s PhD thesis (UNT, 2025); a PyPI package exists (diffnovo-dia v0.1.2) but the GitHub repo is currently empty and no standalone paper has been published.
- Casanovo-DM2 (2025): Diffusion decoder variants (Casanovo-DS / DM1 / DM2) plugged into Casanovo’s spectrum encoder; DM2 + DINOISER loss reported the best amino-acid recall.
- DiffuNovo (2026): Regressor-guided diffusion
- Diffusion spectrum foundation model (2026): Self-supervised diffusion encoder for tandem mass spectra, trained by masked spectrum prediction on unlabelled data to give general-purpose spectrum embeddings transferable to downstream proteomics tasks instead of representations inherited from a model optimised only for de novo sequencing. Reaches an R² of 0.784 for precursor m/z at scale against a 0.923 baseline while learning compact, organised latent representations, a trade-off between compactness and predictive accuracy; precursor-charge and fragmentation-type classification stay near the majority-class baselines. Ablations indicate the masking and multi-loss objective drive the latent structure, and that the centroid loss needs modulation such as AdaLN to avoid penalising the learned embeddings.
- PhysNovo (2026): Discrete-diffusion de novo peptide sequencer that folds in physical mass-constraint terms at inference time, so generated sequences respect the observed precursor mass rather than relying purely on the learned amino-acid prior.
Papers describing them (8)
- De novo peptide sequencing with InstaNovo: Accurate, database-free peptide identification for large scale proteomics experiments (2023, bioRxiv, preprint)
- InstaNovo enables diffusion-powered de novo peptide sequencing in large-scale proteomics experiments (2025, Nature Machine Intelligence, peer-reviewed)
- De Novo Peptide Sequencing for Data-independent Acquisition (DIA) Using Deep Learning (2025, thesis)
- Diffusion Decoding for Peptide De Novo Sequencing (2025, arXiv, preprint)
- Regressor-guided Diffusion Model for De Novo Peptide Sequencing with Explicit Mass Control (2026, arXiv, preprint)
- Regressor-guided Diffusion Model for De Novo Peptide Sequencing with Explicit Mass Control (2026, AAAI 2026, peer-reviewed)
- Diffusion-based Foundation Model For Mass Spectrometry Via Self-Supervised Learning on Unlabelled Data (2026, thesis)
- Discrete Diffusion with Physical Mass Constraints for De Novo Peptide Sequencing (2026, ICML 2026, ML conference)