Nine-species benchmark

benchmark · DDA · 5 versions · 34 papers

The field’s most-cited evaluation set: tryptic DDA runs from nine taxonomically distant organisms, used leave-one-species-out so a model is tested on a proteome it never trained on. Assembled by DeepNovo from nine unrelated public submissions.
Kind benchmark
Acquisition DDA
Organisms Apis mellifera, Bacillus subtilis, Candidatus Thiodiazotropha endoloripes, Homo sapiens, Methanosarcina mazei, Mus musculus, Saccharomyces cerevisiae, Solanum lycopersicum, Vigna mungo
Home https://massive.ucsd.edu/ProteoSAFe/dataset.jsp?accession=MSV000081382

The field’s most-cited evaluation set: tryptic DDA runs from nine taxonomically distant organisms, used leave-one-species-out so a model is tested on a proteome it never trained on. Assembled by DeepNovo from nine unrelated public submissions.

Versions

original (DeepNovo, 2017)

The MGF files DeepNovo curated and deposited. What a paper means by “the nine-species dataset” unless it says otherwise.

released 2017-07-18 · introduced by De novo peptide sequencing by deep learning

Where it lives:

Assembled from 9 third-party submissions:

revised (main)

Re-curated to remove peptide redundancy between species, which leaked test peptides into training in the original, and re-released with a fix for a bug that wrongly detected shared peptides between species. Introduced by “A multi-species benchmark for training and validating mass spectrometry proteomics machine learning models”. 2,844,842 spectra. The MassIVE record carries dated update folders, so that accession is not a single fixed object.

spectra 2,844,842 · released 2024-11-08 · introduced by A multi-species benchmark for training and validating mass spectrometry proteomics machine learning models

Where it lives:

InstaNovo split

A fixed train/validation/test split published as parquet with its own DOI, which makes it the only version of this benchmark that is reproducible by citation alone. Its 499,402 training spectra are the same count NovoBench retrains every architecture on.

spectra 639,286 · train 499,402 · validation 28,572 · test 111,312 · released 2024-12-01

Where it lives:

ProteoBench selection

The selection ProteoBench’s de novo DDA-HCD module scores submissions against.

spectra 779,879

Where it lives:

revised (balanced)

The same re-curation, then randomly thinned so the species are more evenly represented: smaller and balanced rather than larger and skewed. Shipped in the same Zenodo record as the main variant, so “the revised benchmark” is still two different things.

released 2024-11-08 · introduced by A multi-species benchmark for training and validating mass spectrometry proteomics machine learning models

Where it lives:

Deposited by (2)

More than one paper here means one piece of work published twice, a preprint and its version of record; a deposit itself happens once.

Used by (32)

Checkpoints trained on this (2)

Method Checkpoint Version of this dataset Size Get it
Casanovo nine-species not stated 5.1 GB Zenodo
PhysNovo nine-species not stated 1.7 GB Zenodo

Each link rests on stated evidence rather than a text match:

  • Casanovo nine-species: the Zenodo record is titled ‘Casanovo model weights on nine-species benchmark’; it does not say WHICH version of the benchmark, so none is asserted
  • PhysNovo nine-species: the Zenodo record is titled ‘PhysNovo Model Weights Trained on the Nine-Species Benchmark Dataset’; it does not say which version

Methods on these papers (29)

Taken from the describing links only, so a paper that merely ran a tool on this data does not make that tool a method of it.

Seen in the charts

Back to the full map

Back to top