CompOmics PRIDE

training · DDA · 1 version

Eighteen public PRIDE projects reprocessed into one labelled corpus and split for training. Every row keeps its source accession in a PXD_identifier column, so the provenance survives inside the data rather than only in the paper.
Kind training
Acquisition DDA
Organisms mixed
Home https://huggingface.co/datasets/InstaDeepAI/CompOmics_PRIDE

Eighteen public PRIDE projects reprocessed into one labelled corpus and split for training. Every row keeps its source accession in a PXD_identifier column, so the provenance survives inside the data rather than only in the paper.

Versions

as published

Part of InstaNovo’s training data. 18 source projects, split train/validation/test.

spectra 8,289,063 · train 6,992,871 · validation 123,307 · test 1,172,885 · released 2025-05-06 · introduced by InstaNovo enables diffusion-powered de novo peptide sequencing in large-scale proteomics experiments

Where it lives:

Assembled from 18 third-party submissions:

Seen in the charts

Back to the full map

Back to top