CompOmics PRIDE
training · DDA · 1 version
Eighteen public PRIDE projects reprocessed into one labelled corpus and split for training. Every row keeps its source accession in a PXD_identifier column, so the provenance survives inside the data rather than only in the paper.
| Kind | training |
| Acquisition | DDA |
| Organisms | mixed |
| Home | https://huggingface.co/datasets/InstaDeepAI/CompOmics_PRIDE |
Eighteen public PRIDE projects reprocessed into one labelled corpus and split for training. Every row keeps its source accession in a PXD_identifier column, so the provenance survives inside the data rather than only in the paper.
Versions
as published
Part of InstaNovo’s training data. 18 source projects, split train/validation/test.
spectra 8,289,063 · train 6,992,871 · validation 123,307 · test 1,172,885 · released 2025-05-06 · introduced by InstaNovo enables diffusion-powered de novo peptide sequencing in large-scale proteomics experiments
Where it lives:
- Hugging Face · InstaDeepAI/CompOmics_PRIDE
Assembled from 18 third-party submissions: