Sequence similarity‐driven proteomics in organisms with unknown genomes by LC‐MS/MS and automated de novo sequencing

peer-reviewed · PROTEOMICS · 2007

peer-reviewed · PROTEOMICS · 2007. Patrice Waridel et al. LC-MS/MS analysis on a linear ion trap LTQ mass spectrometer, combined with data processing, stringent, and…
Date 2007-07-01
Type peer-reviewed
Venue PROTEOMICS
Publisher Wiley
Contribution algorithm
DOI 10.1002/pmic.200700003
Citations (OpenAlex) 99

Abstract

LC-MS/MS analysis on a linear ion trap LTQ mass spectrometer, combined with data processing, stringent, and sequence-similarity database searching tools, was employed in a layered manner to identify proteins in organisms with unsequenced genomes. Highly specific stringent searches (MASCOT) were applied as a first layer screen to identify either known (i.e. present in a database) proteins, or unknown proteins sharing identical peptides with related database sequences. Once the confidently matched spectra were removed, the remainder was filtered against a nonannotated library of background spectra that cleaned up the dataset from spectra of common protein and chemical contaminants. The rectified spectral dataset was further subjected to rapid batch de novo interpretation by PepNovo software, followed by the MS BLAST sequence-similarity search that used multiple redundant and partially accurate candidate peptide sequences. Importantly, a single dataset was acquired at the uncompromised sensitivity with no need of manual selection of MS/MS spectra for subsequent de novo interpretation. This approach enabled a completely automated identification of novel proteins that were, otherwise, missed by conventional database searches.

Authors

  1. Patrice Waridel · Max Planck Institute of Molecular Cell Biology and Genetics, University of California San Diego
  2. Ari Frank · Affectivon, Inc., Max Planck Institute of Molecular Cell Biology and Genetics, University of California San Diego
  3. Henrik Thomas · Max Planck Institute of Molecular Cell Biology and Genetics, University of California San Diego
  4. Vineeth Surendranath · Max Planck Institute of Molecular Cell Biology and Genetics, University of California San Diego
  5. Shamil Sunyaev · Brigham & Women’s Hospital and Harvard Medical School, European Molecular Biology Laboratory
  6. Pavel Pevzner · University of California San Diego
  7. Andrej Shevchenko · European Molecular Biology Laboratory, Max Planck Institute of Molecular Cell Biology and Genetics, University of California San Diego

Methods and tools

  • Layered sequence-similarity proteomics: Identifies proteins in organisms with unsequenced genomes by layering the search: a stringent database search first, then automated de novo sequencing of what it leaves behind, then sequence-similarity searching of those de novo reads against related species.

Cites (11)

Cited by (7)

Seen in the charts

Back to the full map

Back to top