The Power and the Limitations of Cross-Species Protein Identification by Mass Spectrometry-driven Sequence Similarity Searches

peer-reviewed · Molecular & Cellular Proteomics · 2004

peer-reviewed · Molecular & Cellular Proteomics · 2004. Bianca Habermann et al. Mass spectrometry-driven BLAST (MS BLAST) is a database search protocol for identifying unknown proteins by…
Date 2004-03-01
Type peer-reviewed
Venue Molecular & Cellular Proteomics
Publisher Elsevier BV
Contribution adjacent
DOI 10.1074/mcp.m300073-mcp200
Citations (OpenAlex) 150
Venue 2-year citedness 4.69

Abstract

Mass spectrometry-driven BLAST (MS BLAST) is a database search protocol for identifying unknown proteins by sequence similarity to homologous proteins available in a database. MS BLAST utilizes redundant, degenerate, and partially inaccurate peptide sequence data obtained by de novo interpretation of tandem mass spectra and has become a powerful tool in functional proteomic research. Using computational modeling, we evaluated the potential of MS BLAST for proteome-wide identification of unknown proteins. We determined how the success rate of protein identification depends on the full-length sequence identity between the queried protein and its closest homologue in a database. We also estimated phylogenetic distances between organisms under study and related reference organisms with completely sequenced genomes that allow substantial coverage of unknown proteomes.

Authors

  1. Bianca Habermann · Max Planck Institute of Molecular Cell Biology and Genetics
  2. Jeffrey Oegema
  3. Shamil Sunyaev · Brigham & Women’s Hospital and Harvard Medical School, Brigham and Women’s Hospital, European Molecular Biology Laboratory, Harvard University
  4. Andrej Shevchenko · European Molecular Biology Laboratory, Max Planck Institute of Molecular Cell Biology and Genetics, Max Planck Society, University of California San Diego

Methods and tools

  • MS BLAST: Homology search of error-tolerant de novo sequences against a protein database, introduced for charting the proteomes of organisms with unsequenced genomes and later used to validate borderline identifications.

Seen in the charts

Back to the full map

Back to top