MegaPX: fast and space-efficient peptide assignment method using IBF-based multi-indexing

peer-reviewed · Bioinformatics · 2026

peer-reviewed · Bioinformatics · 2026. Ahmad Lutfi et al. Motivation A central problem for metaproteomic analysis is the often-unknown taxonomic composition of the…
Date 2026-05-03
Type peer-reviewed
Venue Bioinformatics
Publisher Oxford University Press (OUP)
Contribution post-processor
DOI 10.1093/bioinformatics/btag134
Citations (OpenAlex) 0
Venue 2-year citedness 6.16

Abstract

Motivation A central problem for metaproteomic analysis is the often-unknown taxonomic composition of the analyzed microbiomes. Using a database search, the standard approach requires prior knowledge of which proteins and taxa to include in the protein reference database or to use tailored metagenome-derived databases, which are expensive and error-prone in their generation. A possible strategy to circumvent this database search issue is de novo sequencing, where peptide sequences are directly identified from mass spectra. However, these sequences must still be mapped back to potentially extensive databases. Here, alignment-based approaches enable robust and precise results, with the potential drawback of high memory usage and long run times. Results We present MegaPX, a software for rapidly classifying de novo peptide sequences against large protein databases. MegaPX implemented as a C++-based tool, uses an alignment-free, k-mer approach as a taxonomic classification method with the possibility of generating mutated reference databases for error-tolerant searching. It uses various algorithms, including interleaved Bloom filters, to efficiently compute approximate membership queries, ensuring fast processing times while querying and indexing large databases in a multi-indexing fashion. We demonstrate the potential of MegaPX by analyzing different samples, including metaproteomics, against extensive reference databases, highlighting its use as a fast screening tool.

Authors

  1. Ahmad Lutfi · Robert Koch Institute
  2. Tanja Holstein · Centre National de la Recherche Scientifique, Ghent University, Institut Pluridisciplinaire Hubert Curien, Proteomics French Infrastructure Core, Robert Koch Institute, Université de Strasbourg, VIB, VIB-UGent Center for Medical Biotechnology, Vlaams Instituut voor Biotechnologie
  3. Sandro Andreotti · Freie Universität Berlin, International Max Planck Research School
  4. Thilo Muth · Federal Institute for Materials Research and Testing (BAM), Max Planck Institute for Dynamics of Complex Technical Systems, Robert Koch Institute, glyXera GmbH

Methods and tools

  • MegaPX: Alignment-free k-mer and interleaved Bloom filter tool that rapidly assigns de novo peptide sequences to taxa in very large protein databases, for metaproteomics.

Seen in the charts

Back to the full map

Back to top