Hybrid de novo + database search
3 methods · 2024–2026
Hybrid de novo + database search: Two engines run side by side and reconciled, a de novo sequencer and a database search, each keeping the identifications the other cannot make.
Two engines run side by side and reconciled, a de novo sequencer and a database search, each keeping the identifications the other cannot make.
The earliest of its 3 methods is Orthrus (2024); 2 more have followed.
| Methods | 3 |
| Papers describing them | 3 |
| Authors | 10 |
| Active | 2024-11-15 to 2026-06-29 |
| Deep learning | 2 of 3 |
| Kinds | downstream-application (2), adjacent |
| Acquisition | DDA (3) |
Methods (3)
Oldest first, by the paper that describes each one.
- Orthrus (2024): Open-source metaproteomics pipeline combining Casanovo transformer-based de novo sequencing with Sage database search and Mokapot rescoring.
- HDPS (2025): Heuristic two-round sequence assembly strategy combining multi-enzyme and microwave-assisted acid hydrolysis, pNovo de novo peptide sequencing, and pFind homology database search with k-mer graph assembly and majority-vote error correction; achieved 100% sequence coverage and >98% amino-acid accuracy on full-length Ricin toxin A and B chains without a reference sequence, outperforming ALPS.
- INSearch (2026): Prototype AI-native database search that replaces the combinatorial scan with retrieval. A dual encoder projects experimental spectra and theoretical peptide sequences into one shared latent space under a contrastive objective with variance regularisation to prevent feature collapse, and an alternative alignment-and-uniformity loss; approximate nearest-neighbour search then returns the top-k candidates, cutting retrieval from the O(S × ρP) of a conventional engine to O(S(log P + K)). Because retrieval is approximate, candidates are re-ranked by InstaNovo’s decoder run in teacher-forcing mode as a scoring function, aggregating residue log-probabilities by geometric mean. On the nine-species benchmark the held-out yeast split reaches Recall@1/5/100 of 60.0/78.9/91.9%, with neural re-scoring lifting Recall@1 to 82.9%; the harder S. brodae proteome reaches only 55.5% Recall@100, pointing to a need for larger-scale training.
Applied in
Papers describing them (3)
- Orthrus: an AI-powered, cloud-ready, and open-source hybrid approach for metaproteomics (2024, bioRxiv, preprint)
- Identification of Unknown Biological Toxin Proteins Using Mass Spectrometry: A Case Study on De Novo Sequencing of Ricin (2025, Toxins, peer-reviewed)
- INSearch: AI-native framework for large scale proteomic database search via contrastive joint embeddings and transformer-based scoring function (2026, thesis)