Sparse autoencoder interpretability of InstaNovo
adjacent · Sparse autoencoder
Mechanistic-interpretability probe of a de novo sequencer: sparse autoencoders trained on the internal activations at four depths of InstaNovo’s spectrum encoder, paired with a fragment-ion annotation pipeline that labels every peak with chemical concepts derived from first principles, so learned features can be matched to chemical events. Reconstructs the encoder’s representations at over 92% fidelity per layer while preserving more than 99.5% of sequencing loss, and finds the feature-to-concept associations are organised by depth: physical spectral properties sharpest in early layers, cleavage specificity and modification localisation emerging deeper. Causal ablation tests how far those associations reflect genuine use by the model, and exposes a structural limit in correlation-based feature ranking when chemical concepts co-occur.
| Kind | adjacent |
| Family | Sparse autoencoder |
| Deep learning | yes |
| Acquisition | DDA |
Papers
- Sparse Autoencoders for Interpretability of InstaNovo (2026, thesis)