High-resolution dissection of concept acquisition in different families of protein language models

Avatar
Poster
Voice is AI-generated
Connected to paperThis paper is a preprint and has not been certified by peer review

High-resolution dissection of concept acquisition in different families of protein language models

Authors

Whitfield, S. T.; Marty, T.; Vernon, R. M.; Langmead, C. J.; Sridhar, D.; Fournier, Q.

Abstract

Protein language models have been increasingly successful on tasks ranging from fitness prediction to functional design, yet what biological knowledge they acquire and where it is encoded within their internal representations remain underexplored. Through a high-resolution layer-by-layer interpretability analysis of 8 models from the ESM2 and AMPLIFY families on 22 concepts from human proteome annotations, we found that these models encode concepts of increasing levels of complexity along their depth: basic physicochemical properties and linear motifs are best captured by early-layer embeddings, secondary structure from subsequent layers, and domain-level semantics from middle layers. Principal component projections of these embeddings showed that they separate biologically meaningful protein groupings, and molecular-biology-inspired interventions demonstrated that pLM embeddings can discriminate phosphomimic-active from inactive mutants. Perhaps surprisingly, we observed that pretraining data and compute had a greater impact on the linear emergence of biological concepts than scaling up parameters. By revealing where biological knowledge is captured in pLMs and which choices shape its emergence, our work offers insights to develop more robust, biologically grounded protein language models.

Follow Us on

0 comments

Add comment