Glossary · term

ProtoMech

ProtoMech is a named mechanistic-interpretability framework for tracing task-specific computation in protein language models. It trains cross-layer transcoders (CLTs) to approximate ESM2 feed-forward-layer outputs with sparse features from the current and earlier layers, then selects small feature sets as circuits for a probe-defined task. `Protein circuit tracing` describes this application; it is not yet a general standard or a synonym for every method that interprets protein models.

Safety2026-02-12Wave 3 · 2025–26Maturity: 3/5

Origin and context

Darin Tsui, Kunal Talreja, Daniel Saeedi and Amirali Aghazadeh submitted the first ProtoMech preprint on 12 February 2026 and revised it in May; the work was accepted at ICML 2026. The authors released code for training CLTs, finding circuits, steering representations and visualizing results, with documented support for ESM2-8M and ESM2-35M. Those artifacts define ProtoMech more precisely than the record's inherited slash label.

Sources: s1, s2, s3

Why it matters

Protein-model interpretability has often used sparse autoencoders to decompose activations into features associated with motifs, sites or domains. ProtoMech asks a different question: can a sparse replacement approximate computation across layers and retain a supervised task signal? This distinction separates three properties that are easy to conflate: replacement fidelity, circuit sparsity and biological interpretability. A compact circuit can preserve a probe score without proving that its nodes are the biological mechanism used by a protein or even the complete mechanism used by ESM2.

Sources: s1, s5

Example

For family classification, the authors trained a logistic probe on ESM2's final MLP output and added attributed CLT latents until a sparse circuit reached a task-performance target. On ESM2-8M, the full replacement recovered 89% of the original classifier's F1, while selected circuits recovered 79% using about 0.8% of the latent space on average. An independent SJSU study then used a ProtoMech transcoder for beta-lactamase classes and found relevant signals distributed across layers; several strong nodes failed its additional validation stages. That is independent technical use, not a replication of every original result.

Sources: s1, s4

How it differs

Cross-layer transcoders / CLTs

A cross-layer transcoder is the sparse replacement-model component. ProtoMech combines CLTs with task probes, circuit selection, steering and protein-specific visualization.

Sparse Autoencoders (SAEs)

A sparse autoencoder reconstructs the representation it receives. ProtoMech's CLT predicts MLP outputs from sparse features across layers, so its fidelity target and circuit claims differ.

Circuit Tracing

Circuit tracing is the broader interpretability method family. ProtoMech adapts it to protein models and evaluates a hybrid replacement whose attention activations still come from the original ESM2 model.

Maturity and evidence

Maturity is 3. ProtoMech has an ICML-accepted paper, public code and model artifacts, and an organizationally independent academic application using the same named method. The independent study adds a useful validation challenge rather than merely repeating the abstract. Evidence remains narrow, however: one originating study and one independent preprint or master's project do not establish broad adoption, standardization or general performance across protein-model families.

Sources: s1, s2, s3, s4

Limits and open questions

The main experiments cover masked ESM2 models, supervised downstream probes and a replacement that keeps original attention outputs fixed; fully recursive replacement accumulated substantial error. CLT decoder count grows quadratically with layer count, and biological labels were assigned through manual analysis of selected examples. The protein-steering evaluation used a CNN fitness proxy trained from DMS data, not new wet-lab measurements, and generated variants stayed within five mutations of wild type. ProtoMech therefore does not by itself prove a biological mechanism, experimental fitness, safety or readiness for protein-engineering decisions.

Sources: s1, s4

Related terms

References

Last updated: 2026-09-07

In the Skills Atlas

This term is also covered in the Skills Atlas as mechanistic interpretability skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as deep learning skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as model evaluation skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as research to engineering translation skill.