Premium

How scientists got a glimpse of the inner workings of protein language models

A team of researchers based at the Massachusetts Institute of Technology (the United States) has tried to shed light on the inner workings of the language models that predict the structure and function of proteins by using an innovative technique

Typically, proteins are made of a combination of 20 different kinds of amino acids, and the structure and function of each protein are governed by the arrangement of the various amino acids in it.Typically, proteins are made of a combination of 20 different kinds of amino acids, and the structure and function of each protein are governed by the arrangement of the various amino acids in it. (Photo: Wikimedia Commons)
Written by: Alind Chauhan
6 min readNew DelhiSep 8, 2025 01:11 PM IST First published on: Sep 6, 2025 at 02:35 PM IST

The recent emergence of large language models (LLMs) has revolutionised the research on proteins — the microscopic mechanisms that are involved in virtually every important activity happening inside all living things. Scientists use a version of LLM for various tasks, such as predicting the structure or function of a protein, which contributes to the development of drugs and vaccines.

However, little is known about how these models make such predictions, as they work like a “black box”, meaning there is no way to determine what is happening inside them. This poses several kinds of issues. For instance, researchers do not know if the basis for the model’s prediction is meaningful or not, and might waste months or years on dead-end experiments.

Latest Comment
Post Comment
Read Comments