What twenty-five years taught us about protein structure data

20 Jul 2026

How a major pharmaceutical collaboration highlighted the critical need for interrogable data pipelines in the era of machine learning

By Dr Neil Taylor

A number of years ago, our research team began a close collaboration with a major pharmaceutical partner on a question that had been on my mind for some time: how much of the interesting behaviour in a binding site is actually being missed by the way we conventionally look at it?

Structure-based design and lead optimisation has, for decades, focused heavily on near neighbours and close contacts – the direct interactions between a ligand and the residues immediately surrounding it. That focus is well justified, but it is also, in a sense, a narrow lens. Proteins are not collections of independent local contacts. They also exhibit small-world network behaviour, where residues that are distant from one another in sequence, and sometimes in space, can be connected through short paths of interaction that produce cooperative effects well beyond the binding site itself. Looking only at close contacts means missing a great deal of what is actually driving binding behaviour and protein function.

The limitation of isolated structures

The core lesson from our work is this: cooperative effects of this kind are rarely visible from a single structure examined in isolation. A network of distant, weakly coupled interactions only becomes apparent when you can compare the same binding site, or the same fold, across many structures at once. This means noticing:

  • Conserved binding site water molecules
  • Which residue sidechains move together
  • Which contacts appear only when others are present
  • Which features recur across a family of related targets

That kind of comparison is only possible if the underlying structural data is organised well enough to be queried across tens, hundreds, or thousands of structures at once, rather than examined one file at a time. It sounds like a modest requirement. In practice, it is one of the harder problems in the field of structural biology, and one that receives far less attention than it deserves.

Lessons from a high-calibre collaboration

Working alongside a modelling team of the highest calibre, with access to top-quality data and substantial resources, was a genuinely clarifying experience for us. It sharpened something we had always suspected: that the difficulty of quantifying network and cooperative effects has very little to do with the sophistication of the processes within the research organisation involved, and a great deal to do with the levels of detail required.

Even with excellent science and substantial resources on both sides, the underlying challenge does not go away. It is rarely a shortage of data, and rarely a shortage of computational power. Instead, it is the absence of very high quality structures and very fine-detail analyses that allow the right questions to be asked across that data efficiently, and the right comparisons to be drawn.

Proasis: The invisible foundation

This is, in a sense, the problem DesertSci has spent a quarter of a century working on, in one form or another. Proasis was never intended to be the most visible layer of a research organisation’s toolkit, and it was not designed to compete with the modelling and design platforms our clients already use and value.

Its purpose has always been to sit underneath those tools, ensuring that the high-value experimental structural data feeding into them is organised, consistent, and genuinely interrogable, rather than simply archived.

In other words, Proasis can be totally invisible to many scientists. As the saying goes, “nobody gets thanked when nothing goes wrong.”

The AI era makes data discipline acute

I raise this now because the current industry interest in AI-driven approaches to drug discovery has, if anything, made our structural data challenges more acute rather than less. A well-constructed machine learning model can identify patterns across structural data with a speed and subtlety that would have been impractical a decade ago. But a model can only interrogate the data in the form given to it. Where that underlying data is fragmented, inconsistently curated, or difficult to cross-reference, even the most capable model will be working with one hand tied behind its back.

Our recent collaboration was a concrete demonstration of this principle in a real-world research setting rather than an abstract one. It strengthened our conviction that the discipline of organising structural data well is not a solved problem to be taken for granted.

The path forward

At DesertSci, we are continuing to develop the methods that emerged from this work, and in time, we look forward to discussing them in more detail. In the meantime, I wanted to share this broader observation, because I suspect it will resonate with others doing similar work in the life sciences sector:

The organisations that get the most out of the next generation of computational tools are unlikely to be those with access to the most advanced AI models alone. They will be the ones that have also done the less glamorous work of getting their structural data into a state where those models can actually be trusted.

Dr Neil Taylor

For more insights into AI in drug discovery, computational chemistry and research-grade scientific software, follow follow Dr Neil Taylor on LinkedIn. Or, if you’d like to arrange a demonstration of DesertSci’s Proasis, please get in touch with our team.

Posted in: Current

Comments: (0)

Leave a Comment