GSK and Relation Therapeutics Partner in AI Drug Discovery Deal
Global pharmaceutical giant GSK has expanded its collaboration with British biotech firm Relation Therapeutics in a deal worth up to $110 million focusing on biological data generation.

Stock photo for illustration only, not from the actual event
- GSK partners with Relation Therapeutics in a deal worth up to $110 million
- Focuses on generating large-scale datasets to train AI models within the MORGAN platform
- Combines laboratory experimentation with computational analysis
- Highlights the critical role of high-quality, diverse biological datasets
Global pharmaceutical company GSK has entered into a research collaboration with British biotechnology firm Relation Therapeutics worth up to $110 million, expanding the companies' existing work in AI-assisted drug discovery.
Under the agreement, Relation will generate large-scale datasets measuring how human cells respond to genetic changes and drug interventions. The data will be used to train AI models designed to identify potential drug targets, including models within Relation’s MORGAN platform.
The agreement places biological data generation alongside AI model development. Relation’s research approach links computational analysis with experiments that generate new information on human cells. This collaboration builds on earlier agreements between GSK and Relation focused on fibrotic diseases and osteoarthritis, which involved observational studies designed to create functional disease datasets for analysis using Relation’s Lab-in-the-Loop platform.

Stock photo for illustration only, not from the actual event
The massive investments pharmaceutical giants are making into proprietary biological datasets indicate that the primary bottleneck in AI-driven drug discovery has shifted from algorithm design to data quality and biological relevance. Bridging wet-lab experiments directly with computational pipelines is rapidly becoming the defining strategy for next-generation therapeutics.
Public repositories remain an important source of training material for biological foundation models, although combining information produced across different studies can introduce technical challenges. A 2025 review in Experimental & Molecular Medicine noted that repositories including CZ CELLxGENE, the Human Cell Atlas, and NCBI Gene Expression Omnibus give researchers access to large volumes of single-cell data, with CZ CELLxGENE alone providing access to more than 100 million standardised cells.
However, sampling methods, sequencing protocols, and processing pipelines can differ between studies, introducing technical noise that requires careful quality control. Furthermore, research published in Nature Methods in June examined 22.2 million cells across 400 models and 6,400 experiments, finding that current single-cell foundation models tended to reach performance plateaus after training on only a fraction of the available corpus, differing from the clear data-scaling laws seen in large language models.
Source: AI News
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment