Jobs and Careers
GS

Principal Scientist, Oncology Data Science (Translational Science)

GSK
Providence, United StatesRemotefull_timeVerifiedPosted 5 Aug 2026
💰 $202,125/yr($121,275/yr$202,125/yr)

About the role

Business Introduction
At GSK, we have bold ambitions for patients, aiming to positively impact the health of 2.5 billion people by the end of the decade. Our R&D focuses on discovering and delivering vaccines and medicines, combining our understanding of the immune system with cutting-edge technology to transform people’s lives. GSK fosters a culture ambitious for patients, accountable for impact, and committed to doing the right thing, making sure that we focus our efforts on accelerating significant assets that meet patients’ needs and have the highest probability of success. We’re uniting science, technology, and talent to get ahead of disease together.
Find out more:
Our approach to R&D
 

For candidates seeking to be located at our Stevenage site, this role will temporarily be based at Stevenage. However, the Company plans to relocate its offices to Cambridge, UK. The location of this role will therefore subsequently change to Cambridge, UK in accordance with timelines to be set by the Company. The relocation is currently proposed to take effect by early 2029.
 

Position Summary
The GSK Oncology Data Science team in R&D Translational Science is seeking a Translational AI scientist to build ML applications for a long-sought-after problem: if we alter a patient tumor’s molecular state in silico, can we predict how their clinical trajectory will change? To tackle this problem, you will integrate and validate multimodal foundation models; bridge functional genomics, spatial omics, and real-world data; and apply cutting edge causal inference techniques.

We operate with high velocity at the intersection of machine learning, causal inference, functional genomics, spatial biology, and real-world clinical data; your expertise, execution, technical leadership and communication will drive our efforts to bring the right therapies to the right patients.

Responsibilities
This role will provide YOU the opportunity to lead key activities to progress YOUR career. These responsibilities include some of the following:

  • Own the pipeline and develop advanced ML architectures to integrate complex multimodal datasets, including single-cell, spatial omics, histopathology, functional genomics, and real-world clinical data.

  • Partner closely with wet-lab scientists, clinicians, and pathologists to validate machine learning models, including in-silico perturbations within the tumor microenvironment against ground-truth data (counterfactual validation).

  • Develop approaches to extract interpretable features from models to generate testable oncological hypotheses and link insights to clinical pipeline decisions such as asset prioritization and patient subpopulation selection.

  • Contribute clean, reproducible tooling to cross-team frameworks. We enforce good engineering practices in our research—utilizing code architecture planning, clean code and automated testing to build trustworthy, reusable code.

  • Maintain cutting edge knowledge of advancements, share with the team and maintain our team as a thought leader through publications in high-impact venues and engaging with the broader community.


Why You?

Basic Qualification
We are seeking professionals with the following required skills and qualifications to help us achieve our goals:

  • PhD (or equivalent experience) in a quantitative field (Applied ML, Computer Science, Physics, Systems/Computational Biology, or equivalent) with 1+ years of industry or productive post-doctoral academic experience.

  • Experience with deeply embedded in cancer / computational biology, with a strong understanding of tumor microenvironment dynamics and high dimensional datasets.

  • Experience with analytical and modelling skills, including expertise in statistical and machine learning approaches.

  • Experience with the analysis of single cell omics data.

  • Experience in one or more of the following: statistical modelling of functional genomics screening datasets (e.g., bulk CRISPR screens, Perturb-seq) or spatial omics.

  • Experience in Python and deep learning frameworks (PyTorch) for data processing and machine learning model development, with a strong grasp of software engineering fundamentals (e.g., version control, modular design, CI/CD).


Preferred Qualification
If you have the following characteristics, it would be a plus:

  • Experience with multi-modal integration, including spatial transcriptomics/proteomics, histopathology, and single cell omics data.

  • Experience working with longitudinal clinical health record trajectory data.

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

GSK

View company profile →