Jobs and Careers
MS

Assoc Dir , Representation Learning, Data Science

MSD
United StatesRemotefull_timeVerifiedPosted 24 Jun 2025
💰 $242,200/yr($153,800/yr$242,200/yr)

About the role

Job Description

The Data Science and Scientific Informatics Team at Research and Development Sciences IT (RaDS-IT) of our Company is seeking a Lead Data Scientist for Representation Learning.

Our team is a diverse collection of scientists and engineers working towards the same goal – enabling and accelerating the next generation of pharmaceutical sciences.  We collaborate closely with laboratory and in silico scientists, proposing and implementing innovative solutions that enable new organizational capabilities.

An ideal candidate will have a strong background in computational biology, machine learning, and chemistry, with a focus on developing and applying advanced methods for molecular and protein design. This role will involve creating and optimizing foundation models that support the design and evaluation of novel therapeutic candidates, in close collaboration with RaDS-IT product lines supporting the corresponding functionalities under Discovery Chemistry, Discovery Biologics, and IDVAX.

Key Responsibilities:

Molecular Representation Learning:

  • Develop, validate, and implement state-of-the-art machine learning and deep learning algorithms for molecular representation, focusing on capturing complex chemical properties and biological activities.

  • Utilize various techniques, including graph neural networks and transformer architectures, to enhance molecular and protein representations.

  • Collaborate with cross-functional teams to contribute in the design of novel small molecules and protein constructs tailored to specific therapeutic targets.

Protein Design:

  • Apply computational tools and methodologies for de novo protein design and engineering such as RF-Diffusion, ProteinMPNN and AlphaFold, using AI-driven approaches to predict protein stability, function, and interaction.

  • Oversee the integration of structural biology data into machine learning models to improve predictive capabilities.

Foundation Models:

  • Lead initiatives to develop foundation models that enable scalable and efficient molecular and protein design workflows.

  • Conduct research on transfer learning and few-shot learning to maximize model performance on diverse datasets.

Data Management and Collaboration:

  • Manage and curate large-scale datasets relevant to molecular and protein design, ensuring data integrity and accessibility for team members.

  • Collaborate closely with experimental chemists, biologists, data scientists, and other product teams on RaDS-IT to translate computational insights into practical a

Mentorship and Leadership:

  • Provide mentorship to junior scientists and researchers on the team, fostering an environment of creativity and scientific rigor.

  • Contribute to strategic planning and project prioritization within the team and at the higher level of the organization.

Publications and Presentations:

  • Lead efforts in publishing research findings in peer-reviewed journals and presenting at conferences.

  • Stay abreast of advancements in molecular representation learning and related fields to inform ongoing research and development.

Required Skills:

  • Ph.D. in Computational Biology, Bioinformatics, Chemistry, Machine Learning, or a related field, with 3+ years of experience in industry (including full time job and internship/co-op)

  • Proven experience in molecular and/or protein design, with a strong publication record in relevant areas.

  • Proficient in programming languages such as Python and R, and familiarity with ML frameworks (e.g., TensorFlow, PyTorch).

  • Strong understanding of molecular modeling software and tools (e.g., RDKit, OpenMM, AlphaFold, RosettaFold, MPNN, RF-Diffusion).

  • Excellent communication skills and ability to work collaboratively in a multidisciplinary team.

  • Deep knowledge of statistical methods and experimental design as applied to computational biology.

  • Experience with large-scale datasets and big data analytics techniques.

Optional Skills:

  • Experience with cloud computing platforms (e.g., AWS, Google Cloud) for computational modeling and data analysis.

  • Familiarity with cheminformatics and bioinformatics databases and tools (e.g., ChEMBL, UniProt).

  • Knowledge of synthetic chemistry or organic chemistry principles.

  • Experience in project management and leading cross-functional research initiatives.

  • Understan

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

MSD

View company profile →