Jobs and Careers
MS

Data Scientist and AI/ML Engineer – Generative AI and Natural Language Processing (Hybrid - NJ or MA)

MSD
United StatesRemotefull_timeVerifiedPosted 27 Jun 2025
💰 $163,900/yr($104,200/yr$163,900/yr)

About the role

Job Description

The Data Scientist and AI/ML Engineer – Generative AI and Natural Language Processing role involves helping to develop and deploy production-grade NLP products for unstructured and semi-structured data from across our company’s research and development pipeline. These models and workflows will help solve real-world problems and contribute to Artificial Intelligence and Machine Learning (AI/ML) in therapeutic research and development. Key focus areas will include the scalable deployment of ML and Generative AI approaches (such as Large Language Models, or LLMs) for surfacing insights from proprietary unstructured research data and biomedical literature, as well as the integration of structured information from the likes of knowledge graphs. The position is embedded in a cross-disciplinary team of data scientists, bioinformaticians, and engineers that are all focused on using cutting-edge software, AI/ML, and data science techniques to drive drug discovery and development.  

 

You enjoy:  

  • Building novel NLP/AI-enriched software that enables the discovery, development, and delivery of new therapeutics to patients in need  

  • Understanding real-world challenges and developing automated data solutions for them  

  • Opportunities to directly interact with users and stakeholders of your data science, ML, and AI products  

  • Evaluating, developing, testing, and deploying new techniques for natural language understanding and new DevOps and ML/LLMOps frameworks. 

  • Freedom to propose projects that interest you and to collaborate cross-functionally on delivery  

  • Staying updated on the newest methods in NLP, ML, generative AI, and ML/LLMOps 

  • Sharing the approaches you implement and their impact with internal company audiences and externally  

 

You have: 

The following are preferred skills and experience, not strict requirements 

  • Fluency in Python programming, version control and collaboration with git, environment management (e.g., poetry, conda, docker), standard Python packages for data exploration (e.g., pandas, numpy, matplotlib) 

  • Fluency with data science and NLP approaches such as exploratory data analysis, performance metrics and benchmarks, supervised and unsupervised learning, transformers, and LLMs.  

  • Fluency with standard cloud and DevOps tools, such as Infrastructure as Code (IaC) and Github Actions. 

  • Experience with at least one ML framework (e.g., pytorch, tensorflow, fairseq) and with ML model deployment and operations (MLOps/LLMOps) 

  • Experience with scalable data engineering frameworks such as Apache Spark and orchestration frameworks such as Airflow, semantic search and retrieval frameworks (e.g., development and benchmarking of embedding models and retrieval approaches in the context of Retrieval Augmented Generation, RAG), and/or semantic knowledge frameworks (e.g. RDF triplestores, property graphs, ontology management).  

  • Experience with standard operations on non-relational (e.g.,

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

MSD

View company profile →