Scientific Data Engineer
NovartisAbout the role
Band
Level 3
Ā
Job Description Summary
#LI-HybridLocation: Cambridge, MA
Job Description Summary
The Oncology Data Science (OncDS) team in Biomedical Research provides computational biology, AI, and data expertise and brings together projects spanning the entire pre-clinical to clinical development pipeline across a wide and unique range of novel therapeutics. OncDS specializes in using high-throughput genomic and biomarker data for target identification, drug discovery, and clinical development.
We are seeking a motivated and curious Scientific Data Engineer at the beginning of their professional data journey to support the development, maintenance, and modernization of FAIR oncology reference datasets. This role combines hands-on data curation with workflow automation, AI-enabled process improvement, dataset refresh and quality control, and fit-for-purpose data engineering to help ensure that key oncology data assets are reliable, reproducible, and ready for downstream scientific use.
This position will work closely with senior team members and cross-functional collaborators to modernize data curation processes, identify opportunities for AI-enabled automation, and contribute to high-value reference datasets that support Oncology-wide data strategy and FAIR data goals. The role is designed to build practical experience in scalable data practices, operational excellence, and automation-driven continuous improvement.
Why join us: This is an opportunity to work at the intersection of data engineering, computational biology, and oncology research, contributing to data assets that enable important scientific decisions across the drug discovery pipeline. You will join a collaborative team that values curiosity, learning, and continuous improvement, and you will have the chance to grow your technical and scientific skills while helping build reliable, scalable data foundations for cutting-edge oncology research.
Ā
Job Description
Key responsibilities:
Work with senior team members to identify opportunities for simplification, automation, and continuous improvement in operational data processes, including modernization of data curation workflows
Support the development and improvement of automated workflows for data processing, quality control, recurring dataset updates, and AI-enabled process improvement
Collaborate with cross-functional teams to curate, maintain, and improve key oncology reference datasets, with an emphasis on quality, correctness, reproducibility, and operational reliability
Contribute to reference dataset lifecycle management, including documentation, versioning, and traceable update processes
Assess dataset structure, integrity, and fitness for downstream scientific use, including attention to identifiers, primary keys, missing values, consistency, and related data quality issues, and help resolve data issues in collaboration with relevant partners
Contribute to best practices for reproducible workflows, data handling, and scalable dataset operations
Help keep oncology reference datasets current, usable, and responsive to evolving scientific and organizational priorities
Essential Requirements:
Degree in bioinformatics, computational biology, data science, computer science, or a related field, or equivalent relevant experience
Experience working with biological, experimental, or omics datasets in a research environment
Proficiency in one or more programming languages used in data science or bioinformatics, such as Python or R
Experience with data wrangling, data quality control, and reproducible analysis workflows
Familiarity with workflow automation and scripting for recurring data processing tasks
Familiarity with database principles and sound data modeling practices, including normalization, primary keys, joins, and thoughtful handling of missing values and data integrity issues
Knowledge of version control and Unix / Linux-based working environments
Strong organizational skills and attention to detail, with a demonstrated interest in operational excellence
Ability to learn quickly, manage multiple tasks, and work effectively in a collaborative environment
Excellent written and verbal communication skills, and the ability to work effectively across technical and scientific teams
Desirable Requirements:
Experience with oncology datasets, translational research data, or biomedical reference data resources
Familiarity with FAIR data principles, metadata curation, or scientific data stewardship
Experience maintaining r
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights ā in under 60 seconds.
Apply Now āGenerate Application KitFree account required ā sign up in 30s