Data Engineer II
Howard Hughes Medical InstituteAbout the role
Current HHMI Employees, click here to apply via your Workday account.
The Howard Hughes Medical Institute’s Janelia Research Campus is a pioneering research center in Ashburn, Virginia, where scientists pursue fundamental questions in the life sciences. Our integrated teams of biologists, computational scientists, and tool-builders innovate research practices and technologies to solve biology’s deepest mysteries. HHMI launched Janelia in 2006, establishing an intellectually enriching environment for scientists to do creative, collaborative, hands-on work. We share our methods, results, and tools with the scientific community.
Summary:
AI@HHMI: HHMI is investing $500 million over the next 10 years to support AI-driven projects and to embed AI systems throughout every stage of the scientific process in labs across HHMI. The AI initiative will be centered at HHMI’s Janelia Research Campus. Janelia has been at the forefront of AI-driven research in biology for more than 15 years. Its forward-thinking structure, centralized funding, and collaborative culture make it ideally suited to take this bold leap forward. To learn more about the initiative, visit here.
About the role:
We're seeking a skilled Data Engineer to drive scientific innovation through robust data infrastructure. In this role, you’ll design, develop, and optimize scalable data pipelines and tools for the ingestion, transformation, and integration of large, heterogeneous datasets. Your work will directly support computational research initiatives, including machine learning and AI applications. Collaborating closely with multidisciplinary teams of computational and experimental scientists, you’ll help define and implement best practices in data engineering, ensuring data quality, accessibility, and reproducibility. You’ll also be responsible for maintaining detailed documentation, mentoring junior engineers, and automating workflows to streamline the path from raw data to scientific insight.
What we provide:
A competitive compensation package, with comprehensive health and welfare benefits.
A supportive team environment that promotes collaboration and knowledge sharing.
The opportunity to engage with world-class researchers, software engineers and AI/ML experts, contribute to impactful science, and be part of a dynamic community committed to advancing humanity’s understanding of fundamental scientific questions.
Amenities that enhance work-life balance such as on-site childcare, free gyms, available on-campus housing, social and dining spaces, and convenient shuttle bus service to Janelia from the Washington D.C. metro area.
What you’ll do:
Design and implement scalable data solutions, including mining, transformation, storage optimization, and orchestration of structured and unstructured datasets to support AI and computational research.
Apply statistical tools and programming languages (e.g., Python, R) to analyze large datasets, develop custom functions, and extract actionable insights through effective visualization.
Establish and maintain data standards, formats, workflows, and documentation to ensure data quality, accessibility, and reproducibility across projects.
Collaborate closely with interdisciplinary teams, mentor junior engineers, and advise stakeholders on data strategies, source identification, and best practices.
Continuously explore emerging tools and technologies to enhance data infrastructure and analytical capabilities in support of evolving research needs.
What you bring:
A Bachelor’s degree in Computer Science, Data Science, Statistics, Applied Mathematics or related fields with 3 to 5 years of relevant experience. An equivalent combination of education and relevant experience will be considered.
Advanced proficiency in the use of the Linux command line, programming languages and frameworks, and formats for data management (e.g. Python, R, Numpy, Pandas, HDF5).
Advanced proficiency in the application of data mining and data analysis methods and techniques.
Advanced proficiency in utilizing data visualization software (e.g., Matplotlib,R, Jupyter notebooks Plotly, Domo, Looker, Tableau).
Understanding of high-performance computing environments and cloud storage (e.g., AWS, GoogleCloud).
Familiarity with machine learning frameworks (e.g., PyTorch,
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s