Jobs and Careers
AP
Senior Data Engineer – ADMET
ApherisRemote (UTC +/- 2 hrs)Remotefull_timeVerifiedPosted 11 Apr 2025
About the role
About the role
At Apheris, we power federated data networks in life sciences to address the data bottleneck in training highly performant ML models. Publicly available, molecular datasets are insufficient to train high-quality ML models that meet industry requirements. Our product addresses this by hosting networks where biopharma organizations collaboratively train higher quality models on their combined data. The Apheris product is a set of drug discovery applications - enriched with the proprietary data of network participants. Our federated computing infrastructure with built-in governance and privacy controls ensure that the data IP and ownership always stays with the data custodians.As we are doubling down on ADMET (absorption, distribution, metabolism, excretion, and toxicity) use cases as a focus area within our drug discovery work, we are looking for a Senior Data Engineer to help us build great ADMET models. This is a hands-on, high-impact role focused on advancing the state of the art in applying foundational models to drug discovery problems. You’ll work closely with our ADMET team and will serve as the technical authority on data preparation, data harmonization, and data pipelines in this domain.
You should bring deep expertise in data infrastructure and data preparation with domain knowledge in pharmacokinetics and toxicity with a focus on ADMET modelling and related tasks. You must also understand the application of these models within industrial drug discovery workflows.
If you want to be part of a mission-driven team building cutting-edge AI systems for life sciences – and you know what it takes to leverage domain-specific data – this role is for you.
What you will do
- ADMET Data Pipeline Development: Design, build and maintain scalable pipelines for ingesting, processing, and harmonizing diverse ADMET datasets from public sources (e.g., ChEMBL, PubChem) and proprietary assays.
- Data Harmonization: Standardize heterogeneous ADMET data formats (e.g., in vitro assays, in silico predictions) across network participants to enable modelling readiness of the data
- Model-Ready Dataset Curation: Preprocess raw ADMET data (e.g., normalizing units, handling missing values) to support AI/ML model training for a variety of endpoints (like bioavailability, hERG inhibition, or CYP450 interactions)
- Data Quality Assurance: Implement and automate validation checks to ensure ADMET data integrity
- Cross-Functional Integration: Work with computational chemists to optimize data structures for AI-driven ADMET models (e.g., graph-based representations for metabolic pathways)
- Work with our customers and potentially academic partners to define data preprocessing, selection, and benchmarking strategies for novel training tasks involving ADMET data, including leveraging and harmonizing assay data from different sources.
- Collaborate cross-functionally to ensure data and resulting models address real-world drug discovery needs.
- Mentor and guide team members on a content level, supporting the planning and breakdown of complex ADMET data preparation.
- Influence strategic decisions on data infrastructure and data quality assurance
- Contribute to publications or open-source contributions where relevant.
What we expect from you
- By month 3: Develop a deep technical understanding of the Apheris product and how it maps to the current ADMET use-cases we are working on. Take ownership of an ADMET data preparation stream.
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s