Jobs and Careers
N-
AI Data Operations/Labeling Engineer
N-Power MedicineUKRemotefull_timeVerifiedPosted 19 Mar 2025
💰 $183,000/yr($155,000/yr – $183,000/yr)
About the role
About N-Power Medicine
N-Power Medicine aims to establish a new paradigm in drug development by reinventing the ‘how’ and transforming clinical trials through better integration with clinical practice, ensuring broader participation by physicians and patients. We are building an exceptional multi-disciplinary team with diverse expertise spanning healthcare, engineering, technology and regulatory, and with people who share our core value of Empowering Community through generosity, curiosity and humility. We are working with urgency to bring better therapies to patients faster.
Position Overview
We are seeking a highly motivated and skilled AI Data Operations/Labeling Engineer to play a critical role in building and optimizing our data labeling pipeline. You will be instrumental in ensuring the creation of high-quality labeled datasets from diverse healthcare sources, including medical notes, radiology reports, pathology reports, and imaging data. This role involves close collaboration with our team of AI scientists, engineers, and teams of subject matter experts to power the development and validation of advanced AI/ML models, especially Large Language Models (LLMs).
This is a hybrid or remote role. Bay Area is preferred but not required.
Role Objectives and Responsibilities
-Design and implement efficient data labeling workflows that integrate with data processing pipelines.-Write appropriate statistical design and analysis plans in the curation of AI labeling dataset to ensure the translation of labeling effort into improvements in model analytical performance or validation.-Select and manage data labeling platforms and tools.-Collaborate with stakeholders, SMEs to craft detailed labeling guidelines.-Work with TPM to recruit, train, and manage human labelers (internal or external).-Track labeler performance and provide feedback.-Implement robust quality control measures to ensure labeling accuracy and consistency.-Develop data pipelines for data preprocessing and labeled data integration, utilizing scalable data processing frameworks.-Manage data storage and versioning for labeled datasets.-Develop and refine tools to automate labeling tasks and improve efficiency.-Integrate labeling platforms with other AI/ML tools.-Create and implement quality assurance procedures for labeling.-Collaborate closely with AI data scientists and engineers to understand labeling requirements.-Understand the principles of AI models that explicitly learn from human feedback and assist humans in evaluating AI model output accuracy.-Applies safeguards and protections in line with HIPAA and applicable privacy laws and adheres to relevant compliance, quality, security, privacy, legal, and ethical standards when it comes to the use of AI.
Education, Experience, Behavioral Competencies, & Skills
-BS/MSc in the field of engineering, computer science, applied mathematics, physics, statistics, or relevant field of study.-5+ years of relevant experience in AI oriented engineering or data science, including proven experience with data labeling platforms and tools.-Strong proficiency in Python and SQL. -Strong project management and organizational skills, with an understanding of data workflows.-Solid understanding of data engineering principles and practices, particularly within distributed data processing environments.-Experience with scalable data processing frameworks.-Ability to see problems and solutions from a ‘holistic’ point of view and communicate how specific solutions bring business value.-Excellent ability to see the big picture while decomposing complex solutions into incremental steps.-Insight into how to create sustainable, reusable, and properly modular code.-Excellent written, verbal, interpersonal, and communication skills.-Generous, Curious and Humble.
Preferred-Experience and/or knowledge of the healthcare data domain in at least one of: electronic medical records (EHR/EMR), clinical imaging, clinical workflows; Electronic Data captures (EDC), and/or clinical research/trials.-Polyglot coder with hands-on development experience across additional modern programming languages (e.g. R, javascript, C++, …).-Familiarity of and documented experience with one or more deep learning frameworks (e.g. PyTorch, Tensorflow, MLflow, etc.). -Hands-on experience with modern, web-based compute - and storage services (e.g., Databricks, AWS, Google Colab, Microsoft Azure, etc.).-Familiarity with the software development life cycle (SDLC).-Experience in creating detailed labeling guidelines and quality assurance processes.-Ability to effectively manage and motivate human labelers.
Travel Requirements
5% Travel may be required
Pay Information
The expected salary range for this position is $155,000 and $183,000. Actual pay will be determined based on experience, qualifications, geographic location, and other
N-Power Medicine aims to establish a new paradigm in drug development by reinventing the ‘how’ and transforming clinical trials through better integration with clinical practice, ensuring broader participation by physicians and patients. We are building an exceptional multi-disciplinary team with diverse expertise spanning healthcare, engineering, technology and regulatory, and with people who share our core value of Empowering Community through generosity, curiosity and humility. We are working with urgency to bring better therapies to patients faster.
Position Overview
We are seeking a highly motivated and skilled AI Data Operations/Labeling Engineer to play a critical role in building and optimizing our data labeling pipeline. You will be instrumental in ensuring the creation of high-quality labeled datasets from diverse healthcare sources, including medical notes, radiology reports, pathology reports, and imaging data. This role involves close collaboration with our team of AI scientists, engineers, and teams of subject matter experts to power the development and validation of advanced AI/ML models, especially Large Language Models (LLMs).
This is a hybrid or remote role. Bay Area is preferred but not required.
Role Objectives and Responsibilities
-Design and implement efficient data labeling workflows that integrate with data processing pipelines.-Write appropriate statistical design and analysis plans in the curation of AI labeling dataset to ensure the translation of labeling effort into improvements in model analytical performance or validation.-Select and manage data labeling platforms and tools.-Collaborate with stakeholders, SMEs to craft detailed labeling guidelines.-Work with TPM to recruit, train, and manage human labelers (internal or external).-Track labeler performance and provide feedback.-Implement robust quality control measures to ensure labeling accuracy and consistency.-Develop data pipelines for data preprocessing and labeled data integration, utilizing scalable data processing frameworks.-Manage data storage and versioning for labeled datasets.-Develop and refine tools to automate labeling tasks and improve efficiency.-Integrate labeling platforms with other AI/ML tools.-Create and implement quality assurance procedures for labeling.-Collaborate closely with AI data scientists and engineers to understand labeling requirements.-Understand the principles of AI models that explicitly learn from human feedback and assist humans in evaluating AI model output accuracy.-Applies safeguards and protections in line with HIPAA and applicable privacy laws and adheres to relevant compliance, quality, security, privacy, legal, and ethical standards when it comes to the use of AI.
Education, Experience, Behavioral Competencies, & Skills
-BS/MSc in the field of engineering, computer science, applied mathematics, physics, statistics, or relevant field of study.-5+ years of relevant experience in AI oriented engineering or data science, including proven experience with data labeling platforms and tools.-Strong proficiency in Python and SQL. -Strong project management and organizational skills, with an understanding of data workflows.-Solid understanding of data engineering principles and practices, particularly within distributed data processing environments.-Experience with scalable data processing frameworks.-Ability to see problems and solutions from a ‘holistic’ point of view and communicate how specific solutions bring business value.-Excellent ability to see the big picture while decomposing complex solutions into incremental steps.-Insight into how to create sustainable, reusable, and properly modular code.-Excellent written, verbal, interpersonal, and communication skills.-Generous, Curious and Humble.
Preferred-Experience and/or knowledge of the healthcare data domain in at least one of: electronic medical records (EHR/EMR), clinical imaging, clinical workflows; Electronic Data captures (EDC), and/or clinical research/trials.-Polyglot coder with hands-on development experience across additional modern programming languages (e.g. R, javascript, C++, …).-Familiarity of and documented experience with one or more deep learning frameworks (e.g. PyTorch, Tensorflow, MLflow, etc.). -Hands-on experience with modern, web-based compute - and storage services (e.g., Databricks, AWS, Google Colab, Microsoft Azure, etc.).-Familiarity with the software development life cycle (SDLC).-Experience in creating detailed labeling guidelines and quality assurance processes.-Ability to effectively manage and motivate human labelers.
Travel Requirements
5% Travel may be required
Pay Information
The expected salary range for this position is $155,000 and $183,000. Actual pay will be determined based on experience, qualifications, geographic location, and other
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s