Machine Learning Performance Engineer
Northeastern UniversityAbout the role
About the Opportunity
Job Summary:
As part of the Research Computing (RC) team at Northeastern University (NU), the Machine Learning Performance Engineer (MLPE) supports the incredible growth of RC’s user base, computing infrastructure, and ever-increasing need to support Artificial Intelligence/Deep Learning (AI/DL) workloads. The successful candidate will be a key link between the RC team and the research community at NU, including faculty and students across a broad range of departments, as well as outside users and partners, to help them leverage RC resources for their research and teaching.
The MLPE will help improve the overall reliability and efficiency of GPU-based software applications and optimize the per-watt performance of the ML models on the NU High Performance Computing (HPC) resources. The MLPE will provide key support in the full development cycle of ML models to ensure their optimal performance for faculty research groups in addition to helping increase both the adoption and efficient use of the university’s HPC resources in executing AI/DL workloads. This position will work with faculty and researchers, advise members of the research community on best practices, and assist them in getting the most out of NU’s cloud and HPC service offerings. The MLPE will also participate in and provide assistance with research proposals related to computational science and the use of HPC solutions for a diverse range of scientific research.
As an ML scientist, you will also participate in grant funding opportunities, and author research papers and presentations, with faculty members at Northeastern University. Additionally, you will have the opportunity to design and lead projects, working directly with RC Graduate Research Assistant (GRA) and Co-op student workers.
Qualifications:
- Expert knowledge with an object-oriented language (e.g. C++, Java).
- Proficiency in the areas of GPUs and GPU-based software applications; machine learning, modeling, measurement techniques, testing, and statistical methods; and Python.
- Ability to work with faculty to build technically-focused proposals around HPC solutions.
- Excellent time management skills.
- Ability to manage multiple projects simultaneously, plan and implement project specifications, report project status, and identify delays or resource shortages.
- Ability to communicate with team members effectively and work efficiently with team members to achieve daily, weekly, and monthly objectives.
- Excellent verbal and written communication skills with an ability to communicate solutions by providing both technical and non-technical interpretations of models and results
Knowledge and skills required for this role are typically acquired through a combination of formal education and experience: Master’s degree in a computational science or a related field.
- Minimum of 2-3 years of experience in performance optimization.
- Minimum of 1-2 years of experience in high performance computing.
- Minimum of 1-2 years of research experience in a higher education or government setting.
Additionally, experience should include working with: linking the performance of hardware and software components of scientific applications, with emphasis on GPUs and GPU-based applications; optimization techniques including: SIMD (SSE, AVX), vectorization, loop dependencies, multithreading, multi-processor usage, and tensor cores; diverse communities regarding complex computing requirements and capabilities; batch management systems (e.g. Slurm, PBS, SGE, etc), including cluster configuration and management tools; and leveraging open source or commercial cloud technologies (e.g. Open Science Grid, AWS, Azure, GCP).
Key Responsibilities:
- Partner with faculty and research staff to leverage NU’s HPC cluster at MGHPCC. Work with research groups to help strategize, streamline, and implement optimized ML workflows on the cluster. Participate in the research, deployment, and advertising of new ML technologies. Troubleshoot, isolate, and resolve application errors, and other technical issues. Benchmark application codes and deploy application performance tools. 40%
- Help define and deploy a comprehensive scientific Machine Learning vision for NU researchers by engaging in research and authoring/co-authoring papers and research grants. Assist in developing and writing proposals to enhance the research enterprise at NU. 30%
- Enhance communication with the faculty about technology in support of their research. Forecast and anticipate the need for new HPC technologies and services to improve NU’s standing as a state-of-the-art educational and research environment. 15%
- Ensure the maintenance and/or creation of documentation, training (internal and external), and communica
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s