Jobs and Careers
ME
Software Engineer, Systems ML - HPC
MetaBurlingame, United Statesfull_timeVerifiedPosted 13 Jul 2024
💰 $208,000/yr
About the role
Meta is seeking an AI Software Engineer to join our Research & Development teams. The ideal candidate will have industry experience working on AI Infrastructure related topics. The position will involve taking these skills and applying them to solve for some of the most crucial & exciting problems that exist on the web.
Some aspects of this role as an HPC specialist will include using lower precision numeric formats (fp8) for training. In addition, you would be optimizing training workloads on new hardware platforms (Blackwell) and exploring newer paradigms for distributed training and data loading.
We are hiring in multiple locations.Software Engineer, Systems ML - HPC Responsibilities
Individual compensation is determined by skills, qualifications, experience, and location. Compensation details listed in this posting reflect the base hourly rate, monthly rate, or annual salary only, and do not include bonus, equity or sales incentives, if applicable. In addition to base compensation, Meta offers benefits. Learn more about benefits at Meta.
Some aspects of this role as an HPC specialist will include using lower precision numeric formats (fp8) for training. In addition, you would be optimizing training workloads on new hardware platforms (Blackwell) and exploring newer paradigms for distributed training and data loading.
We are hiring in multiple locations.Software Engineer, Systems ML - HPC Responsibilities
- Apply relevant AI and machine learning techniques to build & optimize our intelligent systems that improve Metas products and experiences
- Enable low precision numerics for efficient training.
- Research and deploy techniques for improving training efficiency like data efficient training and efficient optimizers
- Apply techniques for kernel fusion and kernel generation (torch.compile) to improve and optimize training performance
- Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience.
- 2+ years of experience in HPC and parallel computing.
- Track record of bringing research ideas to production to improve training efficiency
- Expertise in pytorch and cuda
- PhD in Computer Science, Computer Engineering, or relevant technical field.
- Publications on conferences like ML Sys, Neurips on the topic of improving training efficiency
Individual compensation is determined by skills, qualifications, experience, and location. Compensation details listed in this posting reflect the base hourly rate, monthly rate, or annual salary only, and do not include bonus, equity or sales incentives, if applicable. In addition to base compensation, Meta offers benefits. Learn more about benefits at Meta.
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s