Jobs and Careers
HE

Senior AI and Machine Learning Engineer

Hewlett Packard Enterprise
All, Texas, United States of America, United Statesfull_timeVerifiedPosted 10 Sept 2024
💰 $295,000/yr($128,000/yr$295,000/yr)

About the role

Senior AI and Machine Learning Engineer

  

This role has been designated as ‘Remote/Teleworker’, which means you will primarily work from home.

Who We Are:

Hewlett Packard Enterprise is the global edge-to-cloud company advancing the way people live and work. We help companies connect, protect, analyze, and act on their data and applications wherever they live, from edge to cloud, so they can turn insights into outcomes at the speed required to thrive in today’s complex world. Our culture thrives on finding new and better ways to accelerate what’s next. We know diverse backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs. We make bold moves, together, and are a force for good. If you are looking to stretch and grow your career our culture will embrace you. Open up opportunities with HPE.

Job Description:

   

Job Description: 

High Performance Computing, AI and Labs is a critical element of HPE. We are focused on delivering innovative solutions that accelerate our customers’ digital transformation, enabling them to tackle their complex, and data-intensive workloads. The next era of computing combines deep learning and machine learning expertise with the development of the world’s most cutting-edge, high-performance supercomputers. Industries are rapidly changing to deliver valuable insight & innovation using ML/DL. Join our team and redefine what’s next for you.

The HPC & AI Performance Engineering team at HPE is building the industry’s highest performing HPC & AI servers and clusters for our customers. We do this by designing, benchmarking, proving, and improving ML/DL application performance on the world’s fastest supercomputers and enabling customers to make quicker and better data-driven decisions.

What you’ll do:

Responsibilities:

  • Studies and improves performance of Large Language Models running on HPE GPU servers
  • Performs system level analysis of HPC & AI workloads on various HPE platforms
  • Runs ML/DL code on accelerated hardware like NVIDIA and AMD GPUs and high-speed networks like InfiniBand
  • Develops software and scripts to automate AI workloads and analyze performance data
  • Installs and configures complex IT infrastructure components (servers, storage, network)
  • Writes white papers and other guidance documents for AI workload and model selection
  • Captures and reviews system performance data, logs, traces to understand workload behavior
  • Communicates technical work well and presents work to non-technical colleagues
  • Works with software and hardware partners in optimizing systems and resolving performance issues
  • Documents and reports issues when testing and evaluating systems
  • Communicates project status and concerns to management in a timely manner
  • Mentors less-experienced staff members

What you need to bring:

Education and Experience Required:

  • Master's degree or PhD in Computer Science, Engineering, Information Technology or Systems, or relevant field.
  • Typically 3+ years of experience.

Knowledge and Skills:

  • 3+ years of experience in Machine Learning/Artificial Intelligence
  • Proficiency in one or more AI & Machine Learning frameworks or libraries (TensorFlow, PyTorch, ONNX, DeepSpeed, Horovod, TensorRT, NeMo)
  • Experience with containers and distributed deep learning and neural networks, including transformers used in generative AI projects
  • Experience with High Performance Computer Servers, High Performance Networking, and associated software
  • Experience with Weka I/O, NTFS and Lustre File Systems
  • Programming experience in Python or C/C++ is strongly desired
  • Strong analytical and critical thinking skills
  • Must be a self-starter, able to work with minimum supervision in a semi-remote setting

Additional Skills:

Artificial Intelligence Technologies and performance benchmarking, Cross Domain Knowledge, Data Engineering, Data Science, Design Thinking, Development Fundamentals, Full Stack Development, IT Performance, Machine Learning Operations, Scalability Testing, Security-First Mindset.

#unitedstates #AIML #frameworks #libraries #TensorFlow #PyTorch, #ONNX #DeepSpeed #Horovod, #TensorRT, #NeMo #hpc #filesystems #python #C #C++ #containers #generativeai

Additional Skills:

Artificial Intelligence Technologies, Cross Domain Knowledge, Data Engineering, Data Science, Design Thinking, Development Fundamentals, Full Stack Development, IT Performance, Machine Learning Operations, Scalability Testing, Security-First Mindset

What We Can Offer You:

Health & Wellbeing

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Hewlett Packard Enterprise

View company profile →