Senior Solution Engineer - AI and ML Storage Architecture
NVIDIAAbout the role
As a Senior Solution Engineer specializing in AI/ML Storage Architecture, you will be an integral part of our dynamic team, contributing to the design, construction, and maintenance of innovative storage solutions tailored for Artificial Intelligence and Machine Learning workloads. This role spans various domains, including software and systems engineering practices, storage, data management, and services. Their responsibilities encompass ensuring reliable storage solutions, managing data efficiently, and providing related services to support the overall stability and performance of the production systems.
Solution Engineer specializing in AI/ML at NVIDIA, your role involves ensuring the reliability and uptime of both our internal and external GPU cloud services, aligning with our commitments to users. Simultaneously, you empower developers to implement system changes through meticulous preparation and planning, with a keen focus on aspects like capacity, latency, and performance. This position embodies a specific attitude and a suite of engineering strategies aimed at enhancing the efficiency of production systems and implementing optimizations. A significant portion of our software development efforts concentrates on automating tasks, fine-tuning performance, and enhancing overall production system efficiency. Given the comprehensive responsibility for understanding how our systems interconnect, you will use a diverse range of tools and approaches to address a wide array of challenges. This role offers engaging and dynamic day-to-day work, emphasizing continual enhancement and ensuring the success of our AI/ML solutions. Solution Engineer's culture of diversity, intellectual curiosity, problem-solving, and openness is important to its success. Our organization brings together people with a wide variety of backgrounds, experiences, and perspectives. We encourage them to collaborate, think big, and take risks in a blame-free environment. We promote self-direction to work on meaningful projects while striving to build an environment that provides the support and mentorship needed to learn and grow.
What You Will Be Doing:
Solution Design: Collaborating with multi-functional teams to design and implement storage architectures optimized for AI/ML workloads, ensuring scalability, performance, and reliability.
Technology Expertise: Apply in-depth knowledge of storage technologies, encompassing Lustre and Cloud storage, to devise and implement solutions tailored to the distinctive requirements of AI/ML training and inference workloads. Stay current with industry trends and advancements in storage technologies to consistently elevate the company's capabilities.
Cloud Infrastructure Integration: Apply proficiency in GCP, AWS, and Azure to incorporate storage solutions emphasizing on efficiency, reliability, and cost-effectiveness.
Demonstrate practical experience in constructing and deploying storage solutions, assuming responsibility for the entire process and troubleshooting as needed.
Technical Consultation: Providing technical expertise and consultation to customers and internal teams on storage solutions, aligning with AI/ML standard methodologies.
Performance Optimization: Analyzing and optimizing storage systems for AI/ML applications to meet performance requirements and enhance overall system efficiency.
Integration: Integrating storage solutions seamlessly with AI/ML frameworks, ensuring compatibility and improving the utilization of storage resources.
Collaboration: Working closely with data scientists, engineers, and stakeholders to understand AI/ML storage requirements and proposing tailored solutions.
Documentation: Creating comprehensive technical documentation for AI/ML storage architectures, guidelines, and best practices.
Emerging Technologies: Staying abreast of industry trends and emerging technologies in AI/ML and storage to drive innovation and continuous improvement.
What We Need To See:
Proven experience in designing and implementing storage architectures for AI/ML workloads, with a focus on scalability and performance.
Strong technical expertise in storage technologies, AI/ML frameworks, and their integration. Proficiency in programming languages commonly used in AI/ML, such as Python, is desirable.
Excellent communication and interpersonal skills to effectively convey complex technical concepts to both technical and non-technical collaborators.
Confirmed ability to analyze and solve complex technical challenges related to AI/ML storage architectures.
A collaborative approach with the ability to work effectively in multi-functional teams and engage with clients to understand their specific AI/ML storage requirements.
<
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s