Staff Engineer, AI Infrastructure Quality and Performance - Lead
Samsung Semiconductor, Inc.About the role
Please Note:
To provide the best candidate experience amidst our high application volumes, each candidate is limited to 10 applications across all open jobs within a 6-month period.
Advancing the World’s Technology Together
Our technology solutions power the tools you use every day--including smartphones, electric vehicles, hyperscale data centers, IoT devices, and so much more. Here, you’ll have an opportunity to be part of a global leader whose innovative designs are pushing the boundaries of what’s possible and powering the future.
We believe innovation and growth are driven by an inclusive culture and a diverse workforce. We’re dedicated to empowering people to be their true selves. Together, we’re building a better tomorrow for our employees, customers, partners, and communities.
The Data Fabric Solutions (weblink) is part of Memory Solution Lab in Samsung Semiconductors, the industry's technology and volume leader in Storage and Memory devices. The Data Fabric Solutions (DFS) group pioneers and creates comprehensive software solutions that optimize data flow efficiency, transformation, and performance for client’s data-intensive applications, leveraging Samsung’s world-leading SSDs, Memory, and Accelerators. As an integral part of Samsung’s strong R&D focus & lab innovation engine, they cater to diverse data fabric deployments.
Specifically, team has opening for seasoned AI Infrastructure Quality & Performance Lead Engineer/Manager, who has experience in managing/leading complex SaaS (Software-as-a-service) product testing to meet quality & performance standards for innovative data fabric technology and solutions for the Cloud or Enterprise markets. He or she will join a team of experts in researching and devising new innovative infrastructure and test tools/solutions that validates performance of existing and emerging AI technologies to add substantial value to Samsung.
The team covers a broad spectrum of topics and is building an ecosystem surrounding HBM, SSD and Non-volatile memory Software. You will be working with the-state-of-art technologies in the context of vertical integration and optimization across H/W, S/W and cloud applications. You will be leading the software products leveraging latest technologies including HBM (High Bandwidth Memory), CXL (Compute Express Link), Flash/SSD, GPUs, high speed networking and others to research/build open-source software framework/eco-systems. In addition, you will be working closely with internal cross-functional teams, partners and customers to drive the products to maturity.
Location: Daily onsite presence at our San Jose, CA office / U.S. headquarters in alignment with our Flexible Work policy.
What You’ll Do
· Lead overall QA testing, performance benchmarking efforts with writing new test plans.
· Hands-on experience in software verification/validation, performance benchmarking of distributed software w/performance tools.
· Develop new test cases and test plans that can exercise the functionality & performance of distributed system software stack
· Expertise in benchmarking and profiling infrastructure components (e.g., CPU, GPU, memory, storage, network).
Develop test automation software/framework using Python, bash and other popular scripting languages
What You Bring
- · A PhD/Master’s in Computer Science, Computer Engineering, or related field.
- · 10+ years of experience of system Testing, Test Plan, Test case writing, software automation, scripting, test development, & debugging
- · Strong software engineering skills with writing efficient, maintainable and testable Python/Shell scripting is required.
- · Understanding of AI/ML workflows, including data preprocessing, model training, and inference pipelines.
- · Familiarity with AI frameworks (TensorFlow, PyTorch, ONNX) to test integration with infrastructure.
- · Deep knowledge of GPU/TPU environments, including CUDA, cuDNN, or other accelerators, and their performance characteristics.
- · Experience with distributed computing frameworks (e.g., Ray, Spark, Dask) for testing large-scale AI workloads
- · Knowledge of latency, throughput, and scalability testing for AI systems.
- · Ability to optimize infrastructure for high-performance AI workloads, including resource allocation and tuning.
- · Experience in system testing, test development and test methodologies. <
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s
Similar roles
Support Staff Substitute Paraprofessional/Secretary/Health Assistant (pool open for the 2026-27 school year)
Mesa County Valley School District 51
Seasonal Park Maintenance Staff - Southern Region Parks Division
The Maryland-National Capital Park and Planning Commission
$39,460/yr
Staff Software Engineer, Gen AI
ServiceNow
$286,300/yr