Jobs and Careers
JP

Software Engineer III - Machine Learning Platform

JPMorgan Chase & Co.
Palo Alto, United Statesfull_timeVerifiedPosted 20 Aug 2026

About the role

We have an exciting and rewarding opportunity for you to take your software engineering career to the next level.

  
As a Software Engineer III at JPMorganChase within the AI/ML data platform team you serve as a seasoned member of an agile team to build and operate scalable, reliable ML training systems and pipelines on AWS and other cloud platforms. You will productionize training workloads (often GPU-based), improve performance and cost efficiency, and enable repeatable, well-governed training across environments

 

Job Responsibilities

 

  • Design, build, and maintain end-to-end ML training platform.

  • Run and optimize GPU training workloads (single-node and distributed), improving throughput, utilization and reproducibility.

  • Build and operate training infrastructure on Kubernetes (e.g., EKS and other manage Kubernetes platforms), including resource management and workload troubleshooting.

  • Enable Gen AI/LLM training and fine-tuning workflows (e.g., supervised fine-tuning), including evaluation harnesses, artifact/version governance, and scalable GPU execution patterns aligned to enterprise controls.

  • Implement observability for training systems: metrics, logs, dashboards, alerting, and operational runbooks.

  • Partner with data engineering and platform teams to define interfaces, standards, and guardrails (security, access, cost controls)

  • Improve developer experience for training: standardized containers, CI/CD, templates, documentation, and self-service workflow

  • Leverages enterprise-authorized AI coding assist tools within the work environment to improve code quality, delivery speed, and productivity across complex deliverables (e.g., code generation/refactoring, unit test creation, documentation), while validating outputs through peer review, automated testing, and secure coding standards; contributes learnings and reusable patterns to improve broader team effectiveness. 

  • Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation. 

 

Required qualifications, capabilities, and skills

 

  • Formal training or certification on software engineering concepts and 3+ years applied experience 

  • Demonstrated experience running ML training in cloud environments and debugging issues across infrastructure & code.

  • Strong Python skills with solid engineering practices (testing, code reviews, modular design, dependency management).

  • Experience building automation/CI for ML codebases (build, test, release, deployment/promotion workflows).

  • Hands on experience with deep learning training workflows and at least one major framework (eg., PyTorch or TensorFlow).

  • Understanding of training performance and stability: data loading bottlenecks, mixed precision, checkpointing, reproducibility, and evaluation methodology.

  • Experience with distributed training and related concepts (e.g., DDP/FSDP/DeepSpeed concepts, collective communication basics, scaling and bottleneck analysis).

  • Ability to profile and optimize training systems (CPU/GPU utilization, memory, I/O throughput, networking, scheduling).

  • Experience with Kubernetes fundamentals for running compute-intensive workloads and AWS (eg., EKS/ECR, S3, IAM, VPC/networking, Cloudwatch, EC2)

  • Hands-on experience using enterprise-authorized AI-assisted software development tools within the work environment (e.g., for coding, test creation, troubleshooting, or documentation) with demonstrated ability to critically evaluate, validate, and refine AI-generated outputs for correctness, performance, and security. 

  • Understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; ability to guide peers on safe and effective usage within team practices. 

 

Preferred qualifications, capabilities, and skills 

 

  • Experience running training workloads across multiple cloud platforms and managing portability, performance, and governance across environments.

  • Familiarity with cloud-native networking/storage patterns for high-throughput training and artifact management.

  • Experience optimizing training input pipelines (sharding, prefetching, caching, format choices such as Parquet/WebDataset) and working with large datasets.

  • Familiarity with distributed compute frameworks (Spark, Ray, Dask) for feature/dataset generation.

  • Familiarity with

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

JPMorgan Chase & Co.

View company profile →