Lead ML Platform Engineer — GPU Training on Cloud

JPMorgan Chase & Co.

Palo Alto (CA)

On-site

USD 180,000 - 270,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

JPMorganChase in Palo Alto is seeking a Lead Software Engineer to drive the design, build, and operation of scalable ML training systems on AWS and other clouds. You will optimize GPU workloads, enable Gen AI/LLM training, and establish governance and observability across environments.

You will partner with data engineering and platform teams to define interfaces, ensure security and cost controls, and improve developer experience with containers, CI/CD, and self-service workflows.

Qualifications

  • Formal training or certification on software engineering concepts and 5+ years applied experience.
  • Demonstrated experience running ML training in cloud environments and debugging issues across infrastructure & code.
  • Strong Python skills with solid engineering practices (testing, code reviews, modular design, dependency management).
  • Experience building automation/CI for ML codebases (build, test, release, deployment/promotion workflows).
  • Hands on experience with deep learning training workflows and at least one major framework (eg., PyTorch or TensorFlow).
  • Understanding of training performance and stability: data loading bottlenecks, mixed precision, checkpointing, reproducibility, and evaluation methodology.
  • Experience with distributed training and related concepts (e.g., DDP/FSDP/DeepSpeed concepts, collective communication basics, scaling and bottleneck analysis).
  • Ability to profile and optimize training systems (CPU/GPU utilization, memory, I/O throughput, networking, scheduling).
  • Experience with Kubernetes fundamentals for running compute-intensive workloads and AWS (eg., EKS/ECR, S3, IAM, VPC/networking, Cloudwatch, EC2)
  • Demonstrated experience leading effective use of approved AI-assisted software development tools (e.g., for coding, code review, test acceleration, troubleshooting) with the ability to set team expectations for validating AI outputs for correctness, performance, and security.
  • Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; experience coaching engineers on safe, compliant adoption within delivery practices.

Responsibilities

  • Design, build, and maintain end-to-end ML training platform.
  • Run and optimize GPU training workloads (single-node and distributed), improving throughput, utilization and reproducibility.
  • Build and operate training infrastructure on Kubernetes (e.g., EKS and other manage Kubernetes platforms), including resource management and workload troubleshooting.
  • Enable Gen AI/LLM training and fine-tuning workflows (e.g., supervised fine-tuning), including evaluation harnesses, artifact/version governance, and scalable GPU execution patterns aligned to enterprise controls.
  • Implement observability for training systems: metrics, logs, dashboards, alerting, and operational runbooks.
  • Partner with data engineering and platform teams to define interfaces, standards, and guardrails (security, access, cost controls)
  • Improve developer experience for training: standardized containers, CI/CD, templates, documentation, and self-service workflow
  • Drives team adoption of enterprise-authorized AI-assisted engineering practices within the work environment to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test strategy acceleration, incident/root-cause analysis support), while establishing consistent validation standards (secure coding, peer review, automated testing) and promoting reuse of effective patterns across the team.
  • Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation.

Skills

Python
ML CI/CD
PyTorch
TensorFlow
Kubernetes
AWS
Distributed training
GPU training
Profiling
Observability
Security & compliance

Tools

Kubernetes (EKS)
S3
IAM
CloudWatch
EC2

Job description

JPMorganChase in Palo Alto is seeking a Lead Software Engineer to drive the design, build, and operation of scalable ML training systems on AWS and other clouds. You will optimize GPU workloads, enable Gen AI/LLM training, and establish governance and observability across environments.

You will partner with data engineering and platform teams to define interfaces, ensure security and cost controls, and improve developer experience with containers, CI/CD, and self-service workflows.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead ML Platform Engineer: GPU Training on Cloud
Lead ML Platform Engineer: GPU Training on Cloud

JPMorganChase • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Health care coverage
On-site health centers
Retirement savings plan
+2
Senior ML Platform Engineer - GPU Training on Cloud
Senior ML Platform Engineer - GPU Training on Cloud

J.P. Morgan • New York (NY)

On-site
USD 180,000 - 240,000
Health insurance
Retirement plan
Tuition reimbursement
+1
Senior ML Platform Engineer: GPU Training & Cloud
Senior ML Platform Engineer: GPU Training & Cloud

JPMorganChase • Palo Alto (CA)

On-site
USD 160,000 - 210,000
Senior ML Platform Engineer - GPU & Cloud
Senior ML Platform Engineer - GPU & Cloud

Socket.dev • Palo Alto (CA)

On-site
USD 150,000 - 190,000
Senior ML Platform Engineer — GPU Training on Cloud
Senior ML Platform Engineer — GPU Training on Cloud

J.P. Morgan • New York (NY)

On-site
USD 140,000 - 200,000
Senior ML Platform Engineer - GPU Training Systems
Senior ML Platform Engineer - GPU Training Systems

JPMorgan Chase & Co. • Palo Alto (CA)

On-site
USD 150,000 - 210,000
Senior ML Platform Engineer (GPU/Cloud)
Senior ML Platform Engineer (GPU/Cloud)

Fairygodboss • Palo Alto (CA)

On-site
USD 180,000 - 250,000
Senior Lead AI/ML Platform Engineer
Senior Lead AI/ML Platform Engineer

Fairygodboss • Jersey City (NJ)

On-site
USD 180,000 - 240,000
Lead Software Engineer - Machine Learning Platform
Lead Software Engineer - Machine Learning Platform

JPMorgan Chase & Co. • Palo Alto (CA)

On-site
USD 180,000 - 270,000
Lead AI/ML Platform Engineer, Cloud & Automation
Lead AI/ML Platform Engineer, Cloud & Automation

JPMorgan Chase & Co. • Wilmington (DE)

On-site
USD 170,000 - 210,000