Lead ML Platform Engineer: GPU Training on Cloud

JPMorganChase

Palo Alto (CA)

On-site

USD 180,000 - 260,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health care coverage
On-site health centers
Retirement savings plan
Tuition reimbursement
Mental health support

Job summary

JPMorganChase in Palo Alto is seeking a Lead Software Engineer to advance scalable ML training systems and governance across cloud platforms. You will productionize training workloads, optimize performance and cost, and enable repeatable, well-governed training across environments.

You will lead design and implementation of end-to-end training pipelines, work with Kubernetes and AWS infra, and drive observability, reproducibility, and secure development practices across teams.

Qualifications

  • Formal training or certification on software engineering concepts and 5+ years applied experience.
  • Demonstrated experience running ML training in cloud environments and debugging issues across infrastructure & code.
  • Strong Python skills with solid engineering practices (testing, code reviews, modular design, dependency management).
  • Experience building automation/CI for ML codebases (build, test, release, deployment/promotion workflows).
  • Hands on experience with deep learning training workflows and at least one major framework (e.g., PyTorch or TensorFlow).
  • Understanding of training performance and stability: data loading bottlenecks, mixed precision, checkpointing, reproducibility, and evaluation methodology.
  • Experience with distributed training and related concepts (DDP/FSDP/DeepSpeed, etc.).
  • Ability to profile and optimize training systems (CPU/GPU utilization, memory, I/O, networking).
  • Experience with Kubernetes fundamentals for running compute-intensive workloads and AWS (EKS/ECR, S3, IAM, VPC, CloudWatch, EC2).
  • Experience leading AI-assisted development tools and ensuring secure, compliant outputs.

Responsibilities

  • Design, build, and maintain end-to-end ML training platform.
  • Run and optimize GPU training workloads (single-node and distributed).
  • Build and operate training infrastructure on Kubernetes (EKS) including resource management and troubleshooting.
  • Enable Gen AI/LLM training and fine-tuning workflows with governance and scalable GPU execution patterns.
  • Implement observability for training systems: metrics, logs, dashboards, alerting, runbooks.
  • Partner with data engineering and platform teams to define interfaces, standards, and guardrails.
  • Improve developer experience for training: containers, CI/CD, templates, documentation, self-service workflows.
  • Drive adoption of enterprise AI-assisted engineering practices to improve quality and efficiency.

Skills

Python
ML training
Kubernetes
AWS
GPU training
CI/CD
Testing and code reviews
Performance optimization
Debugging infrastructure
Distributed training

Tools

Kubernetes (EKS)
AWS (S3, EC2)
PyTorch
TensorFlow

Job description

JPMorganChase in Palo Alto is seeking a Lead Software Engineer to advance scalable ML training systems and governance across cloud platforms. You will productionize training workloads, optimize performance and cost, and enable repeatable, well-governed training across environments.

You will lead design and implementation of end-to-end training pipelines, work with Kubernetes and AWS infra, and drive observability, reproducibility, and secure development practices across teams.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead ML Training Platform Engineer
Lead ML Training Platform Engineer

JPMorgan Chase & Co. • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Senior ML Platform Engineer - GPU Training on Cloud
Senior ML Platform Engineer - GPU Training on Cloud

J.P. Morgan • New York (NY)

On-site
USD 180,000 - 240,000
Health insurance
Retirement plan
Tuition reimbursement
+1
Senior Lead AI/ML Platform Engineer
Senior Lead AI/ML Platform Engineer

Fairygodboss • Jersey City (NJ)

On-site
USD 180,000 - 240,000
Lead Software Engineer - Machine Learning Platform
Lead Software Engineer - Machine Learning Platform

JPMorgan Chase & Co. • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Lead AI/ML Platform Engineer—Cloud & Infra
Lead AI/ML Platform Engineer—Cloud & Infra

JPMorganChase • Wilmington (DE)

On-site
USD 150,000 - 210,000
Senior Lead AI/ML Platform Engineer
Senior Lead AI/ML Platform Engineer

JPMorgan Chase & Co. • Wilmington (DE)

On-site
USD 180,000 - 250,000
Senior ML Platform Engineer & MLOps Lead
Senior ML Platform Engineer & MLOps Lead

Next Frontier Capital • New York (NY)

On-site
USD 180,000 - 240,000
Health care coverage
On-site health and wellness centers
Retirement savings plan
+4
Senior ML Platform Architect
Senior ML Platform Architect

Fairygodboss • New York (NY)

On-site
USD 180,000 - 250,000
Lead Cloud Platform Engineer for AI & ML Data Platforms
Lead Cloud Platform Engineer for AI & ML Data Platforms

Next Frontier Capital • Jersey City (NJ)

On-site
USD 140,000 - 190,000
Lead Software Engineer - Machine Learning Platform
Lead Software Engineer - Machine Learning Platform

J.P. Morgan • New York (NY)

On-site
USD 180,000 - 240,000
Health insurance
Retirement plan
Tuition reimbursement
+1