Lead ML Platform Engineer for Scalable GPU Training

Next Frontier Capital

Palo Alto (CA)

On-site

USD 180,000 - 245,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

JPMorganChase in Palo Alto seeks a Lead Software Engineer to architect and operate scalable ML training platforms on AWS and other clouds. You will optimize GPU workloads, enable Gen AI workflows, and enforce governance with observability and automation across environments.

The role emphasizes enterprise-grade security, reproducibility, and efficient collaboration with data engineering and platform teams. A strong Linux+Python background and experience in cloud ML are essential.

Qualifications

  • Formal training or certification on software engineering concepts and 5+ years applied experience.
  • Demonstrated experience running ML training in cloud environments and debugging issues across infrastructure & code.
  • Strong Python skills with solid engineering practices (testing, code reviews, modular design, dependency management).
  • Experience building automation/CI for ML codebases (build, test, release, deployment/promotion workflows).
  • Hands on experience with deep learning training workflows and at least one major framework (e.g., PyTorch or TensorFlow).
  • Understanding of training performance and stability: data loading bottlenecks, mixed precision, checkpointing, reproducibility, and evaluation methodology.
  • Experience with distributed training and related concepts (e.g., DDP/FSDP/DeepSpeed).
  • Ability to profile and optimize training systems (CPU/GPU utilization, memory, I/O throughput, networking, scheduling).
  • Experience with Kubernetes fundamentals for running compute-intensive workloads and AWS (EKS/ECR, S3, IAM, VPC...).
  • Demonstrated experience leading effective use of approved AI-assisted software development tools with CI/CD and code reviews.
  • Strong understanding of responsible AI use in engineering workflows, including data sensitivity and secure handling.

Responsibilities

  • Design, build, and maintain end-to-end ML training platform.
  • Run and optimize GPU training workloads, improving throughput and reproducibility.
  • Build and operate training infrastructure on Kubernetes and cloud platforms.
  • Enable Gen AI/LLM training and fine-tuning workflows with governance and scalable GPU execution.
  • Implement observability: metrics, logs, dashboards, alerting, and runbooks.
  • Partner with data engineering and platform teams to define interfaces and guardrails.
  • Improve developer experience: containers, CI/CD, templates, documentation, self-service workflows.
  • Drive adoption of enterprise AI-assisted engineering practices for code quality and delivery speed.
  • Apply SDLC toolchain knowledge to improve value from automation.

Skills

Python
ML in cloud
CI/CD for ML
DL frameworks (PyTorch/TensorFlow)
Distributed training
Kubernetes
AWS
Profiling & optimization
AI-assisted tooling
Responsible AI

Tools

Kubernetes
AWS
EKS
S3
IAM

Job description

JPMorganChase in Palo Alto seeks a Lead Software Engineer to architect and operate scalable ML training platforms on AWS and other clouds. You will optimize GPU workloads, enable Gen AI workflows, and enforce governance with observability and automation across environments.

The role emphasizes enterprise-grade security, reproducibility, and efficient collaboration with data engineering and platform teams. A strong Linux+Python background and experience in cloud ML are essential.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead ML Engineer – Scalable GPU Training on Cloud
Lead ML Engineer – Scalable GPU Training on Cloud

JPMorganChase • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Lead ML Platform Engineer — GPU Training on Cloud
Lead ML Platform Engineer — GPU Training on Cloud

JPMorgan Chase & Co. • Palo Alto (CA)

On-site
USD 180,000 - 270,000
Lead ML Platform Engineer - GPU & Cloud
Lead ML Platform Engineer - GPU & Cloud

JPMorgan Chase • Palo Alto (CA)

On-site
USD 157,000 - 215,000
Senior ML Platform Engineer: GPU Training & Cloud Pipelines
Senior ML Platform Engineer: GPU Training & Cloud Pipelines

Next Frontier Capital • Palo Alto (CA)

On-site
USD 170,000 - 250,000
Senior ML Platform Engineer: GPU Training & Cloud
Senior ML Platform Engineer: GPU Training & Cloud

JPMorganChase • Palo Alto (CA)

On-site
USD 160,000 - 210,000
Senior ML Platform Engineer: GPU Training & Cloud Infra
Senior ML Platform Engineer: GPU Training & Cloud Infra

JPMorgan Chase & Co. • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Senior ML Platform Engineer - GPU Training on Cloud
Senior ML Platform Engineer - GPU Training on Cloud

J.P. Morgan • New York (NY)

On-site
USD 180,000 - 240,000
Health insurance
Retirement plan
Tuition reimbursement
+1
Senior ML Platform Engineer (GPU/Cloud)
Senior ML Platform Engineer (GPU/Cloud)

Fairygodboss • Palo Alto (CA)

On-site
USD 180,000 - 250,000
Senior ML Platform Engineer - GPU Training Systems
Senior ML Platform Engineer - GPU Training Systems

JPMorgan Chase & Co. • Palo Alto (CA)

On-site
USD 150,000 - 210,000
Senior ML Platform Engineer - GPU & Cloud Pipelines
Senior ML Platform Engineer - GPU & Cloud Pipelines

Next Frontier Capital • Palo Alto (CA)

On-site
USD 150,000 - 210,000