Lead ML Platform Engineer - GPU & Cloud

JPMorgan Chase

Palo Alto (CA)

On-site

USD 157,000 - 215,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

JPMorganChase is seeking a Lead Software Engineer - Machine Learning in Palo Alto to architect scalable ML training platforms on AWS and other clouds. You will push GPU-based workloads, optimize performance, and own end-to-end pipelines within a fast-moving AI/ML Data Platforms team.

Responsibilities include building training infra on Kubernetes, enabling Gen AI workflows, and delivering governed, reproducible experiments. Strong Python, deep learning, and CI/CD capabilities are required.

Qualifications

  • Formal training or certification on software engineering concepts is required.
  • 5+ years of applied software engineering experience in ML or cloud environments.
  • Deep learning frameworks (PyTorch or TensorFlow) experience.
  • Experience building automation/CI for ML codebases and training workloads.

Responsibilities

  • Design, build, and maintain end-to-end ML training platform.
  • Run and optimize GPU training workloads and distributed training.
  • Operate training infrastructure on Kubernetes and AWS cloud services.
  • Enable GenAI/LLM training workflows with governance and security.
  • Implement observability with metrics, logs, dashboards and runbooks.
  • Collaborate with data engineering to define interfaces and guardrails.
  • Improve developer experience with containers, templates, and self-service workflows.
  • Promote AI-assisted engineering practices and standardized validation.

Skills

Python
ML training in cloud
Kubernetes
AWS
PyTorch
TensorFlow
CI/CD
Observability
Secure coding
AI governance

Education

Software engineering training/certification

Tools

Kubernetes (EKS)
S3
IAM
CI/CD tooling

Job description

JPMorganChase is seeking a Lead Software Engineer - Machine Learning in Palo Alto to architect scalable ML training platforms on AWS and other clouds. You will push GPU-based workloads, optimize performance, and own end-to-end pipelines within a fast-moving AI/ML Data Platforms team.

Responsibilities include building training infra on Kubernetes, enabling Gen AI workflows, and delivering governed, reproducible experiments. Strong Python, deep learning, and CI/CD capabilities are required.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead ML Platform Engineer — GPU Training on Cloud
Lead ML Platform Engineer — GPU Training on Cloud

JPMorgan Chase & Co. • Palo Alto (CA)

On-site
USD 180,000 - 270,000
Lead ML Platform Engineer for Scalable GPU Training
Lead ML Platform Engineer for Scalable GPU Training

Next Frontier Capital • Palo Alto (CA)

On-site
USD 180,000 - 245,000
Lead ML Engineer – Scalable GPU Training on Cloud
Lead ML Engineer – Scalable GPU Training on Cloud

JPMorganChase • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Senior ML Platform Engineer: GPU Training & Cloud Infra
Senior ML Platform Engineer: GPU Training & Cloud Infra

JPMorgan Chase & Co. • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Senior ML Platform Engineer (GPU/Cloud)
Senior ML Platform Engineer (GPU/Cloud)

Fairygodboss • Palo Alto (CA)

On-site
USD 180,000 - 250,000
Senior ML Platform Engineer: GPU Training & Cloud Pipelines
Senior ML Platform Engineer: GPU Training & Cloud Pipelines

Next Frontier Capital • Palo Alto (CA)

On-site
USD 170,000 - 250,000
Senior ML Platform Engineer: GPU Training & Cloud
Senior ML Platform Engineer: GPU Training & Cloud

JPMorganChase • Palo Alto (CA)

On-site
USD 160,000 - 210,000
Senior ML Platform Engineer - GPU Training on Cloud
Senior ML Platform Engineer - GPU Training on Cloud

J.P. Morgan • New York (NY)

On-site
USD 180,000 - 240,000
Health insurance
Retirement plan
Tuition reimbursement
+1
Senior ML Platform Engineer - GPU & Cloud Pipelines
Senior ML Platform Engineer - GPU & Cloud Pipelines

Next Frontier Capital • Palo Alto (CA)

On-site
USD 150,000 - 210,000
Senior ML Platform Engineer — GPU Training on Cloud
Senior ML Platform Engineer — GPU Training on Cloud

J.P. Morgan • New York (NY)

On-site
USD 140,000 - 200,000