Lead ML Engineer – Scalable GPU Training on Cloud

JPMorganChase

Palo Alto (CA)

On-site

USD 180,000 - 240,000

Full time

7 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

JPMorganChase in Palo Alto seeks a Lead Software Engineer to design, build, and operate scalable ML training systems on AWS and cloud platforms. You will drive end-to-end training pipelines, run GPU workloads, and ensure governance and observability across environments.

You will collaborate with data engineering and platform teams to standardize interfaces, improve developer experience, and advance AI-assisted engineering practices while maintaining strong security and compliance.

Qualifications

  • 5+ years of software engineering experience with ML systems.
  • Experience running ML training in cloud environments and debugging across infra.
  • Strong Python and engineering practices (testing, reviews, modular design).
  • Experience building automation/CI for ML codebases and deployment workflows.
  • Hands-on with deep learning training workflows (PyTorch or TensorFlow).
  • Understanding of training performance, stability, data loading, and reproducibility.
  • Experience with distributed training concepts (DDP/FSDP/DeepSpeed).
  • Ability to profile/optimize training systems (CPU/GPU/memory/I/O).
  • Experience with Kubernetes and AWS (EKS/ECR, S3, IAM, VPC, CloudWatch).
  • Demonstrated use of AI-assisted software development tools with governance.

Responsibilities

  • Design, build, and maintain end-to-end ML training platform.
  • Run and optimize GPU training workloads and improve throughput.
  • Build and operate training infra on Kubernetes (EKS and others).
  • Enable Gen AI/LLM training and fine-tuning workflows with governance.
  • Implement observability for training systems: metrics, logs, dashboards, alerting.
  • Partner with data engineering and platform teams to define interfaces and guardrails.
  • Improve developer experience with containers, CI/CD, templates, docs, self-service workflows.
  • Promote enterprise-authorized AI-assisted engineering practices with validation standards.
  • Apply SDLC tooling to improve automation and value.

Skills

Python
ML Training in Cloud
Kubernetes
AWS
GPU Training
CI/CD for ML
Observability
AI Governance

Tools

Kubernetes
AWS
PyTorch
TensorFlow
Docker
CI/CD tooling
S3
EKS
CloudWatch

Job description

JPMorganChase in Palo Alto seeks a Lead Software Engineer to design, build, and operate scalable ML training systems on AWS and cloud platforms. You will drive end-to-end training pipelines, run GPU workloads, and ensure governance and observability across environments.

You will collaborate with data engineering and platform teams to standardize interfaces, improve developer experience, and advance AI-assisted engineering practices while maintaining strong security and compliance.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead ML Platform Engineer — GPU Training on Cloud
Lead ML Platform Engineer — GPU Training on Cloud

JPMorgan Chase & Co. • Palo Alto (CA)

On-site
USD 180,000 - 270,000
Senior ML Platform Engineer: GPU Training & Cloud Infra
Senior ML Platform Engineer: GPU Training & Cloud Infra

JPMorgan Chase & Co. • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Senior ML Platform Engineer - GPU Training on Cloud
Senior ML Platform Engineer - GPU Training on Cloud

J.P. Morgan • New York (NY)

On-site
USD 180,000 - 240,000
Health insurance
Retirement plan
Tuition reimbursement
+1
Senior ML Platform Engineer: GPU Training & Cloud
Senior ML Platform Engineer: GPU Training & Cloud

JPMorganChase • Palo Alto (CA)

On-site
USD 160,000 - 210,000
Senior ML Platform Engineer — GPU Training on Cloud
Senior ML Platform Engineer — GPU Training on Cloud

J.P. Morgan • New York (NY)

On-site
USD 140,000 - 200,000
Senior ML Platform Engineer (GPU/Cloud)
Senior ML Platform Engineer (GPU/Cloud)

Fairygodboss • Palo Alto (CA)

On-site
USD 180,000 - 250,000
Senior ML Platform Engineer - GPU Training Systems
Senior ML Platform Engineer - GPU Training Systems

JPMorgan Chase & Co. • Palo Alto (CA)

On-site
USD 150,000 - 210,000
Lead ML Platform Engineer | MLOps on AWS
Lead ML Platform Engineer | MLOps on AWS

Fairygodboss • Plano (TX)

On-site
USD 140,000 - 210,000
Lead AI/ML Engineer - GPU-Accelerated Serving
Lead AI/ML Engineer - GPU-Accelerated Serving

Socket.dev • Palo Alto (CA), Northern (KY)

Hybrid
USD 180,000 - 230,000
Senior Lead AI/ML Platform Engineer
Senior Lead AI/ML Platform Engineer

Fairygodboss • Jersey City (NJ)

On-site
USD 180,000 - 240,000