Senior ML Platform Engineer - GPU Training on Cloud

J.P. Morgan

New York (NY)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
Retirement plan
Tuition reimbursement
On-site wellness

Job summary

JPMorganChase is seeking a Lead Software Engineer to design, build and operate an end-to-end ML training platform. You will run GPU training workloads, scale training, and manage infrastructure on Kubernetes (EKS) across cloud environments.

You will enable Gen AI/LLM training, implement observability, and collaborate with data engineering to enforce security, cost controls and reusable patterns across teams.

Qualifications

  • Formal training or certification in software engineering concepts with 5+ years applied experience.
  • Demonstrated experience running ML training in cloud environments and debugging across infra & code.
  • Strong Python skills with solid engineering practices (testing, code reviews, modular design).
  • Experience building automation/CI for ML codebases (build, test, release, deployment/promotion).
  • Hands-on experience with deep learning training workflows and at least one major framework (PyTorch or TensorFlow).
  • Understanding training performance and stability: data loading, mixed precision, checkpointing, reproducibility.
  • Experience with distributed training concepts (DDP/FSDP/DeepSpeed), scaling and bottleneck analysis.
  • Ability to profile and optimize training systems (CPU/GPU, memory, I/O, networking, scheduling).
  • Experience with Kubernetes fundamentals for running compute-intensive workloads and AWS (EKS/ECR, S3, IAM, VPC, CloudWatch, EC2).
  • Strong understanding of responsible AI use and secure adoption within delivery practices.

Responsibilities

  • Design, build, and maintain end-to-end ML training platform.
  • Run and optimize GPU training workloads (single-node and distributed).
  • Build and operate training infrastructure on Kubernetes (EKS and others).
  • Enable Gen AI/LLM training and fine-tuning workflows with governance.
  • Implement observability: metrics, logs, dashboards, alerting, runbooks.
  • Collaborate with data engineering and platform teams to define interfaces and guardrails.
  • Improve developer experience with containers, CI/CD, templates, docs, self-service workflow.
  • Promote enterprise-authored AI-assisted engineering practices, code reviews and testing.
  • Apply SDLC tools with AI-assisted development to improve automation.

Skills

Python
ML training
Kubernetes
AWS
Distributed training
GPU
CI/CD
Security/compliance
Performance optimization
Code reviews

Tools

Kubernetes (EKS)
S3
IAM
VPC
CloudWatch

Job description

JPMorganChase is seeking a Lead Software Engineer to design, build and operate an end-to-end ML training platform. You will run GPU training workloads, scale training, and manage infrastructure on Kubernetes (EKS) across cloud environments.

You will enable Gen AI/LLM training, implement observability, and collaborate with data engineering to enforce security, cost controls and reusable patterns across teams.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Platform Engineer — GPU Training on Cloud
Senior ML Platform Engineer — GPU Training on Cloud

J.P. Morgan • New York (NY)

On-site
USD 140,000 - 200,000
Senior ML Platform Engineer: GPU Training & Cloud
Senior ML Platform Engineer: GPU Training & Cloud

JPMorganChase • Palo Alto (CA)

On-site
USD 160,000 - 210,000
Lead ML Platform Engineer: GPU Training on Cloud
Lead ML Platform Engineer: GPU Training on Cloud

JPMorganChase • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Health care coverage
On-site health centers
Retirement savings plan
+2
Senior ML Platform Engineer - GPU Training Systems
Senior ML Platform Engineer - GPU Training Systems

JPMorgan Chase & Co. • Palo Alto (CA)

On-site
USD 150,000 - 210,000
Senior ML Platform Engineer - GPU & Cloud
Senior ML Platform Engineer - GPU & Cloud

Socket.dev • Palo Alto (CA)

On-site
USD 150,000 - 190,000
Lead ML Training Platform Engineer
Lead ML Training Platform Engineer

JPMorgan Chase & Co. • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Senior ML Platform Engineer (GPU/Cloud)
Senior ML Platform Engineer (GPU/Cloud)

Fairygodboss • Palo Alto (CA)

On-site
USD 180,000 - 250,000
Senior Lead AI/ML Platform Engineer
Senior Lead AI/ML Platform Engineer

Fairygodboss • Jersey City (NJ)

On-site
USD 180,000 - 240,000
Senior AI Platform Engineer — Cloud, GPUs & MLOps
Senior AI Platform Engineer — Cloud, GPUs & MLOps

JPMorgan Chase & Co. • Fairfax (DE)

Hybrid
USD 140,000 - 200,000
Lead Software Engineer - Machine Learning Platform
Lead Software Engineer - Machine Learning Platform

JPMorgan Chase & Co. • Palo Alto (CA)

On-site
USD 180,000 - 260,000