Senior ML Platform Engineer: GPU Training & Cloud Infra

JPMorgan Chase & Co.

Palo Alto (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

JPMorganChase is seeking a Lead Software Engineer in the AI/ML Data Platforms group to design, build, and operate scalable ML training systems on AWS and Kubernetes. You will drive GPU-based training pipelines, performance optimizations, and governance across environments.

You will partner with data engineering and platform teams, improve developer experience with reusable containers and CI/CD, and advance Gen AI/LLM workflows while upholding secure, compliant practices.

Qualifications

  • Formal training or certification on software engineering concepts and 5+ years applied experience.
  • Demonstrated experience running ML training in cloud environments and debugging issues across infrastructure & code.
  • Strong Python skills with solid engineering practices (testing, code reviews, modular design, dependency management).
  • Experience building automation/CI for ML codebases (build, test, release, deployment/promotion workflows).
  • Hands on experience with deep learning training workflows and at least one major framework (e.g., PyTorch or TensorFlow).
  • Understanding of training performance and stability: data loading bottlenecks, mixed precision, checkpointing, reproducibility, and evaluation methodology.
  • Experience with distributed training and related concepts (e.g., DDP/FSDP/DeepSpeed concepts, collective communication basics).
  • Ability to profile and optimize training systems (CPU/GPU utilization, memory, I/O throughput).
  • Experience with Kubernetes fundamentals for running compute-intensive workloads and AWS (EKS/ECR, S3, IAM, VPC).
  • Demonstrated experience leading effective use of AI-assisted development tools with ability to set team expectations for correctness, performance, and security.
  • Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations and secure handling.

Responsibilities

  • Design, build, and maintain end-to-end ML training platform.
  • Run and optimize GPU training workloads (single-node and distributed).
  • Build and operate training infrastructure on Kubernetes (EKS and others).
  • Enable Gen AI/LLM training and fine-tuning workflows with governance.
  • Implement observability for training systems: metrics, logs, dashboards, alerts.
  • Partner with data engineering and platform teams to define interfaces and guardrails.
  • Improve developer experience for training: standardized containers, CI/CD, templates, docs, self-service workflows.
  • Drive team adoption of enterprise AI-assisted engineering practices for code quality and automation.
  • Apply SDLC toolchain knowledge to improve automation value.

Skills

Python
ML/AI systems
Distributed training
Kubernetes
AWS
PyTorch
TensorFlow
CI/CD
Profiling/Optimization
Security/Compliance

Tools

Kubernetes
Docker
EKS
CI/CD tooling

Job description

JPMorganChase is seeking a Lead Software Engineer in the AI/ML Data Platforms group to design, build, and operate scalable ML training systems on AWS and Kubernetes. You will drive GPU-based training pipelines, performance optimizations, and governance across environments.

You will partner with data engineering and platform teams, improve developer experience with reusable containers and CI/CD, and advance Gen AI/LLM workflows while upholding secure, compliant practices.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior ML Platform Engineer - GPU Training on Cloud
Senior ML Platform Engineer - GPU Training on Cloud

J.P. Morgan • New York (NY)

On-site
USD 180,000 - 240,000
Health insurance
Retirement plan
Tuition reimbursement
+1
Lead ML Platform Engineer — GPU Training on Cloud
Lead ML Platform Engineer — GPU Training on Cloud

JPMorgan Chase & Co. • Palo Alto (CA)

On-site
USD 180,000 - 270,000
Senior ML Platform Engineer — GPU Training on Cloud
Senior ML Platform Engineer — GPU Training on Cloud

J.P. Morgan • New York (NY)

On-site
USD 140,000 - 200,000
Senior ML Platform Engineer: GPU Training & Cloud
Senior ML Platform Engineer: GPU Training & Cloud

JPMorganChase • Palo Alto (CA)

On-site
USD 160,000 - 210,000
Lead ML Platform Engineer: GPU Training on Cloud
Lead ML Platform Engineer: GPU Training on Cloud

JPMorgan Chase • Palo Alto (CA)

On-site
USD 157,000 - 215,000
Senior ML Platform Engineer - GPU Training Systems
Senior ML Platform Engineer - GPU Training Systems

JPMorgan Chase & Co. • Palo Alto (CA)

On-site
USD 150,000 - 210,000
Lead ML Engineer – Scalable GPU Training on Cloud
Lead ML Engineer – Scalable GPU Training on Cloud

JPMorganChase • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Senior ML Platform Engineer (GPU/Cloud)
Senior ML Platform Engineer (GPU/Cloud)

Fairygodboss • Palo Alto (CA)

On-site
USD 180,000 - 250,000
Senior Lead AI/ML Platform Engineer
Senior Lead AI/ML Platform Engineer

Fairygodboss • Jersey City (NJ)

On-site
USD 180,000 - 240,000
Senior Lead AI/ML Training Infrastructure Engineer
Senior Lead AI/ML Training Infrastructure Engineer

JPMorganChase • Seattle (WA)

On-site
USD 180,000 - 240,000