Senior ML Platform Engineer: GPU Training & Cloud Pipelines

Next Frontier Capital

Palo Alto (CA)

On-site

USD 170,000 - 250,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

JPMorganChase, a leading global financial services firm, seeks a Lead Software Engineer to architect and operate scalable ML training pipelines on AWS and cloud platforms. You will optimize GPU workloads, ensure reproducibility, and drive enterprise-grade AI governance across environments.

You will collaborate with data engineering and platform teams to standardize interfaces, security, and cost controls, while improving developer experience with containers, CI/CD, and self-service workflows.

Qualifications

  • Formal training or certification on software engineering concepts and 5+ years applied experience.
  • Demonstrated experience running ML training in cloud environments and debugging issues across infrastructure & code.
  • Strong Python skills with solid engineering practices (testing, code reviews, modular design, dependency management).
  • Experience building automation/CI for ML codebases (build, test, release, deployment/promotion workflows).
  • Hands on experience with deep learning training workflows and at least one major framework (eg., PyTorch or TensorFlow).
  • Understanding of training performance and stability: data loading bottlenecks, mixed precision, checkpointing, reproducibility, and evaluation methodology.
  • Experience with distributed training and related concepts (e.g., DDP/FSDP/DeepSpeed concepts, collective communication basics, scaling and bottleneck analysis).
  • Ability to profile and optimize training systems (CPU/GPU utilization, memory, I/O throughput, networking, scheduling).
  • Experience with Kubernetes fundamentals for running compute-intensive workloads and AWS (eg., EKS/ECR, S3, IAM, VPC/networking, Cloudwatch, EC2)
  • Demonstrated experience leading effective use of approved AI-assisted software development tools (e.g., for coding, code review, test acceleration, troubleshooting) with the ability to set team expectations for validating AI outputs for correctness, performance, and security.

Responsibilities

  • Design, build, and maintain end-to-end ML training platform.
  • Run and optimize GPU training workloads (single-node and distributed), improving throughput, utilization and reproducibility.
  • Build and operate training infrastructure on Kubernetes (e.g., EKS and other manage Kubernetes platforms), including resource management and workload troubleshooting.
  • Enable Gen AI/LLM training and fine-tuning workflows (e.g., supervised fine-tuning), including evaluation harnesses, artifact/version governance, and scalable GPU execution patterns aligned to enterprise controls.
  • Implement observability for training systems: metrics, logs, dashboards, alerting, and operational runbooks.
  • Partner with data engineering and platform teams to define interfaces, standards, and guardrails (security, access, cost controls)
  • Improve developer experience for training: standardized containers, CI/CD, templates, documentation, and self-service workflow
  • Drives team adoption of enterprise-authorized AI-assisted engineering practices within the work environment to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test strategy acceleration, incident/root-cause analysis support), while establishing consistent validation standards (secure coding, peer review, automated testing) and promoting reuse of effective patterns across the team.
  • Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation.

Skills

Python
ML training in cloud
Kubernetes
AWS
CI/CD
Code review

Tools

Kubernetes
AWS
EKS
CI/CD tooling

Job description

JPMorganChase, a leading global financial services firm, seeks a Lead Software Engineer to architect and operate scalable ML training pipelines on AWS and cloud platforms. You will optimize GPU workloads, ensure reproducibility, and drive enterprise-grade AI governance across environments.

You will collaborate with data engineering and platform teams to standardize interfaces, security, and cost controls, while improving developer experience with containers, CI/CD, and self-service workflows.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior ML Platform Engineer: GPU Training & Cloud Infra
Senior ML Platform Engineer: GPU Training & Cloud Infra

JPMorgan Chase & Co. • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Senior ML Platform Engineer - GPU & Cloud Pipelines
Senior ML Platform Engineer - GPU & Cloud Pipelines

Next Frontier Capital • Palo Alto (CA)

On-site
USD 150,000 - 210,000
Senior ML Platform Engineer - GPU Training on Cloud
Senior ML Platform Engineer - GPU Training on Cloud

J.P. Morgan • New York (NY)

On-site
USD 180,000 - 240,000
Health insurance
Retirement plan
Tuition reimbursement
+1
Lead ML Platform Engineer — GPU Training on Cloud
Lead ML Platform Engineer — GPU Training on Cloud

JPMorgan Chase & Co. • Palo Alto (CA)

On-site
USD 180,000 - 270,000
Senior ML Platform Engineer — GPU Training on Cloud
Senior ML Platform Engineer — GPU Training on Cloud

J.P. Morgan • New York (NY)

On-site
USD 140,000 - 200,000
Lead ML Engineer – Scalable GPU Training on Cloud
Lead ML Engineer – Scalable GPU Training on Cloud

JPMorganChase • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Senior ML Platform Engineer - GPU Training Systems
Senior ML Platform Engineer - GPU Training Systems

JPMorgan Chase & Co. • Palo Alto (CA)

On-site
USD 150,000 - 210,000
Senior ML Platform Engineer: GPU Training & Cloud
Senior ML Platform Engineer: GPU Training & Cloud

JPMorganChase • Palo Alto (CA)

On-site
USD 160,000 - 210,000
Lead ML Platform Engineer - GPU & Cloud
Lead ML Platform Engineer - GPU & Cloud

JPMorgan Chase • Palo Alto (CA)

On-site
USD 157,000 - 215,000
Lead ML Platform Engineer for Scalable GPU Training
Lead ML Platform Engineer for Scalable GPU Training

Next Frontier Capital • Palo Alto (CA)

On-site
USD 180,000 - 245,000