Senior ML Platform Engineer - GPU & Cloud Pipelines

Next Frontier Capital

Palo Alto (CA)

On-site

USD 150,000 - 210,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

JPMorganChase is seeking a Software Engineer III to design and operate an end-to-end ML training platform on AWS and other clouds. You will run GPU workloads, optimize performance, and enable Gen AI workflows within a governed, secure environment.

You will collaborate with data engineering and platform teams, improve developer experience, and implement observability with metrics, logs, and dashboards across the training stack.

Qualifications

  • Formal training or certification in software engineering concepts and 3+ years applied experience.
  • Demonstrated experience running ML training in cloud environments and debugging issues across infrastructure & code.
  • Strong Python skills with solid engineering practices (testing, code reviews, modular design, dependency management).
  • Experience building automation/CI for ML codebases (build, test, release, deployment/promotion workflows).
  • Hands on experience with deep learning training workflows and at least one major framework (e.g. PyTorch or TensorFlow).
  • Understanding of training performance and stability: data loading bottlenecks, mixed precision, checkpointing, reproducibility, and evaluation methodology.
  • Experience with distributed training and related concepts (e.g., DDP/FSDP/DeepSpeed concepts, collective communication basics, scaling and bottleneck analysis).
  • Ability to profile and optimize training systems (CPU/GPU utilization, memory, I/O throughput, networking, scheduling).
  • Experience with Kubernetes fundamentals for running compute-intensive workloads and AWS (eg., EKS/ECR, S3, IAM, VPC/networking, Cloudwatch, EC2)
  • Hands-on experience using enterprise-authorized AI-assisted software development tools within the work environment with demonstrated ability to evaluate, validate, and refine AI-generated outputs for correctness, performance, and security.
  • Understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; ability to guide peers on safe and effective usage within team practices.

Responsibilities

  • Design, build, and maintain end-to-end ML training platform.
  • Run and optimize GPU training workloads (single-node and distributed).
  • Build and operate training infrastructure on Kubernetes (e.g., EKS) including resource management and troubleshooting.
  • Enable Gen AI/LLM training and fine-tuning workflows with evaluation harnesses and governance.
  • Implement observability for training systems: metrics, logs, dashboards, alerting, runbooks.
  • Partner with data engineering and platform teams to define interfaces, guardrails (security, cost controls).
  • Improve developer experience for training: containers, CI/CD, templates, documentation, self-service workflow.
  • Leverage AI-assisted development tools and ensure outputs are correct, secure, and well-tested.

Skills

Python
ML training in cloud
Kubernetes familiarity
AWS cloud concepts
Performance optimization
CI/CD for ML
Debugging infrastructure
GPU training workflows

Education

Formal training or certification in software engineering concepts

Tools

Kubernetes (EKS)
AWS (S3, IAM, EC2, ECR)
GPU compute
PyTorch/TensorFlow
Observability tools

Job description

JPMorganChase is seeking a Software Engineer III to design and operate an end-to-end ML training platform on AWS and other clouds. You will run GPU workloads, optimize performance, and enable Gen AI workflows within a governed, secure environment.

You will collaborate with data engineering and platform teams, improve developer experience, and implement observability with metrics, logs, and dashboards across the training stack.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior ML Platform Engineer - GPU Training Systems
Senior ML Platform Engineer - GPU Training Systems

JPMorgan Chase & Co. • Palo Alto (CA)

On-site
USD 150,000 - 210,000
Senior ML Platform Engineer — GPU Training on Cloud
Senior ML Platform Engineer — GPU Training on Cloud

J.P. Morgan • New York (NY)

On-site
USD 140,000 - 200,000
Senior ML Platform Engineer: GPU Training & Cloud Pipelines
Senior ML Platform Engineer: GPU Training & Cloud Pipelines

Next Frontier Capital • Palo Alto (CA)

On-site
USD 170,000 - 250,000
Senior ML Platform Engineer: GPU Training & Cloud
Senior ML Platform Engineer: GPU Training & Cloud

JPMorganChase • Palo Alto (CA)

On-site
USD 160,000 - 210,000
Senior ML Platform Engineer (GPU/Cloud)
Senior ML Platform Engineer (GPU/Cloud)

Fairygodboss • Palo Alto (CA)

On-site
USD 180,000 - 250,000
Senior ML Platform Engineer: GPU Training & Cloud Infra
Senior ML Platform Engineer: GPU Training & Cloud Infra

JPMorgan Chase & Co. • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Senior ML Platform Engineer - GPU Training on Cloud
Senior ML Platform Engineer - GPU Training on Cloud

J.P. Morgan • New York (NY)

On-site
USD 180,000 - 240,000
Health insurance
Retirement plan
Tuition reimbursement
+1
Lead ML Platform Engineer — GPU Training on Cloud
Lead ML Platform Engineer — GPU Training on Cloud

JPMorgan Chase & Co. • Palo Alto (CA)

On-site
USD 180,000 - 270,000
Lead ML Platform Engineer - GPU & Cloud
Lead ML Platform Engineer - GPU & Cloud

JPMorgan Chase • Palo Alto (CA)

On-site
USD 157,000 - 215,000
Lead ML Engineer – Scalable GPU Training on Cloud
Lead ML Engineer – Scalable GPU Training on Cloud

JPMorganChase • Palo Alto (CA)

On-site
USD 180,000 - 240,000