Senior ML Platform Engineer (GPU/Cloud)

Fairygodboss

Palo Alto (CA)

On-site

USD 180,000 - 250,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

JPMorganChase in Palo Alto is seeking a Software Engineer III to join the AI/ML data platform team. You will design and operate scalable ML training systems on AWS and other clouds, productionizing training workloads, and improving performance and cost efficiency.

You will work on end-to-end pipelines, Kubernetes, EKS, AI tooling, and governance. The role requires strong Python, ML frameworks, and cloud experience with a focus on reliability and secure workflows.

Qualifications

  • Formal training or certification on software engineering concepts and 3+ years applied experience.
  • Demonstrated experience running ML training in cloud environments and debugging issues across infrastructure & code.
  • Strong Python skills with solid engineering practices (testing, code reviews, modular design, dependency management).
  • Experience building automation/CI for ML codebases (build, test, release, deployment/promotion workflows).
  • Hands on experience with deep learning training workflows and at least one major framework (eg., PyTorch or TensorFlow).
  • Understanding of training performance and stability: data loading bottlenecks, mixed precision, checkpointing, reproducibility, and evaluation methodology.
  • Experience with distributed training and related concepts (e.g., DDP/FSDP/DeepSpeed concepts, collective communication basics, scaling and bottleneck analysis).
  • Ability to profile and optimize training systems (CPU/GPU utilization, memory, I/O throughput, networking, scheduling).
  • Experience with Kubernetes fundamentals for running compute-intensive workloads and AWS (eg., EKS/ECR, S3, IAM, VPC/networking, Cloudwatch, EC2)
  • Hands-on experience using enterprise-authorized AI-assisted software development tools within the work environment with demonstrated ability to evaluate and refine outputs for correctness, performance, and security.
  • Understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations.

Responsibilities

  • Design, build, and maintain end-to-end ML training platform.
  • Run and optimize GPU training workloads (single-node and distributed), improving throughput, utilization and reproducibility.
  • Build and operate training infrastructure on Kubernetes (e.g., EKS and other manage Kubernetes platforms), including resource management and workload troubleshooting.
  • Enable Gen AI/LLM training and fine-tuning workflows, including evaluation harnesses, artifact/version governance, and scalable GPU execution patterns aligned to enterprise controls.
  • Implement observability for training systems: metrics, logs, dashboards, alerting, and operational runbooks.
  • Partner with data engineering and platform teams to define interfaces, standards, and guardrails (security, access, cost controls).
  • Improve developer experience for training: standardized containers, CI/CD, templates, documentation, and self-service workflow.
  • Leverages enterprise-authorized AI coding assist tools to improve code quality, delivery speed, and productivity; validates outputs through peer review and testing.

Tools

Kubernetes
AWS
EKS
S3
IAM
VPC
CloudWatch
EC2

Job description

JPMorganChase in Palo Alto is seeking a Software Engineer III to join the AI/ML data platform team. You will design and operate scalable ML training systems on AWS and other clouds, productionizing training workloads, and improving performance and cost efficiency.

You will work on end-to-end pipelines, Kubernetes, EKS, AI tooling, and governance. The role requires strong Python, ML frameworks, and cloud experience with a focus on reliability and secure workflows.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Platform Engineer - GPU & Cloud
Senior ML Platform Engineer - GPU & Cloud

Socket.dev • Palo Alto (CA)

On-site
USD 150,000 - 190,000
Senior ML Platform Engineer: GPU Training & Cloud
Senior ML Platform Engineer: GPU Training & Cloud

JPMorganChase • Palo Alto (CA)

On-site
USD 160,000 - 210,000
Senior ML Platform Engineer - GPU Training Systems
Senior ML Platform Engineer - GPU Training Systems

JPMorgan Chase & Co. • Palo Alto (CA)

On-site
USD 150,000 - 210,000
Senior ML Platform Engineer — GPU Training on Cloud
Senior ML Platform Engineer — GPU Training on Cloud

J.P. Morgan • New York (NY)

On-site
USD 140,000 - 200,000
Lead ML Platform Engineer: GPU Training on Cloud
Lead ML Platform Engineer: GPU Training on Cloud

JPMorganChase • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Health care coverage
On-site health centers
Retirement savings plan
+2
Senior ML Platform Engineer - GPU Training on Cloud
Senior ML Platform Engineer - GPU Training on Cloud

J.P. Morgan • New York (NY)

On-site
USD 180,000 - 240,000
Health insurance
Retirement plan
Tuition reimbursement
+1
Lead ML Training Platform Engineer
Lead ML Training Platform Engineer

JPMorgan Chase & Co. • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Senior AI Platform Engineer — Cloud, GPUs & MLOps
Senior AI Platform Engineer — Cloud, GPUs & MLOps

JPMorgan Chase & Co. • Fairfax (DE)

Hybrid
USD 140,000 - 200,000
Senior Lead AI/ML Platform Engineer
Senior Lead AI/ML Platform Engineer

Fairygodboss • Jersey City (NJ)

On-site
USD 180,000 - 240,000
Senior Python Engineer - AI/ML on AWS
Senior Python Engineer - AI/ML on AWS

JPMorganChase • Houston (TX)

On-site
USD 110,000 - 160,000