Senior GenAI Platform Engineer — Scale Multi-GPU Infra

Quantiphi

United States

Hybrid

USD 150,000 - 190,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Impact at Quantiphi
Growth opportunities
Research-focused environment
Fortune 500 exposure

Job summary

Quantiphi is seeking a Senior Platform Engineer to design, optimize, and scale GenAI and LLM infrastructure across multi-GPU environments. You will profile GPUs, tune distributed training, and support production-grade deployments in collaboration with data science, MLOps, and application teams.

The role emphasizes building reusable infrastructure with Terraform/Helm, and deploying GenAI workloads using Slurm, OpenShift, and Kubernetes in cloud or on-prem environments.

Qualifications

  • Proven experience with Slurm and distributed training environments.
  • Hands-on expertise with Red Hat OpenShift and/or Kubernetes.
  • Deep knowledge of the NVIDIA GPU ecosystem (CUDA, cuDNN, NCCL, Nsight, Triton/TensorRT).
  • Strong foundation in Linux systems, performance tuning, and multi-GPU optimization.
  • Experience deploying GenAI workloads (LLM fine-tuning, RAG pipelines, multi-modal systems).
  • Familiarity with Infrastructure-as-Code tools (Terraform, Ansible).
  • Experience with cloud GPU environments (GCP, Azure, AWS, OCI) and/or on-prem GPU clusters.

Responsibilities

  • Design and implement scalable infrastructure for LLM and GenAI workloads across multi-GPU environments.
  • Perform GPU profiling, benchmarking, and performance optimization for distributed training workloads.
  • Manage and schedule compute-intensive jobs using Slurm-based clusters and OpenShift/Kubernetes environments.
  • Enable and optimize the NVIDIA GPU stack (CUDA, cuDNN, NCCL, Triton, RAPIDS, etc.).
  • Collaborate with cross-functional teams to deploy models in research and production environments.
  • Build and support GenAI pipelines (fine-tuning, RAG, multi-modal inferencing, LLMOps).
  • Develop reusable infrastructure templates using tools like Terraform and Helm.
  • Contribute to internal innovation (PoCs, workshops) and support client-facing delivery engagements.

Skills

Slurm-based training
OpenShift/Kubernetes
NVIDIA GPU stack
Linux performance tuning
GenAI deployment
Terraform/Helm

Tools

Terraform
Helm
Ansible

Job description

Quantiphi is seeking a Senior Platform Engineer to design, optimize, and scale GenAI and LLM infrastructure across multi-GPU environments. You will profile GPUs, tune distributed training, and support production-grade deployments in collaboration with data science, MLOps, and application teams.

The role emphasizes building reusable infrastructure with Terraform/Helm, and deploying GenAI workloads using Slurm, OpenShift, and Kubernetes in cloud or on-prem environments.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior GenAI Platform Engineer: Multi-GPU & Kubernetes
Senior GenAI Platform Engineer: Multi-GPU & Kubernetes

Quantiphi • United States

Hybrid
USD 180,000 - 240,000
GenAI Platform Architect (Multi-GPU & MLOps)
GenAI Platform Architect (Multi-GPU & MLOps)

Quantiphi • Boston (MA)

Hybrid
USD 150,000 - 190,000
Platform Engineer - Senior - US
Platform Engineer - Senior - US

Quantiphi • United States

Hybrid
USD 180,000 - 240,000
GenAI Infra Architect - Multi-Cloud, Kubernetes, Scale
GenAI Infra Architect - Multi-Cloud, Kubernetes, Scale

Scale • San Francisco (CA), New York (NY)

On-site
USD 179,000 - 224,000
Health, dental & vision insurance
Retirement benefits
Learning & development stipend
+2
AI/ML Platform Engineer — Hybrid Cloud & GPU
AI/ML Platform Engineer — Hybrid Cloud & GPU

Madrona Venture Labs • United States

Hybrid
USD 180,000 - 260,000
Senior LLMOps Platform Engineer — GPU AI Infra
Senior LLMOps Platform Engineer — GPU AI Infra

Quantum Technologies. LLC • Jersey City (NJ), Northern (KY)

Hybrid
USD 120,000 - 165,000
Senior Platform Engineer
Senior Platform Engineer

Quantiphi • United States

Hybrid
USD 150,000 - 190,000
Impact at Quantiphi
Growth opportunities
Research-focused environment
+1
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Senior AI Infra Engineer - Scale ML Platforms
Senior AI Infra Engineer - Scale ML Platforms

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Enterprise GenAI Infra Engineer — Multi-Cloud Scale
Enterprise GenAI Infra Engineer — Multi-Cloud Scale

Front Door Defense • New York (NY), Northern (KY)

Hybrid
USD 179,000 - 224,000
Health, dental, vision coverage
Retirement benefits
Learning and development stipend
+2