GenAI Platform Engineer: Multi-GPU & Kubernetes Lead

Quantiphi

United States

On-site

USD 140,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Quantiphi seeks a Senior Platform Engineer to design, optimize, and scale GenAI infrastructure. You will lead multi-GPU environments, GPU profiling, and distributed training, partnering with data science, MLOps, and application teams for production deployments.

Ideal candidates have hands-on OpenShift/Kubernetes, NVIDIA GPU stack expertise, and IaC skills. This remote US/Canada role offers growth, collaboration with Fortune 500 clients, and cutting-edge AI workloads.

Qualifications

  • Experience with Slurm and distributed training environments.
  • Hands-on expertise with Red Hat OpenShift and/or Kubernetes.
  • Deep knowledge of the NVIDIA GPU ecosystem (CUDA, cuDNN, NCCL, TensorRT).
  • Strong foundation in Linux systems, performance tuning, and multi-GPU optimization.
  • Experience deploying GenAI workloads (LLM fine-tuning, RAG pipelines, multi-modal systems).
  • Familiarity with Infrastructure-as-Code tools (Terraform, Ansible).
  • Experience with cloud GPU environments (GCP, Azure, AWS, OCI) and/or on-prem GPU clusters.
  • Experience with NVIDIA NIMs, DGX systems, or GPU-accelerated containers.
  • Knowledge of LLMOps frameworks and MLOps integration.
  • Familiarity with vector databases and retrieval systems for RAG architectures.

Responsibilities

  • Design and implement scalable infrastructure for LLM and GenAI workloads across multi-GPU environments.
  • Perform GPU profiling, benchmarking, and performance optimization for distributed training workloads.
  • Manage and schedule compute-intensive jobs using Slurm-based clusters and OpenShift/Kubernetes environments.
  • Enable and optimize the NVIDIA GPU stack (CUDA, cuDNN, NCCL, Triton, RAPIDS, etc.).
  • Collaborate with cross-functional teams to deploy models in research and production environments.
  • Build and support GenAI pipelines (fine-tuning, RAG, multi-modal inferencing, LLMOps).
  • Develop reusable infrastructure templates using tools like Terraform and Helm.
  • Contribute to internal innovation (PoCs, workshops) and support client-facing delivery engagements.

Skills

Slurm distributed training
LLMOps frameworks
MLOps integration
Multi-GPU optimization
Linux performance tuning

Tools

Slurm
OpenShift
Kubernetes
CUDA
cuDNN
NCCL
TensorRT
Nsight
Triton
Terraform
Ansible
GCP
AWS
Azure
OCI
NVIDIA NIMs/DGX
Vector databases

Job description

Quantiphi seeks a Senior Platform Engineer to design, optimize, and scale GenAI infrastructure. You will lead multi-GPU environments, GPU profiling, and distributed training, partnering with data science, MLOps, and application teams for production deployments.

Ideal candidates have hands-on OpenShift/Kubernetes, NVIDIA GPU stack expertise, and IaC skills. This remote US/Canada role offers growth, collaboration with Fortune 500 clients, and cutting-edge AI workloads.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GenAI Platform Engineer — Multi-GPU & MLOps
GenAI Platform Engineer — Multi-GPU & MLOps

quantiphi • Boston (MA)

Hybrid
USD 140,000 - 190,000
GenAI Platform Architect (Multi-GPU & MLOps)
GenAI Platform Architect (Multi-GPU & MLOps)

Quantiphi • Boston (MA)

Hybrid
USD 150,000 - 190,000
Platform Engineer - Senior - US
Platform Engineer - Senior - US

quantiphi • Boston (MA)

On-site
USD 140,000 - 190,000
Senior Platform Engineer
Senior Platform Engineer

Quantiphi • United States

On-site
USD 140,000 - 210,000
Senior Tech Architect - PE
Senior Tech Architect - PE

Quantiphi • Boston (MA)

Hybrid
USD 150,000 - 190,000
Compute Platform Lead: Multi-Cloud, GPU & Kubernetes
Compute Platform Lead: Multi-Cloud, GPU & Kubernetes

B Capital • San Francisco (CA)

On-site
USD 210,000 - 290,000
Top-tier compensation
Stock options
Comprehensive health/dental/vision
+5
GenAI Solutions Architect — GPU Cloud & AI Infra
GenAI Solutions Architect — GPU Cloud & AI Infra

NVIDIA • Town of Texas (WI)

On-site
USD 184,000 - 288,000
Equity
Benefits
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
Lead - Machine Learning
Lead - Machine Learning

Quantiphi • United States

Remote
USD 50,000 - 80,000
Access to state-of-the-art GPU infrastructure
Peer learning opportunities
Exposure to Fortune 500 leaders
Senior ML Platform Engineer — GPU, Kubernetes Infra
Senior ML Platform Engineer — GPU, Kubernetes Infra

WorkGenius Group • Los Angeles (CA)

On-site
USD 117,000 - 186,000