Platform Engineer - Senior - US

quantiphi

Boston (MA)

Hybrid

USD 140,000 - 190,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Quantiphi is seeking a Senior - Platform Engineer to design, optimize, and scale GenAI infrastructure. You’ll work across multi-GPU environments, profiling GPU performance and supporting production-grade deployments.

Collaboration with data science, MLOps, and application teams is key to delivering cutting-edge AI solutions. Ideal candidates have deep Linux expertise, hands-on with Slurm, OpenShift/Kubernetes, and NVIDIA GPU ecosystems, plus IaC skills with Terraform and Helm.

Qualifications

  • Strong experience with Slurm and distributed training environments.
  • Hands-on expertise with Red Hat OpenShift and/or Kubernetes.
  • Deep knowledge of the NVIDIA GPU ecosystem (CUDA, cuDNN, NCCL, Nsight, Triton/TensorRT).
  • Strong foundation in Linux systems, performance tuning, and multi-GPU optimization.
  • Experience deploying GenAI workloads (LLM fine-tuning, RAG pipelines, multi-modal systems).

Responsibilities

  • Design scalable infrastructure for LLM and GenAI workloads across multi-GPU environments.
  • Perform GPU profiling, benchmarking, and performance optimization for distributed training workloads.
  • Manage compute resources using Slurm clusters and OpenShift/Kubernetes environments.
  • Enable and optimize the NVIDIA GPU stack (CUDA, cuDNN, NCCL, Triton, TensorRT).
  • Collaborate with data science, MLOps and engineering teams to deploy models.
  • Build reusable infrastructure templates with Terraform and Helm.
  • Support PoCs and client-facing delivery engagements.

Skills

Slurm
OpenShift
Kubernetes
NVIDIA CUDA
GPU profiling
Distributed training

Tools

Terraform
Ansible
Helm
OpenShift tooling

Job description

While technology is the heart of our business, a global and diverse culture is the heart of our success. We love our people and we take pride in catering them to a culture built on transparency, diversity, integrity, learning and growth.

If working in an environment that encourages you to innovate and excel, not just in professional but personal life, interests you- you would enjoy your career with Quantiphi!

About Quantiphi:

Quantiphi is an award-winning, AI-First global digital engineering company that helps the world's leading Fortune 1000 organizations transform bold ideas into measurable business impact. We go beyond building innovative AI technologies, we solve the problems that matter most to our clients.

  • 21 Google Cloud Partner of the Year awards in the past 10 years
  • 3 AWS AI/ML Partner of the Year awards
  • 3 NVIDIA Partner of the Year awards
  • 3 Snowflake Partner of the Year awards

Rated Leaders by Gartner, Forrester, IDC, ISG, Everest Group and other leading analyst firms

Quantiphi delivers First-in-class AI solutions across Life Sciences, Healthcare, Banking, Financial Services, CPG, Manufacturing, Energy, High-Tech, Telecommunications, etc., powered by cutting-edge Generative AI and Agentic AI accelerators.

We are also proud to be certified as a Great Place to Work, reflecting our commitment to our people and our culture.

For more details, visit: Website or Linkedin Page

Role: Senior - Platform Engineer

Experience Level: 5 + yrs

Work Location: East/Canada [ET & CT]

Role Overview:

We are looking for a highly skilled Architect - Platform Engineer to design, optimize, and scale infrastructure for GenAI and LLM workloads . This role is ideal for someone with deep hands-on experience in GPU profiling, distributed training, and high-performance compute environments . You'll play a key role in building out GenAI platform foundations , supporting production-grade deployments, and partnering closely with data science, MLOps, and application teams to bring cutting-edge AI solutions to life.

Key Responsibilities:
  • Design and implement scalable infrastructure for LLM and GenAI workloads across multi-GPU environments
  • Perform GPU profiling, benchmarking, and performance optimization for distributed training workloads
  • Manage and schedule compute-intensive jobs using Slurm-based clusters and OpenShift/Kubernetes environments
  • Enable and optimize the NVIDIA GPU stack (CUDA, cuDNN, NCCL, Triton, RAPIDS, etc.)
  • Collaborate with cross-functional teams to deploy models in research and production environments
  • Build and support GenAI pipelines (fine-tuning, RAG, multi-modal inferencing, LLMOps)
  • Develop reusable infrastructure templates using tools like Terraform and Helm
  • Contribute to internal innovation (PoCs, workshops) and support client-facing delivery engagements
Basic Qualifications:
  • Strong experience with Slurm and distributed training environments
  • Hands-on expertise with Red Hat OpenShift and/or Kubernetes
  • Deep knowledge of the NVIDIA GPU ecosystem (CUDA, cuDNN, NCCL, Nsight, Triton/TensorRT)
  • Strong foundation in Linux systems, performance tuning, and multi-GPU optimization
  • Experience deploying GenAI workloads (LLM fine-tuning, RAG pipelines, multi-modal systems)
  • Familiarity with Infrastructure-as-Code tools (Terraform, Ansible)
  • Experience with cloud GPU environments (GCP, Azure, AWS, OCI) and/or on-prem GPU clusters
Other Qualifications (OQs):
  • Experience with NVIDIA NIMs, DGX systems, or GPU-accelerated containers
  • Knowledge of LLMOps frameworks and MLOps integration
  • Familiarity with vector databases and retrieval systems for RAG architectures
  • Comfortable working in client-facing environments and collaborating with AI solution teams
Healthcare Domain Experience (Nice to Have):
  • Experience working with FHIR R4, HL7 v2, or SMART on FHIR
  • Integration with EHR systems (e.g., Epic)
  • Understanding of HIPAA compliance and healthcare data privacy
  • Exposure to clinical workflows, CDS Hooks, or patient-facing applications
  • Experience building clinical decision support systems or healthcare interoperability solutions
What's in it for YOU at Quantiphi:

Make an impact at one of the world's fastest-growing AI-first digital engineering companies.

Upskill and discover your potential as you solve complex challenges in cutting-edge areas of technology alongside passionate, talented colleagues.

Work where innovation happens - work with disruptive innovators in a research-focused organization wi

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Platform Engineer - Senior - US
Platform Engineer - Senior - US

Quantiphi, Inc. • Boston (MA)

Hybrid
USD 150,000 - 210,000
Senior Tech Architect - PE
Senior Tech Architect - PE

Quantiphi, Inc. • United States

Hybrid
USD 150,000 - 190,000
Senior Platform Engineer
Senior Platform Engineer

Quantiphi • United States

On-site
USD 140,000 - 210,000
Senior Tech Architect - PE
Senior Tech Architect - PE

Quantiphi • Boston (MA)

Hybrid
USD 150,000 - 190,000
Architect - Software Developer
Architect - Software Developer

Quantiphi • Boston (MA)

On-site
CAD 100,000 - 130,000
Access to advanced GPU infrastructure
Strong peer learning opportunities
Exposure to Fortune 500 companies
Senior Software Engineer
Senior Software Engineer

Quantiphi • Boston (MA)

On-site
USD <1,000
Impactful AI projects
Growth opportunities
Competitive compensation
+1
GenAI Platform Engineer — Multi-GPU & MLOps
GenAI Platform Engineer — Multi-GPU & MLOps

quantiphi • Boston (MA)

Hybrid
USD 140,000 - 190,000
GenAI Platform Engineer: Multi-GPU & Kubernetes Lead
GenAI Platform Engineer: Multi-GPU & Kubernetes Lead

Quantiphi • United States

On-site
USD 140,000 - 210,000
Lead - Machine Learning
Lead - Machine Learning

Quantiphi • United States

Remote
USD 50,000 - 80,000
Access to state-of-the-art GPU infrastructure
Peer learning opportunities
Exposure to Fortune 500 leaders
Generative AI Architect
Generative AI Architect

Quantiphi • Boston (MA)

On-site
GBP 90,000 - 130,000