Hybrid AI HPC Infrastructure Engineer (GPU/ML)

Analysis Group, Inc.

Boston (MA)

On-site

USD 150,000 - 170,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Discretionary annual bonus
Benefits package

Job summary

Analysis Group, Inc. seeks an experienced AI HPC Infrastructure Engineer to own and operate a hybrid HPC/AI computing environment. You will manage Linux clusters, GPU fleets, and multi-tenant workloads to support researchers and data scientists.

You will optimize performance with MPI/OpenMP, implement scalable GPU training/inference, and build MLOps pipelines. Strong scripting, collaboration, and problem-solving are essential in a fast-paced research setting.

Qualifications

  • Bachelor's degree in computer science, electrical engineering, or related field.
  • 5+ years hands-on Linux systems administration in research/HPC/production.
  • Experience with Posit Workbench (RStudio Server Pro) and Python/R environments.
  • Experience with SLURM, Platform LSF, or other job schedulers; GPU scheduling preferred.
  • Hands-on NVIDIA GPU stack (CUDA, cuDNN, NCCL) and containerization knowledge.
  • Familiarity with ML frameworks (PyTorch, TensorFlow) and distributed training.
  • Experience with MLOps tools (MLflow, Kubeflow).
  • Proficiency in remote access (SSH/RDP) and GPFS.
  • Excellent troubleshooting, documentation, and collaboration skills.

Responsibilities

  • Maintain and expand HPC/AI compute environment for researchers and data scientists.
  • Optimize performance using MPI/OpenMP and distributed multi-GPU training.
  • Design and maintain GPU-accelerated infrastructure for large-scale model training.
  • Manage GPU scheduling across SLURM/Kubernetes multi-tenant clusters.
  • Administer NVIDIA stack and health monitoring for GPU fleets.
  • Tune LLM training/inference (batching, quantization, KV-cache).
  • Build MLOps pipelines (MLflow, Kubeflow) for training and deployment.
  • Manage containers (Docker, Kubernetes, Singularity) for HPC/ML workloads.
  • Develop automation scripts and usage reporting across resources.
  • Ensure storage and data pipelines on GPFS are robust and performant.
  • Troubleshoot hardware/software/network issues and perform root-cause analysis.
  • Participate in 24x7 on-call rotation and provide remote/on-site support.

Education

Bachelor's degree in CS/EE

Tools

Posit Workbench
GPFS
Docker
Kubernetes
Singularity/Apptainer
MLflow
Kubeflow
Ansible

Job description

Analysis Group, Inc. seeks an experienced AI HPC Infrastructure Engineer to own and operate a hybrid HPC/AI computing environment. You will manage Linux clusters, GPU fleets, and multi-tenant workloads to support researchers and data scientists.

You will optimize performance with MPI/OpenMP, implement scalable GPU training/inference, and build MLOps pipelines. Strong scripting, collaboration, and problem-solving are essential in a fast-paced research setting.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI HPC Infrastructure Engineer
AI HPC Infrastructure Engineer

Analysis Group, Inc. • Boston (MA)

On-site
USD 150,000 - 170,000
Discretionary annual bonus
Benefits package
HPC AI Systems Architect (On-Prem GPU Cluster)
HPC AI Systems Architect (On-Prem GPU Cluster)

MRE Consulting • Houston (TX)

On-site
USD 95,000 - 140,000
Hybrid HPC Systems Architect - GPU Cloud for AI
Hybrid HPC Systems Architect - GPU Cloud for AI

The Consensus • San Jose (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Cash compensation
Equity compensation
Health, dental and vision coverage
+1
AI/ML Engineer for HPC & Distributed Systems (Semi-Remote)
AI/ML Engineer for HPC & Distributed Systems (Semi-Remote)

Hewlett Packard Enterprise • Spring (TX)

On-site
USD 121,000 - 277,000
Senior AI Factory Architect — Multi-GPU HPC, NCCL, Equity
Senior AI Factory Architect — Multi-GPU HPC, NCCL, Equity

NVIDIA • California (MO)

On-site
USD 152,000 - 288,000
Equity
Benefits
AI Systems Engineer: HPC & GPU Clusters
AI Systems Engineer: HPC & GPU Clusters

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 180,000 - 260,000
Hybrid AI & HPC ML Engineer
Hybrid AI & HPC ML Engineer

Hewlett Packard Enterprise Company • Springs (NY)

Hybrid
USD 121,000 - 277,000
Health & Wellbeing
Professional development
Unconditional inclusion
Senior AI Infrastructure & HPC Ops Manager
Senior AI Infrastructure & HPC Ops Manager

Allen Institute for Artificial Intelligence • Seattle (WA)

On-site
USD 146,000 - 221,000
Medical, dental, vision
401k plan
Commuting stipend
+3
AI Infrastructure Engineer - GPU & HPC Expert
AI Infrastructure Engineer - GPU & HPC Expert

Socket.dev • Town of Florida (NY)

On-site
USD 87,000 - 266,000
Lead GPU Systems Engineer - HPC & AI Infrastructure
Lead GPU Systems Engineer - HPC & AI Infrastructure

Socket.dev • New York (NY)

Hybrid
USD 200,000 - 300,000
Hybrid working opportunities
Generous PTO
Wellness programs
+2