Hybrid AI HPC Infrastructure Engineer (GPU/ML)

Analysis Group, Inc.

Boston (MA)

On-site

USD 150,000 - 170,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Discretionary annual bonus
Benefits package

Job summary

Analysis Group, Inc. seeks an experienced AI HPC Infrastructure Engineer to own and operate a hybrid HPC/AI computing environment. You will manage Linux clusters, GPU fleets, and multi-tenant workloads to support researchers and data scientists.

You will optimize performance with MPI/OpenMP, implement scalable GPU training/inference, and build MLOps pipelines. Strong scripting, collaboration, and problem-solving are essential in a fast-paced research setting.

Qualifications

  • Bachelor's degree in computer science, electrical engineering, or related field.
  • 5+ years hands-on Linux systems administration in research/HPC/production.
  • Experience with Posit Workbench (RStudio Server Pro) and Python/R environments.
  • Experience with SLURM, Platform LSF, or other job schedulers; GPU scheduling preferred.
  • Hands-on NVIDIA GPU stack (CUDA, cuDNN, NCCL) and containerization knowledge.
  • Familiarity with ML frameworks (PyTorch, TensorFlow) and distributed training.
  • Experience with MLOps tools (MLflow, Kubeflow).
  • Proficiency in remote access (SSH/RDP) and GPFS.
  • Excellent troubleshooting, documentation, and collaboration skills.

Responsibilities

  • Maintain and expand HPC/AI compute environment for researchers and data scientists.
  • Optimize performance using MPI/OpenMP and distributed multi-GPU training.
  • Design and maintain GPU-accelerated infrastructure for large-scale model training.
  • Manage GPU scheduling across SLURM/Kubernetes multi-tenant clusters.
  • Administer NVIDIA stack and health monitoring for GPU fleets.
  • Tune LLM training/inference (batching, quantization, KV-cache).
  • Build MLOps pipelines (MLflow, Kubeflow) for training and deployment.
  • Manage containers (Docker, Kubernetes, Singularity) for HPC/ML workloads.
  • Develop automation scripts and usage reporting across resources.
  • Ensure storage and data pipelines on GPFS are robust and performant.
  • Troubleshoot hardware/software/network issues and perform root-cause analysis.
  • Participate in 24x7 on-call rotation and provide remote/on-site support.

Education

Bachelor's degree in CS/EE

Tools

Posit Workbench
GPFS
Docker
Kubernetes
Singularity/Apptainer
MLflow
Kubeflow
Ansible

Job description

Analysis Group, Inc. seeks an experienced AI HPC Infrastructure Engineer to own and operate a hybrid HPC/AI computing environment. You will manage Linux clusters, GPU fleets, and multi-tenant workloads to support researchers and data scientists.

You will optimize performance with MPI/OpenMP, implement scalable GPU training/inference, and build MLOps pipelines. Strong scripting, collaboration, and problem-solving are essential in a fast-paced research setting.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI/HPC Infrastructure Engineer: GPU Compute & Hybrid Cloud
AI/HPC Infrastructure Engineer: GPU Compute & Hybrid Cloud

Saigepartners • San Jose (CA)

Hybrid
USD 120,000 - 180,000
AI HPC Infrastructure Engineer
AI HPC Infrastructure Engineer

Analysis Group, Inc. • Boston (MA)

On-site
USD 150,000 - 170,000
Discretionary annual bonus
Benefits package
Senior GPU HPC Infrastructure Engineer for ML Pipelines
Senior GPU HPC Infrastructure Engineer for ML Pipelines

Jaide Health • United States

Hybrid
USD 180,000 - 250,000
Weekly lunch stipend
Health and dental benefits
Parental leave
+3
Hybrid HPC Systems Architect - GPU Cloud for AI
Hybrid HPC Systems Architect - GPU Cloud for AI

The Consensus • San Jose (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Cash compensation
Equity compensation
Health, dental and vision coverage
+1
AI HPC Systems Engineer: GPU Clusters & ML Platforms
AI HPC Systems Engineer: GPU Clusters & ML Platforms

AMD • San Jose (CA)

On-site
USD 140,000 - 210,000
Hybrid HPC Systems Administrator for AI/ML
Hybrid HPC Systems Administrator for AI/ML

Eli Lilly and Company • San Francisco (CA)

Hybrid
USD 141,000 - 231,000
Bonus program
401(k) plan
Health, dental, vision insurance
+1
Hybrid AI/HPC Systems Administrator - GPU & Cloud
Hybrid AI/HPC Systems Administrator - GPU & Cloud

Eli Lilly and Company • South San Francisco (CA)

Hybrid
USD 141,000 - 231,000
Hybrid work schedule
Comprehensive benefits
HPC & AI Systems Administrator for GPU-Driven ML
HPC & AI Systems Administrator for GPU-Driven ML

Scorpion Therapeutics • South San Francisco (CA)

Hybrid
USD 140,000 - 200,000
AI & HPC GPU Compute Performance Engineer
AI & HPC GPU Compute Performance Engineer

engineeringjobs.net, Inc. • San Jose (CA)

On-site
USD 150,000 - 190,000
Medical Insurance
Dental Insurance
Vision Insurance
+16
Senior GPU Infrastructure Engineer — HPC & Clusters
Senior GPU Infrastructure Engineer — HPC & Clusters

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000