HPC (High-Performance Computing) Consultant @ Remote

BURGEON IT SERVICES LLC

United States

Remote

USD 120,000 - 180,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

BURGEON IT SERVICES LLC is seeking an experienced HPC Consultant to design and operate large-scale Kubernetes and HPC/GPU infrastructure in a remote, full-time role.

The ideal candidate will own infrastructure automation, cloud deployments, and SRE practices for AI/ML workloads, with strong experience in Slurm, NVIDIA GPUs, and multi-cloud environments.

Qualifications

  • 10+ years of hands-on HPC/Kubernetes experience required.
  • Proven ability to manage large Kubernetes clusters and upgrades.
  • Deep knowledge of Slurm, GPU scheduling, and HPC/GPU workloads.

Responsibilities

  • Operate and scale large Kubernetes clusters (up to 500 nodes).
  • Design and maintain HPC/GPU infrastructure for AI/ML workloads.
  • Automate infrastructure with Terraform and cloud tools.
  • Ensure SRE practices: monitoring, alerting, incident response, RCA.

Skills

Kubernetes large-scale production
HPC / High-Performance Computing
Slurm
NVIDIA GPU infrastructure
AI/ML infrastructure
Python automation
Bash/Shell scripting
SRE monitoring & incident response

Tools

Terraform
AWS (EKS, EC2, S3, FSx Lustre)
Prometheus/Grafana

Job description

Position

HPC (High-Performance Computing) Consultant

Location

Remote

Employment

Full-Time | No Contractors

Experience

10 Years

Role Summary

We are seeking an experienced HPC / Kubernetes / Cloud Infrastructure Engineer with strong hands-on experience supporting large-scale Kubernetes environments, HPC/GPU infrastructure, AI/ML workloads, cloud platforms, infrastructure automation, and SRE operations.

Mandatory Skills
  • Kubernetes large-scale production environments, cluster lifecycle, node management, upgrades, troubleshooting
  • Kubernetes Troubleshooting Scheduler, CNI/networking, nodes, storage, and cluster issues
  • HPC / High-Performance Computing
  • Slurm
  • NVIDIA GPU Infrastructure and GPU scheduling
  • AI/ML Infrastructure and GPU-based workloads
  • AWS especially EKS, EC2, VPC, IAM, S3, and FSx for Lustre
  • Terraform / Infrastructure as Code (IaC)
  • Python automation and coding
  • Bash/Shell scripting
  • Monitoring & SRE PrometheGrafana or equivalent, alerting, SLI/SLO, incident response, RCA
  • Kubernetes Networking Calico or Cilium
  • Helm, RBAC, autoscaling, and rolling upgrades
  • Experience supporting large-scale Kubernetes clusters (500 nodes) is strongly preferred.
Preferred Skills
  • Multi-cloud: AWS, Google Cloud Platform, CoreWeave
  • Karpenter
  • CUDA
  • AWS ParallelCluster
  • Lustre / WekaFS
  • InfiniBand / RDMA
  • Kubeflow, KServe, Ray, MLflow, or vLLM
  • Distributed AI/ML training and inference platforms
Ideal Candidate

Strong hands-on Kubernetes Slurm NVIDIA GPU AI/ML Infrastructure Terraform Python AWS/EKS/FSx for Lustre SRE/Monitoring experience.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Platform Architect - HPC, Kubernetes
Platform Architect - HPC, Kubernetes

EPAM Systems • United States

On-site
USD 140,000 - 230,000
Lead Platform Engineer/Architect - HPC, Kubernetes
Lead Platform Engineer/Architect - HPC, Kubernetes

EPAM Systems • United States

On-site
USD 150,000 - 230,000
HPC Infrastructure Engineer
HPC Infrastructure Engineer

Arcadia • San Francisco (CA)

On-site
USD 180,000 - 260,000
Heavy AWS + HPC
Heavy AWS + HPC

Zeal Solutions Inc • Charlotte (NC)

Remote
USD 120,000 - 180,000
HPC Solution Architect
HPC Solution Architect

Coda Search│Staffing • Dallas (TX)

On-site
USD 120,000 - 160,000
GPU / HPC Consultant
GPU / HPC Consultant

Arke • United States

On-site
USD 120,000 - 180,000
HPC Observability Engineer
HPC Observability Engineer

EIT Professionals Corp • United States

On-site
USD 100,000 - 130,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda Innovation • California (MO)

Hybrid
USD 180,000 - 240,000
HPC Customer Solutions Engineer
HPC Customer Solutions Engineer

GTN Technical Staffing • United States

On-site
USD 120,000 - 180,000
HPC AI Systems Administrator
HPC AI Systems Administrator

MRE Consulting • Houston (TX)

On-site
USD 120,000 - 180,000
Competitive salary
Comprehensive benefits
Professional development support