Senior GPU HPC Engineer for AI Cloud

Verda

United States

Remote

USD 120,000 - 170,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Cash and equity compensation
Healthcare
Lunch

Job summary

Verda is building a full-stack AI cloud, from data centers to our cloud platform, used by leading AI teams worldwide. As an HPC Engineer you will own baremetal GPU/HPC clusters, InfiniBand fabric, and workload orchestration to keep systems healthy and fast.

You'll collaborate with platform, network, and storage teams, design scalable fabrics, deploy shared filesystems, and participate in on-call rotations to ensure reliability for researchers and customers.

Qualifications

  • Solid Linux systems experience with memory management and virtualization basics.
  • Deep InfiniBand knowledge for fabric design, topology planning, and performance tuning.
  • Experience with HPC clusters, benchmarking and day-two operations.

Responsibilities

  • Administer baremetal GPU/HPC clusters end to end, from provisioning to day-two ops.
  • Administer virtualized clusters, including hypervisor and GPU virtualization.
  • Design, deploy, and tune InfiniBand fabrics and topology.
  • Deploy and operate shared filesystems for training/inference workloads.
  • Troubleshoot across the full fabric stack: fibers, NICs, switches, drivers, firmware.
  • Collaborate with remote-hands and data center teams for hardware fixes.
  • Operate Slurm/workload scheduling deployments and keep records up to date.
  • Participate in on-call rotations and incident response for cluster issues.
  • Coordinate with platform, network, and storage teams to integrate new clusters.

Skills

Linux systems
InfiniBand
HPC clustering
Shared filesystems
Slurm
NCCL/CUDA/DOCA
Scripting (Python/Bash)
Remote hands coordination
PCIe/topology
Virtualization basics

Tools

SLURM
Lustre
GPFS/Spectrum Scale
WekaFS
KVM/QEMU with GPU passthrough
MAAS/Foreman

Job description

Verda is building a full-stack AI cloud, from data centers to our cloud platform, used by leading AI teams worldwide. As an HPC Engineer you will own baremetal GPU/HPC clusters, InfiniBand fabric, and workload orchestration to keep systems healthy and fast.

You'll collaborate with platform, network, and storage teams, design scalable fabrics, deploy shared filesystems, and participate in on-call rotations to ensure reliability for researchers and customers.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

On-Site Junior Data Center Technician: GPU & Cabling
On-Site Junior Data Center Technician: GPU & Cabling

Verda • Santa Clara (UT)

On-site
USD 60,000 - 76,000
Healthcare
Lunch
Wellbeing benefits
+1
GPU Platform Engineer for AI/ML Infra
GPU Platform Engineer for AI/ML Infra

Vero • United States

On-site
USD 136,000 - 160,000
Medical, dental, and vision insurance
Equity Scheme
401(k) with employer match
+3
Senior HPC Engineer - GPU Compute & InfiniBand
Senior HPC Engineer - GPU Compute & InfiniBand

Nebius • United States

Remote
USD 150,000 - 230,000
Senior HPC Cloud Engineer — GPU/InfiniBand AI Infra
Senior HPC Cloud Engineer — GPU/InfiniBand AI Infra

Jobgether • Germany (OH)

On-site
USD 81,000 - 105,000
Career development opportunities
Flexible working arrangements
Collaborative engineering environment
+1
AI Infrastructure Engineer: HPC GPU Clusters
AI Infrastructure Engineer: HPC GPU Clusters

Veeda AI • Seattle (WA)

On-site
USD 180,000 - 240,000
AI Infrastructure Architect for HPC & GPU Clusters
AI Infrastructure Architect for HPC & GPU Clusters

Veeda • California (MO)

On-site
USD 150,000 - 200,000
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm

Blue Signal Search • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual bonus
Equity participation
Comprehensive benefits
+1
Senior GPU Cloud Infrastructure Engineer
Senior GPU Cloud Infrastructure Engineer

Hyperbolic • San Francisco (CA)

On-site
USD 180,000 - 240,000
Founding HPC Engineer - GPU Cloud Infra & AI Orchestration
Founding HPC Engineer - GPU Cloud Infra & AI Orchestration

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 235,000 - 315,000
Founding engineer equity
Full benefits package
Senior Data Center Engineer - GPU/AI Infra
Senior Data Center Engineer - GPU/AI Infra

Nebius B.V. • Vineland (NJ)

On-site
USD 85,000 - 140,000
Career growth
Flexibility and ownership
Collaborative culture
+2