GPU Cluster Infrastructure Engineer

Smartshare, Inc.

New York, Northern (NY, KY)

Hybrid

USD 165,000 - 303,000

Full time

6 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Competitive salary and meaningfulequiy
Health, dental, vision benefits
Fitness stipend and learning budget
Industry events in cloud native
Opportunities for equity

Job summary

Beam is seeking an experienced contractor to help stand up high-performance GPU clusters. The role spans design review through bring-in, establishing the operational foundation our team needs to run them.

You will review designs, lead acceptance testing, stand up storage, secure the management plane, and integrate telemetry into our observability stack. You will also document runbooks and provide post-go-live escalation support.

Qualifications

  • Experience building and operating NVIDIA HGX/DGX clusters in production.
  • Hands-on InfiniBand subnet management, fabric bring-up, and diagnosing degraded links.
  • GPU node bring-up including firmware, BMC/Redfish, DCGM, and NCCL testing.
  • Parallel storage experience: WEKA, VAST, GPFS, Lustre, or similar.
  • Ability to work on-site and remotely with strong documentation.
  • Bonus: familiarity with automation and monitoring tools (Prometheus/Grafana).

Responsibilities

  • Review cluster designs and BOMs across compute, networking, and storage.
  • Lead acceptance testing: cabling, optics, InfiniBand bring-up, burn-in.
  • Stand up and validate high-performance storage with vendor teams.
  • Build the out-of-band management layer and firmware baselines.
  • Secure the management plane for customer environments.
  • Integrate telemetry into observability stack with alerts and health checks.
  • Write runbooks, as-builts, and remote-hands procedures.
  • Provide escalation support after go-live.

Skills

NVIDIA HGX/DGX
InfiniBand
BMC/Redfish
PXE imaging
NCCL testing
Parallel storage tech

Tools

Ansible
Prometheus
Grafana
NVIDIA certifications

Job description

Beam is an ultrafast AI inference platform. We built a serverless runtime that launches GPU-backed containers in less than 1 second and quickly scales out to thousands of GPUs. Developers use our platform to serve apps to millions of users around the globe. We're backed by Y Combinator, Tiger Global, and prominent developer-tool founders, including the founder of Snyk and former CTO of GitHub.

About the Role

We're building out our own GPU capacity and we're looking for an experienced contractor to help us stand up high-performance GPU clusters. The work runs from design review through bring-in, and you'll leave behind the operational foundation our team needs to run them.

  • Review cluster designs and bills of materials across compute, networking, and storage, and catch gaps before hardware is ordered.
  • Lead acceptance testing: validate cabling and optics, bring up the InfiniBand fabric, run burn-in, and hold vendors to their deliverables.
  • Stand up and validate high-performance storage alongside vendor teams.
  • Build the out-of-band management layer and firmware baselines, and secure the management plane for customer-facing environments.
  • Integrate hardware, fabric, and storage telemetry into our observability stack, with alerting and automated health checks.
  • Write runbooks, as-builts, and remote-hands procedures.
  • Provide escalation support after go-live and help our team ramp up.

Skills & Experience

  • You've built and operated NVIDIA HGX or DGX clusters in production at a GPU cloud, HPC center, or AI lab.
  • Hands-on experience with InfiniBand: subnet management and UFM, fabric bring-up, and diagnosing degraded links and optics. NDR or newer.
  • GPU node bring-up and burn-in: firmware, BMC/Redfish, DCGM, NCCL testing, PXE and imaging, and XID error triage.
  • Parallel storage experience: WEKA, VAST, GPFS, Lustre, or similar.
  • Equally effective on the data center floor and remotely, including directing colo remote hands.
  • You troubleshoot methodically across hardware, fabric, and software, document as you go, and communicate clearly with technical and non-technical people.
  • Bonus: recent-generation NVIDIA platforms, bare-metal cloud operations, Ansible or similar automation, Prometheus/Grafana, NVIDIA certifications.
Benefits
  • Competitive salary and meaningful equity
  • Join a fast-growing pre-series A company at the ground floor
  • Health, dental, and vision benefits with 90% coverage for you and 50% for dependents
  • Opportunities to participate in events across the cloud native community
  • Fitness stipend, learning budget, and much, much more
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior GPU Cluster Engineer — HPC Infra, InfiniBand
Senior GPU Cluster Engineer — HPC Infra, InfiniBand

Smartshare, Inc. • New York (NY), Northern (KY)

Hybrid
USD 165,000 - 303,000
Competitive salary and meaningfulequiy
Health, dental, vision benefits
Fitness stipend and learning budget
+2
Site Reliability Engineer
Site Reliability Engineer

Beam • San Francisco (CA)

On-site
USD 140,000 - 180,000
Competitive salary
Meaningful equity
Health, dental, vision benefits
+3
Network Engineer
Network Engineer

Beam • New York (NY)

On-site
USD 120,000 - 180,000
Health, dental, vision benefits
Equity
Learning budget
Network Engineer
Network Engineer

Beam • San Francisco (CA)

On-site
USD 120,000 - 180,000
Competitive salary
Equity
Health/dental/vision
+2
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 300,000
Equity incentives
Distributed Systems Engineer
Distributed Systems Engineer

Beam • New York (NY)

On-site
USD 120,000 - 170,000
Competitive salary
Equity
Health, dental, vision
+2
Distributed Systems Engineer
Distributed Systems Engineer

Beam • San Francisco (CA)

On-site
USD 140,000 - 200,000
Competitive compensation
Equity
Health, dental, vision benefits
+3
HPC Infrastructure Engineer
HPC Infrastructure Engineer

Arcadia • San Francisco (CA)

On-site
USD 180,000 - 260,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Remote - Lead GPU Cluster Solutions Architect
Remote - Lead GPU Cluster Solutions Architect

Orion Placement • United States

Remote
USD 150,000 - 210,000
Bonus
Equity