Hybrid GPU Fleet Operations Engineer

Crusoe

San Francisco (CA)

Hybrid

USD 215,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Hybrid work schedule
Industry competitive pay
Restricted Stock Units
Health insurance
401(k) with match

Job summary

Crusoe is seeking a GPU Fleet Operations Engineer to join our Fleet Operations team. You will diagnose, maintain and repair high-density GPU clusters across our data center fleet, ensuring uptime and performance.

The role requires hands-on hardware troubleshooting, firmware updates, and close collaboration with data center operations and engineering. The ideal candidate has strong Linux skills, experience with NVIDIA A100/H200 GPUs, and proficiency in Go, with a track record of working in

Qualifications

  • Hands-on experience diagnosing and repairing GPU rack-mounted hardware.
  • Experience with NVIDIA A100, H200, GB200, B200, B300, and AMD 350X/355X.
  • Strong Linux experience and command-line diagnostics.
  • Ability to work in fast-paced data center environments.

Responsibilities

  • Automate deep-level diagnosis and troubleshooting of hardware faults within GPU racks and high-density compute systems.
  • Develop software to troubleshoot and support GPU platforms including NVIDIA A100, H200, GB200, B200 and AMD 350X / 355X.
  • Execute component-level diagnosis and remediation for failed or degraded hardware.
  • Partner with data center operations to manage and perform FRU repairs for GPUs, power supplies, cooling systems, interconnects, and networking hardware.
  • Conduct post-repair validation, burn-in testing, torch testing, and NVIDIA NCCL testing to ensure system stability and performance.
  • Implement preventative maintenance procedures to improve fleet reliability and extend hardware lifespan.
  • Perform firmware and BIOS upgrades across the GPU fleet.
  • Maintain detailed documentation of maintenance activities, failures, and resolutions in ticketing and asset management systems.
  • Develop and update standard operating procedures (SOPs) for troubleshooting, repair, and validation workflows.
  • Collaborate with engineering, software, and data center operations teams to identify root causes of systemic failures and implement preventative solutions.
  • Participate in a rotating infrastructure on-call schedule (about one week every 4–6 weeks) with daytime coverage and handoff to the Europe team.

Skills

Golang
Linux administration
GPU architectures
Diagnostics & repair
Communication

Education

Bachelor's degree in Electrical Engineering or Computer Science

Tools

NVIDIA DCGM
NVIDIA field diagnostic utilities
InfiniBand
NVLink
RoCE

Job description

Crusoe is seeking a GPU Fleet Operations Engineer to join our Fleet Operations team. You will diagnose, maintain and repair high-density GPU clusters across our data center fleet, ensuring uptime and performance.

The role requires hands-on hardware troubleshooting, firmware updates, and close collaboration with data center operations and engineering. The ideal candidate has strong Linux skills, experience with NVIDIA A100/H200 GPUs, and proficiency in Go, with a track record of working in

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Infrastructure Software Engineer I
GPU Infrastructure Software Engineer I

crusoe • San Francisco (CA)

On-site
USD 117,000 - 135,000
Health insurance
401(k) with match
Stock options/RSUs
+3
Senior GPU Infra Engineer - Automation & Diagnostics
Senior GPU Infra Engineer - Automation & Diagnostics

Crusoe • United States

On-site
USD 250,000 - 300,000
Industry competitive pay
RSUs in a fast-growing tech company
Health insurance with family options
+3
Staff Software Engineer (Cloud Infrastructure)
Staff Software Engineer (Cloud Infrastructure)

Crusoe • San Francisco (CA)

Hybrid
USD 215,000 - 260,000
Hybrid work schedule
Industry competitive pay
Restricted Stock Units
+2
Senior GPU DC Infra Engineer — Automation & Diagnostics
Senior GPU DC Infra Engineer — Automation & Diagnostics

Crusoe • San Francisco (CA)

On-site
USD 215,000 - 260,000
Industry competitive pay
Restricted Stock Units
Health insurance options (HDHP/PPO)
+11
Senior Software Engineer — GPU Data Center Automation
Senior Software Engineer — GPU Data Center Automation

Crusoe • San Francisco (CA)

On-site
USD 170,000 - 205,000
Health insurance package options
Restricted Stock Units
401(k) with match up to 4%
+2
Senior GPU Data Center Operations Engineer
Senior GPU Data Center Operations Engineer

Crusoe • Denver (CO)

On-site
USD 150,000 - 170,000
Equity
Paid time off
Health insurance
+1
Senior GPU Data Center Infrastructure Engineer
Senior GPU Data Center Infrastructure Engineer

CV in • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 220,000
Health insurance
RSUs
401(k) match
+2
Senior GPU Infra Engineer — AI Data Center Automation
Senior GPU Infra Engineer — AI Data Center Automation

Crusoe Energy Systems LLC • San Francisco (CA), Northern (KY)

Hybrid
USD 215,000 - 260,000
Industry competitive pay
RSUs
Health insurance
+12
GPU Fleet Automation Engineer (Hybrid)
GPU Fleet Automation Engineer (Hybrid)

Tower Research Capital • New York (NY)

Hybrid
USD 200,000 - 300,000
Generous PTO
Hybrid work
Free meals
Software Engineer - GPU Fleet
Software Engineer - GPU Fleet

Iceberg • New York (NY)

On-site
USD 120,000 - 170,000