Senior GPU Data Center Operations Engineer

Crusoe

Denver (CO)

On-site

USD 150,000 - 170,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Paid time off
Health insurance
401(k)

Job summary

Crusoe is building next-generation AI infrastructure with GPU-powered compute at scale. We are seeking a Staff Data Center Operations Engineer to own site-level escalations, drive platform-wide improvements, and partner with OEM/ODM teams for reliable fleet operations.

The role is based at Crusoe Cloud's Denver headquarters with cross-site travel across the US and internationally. You will own platform standards, SOPs, and training to scale SiteOps capabilities, while traveling up to 30% for

Qualifications

  • 7+ years in data center operations, field engineering, or OEM/ODM technical support with hands-on GPU infrastructure experience.
  • Direct hands-on experience deploying and supporting GPU platforms at scale across one or more major OEMs or ODMs; familiarity with SuperMicro and HPE platforms.
  • Deep familiarity with server platform architecture and OEM escalation and RMA processes.
  • Experience leading or contributing to large-scale GPU cluster bring-ups including rack staging and production handoff.
  • Demonstrated ability to build technical relationships with OEM and ODM engineering teams and drive platform-level issue resolution.
  • Experience developing SOPs, runbooks, or field troubleshooting procedures and delivering technical training to data center technician teams.
  • Strong written communication — comfortable producing escalation documentation, platform runbooks, and leadership reporting.
  • Willingness to travel domestically and internationally to Crusoe sites as needed (target: up to 30%)

Responsibilities

  • Own Tier 2/3 hardware escalations across all Crusoe sites for issues that exceed local site capability, engaging directly with OEM and ODM engineering teams to drive resolution.
  • Travel to sites as needed for complex platform issues, new hardware bring-ups, and deployment support.
  • Identify recurring failure patterns across sites and translate them into platform feedback, sparing strategy inputs, or OEM improvement requests.
  • Root‑cause complex hardware issues — PCIe, BMC, thermal, fabric — and produce resolution documentation reusable across the SiteOps org.
  • Hand off platform-level findings to the appropriate internal engineering teams with clear, well-documented escalation packages.
  • Develop and maintain deep technical relationships with Crusoe's hardware partners — currently SuperMicro and HPE — at the engineering and field escalation level.
  • Serve as Crusoe's technical voice in OEM/ODM partner conversations, surfacing field observations, influencing hardware roadmaps, and driving platform improvements.
  • Build familiarity with new ODM platform architecture, tooling, and escalation processes as Crusoe expands its ODM footprint.
  • Support vendor evaluations and new platform qualifications in partnership with SiteOps and engineering leadership.
  • Own OEM platform technical knowledge at the SiteOps org level — escalation playbooks, failure pattern analysis, and OEM relationship inputs across all sites.
  • Own the development and maintenance of platform-specific SOPs, runbooks, and field troubleshooting procedures for the SiteOps org, ensuring site teams have current, actionable documentation across all active hardware platforms.
  • Design and deliver technical training for SiteOps technicians covering hardware architecture, platform-specific troubleshooting, and field procedures — both for new hire onboarding and ongoing skill development.
  • Contribute to the technician certification program and technical leveling standards across the org.
  • Support new site bring-up efforts providing platform readiness and deployment execution expertise.
  • Serve as SiteOps' senior technical representative at Crusoe HQ, participating in platform, engineering, and procurement discussions that affect site operations.
  • Partner with engineering and procurement teams on sparing strategy, RMA lifecycle management, and OEM/ODM support contract structures.
  • Provide operational input into next-generation GPU platform evaluations (GB300, VR200, and beyond).
  • Produce escalation reporting, platform health analysis, and operational insights for SiteOps leadership.

Skills

GPU infrastructure
Data center operations
OEM/ODM collaboration
Technical writing
Travel readiness

Tools

SuperMicro hardware
HPE platforms

Job description

Crusoe is building next-generation AI infrastructure with GPU-powered compute at scale. We are seeking a Staff Data Center Operations Engineer to own site-level escalations, drive platform-wide improvements, and partner with OEM/ODM teams for reliable fleet operations.

The role is based at Crusoe Cloud's Denver headquarters with cross-site travel across the US and internationally. You will own platform standards, SOPs, and training to scale SiteOps capabilities, while traveling up to 30% for

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Data Center Infrastructure Engineer
Senior GPU Data Center Infrastructure Engineer

CV in • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 220,000
Health insurance
RSUs
401(k) match
+2
Senior Software Engineer — GPU Data Center Automation
Senior Software Engineer — GPU Data Center Automation

Crusoe • San Francisco (CA)

On-site
USD 170,000 - 205,000
Health insurance package options
Restricted Stock Units
401(k) with match up to 4%
+2
Senior GPU Infra Engineer - Automation & Diagnostics
Senior GPU Infra Engineer - Automation & Diagnostics

Crusoe • United States

On-site
USD 250,000 - 300,000
Industry competitive pay
RSUs in a fast-growing tech company
Health insurance with family options
+3
Staff Data Center Operations Engineer
Staff Data Center Operations Engineer

Crusoe • Denver (CO)

On-site
USD 150,000 - 170,000
Equity
Paid time off
Health insurance
+1
Senior GPU Infra Engineer — AI Data Center Automation
Senior GPU Infra Engineer — AI Data Center Automation

Crusoe Energy Systems LLC • San Francisco (CA), Northern (KY)

Hybrid
USD 215,000 - 260,000
Industry competitive pay
RSUs
Health insurance
+12
Senior GPU DC Infra Engineer — Automation & Diagnostics
Senior GPU DC Infra Engineer — Automation & Diagnostics

Crusoe • San Francisco (CA)

On-site
USD 215,000 - 260,000
Industry competitive pay
Restricted Stock Units
Health insurance options (HDHP/PPO)
+11
Data Center Ops Manager: GPU AI Infra (On-Site)
Data Center Ops Manager: GPU AI Infra (On-Site)

crusoe • Town of Springfield (WI)

On-site
USD 135,000 - 175,000
Competitive compensation and equity
Paid time off, holidays, and leave programs
Medical, dental, and vision insurance
+2
GPU Infrastructure Software Engineer I
GPU Infrastructure Software Engineer I

crusoe • San Francisco (CA)

On-site
USD 117,000 - 135,000
Health insurance
401(k) with match
Stock options/RSUs
+3
Senior HPC Cloud Support Engineer
Senior HPC Cloud Support Engineer

Crusoe • Dallas (TX)

On-site
USD 130,000 - 155,000
Equity
Paid time off
Health insurance
+2
Senior Software Engineer, GPU Fleet Reliability
Senior Software Engineer, GPU Fleet Reliability

Crusoe Energy Systems • San Francisco (CA)

On-site
USD 180,000 - 300,000
Health benefits
Paid time off
401(k) match
+1