Staff Data Center Operations Engineer

Crusoe

Denver (CO)

On-site

USD 150,000 - 170,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Paid time off
Health insurance
401(k)

Job summary

Crusoe is building next-generation AI infrastructure with GPU-powered compute at scale. We are seeking a Staff Data Center Operations Engineer to own site-level escalations, drive platform-wide improvements, and partner with OEM/ODM teams for reliable fleet operations.

The role is based at Crusoe Cloud's Denver headquarters with cross-site travel across the US and internationally. You will own platform standards, SOPs, and training to scale SiteOps capabilities, while traveling up to 30% for

Qualifications

  • 7+ years in data center operations, field engineering, or OEM/ODM technical support with hands-on GPU infrastructure experience.
  • Direct hands-on experience deploying and supporting GPU platforms at scale across one or more major OEMs or ODMs; familiarity with SuperMicro and HPE platforms.
  • Deep familiarity with server platform architecture and OEM escalation and RMA processes.
  • Experience leading or contributing to large-scale GPU cluster bring-ups including rack staging and production handoff.
  • Demonstrated ability to build technical relationships with OEM and ODM engineering teams and drive platform-level issue resolution.
  • Experience developing SOPs, runbooks, or field troubleshooting procedures and delivering technical training to data center technician teams.
  • Strong written communication — comfortable producing escalation documentation, platform runbooks, and leadership reporting.
  • Willingness to travel domestically and internationally to Crusoe sites as needed (target: up to 30%)

Responsibilities

  • Own Tier 2/3 hardware escalations across all Crusoe sites for issues that exceed local site capability, engaging directly with OEM and ODM engineering teams to drive resolution.
  • Travel to sites as needed for complex platform issues, new hardware bring-ups, and deployment support.
  • Identify recurring failure patterns across sites and translate them into platform feedback, sparing strategy inputs, or OEM improvement requests.
  • Root‑cause complex hardware issues — PCIe, BMC, thermal, fabric — and produce resolution documentation reusable across the SiteOps org.
  • Hand off platform-level findings to the appropriate internal engineering teams with clear, well-documented escalation packages.
  • Develop and maintain deep technical relationships with Crusoe's hardware partners — currently SuperMicro and HPE — at the engineering and field escalation level.
  • Serve as Crusoe's technical voice in OEM/ODM partner conversations, surfacing field observations, influencing hardware roadmaps, and driving platform improvements.
  • Build familiarity with new ODM platform architecture, tooling, and escalation processes as Crusoe expands its ODM footprint.
  • Support vendor evaluations and new platform qualifications in partnership with SiteOps and engineering leadership.
  • Own OEM platform technical knowledge at the SiteOps org level — escalation playbooks, failure pattern analysis, and OEM relationship inputs across all sites.
  • Own the development and maintenance of platform-specific SOPs, runbooks, and field troubleshooting procedures for the SiteOps org, ensuring site teams have current, actionable documentation across all active hardware platforms.
  • Design and deliver technical training for SiteOps technicians covering hardware architecture, platform-specific troubleshooting, and field procedures — both for new hire onboarding and ongoing skill development.
  • Contribute to the technician certification program and technical leveling standards across the org.
  • Support new site bring-up efforts providing platform readiness and deployment execution expertise.
  • Serve as SiteOps' senior technical representative at Crusoe HQ, participating in platform, engineering, and procurement discussions that affect site operations.
  • Partner with engineering and procurement teams on sparing strategy, RMA lifecycle management, and OEM/ODM support contract structures.
  • Provide operational input into next-generation GPU platform evaluations (GB300, VR200, and beyond).
  • Produce escalation reporting, platform health analysis, and operational insights for SiteOps leadership.

Skills

GPU infrastructure
Data center operations
OEM/ODM collaboration
Technical writing
Travel readiness

Tools

SuperMicro hardware
HPE platforms

Job description

Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.

We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.

We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.

If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.

About the Role

Crusoe Cloud operates GPU infrastructure across six production sites globally, with a fleet that spans SuperMicro, HPE, and next-generation ODM platforms as we scale. We're looking for a Staff Data Center Operations Engineer to serve as the senior technical operations resource for the SiteOps org — based at our Denver headquarters, with cross-site scope and travel authority across our full portfolio.

This role is the bridge between Crusoe's distributed site teams and our OEM and ODM hardware partners. You'll own platform-level escalations that exceed site-level capability, drive hardware decisions at the org level, and serve as SiteOps' technical presence at headquarters — visible to engineering, procurement, and leadership in a way that a field-based role cannot be.

You'll be hands-on when the situation calls for it, traveling to sites for complex escalations, new platform bring‑ups, and deployment support. But your primary leverage is organizational: building the technical standards, OEM relationships, and institutional knowledge that keeps Crusoe's GPU fleet reliable across every site.

What You'll Do

Cross‑Site Platform Operations & Escalation

  • Own Tier 2/3 hardware escalations across all Crusoe sites for issues that exceed local site capability, engaging directly with OEM and ODM engineering teams to drive resolution

  • Travel to sites as needed for complex platform issues, new hardware bring‑ups, and deployment support

  • Identify recurring failure patterns across sites and translate them into platform feedback, sparing strategy inputs, or OEM improvement requests

  • Root‑cause complex hardware issues — PCIe, BMC, thermal, fabric — and produce resolution documentation reusable across the SiteOps org

  • Hand off platform‑level findings to the appropriate internal engineering teams with clear, well‑documented escalation packages

OEM & ODM Technical Partnership

  • Develop and maintain deep technical relationships with Crusoe's primary hardware partners — currently SuperMicro and HPE, with upcoming ODM’s as growing platforms — at the engineering and field escalation level

  • Serve as Crusoe's technical voice in OEM/ODM partner conversations, surfacing field observations, influencing hardware roadmaps, and driving platform improvements that benefit the full fleet

  • Build familiarity with new ODM platform architecture, tooling, and escalation processes as Crusoe expands its ODM footprint

  • Support vendor evaluations and new platform qualifications in partnership with SiteOps and engineering leadership

Platform Standards & Org Development

  • Own OEM platform technical knowledge at the SiteOps org level — escalation playbooks, failure pattern analysis, and OEM relationship inputs across all sites

  • Own the development and maintenance of platform‑specific SOPs, runbooks, and field troubleshooting procedures for the SiteOps org, ensuring site teams have current, actionable documentation across all active hardware platforms

  • Design and deliver technical training for SiteOps technicians covering hardware architecture, platform‑specific troubleshooting, and field procedures — both for new hire onboarding and ongoing skill development as the fleet and team evolve

  • Contribute to the technician certification program and technical leveling standards across the org

  • Support new site bring‑up efforts providing platform readiness and deployment execution expertise

HQ Presence & Cross‑Functional Collaboration

  • Serve as SiteOps' senior technical representative at Crusoe HQ, participating in platform, engineering, and procurement discussions that affect site operations

  • Partner with engineering and procurement teams on sparing strategy, RMA lifecycle management, and OEM/ODM support contract structures

  • Provide operational input into next‑generation GPU platform evaluations (GB300, VR200, and beyond)

  • Produce escalation reporting, platform health analysis, and operational insights for SiteOps leadership

What We're Looking For

Required

  • 7+ years in data center operations, field engineering, or OEM/ODM technical support with hands‑on GPU infrastructure experience

  • Direct hands‑on experience deploying and supporting GPU platforms at scale across one or more major OEMs or ODMs; familiarity with SuperMicro and HPE platforms required

  • Deep familiarity with server platform architecture and OEM escalation and RMA processes

  • Experience leading or contributing to large‑scale GPU cluster bring‑ups including rack staging and production handoff

  • Demonstrated ability to build technical relationships with OEM and ODM engineering teams and drive platform‑level issue resolution

  • Experience developing SOPs, runbooks, or field troubleshooting procedures and delivering technical training to data center technician teams

  • Strong written communication — comfortable producing escalation documentation, platform runbooks, and leadership reporting

  • Willingness to travel domestically and internationally to Crusoe sites as needed (target: up to 30%)

Preferred

  • Direct experience with SuperMicro GPU platforms (B200, GB200, or newer); SuperMicro Certified Engineer credentials a plus

  • Familiarity with ASUS or Quanta server platforms and ODM engagement models

  • Experience with liquid‑cooled GPU platforms and CDU integration

  • Familiarity with AMD Instinct GPU platforms (MI300X/MI350X/MI355X)

  • Prior experience at an AI cloud provider, hyperscaler, or GPU‑first infrastructure operator

  • Experience contributing to technician certification programs or IC leveling standards within a DC ops organization

Location

This role is based at Crusoe Cloud's headquarters in Denver, CO, with regular travel to our data center sites across the US and internationally. Domestic relocation support is available.

Benefits:

  • Competitive compensation and equity packages
  • Restricted Stock Units
  • Paid time off, paid holidays & leave of absence programs
  • Comprehensive health, dental & vision insurance
  • Employer contributions to HSA account
  • Paid parental leave
  • Paid life insurance, short‑term and long‑term disability
  • Professional development & tuition reimbursement
  • Mental health & wellness support
  • Commuter benefits (parking & transit)
  • Cell phone stipend
  • 401(k) Retirement plan with company match up to 4% of salary
  • Volunteer time off
  • Global travel insurance & emergency assistance
  • Daily meals allowanceAdditional perks & programs specific to location

Compensation Range

Compensation will be paid in the range of up to $150,000 -$170,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicants knowledge, education, and abilities, as well as internal equity and alignment with market data.

Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Manager, Data Center Operations
Manager, Data Center Operations

crusoe • Town of Springfield (WI)

On-site
USD 135,000 - 175,000
Competitive compensation and equity
Paid time off, holidays, and leave programs
Medical, dental, and vision insurance
+2
Senior Staff Software Engineer, DC Infrastructure
Senior Staff Software Engineer, DC Infrastructure

Crusoe • United States

On-site
USD 250,000 - 300,000
Industry competitive pay
RSUs in a fast-growing tech company
Health insurance with family options
+3
Staff Software Engineer, DC Infrastructure
Staff Software Engineer, DC Infrastructure

Crusoe • San Francisco (CA)

On-site
USD 215,000 - 260,000
Industry competitive pay
Restricted Stock Units
Health insurance options (HDHP/PPO)
+11
Senior Software Engineer (DCIE)
Senior Software Engineer (DCIE)

Crusoe • San Francisco (CA)

On-site
USD 170,000 - 205,000
Health insurance package options
Restricted Stock Units
401(k) with match up to 4%
+2
Staff Software Engineer, DC Infrastructure
Staff Software Engineer, DC Infrastructure

Crusoe Energy Systems LLC • San Francisco (CA), Northern (KY)

Hybrid
USD 215,000 - 260,000
Industry competitive pay
RSUs
Health insurance
+12
Software Engineer I (DCIE)
Software Engineer I (DCIE)

crusoe • San Francisco (CA)

On-site
USD 117,000 - 135,000
Health insurance
401(k) with match
Stock options/RSUs
+3
Principal Engineer, Conductor Platform (CAPE)
Principal Engineer, Conductor Platform (CAPE)

Crusoe • San Francisco (CA)

On-site
USD 285,000 - 335,000
Competitive compensation
Equity packages
Paid time off & holidays
+3
Staff Production Engineer, Core PE
Staff Production Engineer, Core PE

Crusoe • San Francisco (CA)

On-site
USD 209,000 - 253,000
Health insurance package options
401(k) with 100% match up to 4%
Generous paid time off
Senior Staff Deployment Automation Engineer
Senior Staff Deployment Automation Engineer

Crusoe • San Francisco (CA)

On-site
USD 250,000 - 300,000
Competitive compensation and equity
Paid time off
Health, dental & vision insurance
+3
Senior Staff Deployment Automation Engineer
Senior Staff Deployment Automation Engineer

ProducePay • United States

On-site
USD 250,000 - 300,000
Competitive compensation
Equity packages
Paid time off
+5