Senior Software Engineer (DCIE)

Crusoe Energy Systems

San Francisco (CA)

On-site

USD 180,000 - 300,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health benefits
Paid time off
401(k) match
Mental wellness resources

Job summary

Crusoe Energy Systems is seeking a Software Engineer to join the Data Center Infrastructure Engineering team. You will develop software for managing a fleet of GPU servers and the data centers housing them, building diagnostics, observability, automation, and repair tooling for high-density compute clusters.

The role emphasizes hands-on problem solving, independence, and helping scale Crusoe's GPU fleet with robust tooling and AI-assisted diagnostics across NVIDIA and AMD platforms.

Qualifications

  • 4–6 years of software engineering experience.
  • Strong in at least one programming language (Go, Python, Java, Rust).
  • Experience with distributed systems, reliability, and cloud platforms.
  • Experience with Kubernetes and Temporal.
  • Experience working with hardware vendors and GPU fleet operations.

Responsibilities

  • Develop software for managing a fleet of GPU servers and data centers.
  • Build advanced diagnostic, observability, automation, and repair tooling for GPU clusters.
  • Own deployment, monitoring, and operational support of tooling to maximize fleet availability and performance.
  • Collaborate with data center operations to create tooling for facilities management including cooling.

Skills

Go
Python
Java
Rust
Distributed systems
Kubernetes
GCP

Tools

Temporal
Kubernetes
Infrastructure as Code
NVIDIA NCCL

Job description

  • We are seeking a highly skilled and motivated Software Engineer to join Crusoe’s Data Center Infrastructure Engineering team
  • This position is focused on the development of software for the management of a fleet of GPU servers as well as the data centers that house those systems
  • The role focuses on the developing and implementing advanced diagnostic, observability, automation and repair tooling for high-performance GPU compute clusters
  • The ideal new team member will be a hands-on problem solver who is comfortable working independently
  • The new team member will play a critical role in maintaining the health and scalability of Crusoe’s rapidly growing GPU fleet
  • Developing and implementing deep-level diagnostics and troubleshooting of hardware faults within GPU racks and high-density compute systems
  • Developing troubleshooting and automation tooling for GPU platforms including NVIDIA A100, H200, GB200, B200 and AMD 350X / 355X
  • Developing automation and AI agents for executing component-level diagnosis and remediation for failed or degraded hardware
  • In conjunction with data center operations develop innovative tooling and AI agents for managing the critical environment
  • Developing tooling for post-repair validation and testing tools such as burn-in, Pytorch, and NVIDIA NCCL to ensure system stability and performance
  • Own the deployment, monitoring, and operational support of developed tooling, ensuring solutions maximize GPU fleet availability and performance to drive customer success
  • Developing automation and operational tooling for facilities management power as well as direct liquid cooling hardware systems
Benefits
  • Health & wellbeing: Comprehensive health benefits designed to support your overall wellness
  • Time away: Paid time off for vacations, family bonding, and unexpected needs
  • 401(k) match: Build your financial future with our 401(k) matching program
  • Mental wellness: Resources and support for your emotional wellbeing and navigating life’s challenges
  • Ability to lean in and assist team members working on critical or complex technical initiatives
  • Ability to work independently and within a team
  • Strong analytical and problem-solving skills
  • Ability to set the technical direction for a specific project and execute
  • The ability to identify a problem, rapidly develop a scalable solution and ship it
  • 4-6 years of software engineering experience
  • Strength in at least one programming language - Go, Python, Java, Rust
  • Expertise in distributed systems, reliability, and cloud platforms (Kubernetes, IaC, GCP etc.)
  • Excellent communication and collaboration skills
  • Experience with Temporal and Kubernetes
  • Experience working directly with hardware vendors
  • Background in large-scale GPU fleet operations or hyperscale data center environments
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer (DCIE)
Senior Software Engineer (DCIE)

Crusoe • San Francisco (CA)

On-site
USD 170,000 - 205,000
Health insurance package options
Restricted Stock Units
401(k) with match up to 4%
+2
Software Engineer I (DCIE)
Software Engineer I (DCIE)

crusoe • San Francisco (CA)

On-site
USD 117,000 - 135,000
Health insurance
401(k) with match
Stock options/RSUs
+3
Staff Software Engineer, DC Infrastructure
Staff Software Engineer, DC Infrastructure

Crusoe • San Francisco (CA)

On-site
USD 215,000 - 260,000
Industry competitive pay
Restricted Stock Units
Health insurance options (HDHP/PPO)
+11
Senior Staff Software Engineer, DC Infrastructure
Senior Staff Software Engineer, DC Infrastructure

Crusoe • United States

On-site
USD 250,000 - 300,000
Industry competitive pay
RSUs in a fast-growing tech company
Health insurance with family options
+3
Staff Software Engineer, DC Infrastructure
Staff Software Engineer, DC Infrastructure

Crusoe Energy Systems LLC • San Francisco (CA), Northern (KY)

Hybrid
USD 215,000 - 260,000
Industry competitive pay
RSUs
Health insurance
+12
Senior Staff Software Engineer, DC Infrastructure
Senior Staff Software Engineer, DC Infrastructure

CV in • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 220,000
Health insurance
RSUs
401(k) match
+2
Senior Software Engineer, GPU Fleet Reliability
Senior Software Engineer, GPU Fleet Reliability

Crusoe Energy Systems • San Francisco (CA)

On-site
USD 180,000 - 300,000
Health benefits
Paid time off
401(k) match
+1
Senior Software Engineer — GPU Data Center Automation
Senior Software Engineer — GPU Data Center Automation

Crusoe • San Francisco (CA)

On-site
USD 170,000 - 205,000
Health insurance package options
Restricted Stock Units
401(k) with match up to 4%
+2
Senior GPU Infra Engineer - Automation & Diagnostics
Senior GPU Infra Engineer - Automation & Diagnostics

Crusoe • United States

On-site
USD 250,000 - 300,000
Industry competitive pay
RSUs in a fast-growing tech company
Health insurance with family options
+3
Staff Software Engineer (Cloud Infrastructure)
Staff Software Engineer (Cloud Infrastructure)

Crusoe • San Francisco (CA)

Hybrid
USD 215,000 - 260,000
Hybrid work schedule
Industry competitive pay
Restricted Stock Units
+2