Senior Infrastructure Software Engineer, GPU Fleet

Crusoe

Sunnyvale (CA)

On-site

USD 250,000 - 300,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) with match
Tuition reimbursement
Cell phone reimbursement
Paid parental leave

Job summary

Crusoe is seeking a highly skilled Software Engineer to join the Data Center Infrastructure Engineering team. This role focuses on software for managing a fleet of GPU servers and the data centers that house them.

You will develop advanced diagnostics, observability, automation, and repair tooling for high‑performance GPU clusters, helping maintain health and scalability while maximizing uptime. The ideal candidate thrives on hands‑on problem solving and works well both independently and with

Qualifications

  • Experience in software engineering for managing GPU compute clusters.
  • Strong background in distributed systems and cloud platforms.
  • Proficiency in Go, Python, Java, or Rust.

Responsibilities

  • Develop deep-level diagnostics and troubleshooting of hardware faults in GPU racks.
  • Developing troubleshooting and automation tooling for GPU platforms including NVIDIA A100, H200, GB200, B200 and AMD 350X / 355X.
  • Developing automation and AI agents for executing component-level diagnosis and remediation for failed or degraded hardware.
  • In conjunction with data center operations develop innovative tooling and AI agents for managing the critical environment.
  • Developing tooling for post-repair validation and testing tools such as burn-in, Pytorch, and NVIDIA NCCL to ensure system stability and performance.
  • Own the deployment, monitoring, and operational support of developed tooling, ensuring solutions maximize GPU fleet availability and performance to drive customer success.
  • Developing automation and operational tooling for facilities management power as well as direct liquid cooling hardware systems

Skills

Go
Python
Java
Rust

Tools

Kubernetes
IaC
GCP

Job description

Crusoe is seeking a highly skilled Software Engineer to join the Data Center Infrastructure Engineering team. This role focuses on software for managing a fleet of GPU servers and the data centers that house them.

You will develop advanced diagnostics, observability, automation, and repair tooling for high‑performance GPU clusters, helping maintain health and scalability while maximizing uptime. The ideal candidate thrives on hands‑on problem solving and works well both independently and with

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Infrastructure & Automation Engineer
Senior GPU Infrastructure & Automation Engineer

Crusoe • San Francisco (CA)

On-site
USD 250,000 - 300,000
Industry competitive pay
Restricted Stock Units
Health insurance
+7
Senior Software Engineer — GPU Data Center Automation
Senior Software Engineer — GPU Data Center Automation

Crusoe • San Francisco (CA)

On-site
USD 170,000 - 205,000
Health insurance package options
Restricted Stock Units
401(k) with match up to 4%
+2
Senior GPU Infra Engineer - Automation & Diagnostics
Senior GPU Infra Engineer - Automation & Diagnostics

Crusoe • United States

On-site
USD 250,000 - 300,000
Industry competitive pay
RSUs in a fast-growing tech company
Health insurance with family options
+3
Hybrid GPU Fleet Operations Engineer
Hybrid GPU Fleet Operations Engineer

Crusoe • San Francisco (CA)

Hybrid
USD 215,000 - 260,000
Hybrid work schedule
Industry competitive pay
Restricted Stock Units
+2
Senior GPU Infra Engineer — AI Data Center Automation
Senior GPU Infra Engineer — AI Data Center Automation

Crusoe Energy Systems LLC • San Francisco (CA), Northern (KY)

Hybrid
USD 215,000 - 260,000
Industry competitive pay
RSUs
Health insurance
+12
Senior GPU DC Infra Engineer — Automation & Diagnostics
Senior GPU DC Infra Engineer — Automation & Diagnostics

Crusoe • San Francisco (CA)

On-site
USD 215,000 - 260,000
Industry competitive pay
Restricted Stock Units
Health insurance options (HDHP/PPO)
+11
Senior GPU Data Center Operations Engineer
Senior GPU Data Center Operations Engineer

Crusoe • Denver (CO)

On-site
USD 150,000 - 170,000
Equity
Paid time off
Health insurance
+1
Software Engineer - GPU Fleet
Software Engineer - GPU Fleet

Iceberg • New York (NY)

On-site
USD 120,000 - 170,000
Staff Software Engineer (Cloud Infrastructure)
Staff Software Engineer (Cloud Infrastructure)

Crusoe • San Francisco (CA)

Hybrid
USD 215,000 - 260,000
Hybrid work schedule
Industry competitive pay
Restricted Stock Units
+2
Senior Staff Software Engineer, DC Infrastructure
Senior Staff Software Engineer, DC Infrastructure

Crusoe • San Francisco (CA)

On-site
USD 250,000 - 300,000
Industry competitive pay
Restricted Stock Units
Health insurance
+7