Senior GPU Infra Engineer — AI Data Center Automation

Crusoe Energy Systems LLC

San Francisco, Northern (CA, KY)

Hybrid

USD 215,000 - 260,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Industry competitive pay
RSUs
Health insurance
HSA contributions
Parental leave
Life insurance
Disability
Teladoc
401(k) match
PTO
Cell phone reimbursement
Tuition reimbursement
Calm app
Legal services
Commuter benefit

Job summary

Crusoe Energy Systems LLC in San Francisco is seeking a hands-on Software Engineer to join Crusoe’s Data Center Infrastructure Engineering team. You will develop software to manage a fleet of GPU servers and the data centers that house them, building advanced diagnostics, observability, automation, and repair tooling for high-performance GPU compute clusters.

The role emphasizes independent problem solving, scalability, and ownership of deployment, monitoring, and tooling that maximizes GPU

Qualifications

  • Experience with distributed systems and reliability in cloud environments.
  • Ability to design scalable tooling for GPU compute clusters.
  • Familiarity with high-density GPU racks and data center operations.

Responsibilities

  • Developing and implementing deep-level diagnostics and troubleshooting of hardware faults within GPU racks and high-density compute systems.
  • Developing troubleshooting and automation tooling for GPU platforms including NVIDIA A100, H200, GB200, B200 and AMD 350X / 355X.
  • Developing automation and AI agents for executing component-level diagnosis and remediation for failed or degraded hardware.
  • In conjunction with data center operations develop innovative tooling and AI agents for managing the critical environment.
  • Developing tooling for post-repair validation and testing tools such as burn-in, Pytorch, and NVIDIA NCCL to ensure system stability and performance.
  • Own the deployment, monitoring, and operational support of developed tooling, ensuring solutions maximize GPU fleet availability and performance to drive customer success.
  • Developing automation and operational tooling for facilities management power as well as direct liquid cooling hardware systems

Skills

Distributed systems
Kubernetes
Go
Python
Java
Rust
Cloud platforms
Analytical skills

Tools

Terraform
NVIDIA NCCL
Kubernetes

Job description

Crusoe Energy Systems LLC in San Francisco is seeking a hands-on Software Engineer to join Crusoe’s Data Center Infrastructure Engineering team. You will develop software to manage a fleet of GPU servers and the data centers that house them, building advanced diagnostics, observability, automation, and repair tooling for high-performance GPU compute clusters.

The role emphasizes independent problem solving, scalability, and ownership of deployment, monitoring, and tooling that maximizes GPU

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Infra Engineer - Automation & Diagnostics
Senior GPU Infra Engineer - Automation & Diagnostics

Crusoe • United States

On-site
USD 250,000 - 300,000
Industry competitive pay
RSUs in a fast-growing tech company
Health insurance with family options
+3
Senior GPU Infrastructure & Automation Engineer
Senior GPU Infrastructure & Automation Engineer

Crusoe • San Francisco (CA)

On-site
USD 250,000 - 300,000
Industry competitive pay
Restricted Stock Units
Health insurance
+7
Senior Software Engineer — GPU Data Center Automation
Senior Software Engineer — GPU Data Center Automation

Crusoe • San Francisco (CA)

On-site
USD 170,000 - 205,000
Health insurance package options
Restricted Stock Units
401(k) with match up to 4%
+2
Senior GPU DC Infra Engineer — Automation & Diagnostics
Senior GPU DC Infra Engineer — Automation & Diagnostics

Crusoe • San Francisco (CA)

On-site
USD 215,000 - 260,000
Industry competitive pay
Restricted Stock Units
Health insurance options (HDHP/PPO)
+11
Senior Infrastructure Software Engineer, GPU Fleet
Senior Infrastructure Software Engineer, GPU Fleet

Crusoe • Sunnyvale (CA)

On-site
USD 250,000 - 300,000
Health insurance
401(k) with match
Tuition reimbursement
+2
Senior Hardware Systems Engineer, AI Compute Infra
Senior Hardware Systems Engineer, AI Compute Infra

Crusoe • San Francisco (CA)

On-site
USD 215,000 - 260,000
RSUs
Health insurance
401(k) with match
+3
Senior GPU Data Center Operations Engineer
Senior GPU Data Center Operations Engineer

Crusoe • Denver (CO)

On-site
USD 150,000 - 170,000
Equity
Paid time off
Health insurance
+1
Senior Deployment Automation Engineer - AI Cloud GPUs
Senior Deployment Automation Engineer - AI Cloud GPUs

Crusoe • Bellevue (WA)

On-site
USD 250,000 - 300,000
Stock options
Paid time off
Health insurance
+5
Senior Software Engineer (DCIE)
Senior Software Engineer (DCIE)

Crusoe • San Francisco (CA)

On-site
USD 170,000 - 205,000
Health insurance package options
Restricted Stock Units
401(k) with match up to 4%
+2
Senior Staff Software Engineer, DC Infrastructure
Senior Staff Software Engineer, DC Infrastructure

Crusoe • San Francisco (CA)

On-site
USD 250,000 - 300,000
Industry competitive pay
Restricted Stock Units
Health insurance
+7