Senior GPU Data Center Infrastructure Engineer

CV in

San Francisco, Northern (CA, KY)

Hybrid

USD 170,000 - 220,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
RSUs
401(k) match
Parental leave
Life insurance

Job summary

Crusoe is hiring a Senior Staff Software Engineer to join the DC Infrastructure Engineering team in San Francisco. You will own software for managing a fleet of GPU servers and the data centers that house them, building tooling for deployment, observability, and remediation to maximize fleet availability and performance.

The role emphasizes hands‑on development, independent problem solving, and mentoring junior engineers.

Qualifications

  • Experience in software engineering with ability to ship scalable solutions.
  • Experience with distributed systems and cloud platforms.
  • Strong problem-solving and collaboration skills.
  • Ability to work independently and in a team.

Responsibilities

  • Developing and implementing deep‑level diagnostics and troubleshooting of hardware faults within GPU racks and high‑density compute systems.
  • Developing troubleshooting and automation tooling for GPU platforms including NVIDIA A100, H200, GB200, B200 and AMD 350X / 355X.
  • Developing automation and AI agents for executing component‑level diagnosis and remediation for failed or degraded hardware.
  • In conjunction with data center operations develop innovative tooling and AI agents for managing the critical environment.
  • Developing tooling for post‑repair validation and testing tools such as burn‑in, Pytorch, and NVIDIA NCCL to ensure system stability and performance.
  • Owning the deployment, monitoring, and operational support of developed tooling, ensuring solutions maximize GPU fleet availability and performance to drive customer success.
  • Developing automation and operational tooling for facilities management power as well as direct liquid cooling hardware systems.
  • Writing and maintaining scalable, observable, and resilient software that integrates with existing cloud and on‑premise infrastructure.
  • Collaborating closely with cross‑functional teams to define, build, and deliver features that meet strict reliability and performance targets.
  • Contributing to on‑call rotations to support critical production issues and driving rapid resolution.
  • Implementing observability pipelines, metrics, and alerting to provide deep insight into system health and performance.
  • Refactoring legacy scripts and workflows into robust, maintainable services that improve reliability and developer velocity.
  • Participating in design reviews and ensuring best practices for security, scalability, and maintainability are followed.
  • Mentoring junior engineers through code reviews, pair programming, and knowledge sharing sessions.

Skills

Go
Python
Java
Rust
Distributed systems
Cloud platforms
Kubernetes
NVIDIA NCCL

Tools

Kubernetes
Terraform
GCP

Job description

Crusoe is hiring a Senior Staff Software Engineer to join the DC Infrastructure Engineering team in San Francisco. You will own software for managing a fleet of GPU servers and the data centers that house them, building tooling for deployment, observability, and remediation to maximize fleet availability and performance.

The role emphasizes hands‑on development, independent problem solving, and mentoring junior engineers.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Infra Engineer — AI Data Center Automation
Senior GPU Infra Engineer — AI Data Center Automation

Crusoe Energy Systems LLC • San Francisco (CA), Northern (KY)

Hybrid
USD 215,000 - 260,000
Industry competitive pay
RSUs
Health insurance
+12
Senior GPU Infra Engineer - Automation & Diagnostics
Senior GPU Infra Engineer - Automation & Diagnostics

Crusoe • United States

On-site
USD 250,000 - 300,000
Industry competitive pay
RSUs in a fast-growing tech company
Health insurance with family options
+3
Senior Software Engineer — GPU Data Center Automation
Senior Software Engineer — GPU Data Center Automation

Crusoe • San Francisco (CA)

On-site
USD 170,000 - 205,000
Health insurance package options
Restricted Stock Units
401(k) with match up to 4%
+2
Senior GPU DC Infra Engineer — Automation & Diagnostics
Senior GPU DC Infra Engineer — Automation & Diagnostics

Crusoe • San Francisco (CA)

On-site
USD 215,000 - 260,000
Industry competitive pay
Restricted Stock Units
Health insurance options (HDHP/PPO)
+11
Senior GPU Data Center Operations Engineer
Senior GPU Data Center Operations Engineer

Crusoe • Denver (CO)

On-site
USD 150,000 - 170,000
Equity
Paid time off
Health insurance
+1
Senior Hardware Systems Engineer, AI Compute Infra
Senior Hardware Systems Engineer, AI Compute Infra

Crusoe • San Francisco (CA)

On-site
USD 215,000 - 260,000
RSUs
Health insurance
401(k) with match
+3
Senior Deployment Automation Engineer - AI Cloud GPUs
Senior Deployment Automation Engineer - AI Cloud GPUs

Crusoe • Bellevue (WA)

On-site
USD 250,000 - 300,000
Stock options
Paid time off
Health insurance
+5
Senior Staff Software Engineer, DC Infrastructure
Senior Staff Software Engineer, DC Infrastructure

CV in • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 220,000
Health insurance
RSUs
401(k) match
+2
Senior Staff Software Engineer, DC Infrastructure
Senior Staff Software Engineer, DC Infrastructure

Crusoe • United States

On-site
USD 250,000 - 300,000
Industry competitive pay
RSUs in a fast-growing tech company
Health insurance with family options
+3
Senior GPU Capacity & Optimization Architect
Senior GPU Capacity & Optimization Architect

Crusoe • San Francisco (CA)

On-site
USD 160,000 - 195,000
Competitive compensation and equity packages
Comprehensive health, dental & vision insurance
401(k) Retirement plan with company match
+2