Senior Software Engineer (DCIE)

Crusoe

San Francisco (CA)

On-site

USD 170,000 - 205,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance package options
Restricted Stock Units
401(k) with match up to 4%
Paid parental leave
Teladoc

Job summary

Crusoe is on a mission to accelerate the abundance of energy and intelligence. We are seeking a Software Engineer to join Crusoe’s Data Center Infrastructure Engineering team, focusing on software for managing a fleet of GPU servers and the data centers that house them.

You will build advanced diagnostics, observability, automation and repair tooling for high‑performance GPU compute clusters. The successful candidate will be hands-on, independent, and pivotal in maintaining health and

Qualifications

  • 4-6 years of software engineering experience.
  • Experience with distributed systems, reliability and cloud platforms.
  • Proficient in Go, Python, Java or Rust.
  • Strong analytical and problem-solving skills.
  • Excellent communication and collaboration skills.
  • Ability to work independently and within a team.

Responsibilities

  • Develop deep-level diagnostics and troubleshooting for GPU racks and high-density compute systems.
  • Build automation tooling for GPU platforms (NVIDIA A100, H200, GB200, B200; AMD 350X/355X).
  • Create automation and AI agents for component-level diagnosis and remediation.
  • Collaborate with data center operations to develop tooling and AI agents for critical environments.
  • Develop post-repair validation and testing tools (burn-in, PyTorch, NVIDIA NCCL).
  • Own deployment, monitoring and operational support of tooling to maximize GPU fleet availability.
  • Develop automation and operational tooling for facilities management power and liquid cooling systems.

Skills

Distributed systems
Cloud platforms
Kubernetes
Go
Python
Java
Rust
Teamwork
Problem solving

Job description

Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world’s most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.

We’re in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We’re solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.

We’re looking for problem‑solving, opportunity‑finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.

If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high‑performing team that believes in each other, come build with us at Crusoe.

About the Role:

We are seeking a highly skilled and motivated Software Engineer to join Crusoe’s Data Center Infrastructure Engineering team. This position is focused on the development of software for the management of a fleet of GPU servers as well as the data centers that house those systems. The role focuses on the developing and implementing advanced diagnostic, observability, automation and repair tooling for high‑performance GPU compute clusters.

The ideal new team member will be a hands‑on problem solver who is comfortable working independently. The new team member will play a critical role in maintaining the health and scalability of Crusoe’s rapidly growing GPU fleet.

What You’ll Be Doing:
  • Developing and implementing deep‑level diagnostics and troubleshooting of hardware faults within GPU racks and high‑density compute systems.
  • Developing troubleshooting and automation tooling for GPU platforms including NVIDIA A100, H200, GB200, B200 and AMD 350X / 355X.
  • Developing automation and AI agents for executing component‑level diagnosis and remediation for failed or degraded hardware.
  • In conjunction with data center operations develop innovative tooling and AI agents for managing the critical environment.
  • Developing tooling for post‑repair validation and testing tools such as burn‑in, Pytorch, and NVIDIA NCCL to ensure system stability and performance.
  • Own the deployment, monitoring, and operational support of developed tooling, ensuring solutions maximize GPU fleet availability and performance to drive customer success.
  • Developing automation and operational tooling for facilities management power as well as direct liquid cooling hardware systems
What You’ll Bring to the Team:
  • 4-6 years of software engineering experience.
  • The ability to identify a problem, rapidly develop a scalable solution and ship it.
  • Ability to lean in and assist team members working on critical or complex technical initiatives.
  • Ability to set the technical direction for a specific project and execute.
  • Expertise in distributed systems, reliability, and cloud platforms (Kubernetes, IaC, GCP etc.)
  • Strength in at least one programming language - Go, Python, Java, Rust.
  • Strong analytical and problem‑solving skills.
  • Excellent communication and collaboration skills.
  • Ability to work independently and within a team.
Nice to Have:
  • Experience with Temporal and Kubernetes.
  • Experience working directly with hardware vendors.
  • Background in large‑scale GPU fleet operations or hyperscale data center environments.
Benefits:
  • Industry competitive pay
  • Restricted Stock Units in a fast‑growing, well‑funded technology company
  • Health insurance package options that include HDHP and PPO, vision, and dental for you and your dependents
  • Employer contributions to HSA accounts
  • Paid Parental Leave
  • Paid life insurance, short‑term and long‑term disability
  • Teladoc
  • 401(k) with a 100% match up to 4% of salary
  • Generous paid time off and holiday schedule
  • Cell phone reimbursement
  • Tuition reimbursement
  • Subscription to the Calm app
  • MetLife Legal
  • Company paid commuter benefit; $50 per pay period
Compensation Range

Compensation will be paid in the range of up to $170,000 - $205,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicants knowledge, education, and abilities, as well as internal equity and alignment with market data.

Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer I (DCIE)
Software Engineer I (DCIE)

crusoe • San Francisco (CA)

On-site
USD 117,000 - 135,000
Health insurance
401(k) with match
Stock options/RSUs
+3
Senior Staff Software Engineer, DC Infrastructure
Senior Staff Software Engineer, DC Infrastructure

Crusoe • United States

On-site
USD 250,000 - 300,000
Industry competitive pay
RSUs in a fast-growing tech company
Health insurance with family options
+3
Staff Software Engineer, DC Infrastructure
Staff Software Engineer, DC Infrastructure

Crusoe • San Francisco (CA)

On-site
USD 215,000 - 260,000
Industry competitive pay
Restricted Stock Units
Health insurance options (HDHP/PPO)
+11
Staff Software Engineer, DC Infrastructure
Staff Software Engineer, DC Infrastructure

Crusoe Energy Systems LLC • San Francisco (CA), Northern (KY)

Hybrid
USD 215,000 - 260,000
Industry competitive pay
RSUs
Health insurance
+12
Senior Staff Software Engineer, DC Infrastructure
Senior Staff Software Engineer, DC Infrastructure

CV in • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 220,000
Health insurance
RSUs
401(k) match
+2
Staff Software Engineer (Cloud Infrastructure)
Staff Software Engineer (Cloud Infrastructure)

Crusoe • San Francisco (CA)

Hybrid
USD 215,000 - 260,000
Hybrid work schedule
Industry competitive pay
Restricted Stock Units
+2
Senior Hardware Systems Engineer
Senior Hardware Systems Engineer

Crusoe • Sunnyvale (CA)

On-site
USD 170,000 - 205,000
Health insurance
401(k) with match
Employee stock options
+2
Staff Hardware Systems Engineer, Performance
Staff Hardware Systems Engineer, Performance

Crusoe • San Francisco (CA)

On-site
USD 215,000 - 260,000
RSUs
Health insurance
401(k) with match
+3
Staff Production Engineer, Core PE
Staff Production Engineer, Core PE

Crusoe • San Francisco (CA)

On-site
USD 209,000 - 253,000
Health insurance package options
401(k) with 100% match up to 4%
Generous paid time off
Senior Production Engineer, Core PE
Senior Production Engineer, Core PE

Crusoe • San Francisco (CA)

On-site
USD 172,000 - 209,000
Health insurance
401(k) with match
Paid parental leave
+4