Senior Production Engineer, Reliability & Automation

Crusoe

San Francisco (CA)

On-site

USD 172,000 - 209,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) with 100% match up to 4%
Parental Leave
Stock options

Job summary

Crusoe is building the most reliable, energy-efficient, AI-optimized cloud platform and Production Engineering sits at the heart of that mission. As a Production Engineer focused on Operational Excellence, you will help ensure the reliability, scalability, and performance of Crusoe’s GPU cloud that powers next-generation AI workloads.

This role is ideal for engineers who enjoy solving complex production problems, improving large-scale distributed systems, and building automation that keeps

Qualifications

  • 5+ years of experience in Production Engineering, SRE, or large-scale infrastructure operations.
  • Experience supporting GPU workloads, HPC environments, or latency/throughput-sensitive distributed systems.
  • Strong knowledge of Linux/Unix systems including debugging kernel/user space issues.
  • Experience in infrastructure roles building or managing compute, storage or networking platforms.
  • Understanding of modern cloud infrastructure fundamentals including Kubernetes, distributed systems, virtualization, and cloud platforms (AWS/GCP).
  • Familiarity with incident management practices and reliability frameworks (SRE, ITIL, or similar).
  • Experience with monitoring/observability tools such as Prometheus and Grafana; knowledge of OpenTelemetry preferred.
  • Familiarity with infrastructure-as-code and configuration management tools such as Terraform or Ansible.
  • Scripting or programming experience with Go, Python, C, or C++.

Responsibilities

  • Collaborate with cross-functional teams to define and evolve availability metrics for Crusoe’s cloud platform, including SLIs/SLOs.
  • Participate in production incident response, diagnosing and resolving service disruptions; contribute to post-incident reviews.
  • Build, operate, and improve observability across Crusoe’s infrastructure using Prometheus, Grafana, Alertmanager, and OpenTelemetry.
  • Identify reliability risks, performance bottlenecks, and early indicators of production issues across distributed systems.
  • Develop automation and tooling that reduces toil, improves recovery times, and enables self-healing infrastructure.
  • Partner with compute, networking, storage and platform teams to strengthen service resilience and disaster recovery capabilities.
  • Contribute to improving operational processes, knowledge sharing, and reliability best practices across the engineering organization.
  • Continue growing technical depth through mentorship, training, and hands-on work operating large-scale AI infrastructure.

Skills

Production Engineering
SRE
Large-scale infra
Linux/Unix
Kubernetes
Cloud platforms
Terraform/Ansible
Go/Python/C/C++
Incident management
Observability
Communication

Tools

Prometheus
Grafana
OpenTelemetry
Terraform
Ansible

Job description

Crusoe is building the most reliable, energy-efficient, AI-optimized cloud platform and Production Engineering sits at the heart of that mission. As a Production Engineer focused on Operational Excellence, you will help ensure the reliability, scalability, and performance of Crusoe’s GPU cloud that powers next-generation AI workloads.

This role is ideal for engineers who enjoy solving complex production problems, improving large-scale distributed systems, and building automation that keeps

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Production Engineer, AI GPU Cloud Reliability & Automation
Production Engineer, AI GPU Cloud Reliability & Automation

Crusoe • San Francisco (CA)

On-site
USD 172,000 - 209,000
Health insurance
401(k) with match
Paid parental leave
+4
Senior AI Cloud Platform Engineer - Reliability & Scale
Senior AI Cloud Platform Engineer - Reliability & Scale

Crusoe • San Francisco (CA)

On-site
USD 170,000 - 205,000
Health insurance
RSUs
401(k) match
+2
Senior Hardware Systems Engineer, AI Compute & Sustaining
Senior Hardware Systems Engineer, AI Compute & Sustaining

Crusoe Energy Systems LLC • San Francisco (CA)

On-site
USD 215,000 - 260,000
Senior AI Hardware Production Engineer
Senior AI Hardware Production Engineer

Crusoe Energy Systems LLC • Sunnyvale (CA)

On-site
USD 170,000 - 205,000
RSUs included
Senior Production Engineer, Core PE
Senior Production Engineer, Core PE

Crusoe • San Francisco (CA)

On-site
USD 172,000 - 209,000
Health insurance
401(k) with match
Paid parental leave
+4
Senior Production Engineer, Operational Excellence
Senior Production Engineer, Operational Excellence

Crusoe • San Francisco (CA)

On-site
USD 172,000 - 209,000
Health insurance
401(k) with 100% match up to 4%
Parental Leave
+1
Staff Network Production & Reliability Engineer
Staff Network Production & Reliability Engineer

ProducePay • United States

On-site
USD 195,000 - 235,000
Competitive compensation
Equity
Health insurance
+2
Senior Production Engineer, Managed Cloud
Senior Production Engineer, Managed Cloud

Crusoe • San Francisco (CA)

On-site
USD 170,000 - 205,000
Health insurance
RSUs
401(k) match
+2
Staff Production Engineer, Core PE
Staff Production Engineer, Core PE

Crusoe • San Francisco (CA)

On-site
USD 209,000 - 253,000
Health insurance package options
401(k) with 100% match up to 4%
Generous paid time off
Senior Deployment Automation Engineer - AI Cloud GPUs
Senior Deployment Automation Engineer - AI Cloud GPUs

Crusoe • Bellevue (WA)

On-site
USD 250,000 - 300,000
Stock options
Paid time off
Health insurance
+5