Production Engineer, AI GPU Cloud Reliability & Automation

Crusoe

San Francisco (CA)

On-site

USD 172,000 - 209,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
401(k) with match
Paid parental leave
Life insurance
Disability insurance
Teladoc
Commuter benefit

Job summary

Crusoe is building a high-performance GPU-focused cloud platform and is seeking a Production Engineer to ensure reliability and scalability of the AI infrastructure. You will lead efforts in incident response, observability, and automation across large-scale systems.

You will collaborate with cross-functional teams to reduce toil, improve SLIs/SLOs, and implement self-healing techniques while growing your technical depth in a dynamic, energy‑driven environment.

Qualifications

  • 5+ years of experience in Production Engineering, SRE, or large-scale infrastructure operations.
  • Experience supporting GPU workloads, HPC environments, or latency/throughput-sensitive distributed systems.
  • Strong knowledge of Linux/Unix systems, including debugging complex issues across kernel and user space.
  • Previous experience in Infrastructure roles building or managing compute, storage or networking platforms.
  • Understanding of modern cloud infrastructure fundamentals including Kubernetes, distributed systems, virtualization, and cloud platforms (AWS/GCP).
  • Familiarity with incident management practices and reliability frameworks (SRE, ITIL, or similar).
  • Experience with monitoring and observability tools such as Prometheus and Grafana, or a strong desire to deepen expertise in this area.
  • Familiarity with infrastructure-as-code and configuration management tools such as Terraform or Ansible.
  • Scripting or programming experience with languages such as Go, Python, C, or C++.
  • Strong communication skills and the ability to collaborate across engineering teams.
  • Ability to remain calm and effective while troubleshooting complex issues in high-impact production environments.
  • A growth mindset and strong interest in reliability engineering, automation, and operational excellence.

Responsibilities

  • Collaborate with cross-functional teams on availability metrics for Crusoe’s cloud platform, including SLIs and SLOs.
  • Participate in production incident response and perform post-incident reviews.
  • Build, operate, and improve observability across Crusoe’s infrastructure using monitoring tools.
  • Identify reliability risks and performance bottlenecks in distributed systems.
  • Develop automation that reduces operational toil and enables self-healing infrastructure.
  • Partner with compute, networking, storage, and platform teams to strengthen service resilience.
  • Contribute to improving operational processes and reliability best practices across the engineering org.
  • Grow technical depth through mentorship, training, and hands-on work on large-scale AI infrastructure.

Skills

Production Engineering
GPU workloads
Linux/Unix
Infrastructure roles
Kubernetes
Cloud fundamentals (AWS/GCP)
SRE/ITIL familiarity
Monitoring (Prometheus/Grafana)
Infrastructure as code (Terraform/Ansb

Tools

Prometheus
Grafana
Terraform
Ansible
Kubernetes

Job description

Crusoe is building a high-performance GPU-focused cloud platform and is seeking a Production Engineer to ensure reliability and scalability of the AI infrastructure. You will lead efforts in incident response, observability, and automation across large-scale systems.

You will collaborate with cross-functional teams to reduce toil, improve SLIs/SLOs, and implement self-healing techniques while growing your technical depth in a dynamic, energy‑driven environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Cloud Platform Engineer - Reliability & Scale
Senior AI Cloud Platform Engineer - Reliability & Scale

Crusoe • San Francisco (CA)

On-site
USD 170,000 - 205,000
Health insurance
RSUs
401(k) match
+2
Principal Engineer - AI Infrastructure Platform
Principal Engineer - AI Infrastructure Platform

Crusoe • San Francisco (CA)

On-site
USD 285,000 - 335,000
Competitive compensation
Equity packages
Paid time off & holidays
+3
Senior GPU Infra Engineer - Automation & Diagnostics
Senior GPU Infra Engineer - Automation & Diagnostics

Crusoe • United States

On-site
USD 250,000 - 300,000
Industry competitive pay
RSUs in a fast-growing tech company
Health insurance with family options
+3
Senior GPU DC Infra Engineer — Automation & Diagnostics
Senior GPU DC Infra Engineer — Automation & Diagnostics

Crusoe • San Francisco (CA)

On-site
USD 215,000 - 260,000
Industry competitive pay
Restricted Stock Units
Health insurance options (HDHP/PPO)
+11
Senior Hardware Systems Engineer, AI Compute & Sustaining
Senior Hardware Systems Engineer, AI Compute & Sustaining

Crusoe Energy Systems LLC • San Francisco (CA)

On-site
USD 215,000 - 260,000
Senior Software Engineer — GPU Data Center Automation
Senior Software Engineer — GPU Data Center Automation

Crusoe • San Francisco (CA)

On-site
USD 170,000 - 205,000
Health insurance package options
Restricted Stock Units
401(k) with match up to 4%
+2
Senior Hardware Systems Engineer, AI Compute & Production
Senior Hardware Systems Engineer, AI Compute & Production

Crusoe • San Francisco (CA)

On-site
USD 215,000 - 260,000
Restricted Stock Units
Senior Deployment Automation Engineer - AI Cloud GPUs
Senior Deployment Automation Engineer - AI Cloud GPUs

Crusoe • Bellevue (WA)

On-site
USD 250,000 - 300,000
Stock options
Paid time off
Health insurance
+5
Senior Cloud Support Engineer - AI HPC Infra
Senior Cloud Support Engineer - AI HPC Infra

Crusoe • Dallas (TX)

On-site
USD 105,000 - 125,000
Competitive compensation
Restricted Stock Units
Paid time off
+12
Senior GPU Data Center Operations Engineer
Senior GPU Data Center Operations Engineer

Crusoe • Denver (CO)

On-site
USD 150,000 - 170,000
Equity
Paid time off
Health insurance
+1