Senior AI Cloud Platform Engineer - Reliability & Scale

Crusoe

San Francisco (CA)

On-site

USD 170,000 - 205,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health insurance
RSUs
401(k) match
Parental leave
Cell phone reimbursement

Job summary

Crusoe is hiring a Senior Production Engineer to ensure reliability and scalability of our AI-optimized cloud platform. You will help deliver highly available, cost-efficient infrastructure for latency-sensitive AI workloads and large-scale training and inference clusters.

You will design and operate managed AI services, define SLIs/SLOs, collaborate with AI and platform teams, and drive observability through telemetry and performance tuning.

Qualifications

  • Strong software engineering background and production-grade systems experience.
  • Experience designing distributed systems and reliability improvements.
  • Familiarity with modern programming languages (Python/Go/Java/C++) and container orchestration.

Responsibilities

  • Design and operate reliable managed AI services serving and scaling LLM workloads.
  • Define, measure, and improve SLIs/SLOs to meet performance and reliability targets.
  • Collaborate with AI, platform, and infrastructure teams to optimize large-scale training and inference clusters.
  • Automate observability by building telemetry and tuning strategies for latency-sensitive services.
  • Investigate and resolve reliability issues in distributed AI systems using telemetry and logs.
  • Contribute to architecture of next-generation distributed systems for AI-first environments.

Skills

Python
Go
Java
C++

Tools

Kubernetes

Job description

Crusoe is hiring a Senior Production Engineer to ensure reliability and scalability of our AI-optimized cloud platform. You will help deliver highly available, cost-efficient infrastructure for latency-sensitive AI workloads and large-scale training and inference clusters.

You will design and operate managed AI services, define SLIs/SLOs, collaborate with AI and platform teams, and drive observability through telemetry and performance tuning.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Cloud Infrastructure Engineer — AI Workloads
Senior Cloud Infrastructure Engineer — AI Workloads

US Health Partners, LLC • Sunnyvale (CA)

On-site
USD 170,000 - 205,000
Restricted Stock Units
Health insurance
HSA contributions
+10
Senior Cloud Infrastructure Engineer — AI-First Platform
Senior Cloud Infrastructure Engineer — AI-First Platform

Crusoe • San Francisco (CA)

On-site
USD 172,000 - 209,000
Restricted Stock Units
Health insurance
401(k) with match
+4
Senior Cloud Infrastructure Engineer for Scalable AI
Senior Cloud Infrastructure Engineer for Scalable AI

Engg • San Francisco (CA)

On-site
USD 170,000 - 205,000
RSUs
Health insurance
401(k) match
+3
Senior Cloud Infra Engineer - AI Platform
Senior Cloud Infra Engineer - AI Platform

Crusoe Energy Systems • Sunnyvale (CA)

On-site
USD 190,000 - 235,000
Health & wellbeing benefits
Paid time off
401(k) match
+1
Staff Cloud Availability Platform Engineer (AI Infra)
Staff Cloud Availability Platform Engineer (AI Infra)

Crusoe Energy Systems • Sunnyvale (CA)

On-site
USD 190,000 - 320,000
Health & wellbeing
Time away
401(k) match
+1
Senior AI Network Reliability Engineer - Automation-Driven
Senior AI Network Reliability Engineer - Automation-Driven

Crusoe • San Francisco (CA)

On-site
USD 165,000 - 200,000
Competitive compensation
Restricted Stock Units
Paid time off & holidays
+10
Senior Network Reliability Engineer — Production & Automation
Senior Network Reliability Engineer — Production & Automation

ProducePay • San Francisco (CA)

On-site
USD 165,000 - 200,000
Competitive compensation
RSUs
Paid time off
+3
Senior Production Engineer, Managed Cloud
Senior Production Engineer, Managed Cloud

Crusoe • San Francisco (CA)

On-site
USD 170,000 - 205,000
Health insurance
RSUs
401(k) match
+2
Senior Cloud Storage Engineer for AI/ML Workloads
Senior Cloud Storage Engineer for AI/ML Workloads

DCYB • San Francisco (CA), Northern (KY)

Hybrid
USD 166,000 - 201,000
Competitive compensation
Restricted Stock Units
Paid time off & holidays
+9
Senior Data Platform Engineer — AI Infrastructure
Senior Data Platform Engineer — AI Infrastructure

Crusoe • San Francisco (CA)

On-site
USD 170,000 - 205,000
Restricted Stock Units
Paid time off & holidays
Comprehensive health, dental & vision