Senior AI Cloud Platform Engineer - Reliability & Scale

Crusoe

San Francisco (CA)

On-site

USD 170,000 - 205,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
RSUs
401(k) match
Parental leave
Cell phone reimbursement

Job summary

Crusoe is hiring a Senior Production Engineer to ensure reliability and scalability of our AI-optimized cloud platform. You will help deliver highly available, cost-efficient infrastructure for latency-sensitive AI workloads and large-scale training and inference clusters.

You will design and operate managed AI services, define SLIs/SLOs, collaborate with AI and platform teams, and drive observability through telemetry and performance tuning.

Qualifications

  • Strong software engineering background and production-grade systems experience.
  • Experience designing distributed systems and reliability improvements.
  • Familiarity with modern programming languages (Python/Go/Java/C++) and container orchestration.

Responsibilities

  • Design and operate reliable managed AI services serving and scaling LLM workloads.
  • Define, measure, and improve SLIs/SLOs to meet performance and reliability targets.
  • Collaborate with AI, platform, and infrastructure teams to optimize large-scale training and inference clusters.
  • Automate observability by building telemetry and tuning strategies for latency-sensitive services.
  • Investigate and resolve reliability issues in distributed AI systems using telemetry and logs.
  • Contribute to architecture of next-generation distributed systems for AI-first environments.

Skills

Python
Go
Java
C++

Tools

Kubernetes

Job description

Crusoe is hiring a Senior Production Engineer to ensure reliability and scalability of our AI-optimized cloud platform. You will help deliver highly available, cost-efficient infrastructure for latency-sensitive AI workloads and large-scale training and inference clusters.

You will design and operate managed AI services, define SLIs/SLOs, collaborate with AI and platform teams, and drive observability through telemetry and performance tuning.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Production Engineer, AI GPU Cloud Reliability & Automation
Production Engineer, AI GPU Cloud Reliability & Automation

Crusoe • San Francisco (CA)

On-site
USD 172,000 - 209,000
Health insurance
401(k) with match
Paid parental leave
+4
Senior AI Infra Engineer - Scalable LLM Platforms
Senior AI Infra Engineer - Scalable LLM Platforms

Crusoe Energy Systems LLC • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 205,000
Health insurance
401(k) with company match
Paid parental leave
+6
Principal Engineer - AI Infrastructure Platform
Principal Engineer - AI Infrastructure Platform

Crusoe • San Francisco (CA)

On-site
USD 285,000 - 335,000
Competitive compensation
Equity packages
Paid time off & holidays
+3
Senior Cloud Support Engineer - AI HPC Infra
Senior Cloud Support Engineer - AI HPC Infra

Crusoe • Dallas (TX)

On-site
USD 105,000 - 125,000
Competitive compensation
Restricted Stock Units
Paid time off
+12
Senior Cloud Support Engineer - HPC & AI Compute
Senior Cloud Support Engineer - HPC & AI Compute

Crusoe • New York (NY)

On-site
USD 125,000 - 145,000
Competitive compensation and equity p 
Restricted Stock Units
Paid time off, paid holidays & leave 
+13
Engineering Manager, Scalable AI Platform
Engineering Manager, Scalable AI Platform

ProducePay • United States

On-site
USD 215,000 - 260,000
Competitive compensation
Equity packages
Health, dental & vision insurance
+5
Senior AI Customer Success Manager: Production Inference
Senior AI Customer Success Manager: Production Inference

Crusoe • New York (NY)

On-site
USD 190,000 - 215,000
Competitive compensation and equity
Restricted Stock Units
Paid time off, holidays & leave
+11
Senior Solutions Engineer: AI Infra & Kubernetes
Senior Solutions Engineer: AI Infra & Kubernetes

ProducePay • Denver (CO)

On-site
USD 175,000 - 250,000
Restricted Stock Units
Health, dental & vision insurance
401(k) with company match
+1
Senior Cloud Support Engineer, AI/ML HPC Infra
Senior Cloud Support Engineer, AI/ML HPC Infra

Crusoe • Denver (CO)

On-site
USD 105,000 - 125,000
Equity
Paid time off
Health insurance
+8
Senior Production Engineer, Managed Cloud
Senior Production Engineer, Managed Cloud

Crusoe • San Francisco (CA)

On-site
USD 170,000 - 205,000
Health insurance
RSUs
401(k) match
+2