Senior AI Cloud Reliability Engineer

ProducePay

Dublin

On-site

EUR 110,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Pension contributions
Private health & dental
Income protection
Life assurance

Job summary

Crusoe is building a high‑performing AI infrastructure and seeks a Production Engineer focused on Operational Excellence to ensure reliability of our GPU cloud. You will drive observability, incident response, and automation across distributed systems, partnering with multiple engineering teams to reduce toil and improve resilience.

Ideal candidates have 5+ years in production engineering or SRE, strong Linux fundamentals, and hands‑on experience with Prometheus, Grafana, and cloud platforms.

Qualifications

  • 5+ years in Production Engineering, SRE, or large-scale infrastructure operations.
  • Experience supporting GPU workloads, HPC environments, or latency/throughput-sensitive distributed systems.
  • Strong knowledge of Linux/Unix systems, including debugging complex issues across kernel and user space.
  • Experience with cloud infrastructure fundamentals including Kubernetes, virtualization, and cloud platforms (AWS/GCP).
  • Familiarity with incident management practices and reliability frameworks (SRE, ITIL, or similar).
  • Experience with monitoring and observability tools such as Prometheus and Grafana.

Responsibilities

  • Collaborate with cross-functional teams to define and evolve availability metrics for Crusoe’s cloud platform, including SLIs and SLOs.
  • Participate in production incident response, diagnosing and resolving service disruptions and performing post-incident reviews.
  • Build, operate, and improve observability across Crusoe’s infrastructure using Prometheus, Grafana, Alertmanager, and OpenTelemetry.
  • Identify reliability risks and performance bottlenecks across distributed systems.
  • Develop automation and tooling to reduce toil and enable self-healing infrastructure.
  • Partner with compute, networking, storage, and platform teams to strengthen service resilience and disaster recovery.
  • Contribute to improving operational processes and reliability best practices across the engineering org.
  • Grow technical depth through mentorship and hands-on work operating large-scale AI infrastructure.

Skills

Production Engineering
SRE
Linux/Unix
Kubernetes
Terraform
Python
Go
Observability
Incident Management

Tools

Prometheus
Grafana
OpenTelemetry
Terraform
Ansible
AWS/GCP

Job description

Crusoe is building a high‑performing AI infrastructure and seeks a Production Engineer focused on Operational Excellence to ensure reliability of our GPU cloud. You will drive observability, incident response, and automation across distributed systems, partnering with multiple engineering teams to reduce toil and improve resilience.

Ideal candidates have 5+ years in production engineering or SRE, strong Linux fundamentals, and hands‑on experience with Prometheus, Grafana, and cloud platforms.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Production Engineer - AI Infra Reliability & Automation
Senior Production Engineer - AI Infra Reliability & Automation

Crusoe • Dublin

On-site
EUR 90,000 - 130,000
Health insurance
Pension plan
Parental leave
+1
Senior Production Engineer — AI Infra Reliability
Senior Production Engineer — AI Infra Reliability

Crusoe Energy Systems LLC • Dublin

On-site
EUR 90,000 - 130,000
Pension contributions
Private health insurance
Life assurance
+1
Senior Production Engineer, AI Infra & Reliability
Senior Production Engineer, AI Infra & Reliability

Crusoe Energy Systems LLC • Dublin

Hybrid
EUR 90,000 - 130,000
Private health insurance
Generous leave policies (maternity, p?
Pension funds
SRE for AI-First Cloud Infrastructure (Hybrid)
SRE for AI-First Cloud Infrastructure (Hybrid)

Dormont Manufacturing Co • Dublin

Hybrid
EUR 75,000 - 105,000
Hybrid work schedule
Competitive Paid Time Off
Retirement benefits
+5
Senior Production Engineer
Senior Production Engineer

Crusoe Energy Systems LLC • Dublin

On-site
EUR 90,000 - 130,000
Pension contributions
Private health insurance
Life assurance
+1
Senior AI Infrastructure Solutions Engineer (Kubernetes/MLOps)
Senior AI Infrastructure Solutions Engineer (Kubernetes/MLOps)

Crusoe • Dublin

On-site
EUR 80,000 - 120,000
Pension contributions
Private health and dental insurance
Income protection
+1
Solutions Engineer
Solutions Engineer

Crusoe Energy Systems • Ireland

On-site
EUR 65,000 - 85,000
Comprehensive health benefits
Paid time off
401(k) matching
+1
Senior Cloud Support Engineer
Senior Cloud Support Engineer

Crusoe • Dublin

On-site
EUR 60,000 - 90,000
Pension contributions
Private health insurance
Dental insurance
+2
Production Engineer (Kubernetes)
Production Engineer (Kubernetes)

Crusoe Energy Systems LLC • Dublin

On-site
EUR 70,000 - 110,000
Pension contributions
Private health insurance
Income protection
+1
Senior Cloud Support Engineer
Senior Cloud Support Engineer

Crusoe Energy Systems • Ireland

On-site
EUR 75,000 - 110,000
Health benefits
Paid time off
401(k) matching