Senior Production Engineer (Core PE)

Crusoe Energy Systems

Ireland

On-site

EUR 90,000 - 130,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health benefits and wellness
Paid time off
401(k) match
Mental health resources

Job summary

Crusoe Energy Systems is seeking a Production Engineer focused on Operational Excellence to ensure reliability and performance of our GPU cloud powering AI workloads. You will reduce toil, automate tasks, and partner with compute, networking, and platform teams to strengthen resilience and disaster recovery.

The role emphasizes building scalable infrastructure, improving incident response, and growing technical depth through mentoring and hands-on work in large-scale AI infrastructure.

Qualifications

  • Understanding cloud infra fundamentals incl. Kubernetes, distributed systems, virtualization, and cloud platforms (AWS/GCP).
  • Experience with monitoring/observability tools such as Prometheus and Grafana; desire to deepen expertise.
  • Experience supporting GPU workloads, HPC, or latency/throughput-sensitive distributed systems.
  • Familiarity with infrastructure-as-code and configuration management tools like Terraform or Ansible.
  • Growth mindset with focus on reliability engineering, automation, and operational excellence.
  • 5+ years in Production Engineering, SRE, or large-scale infrastructure operations.

Responsibilities

  • Improve reliability, scalability, and performance of Crusoe's GPU cloud platform.
  • Collaborate with cross-functional teams to define and evolve availability metrics, SLIs, and SLOs.
  • Participate in incident response, diagnose disruptions, and contribute to root cause analysis.
  • Build, operate, and improve observability—Prometheus, Grafana, Alertmanager, OpenTelemetry.
  • Develop automation and self-healing tooling to reduce toil and improve recovery times.

Skills

Reliability engineering
SRE practices
Scripting (Go/Python/C)
Linux systems

Tools

Kubernetes
Prometheus
Grafana
OpenTelemetry
Terraform
Ansible

Job description

  • Crusoe's building the most reliable, energy-efficient, AI-optimized cloud platform - and Production Engineering sits at the heart of that mission. As a Production Engineer focused on Operational Excellence, you will help ensure the reliability, scalability, and performance of Crusoe's GPU cloud that powers next-generation AI workloads
  • This role is ideal for engineers who enjoy solving complex production problems, improving large-scale distributed systems, and building automation that keeps infrastructure running smoothly. You'll play a key role in strengthening the operational foundation of Crusoe's cloud while helping scale infrastructure that supports demanding AI and HPC workloads
  • You'll partner closely with Production Engineers, infrastructure teams, and platform engineers to improve system reliability, reduce operational toil, and drive continuous improvements across Crusoe's rapidly growing GPU cloud
  • Collaborate with cross-functional teams to define and evolve availability metrics for Crusoe's cloud platform, including establishing, measuring, and improving SLIs and SLOs
  • Participate in production incident response, diagnosing and resolving service disruptions while contributing to post-incident reviews and root cause analysis
  • Build, operate, and improve observability across Crusoe's infrastructure using tools such as Prometheus, Grafana, Alertmanager, and OpenTelemetry
  • Identify reliability risks, performance bottlenecks, and early indicators of potential production issues across distributed systems
  • Develop automation and tooling that reduces operational toil, improves recovery times, and enables self-healing infrastructure
  • Partner with compute, networking, storage, and platform teams to strengthen service resilience and disaster recovery capabilities
  • Contribute to improving operational processes, knowledge sharing, and reliability best practices across the engineering organization
  • Continue growing technical depth through mentorship, training, and hands-on work operating large-scale AI infrastructure
Benefits
  • Health & wellbeing: Comprehensive health benefits designed to support your overall wellness
  • Time away: Paid time off for vacations, family bonding, and unexpected needs
  • 401(k) match: Build your financial future with our 401(k) matching program
  • Mental wellness: Resources and support for your emotional wellbeing and navigating life’s challenges

Understanding of modern cloud infrastructure fundamentals including Kubernetes, distributed systems, virtualization, and cloud platforms (AWS/GCP)Experience with monitoring and observability tools such as Prometheus and Grafana, or a strong desire to deepen expertise in this areaExperience supporting GPU workloads, HPC environments, or latency/throughput-sensitive distributed systemsFamiliarity with infrastructure-as-code and configuration management tools such as Terraform or AnsibleA growth mindset and strong interest in reliability engineering, automation, and operational excellence5+ years of experience in Production Engineering, SRE, or large-scale infrastructure operationsPrevious experience in Infrastructure roles building or managing compute, storage or networking platformsAbility to remain calm and effective while troubleshooting complex issues in high-impact production environmentsFamiliarity with incident management practices and reliability frameworks (SRE, ITIL, or similar)Strong communication skills and the ability to collaborate across engineering teamsScripting or programming experience with languages such as Go, Python, C, or C++Strong knowledge of Linux/Unix systems, including debugging complex issues across kernel and user spaceInterest in scaling AI or HPC infrastructure and solving reliability challenges in GPU-heavy environmentsPassion for mentorship, learning, and developing deeper expertise in Production EngineeringExperience designing self-healing systems, automated remediation, or event-driven operational toolingExperience working with Kubernetes or container orchestration platforms at scaleExposure to change management processes, operational readiness reviews, or structured root cause analysis

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Production Engineer (Kubernetes)
Production Engineer (Kubernetes)

Crusoe Energy Systems LLC • Dublin

On-site
EUR 70,000 - 110,000
Pension contributions
Private health insurance
Income protection
+1
Senior Production Engineer
Senior Production Engineer

Crusoe Energy Systems LLC • Dublin

On-site
EUR 90,000 - 130,000
Pension contributions
Private health insurance
Life assurance
+1
Senior AI Cloud Reliability Engineer
Senior AI Cloud Reliability Engineer

ProducePay • Dublin

On-site
EUR 110,000 - 150,000
Pension contributions
Private health & dental
Income protection
+1
Senior GPU Cloud Reliability Engineer
Senior GPU Cloud Reliability Engineer

Crusoe Energy Systems • Ireland

On-site
EUR 90,000 - 130,000
Health benefits and wellness
Paid time off
401(k) match
+1
Staff Production Engineer
Staff Production Engineer

crusoe • Dublin

On-site
EUR 90,000 - 150,000
Social security coverage
Pension funds
Private health insurance
+3
Senior Cloud Support Engineer
Senior Cloud Support Engineer

Crusoe • Dublin

On-site
EUR 60,000 - 90,000
Pension contributions
Private health insurance
Dental insurance
+2
Senior Network Production Operations Engineer
Senior Network Production Operations Engineer

AI Chopping Block • Dublin

Hybrid
EUR 90,000 - 130,000
Pension contributions
Private health insurance
Dental insurance
+2
Staff Production Engineer
Staff Production Engineer

Linuxconfig • Dublin

Hybrid
EUR 90,000 - 120,000
Social security coverage
Generous leave policies
Private health insurance
+2
Senior Network Production Engineer, Network Ops
Senior Network Production Engineer, Network Ops

ProducePay • Dublin

On-site
EUR 120,000 - 180,000
Pension contributions
Private health insurance
Dental insurance
+2
Senior Network Production Engineer, Network Ops
Senior Network Production Engineer, Network Ops

Crusoe • Ireland

On-site
EUR 90,000 - 130,000
Pension plan
Private health insurance
Dental insurance
+2