Senior Site Reliability Engineer — Token Factory (Inference Platform)

Jobgether

Ireland

On-site

EUR 120,000 - 180,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Competitive compensation
Learning opportunities
Ownership in your work
International team

Job summary

Token Factory in Ireland seeks a Senior Site Reliability Engineer to own the reliability, performance, and observability of a large-scale AI inference platform. You will optimize GPU-heavy workloads and ensure high-throughput APIs meet reliability and cost targets.

You'll design telemetry pipelines, build self-healing systems, and work closely with software and infrastructure teams to scale the platform. The role demands deep Kubernetes, Terraform, Prometheus, and Grafana expertise in a

Qualifications

  • Significant experience in Site Reliability Engineering, Production Engineering, DevOps, or related infra discipline.
  • Deep practical knowledge of Kubernetes in production environments.
  • Strong experience with Prometheus and Grafana for monitoring and observability.
  • Advanced experience with Terraform and IaC practices.
  • Scripting skills with Python and/or Bash.
  • Solid understanding of distributed systems and failure modes.
  • Experience designing alerts, monitoring strategies, and SLOs for high-throughput services.

Responsibilities

  • Own the reliability, performance, and observability of the inference platform and its infra.
  • Design and improve telemetry pipelines for metrics, logs, and traces.
  • Build monitoring solutions for large production signals.
  • Configure Kubernetes for high availability and efficient GPUs.
  • Tune autoscaling to optimize GPU resource utilization.
  • Develop Terraform modules and IaC patterns for clusters and services.
  • Improve request routing, retry, and failure handling mechanisms.
  • Create automation and tooling to detect and remediate incidents.

Skills

Kubernetes
Prometheus & Grafana
Terraform
Python Bash scripting
GPU-heavy workloads
Incident management
Observability
SLOs & alerts

Tools

Kubernetes
Terraform
Prometheus
Grafana
vLLM/Triton/Ray

Job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer — Token Factory (Inference Platform) based in Ireland.

This is a senior engineering role focused on the reliability, performance, and observability of a large-scale AI inference platform.
You will help operate infrastructure serving foundation models across text, vision, audio, and emerging multimodal workloads.
The role combines Kubernetes, infrastructure-as-code, observability, automation, and production incident management at significant scale.
You will optimize GPU-heavy workloads, strengthen resilience, and ensure high-throughput APIs meet demanding reliability and cost targets.
You will work closely with software engineers and infrastructure teams to build self-healing systems and robust operational processes.
The environment is fast-moving, highly technical, international, and focused on solving complex infrastructure challenges for the AI ecosystem.
This is an opportunity to have a direct impact on the infrastructure powering next-generation AI applications.

Accountabilities
  • Own the reliability, performance, and observability of the inference platform and its supporting infrastructure.
  • Design, implement, and continuously improve telemetry pipelines covering metrics, logs, and traces.
  • Build monitoring and observability solutions capable of processing large volumes of production signals and converting them into actionable insights.
  • Configure and optimize Kubernetes infrastructure for high availability, scalability, and efficient resource utilization.
  • Tune Kubernetes autoscaling mechanisms to improve the efficiency and utilization of GPU resources.
  • Develop and maintain Terraform modules and infrastructure-as-code patterns that embed resilience and reliability into new clusters and services.
  • Design and improve request‑routing, retry, and failure‑handling mechanisms to minimize the impact of transient infrastructure or service failures.
  • Develop automation and operational tooling to detect, isolate, and remediate incidents quickly.
  • Create, maintain, and improve runbooks for incident response and operational procedures.
  • Participate in production incident management, troubleshooting issues and restoring services within demanding reliability objectives.
  • Lead or contribute to post‑mortem processes and implement corrective actions to prevent recurring incidents.
  • Define and improve reliability practices for high‑throughput APIs, including alerting strategies and Service Level Objectives (SLOs).
  • Investigate distributed‑system failures and performance issues across infrastructure and application layers.
  • Optimize systems from the kernel and infrastructure layer through to the application layer.
  • Support and improve the operation of GPU‑intensive inference workloads and accelerator‑based infrastructure.
  • Contribute to scaling the inference platform while balancing performance, reliability, and infrastructure costs.
  • Collaborate closely with software engineers to incorporate reliability and operational excellence into product and platform development.
  • Promote automation, self‑healing capabilities, and engineering practices that reduce operational overhead and improve system resilience.
Requirements:
  • Significant experience in Site Reliability Engineering, Production Engineering, DevOps, or a closely related infrastructure discipline.
  • Deep practical knowledge of Kubernetes in production environments.
  • Strong experience with Prometheus and Grafana for monitoring, metrics, dashboards, and observability.
  • Advanced experience with Terraform and infrastructure-as-code practices.
  • Strong scripting and automation skills using Python and/or Bash.
  • Solid understanding of distributed systems and the ways production backends can fail under real-world conditions.
  • Experience designing effective alerts, monitoring strategies, and SLOs for high‑throughput services or APIs.
  • Strong troubleshooting and debugging skills across infrastructure, networking, operating systems, and application layers.
  • Experience designing systems for high availability, resilience, scalability, and graceful failure recovery.
  • Hands‑on experience with GPU‑heavy workloads or accelerator-based infrastructure is highly valuable.
  • Familiarity with GPU inference technologies such as vLLM, Triton, Ray, or comparable accelerator and model‑serving stacks.
  • Experience with MLOps, model hosting, AI infrastructure, or machine‑learning platforms is advantageous.
  • Strong understanding of infrastructure automation, deployment, configuration management, and operational tooling.
  • Ability to analyze complex performance and reliability problems and translate findings into practical engineering improvements.
  • Strong incident‑management and root‑cause‑analysis capabilities.
  • Ability to collaborate effectively with software engineers and other technical teams to integrate reliability into platform development.
  • Proactive mindset with a strong focus on automation, self‑healing systems, and continuous improvement.
  • Comfortable working independently, taking ownership of critical infrastructure, and operating effectively in a fast‑paced technical environment.
Benefits:
  • Competitive compensation.
  • Career growth and continuous learning opportunities.
  • Flexibility and significant ownership in your work.
  • Collaborative and innovative international working environment.
  • Opportunity to work on high‑impact AI infrastructure and inference technologies.
  • Exposure to large‑scale GPU infrastructure and complex distributed systems.
  • Opportunity to contribute to infrastructure supporting next‑generation multimodal AI applications.
  • Diverse and highly technical international teams.
  • Inclusive workplace committed to equal employment opportunities.
  • Workplace accommodations available throughout the application process where required.
  • Employment is subject to authorization to work in the country where the position is based.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Systems Engineer (Inference
Senior ML Systems Engineer (Inference

TensorX • Dublin

Hybrid
EUR 120,000 - 180,000
Hybrid work
25 days paid annual leave
Central Dublin office
+1
Senior ML Systems Engineer (Inference)
Senior ML Systems Engineer (Inference)

Uniting Holding • Dublin

Hybrid
EUR 80,000 - 120,000
25 days paid annual leave
Free inference tokens
Remote work flexibility
Senior Machine Learning Engineer (Inference)
Senior Machine Learning Engineer (Inference)

TensorX • Dublin

Hybrid
EUR 120,000 - 180,000
Hybrid working from Dublin office
25 days paid annual leave
Free inference tokens
Senior Infrastructure Engineer (GPU Cloud)
Senior Infrastructure Engineer (GPU Cloud)

Uniting Holding • Dublin

Hybrid
EUR 75,000 - 95,000
25 days paid annual leave
Free inference tokens
Remote flexibility
Senior Software Engineer (Token Factory)
Senior Software Engineer (Token Factory)

Jobgether • Ireland

On-site
EUR 90,000 - 120,000
Competitive compensation
Career growth and learning
Ownership over your work
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

United States Digital Space LLC • Dublin

On-site
EUR 76,000 - 105,000
Equity
Bonus
Healthcare
Engineering Manager
Engineering Manager

TensorX • Dublin

Hybrid
EUR 120,000 - 180,000
Hybrid working from Dublin office
25 days paid annual leave
Free inference tokens
Engineering Manager
Engineering Manager

Uniting Holding • Dublin

Hybrid
EUR 120,000 - 180,000
Hybrid Dublin office
Free inference tokens
25 days annual leave
Staff Site Reliability Engineer
Staff Site Reliability Engineer

United States Digital Space LLC • Dublin

On-site
EUR 92,000 - 127,000
Senior Staff Embedded Software Engineer – AI Accelerator
Senior Staff Embedded Software Engineer – AI Accelerator

Qualcomm • Cork

On-site
EUR 90,000 - 120,000
Salary, stock and performance-related bonus
Maternity/Paternity Leave
Employee stock purchase scheme
+7