Senior SRE — AI Inference Platform (GPU/Kubernetes)

Jobgether

Ireland

On-site

EUR 120,000 - 180,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Competitive compensation
Learning opportunities
Ownership in your work
International team

Job summary

Token Factory in Ireland seeks a Senior Site Reliability Engineer to own the reliability, performance, and observability of a large-scale AI inference platform. You will optimize GPU-heavy workloads and ensure high-throughput APIs meet reliability and cost targets.

You'll design telemetry pipelines, build self-healing systems, and work closely with software and infrastructure teams to scale the platform. The role demands deep Kubernetes, Terraform, Prometheus, and Grafana expertise in a

Qualifications

  • Significant experience in Site Reliability Engineering, Production Engineering, DevOps, or related infra discipline.
  • Deep practical knowledge of Kubernetes in production environments.
  • Strong experience with Prometheus and Grafana for monitoring and observability.
  • Advanced experience with Terraform and IaC practices.
  • Scripting skills with Python and/or Bash.
  • Solid understanding of distributed systems and failure modes.
  • Experience designing alerts, monitoring strategies, and SLOs for high-throughput services.

Responsibilities

  • Own the reliability, performance, and observability of the inference platform and its infra.
  • Design and improve telemetry pipelines for metrics, logs, and traces.
  • Build monitoring solutions for large production signals.
  • Configure Kubernetes for high availability and efficient GPUs.
  • Tune autoscaling to optimize GPU resource utilization.
  • Develop Terraform modules and IaC patterns for clusters and services.
  • Improve request routing, retry, and failure handling mechanisms.
  • Create automation and tooling to detect and remediate incidents.

Skills

Kubernetes
Prometheus & Grafana
Terraform
Python Bash scripting
GPU-heavy workloads
Incident management
Observability
SLOs & alerts

Tools

Kubernetes
Terraform
Prometheus
Grafana
vLLM/Triton/Ray

Job description

Token Factory in Ireland seeks a Senior Site Reliability Engineer to own the reliability, performance, and observability of a large-scale AI inference platform. You will optimize GPU-heavy workloads and ensure high-throughput APIs meet reliability and cost targets.

You'll design telemetry pipelines, build self-healing systems, and work closely with software and infrastructure teams to scale the platform. The role demands deep Kubernetes, Terraform, Prometheus, and Grafana expertise in a

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Platform SRE — Kubernetes & AI Infra Leader
Senior Platform SRE — Kubernetes & AI Infra Leader

United States Digital Space LLC • Dublin

On-site
EUR 92,000 - 127,000
Senior SRE: Kubernetes Platform & AI Workflows
Senior SRE: Kubernetes Platform & AI Workflows

United States Digital Space LLC • Dublin

On-site
EUR 76,000 - 105,000
Equity
Bonus
Healthcare
Senior Site Reliability Engineer — Token Factory (Inference Platform)
Senior Site Reliability Engineer — Token Factory (Inference Platform)

Jobgether • Ireland

On-site
EUR 120,000 - 180,000
Competitive compensation
Learning opportunities
Ownership in your work
+1
Senior SRE: AI Platform, DevOps & Automation Lead
Senior SRE: AI Platform, DevOps & Automation Lead

RECRUITERS • Dublin

On-site
EUR 111,000 - 150,000
Senior SRE: Scale Global Kubernetes & Cloud Reliability
Senior SRE: Scale Global Kubernetes & Cloud Reliability

Shutterstock, Inc • Dublin

On-site
EUR 95,000 - 150,000
Lead AI/ML SRE — Scale & Production Excellence
Lead AI/ML SRE — Scale & Production Excellence

Mastercard • Greystones

On-site
EUR 110,000 - 140,000
Senior SRE Lead: Reliability, Observability & Automation
Senior SRE Lead: Reliability, Observability & Automation

GCS Recruitment • Ireland

On-site
EUR 90,000 - 130,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

United States Digital Space LLC • Dublin

On-site
EUR 92,000 - 127,000
Senior SRE: AI-Driven Reliability & Platform Leadership
Senior SRE: AI-Driven Reliability & Platform Leadership

JPMorganChase • Dublin

On-site
EUR 90,000 - 150,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

United States Digital Space LLC • Dublin

On-site
EUR 76,000 - 105,000
Equity
Bonus
Healthcare