Senior SRE - AI Inference Platform (GPU, K8s)

Jobgether

Lavamünd

Vor Ort

EUR 127.000 - 190.000

Vollzeit

Vor 3 Tagen
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Hebe dich für diese Rolle von der Masse ab — erstelle in etwa einer Minute einen maßgeschneiderten Lebenslauf und ein Anschreiben.

Schaffe es an den ATS-Filtern vorbei

Benefits dieser Stelle

Competitive compensation
Career growth
Learning opportunities
Ownership in work
International teams

Zusammenfassung

Token Factory is seeking a Senior Site Reliability Engineer to own reliability, performance, and observability for a large-scale AI inference platform. You will design telemetry pipelines, monitor production signals, and optimize GPU-heavy workloads with Kubernetes and IaC tooling.

You will collaborate with software and infra teams to build self-healing systems, automate incidents, and improve runbooks and post-mortems to prevent incidents.

Qualifikationen

  • Significant experience in Site Reliability Engineering or Production Engineering.
  • Deep practical knowledge of Kubernetes in production environments.
  • Strong experience with Prometheus and Grafana for monitoring and observability.
  • Advanced experience with Terraform and infrastructure-as-code practices.
  • Strong scripting and automation skills using Python and Bash.
  • Experience designing alerts, SLOs, and monitoring strategies for high-throughput services.
  • Hands-on experience with GPU-heavy workloads or accelerator-based infra is valuable.

Aufgaben

  • Own reliability, performance, and observability of the inference platform.
  • Design telemetry pipelines covering metrics, logs, and traces.
  • Build monitoring solutions for large production signal volumes.
  • Configure Kubernetes for high availability, scalability, and resource efficiency.
  • Tune autoscaling to optimize GPU resources.
  • Develop and maintain Terraform modules and IaC patterns.
  • Improve request-routing, retries, and failure handling.
  • Develop automation to detect, isolate, and remediate incidents.
  • Create runbooks for incident response and operations.
  • Lead post-mortems and implement corrective actions.

Kenntnisse

SRE
Kubernetes
Prometheus
Grafana
Terraform
Python
Bash
Incident management
GPU workloads
MLOps

Tools

Kubernetes tooling
Terraform modules

Jobbeschreibung

Token Factory is seeking a Senior Site Reliability Engineer to own reliability, performance, and observability for a large-scale AI inference platform. You will design telemetry pipelines, monitor production signals, and optimize GPU-heavy workloads with Kubernetes and IaC tooling.

You will collaborate with software and infra teams to build self-healing systems, automate incidents, and improve runbooks and post-mortems to prevent incidents.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior Site Reliability Engineer — Token Factory (Inference Platform)
Senior Site Reliability Engineer — Token Factory (Inference Platform)

Jobgether • Lavamünd

Vor Ort
EUR 127.000 - 190.000
Competitive compensation
Career growth
Learning opportunities
+2
Senior SRE: Compute Nodes & Linux Infra
Senior SRE: Compute Nodes & Linux Infra

Jobgether • Lavamünd

Vor Ort
EUR 127.000 - 190.000
Competitive compensation
Career growth
Ownership of meaningful projects
+1
Senior SRE - Kubernetes & Cloud Infra (Remote)
Senior SRE - Kubernetes & Cloud Infra (Remote)

Jobgether • Lavamünd

Vor Ort
EUR 148.000 - 201.000
100% remote work within EU time zones
Flexible working hours
High ownership and autonomy
+4
Senior Site Reliability Engineer (SRE, Compute Node Team)
Senior Site Reliability Engineer (SRE, Compute Node Team)

Jobgether • Lavamünd

Vor Ort
EUR 127.000 - 190.000
Competitive compensation
Career growth
Ownership of meaningful projects
+1
Remote Senior SRE: Lead Reliability & Observability
Remote Senior SRE: Lead Reliability & Observability

Jobgether • Lavamünd

Remote
EUR 154.000 - 209.000
Fully remote working environment.
Be the first SRE; establish org-wide S
Autonomy and influence over practices
Senior Software Engineer, Site Reliability Engineering, GenAI
Senior Software Engineer, Site Reliability Engineering, GenAI

Google • Lavamünd

Vor Ort
EUR 90.000 - 130.000
Remote SRE Engineering Manager | Lead Reliability & Platform
Remote SRE Engineering Manager | Lead Reliability & Platform

Jobgether • Lavamünd

Vor Ort
EUR 65.000 - 146.000
Fully remote environment
Parental leave
Home office budget
Senior AI Inference Engineer - High-Performance GPU Systems
Senior AI Inference Engineer - High-Performance GPU Systems

NVIDIA • Lavamünd

Vor Ort
EUR 120.000 - 160.000
Kubernetes Reliability Engineer
Kubernetes Reliability Engineer

JobsinAustria • Wien

Hybrid
EUR 70.000 - 110.000
Senior AI Inference Architect — Multi-Node GPU Scale
Senior AI Inference Architect — Multi-Node GPU Scale

NVIDIA • Lavamünd

Vor Ort
EUR 110.000 - 170.000