Senior SRE — AI Inference Platform (GPU/Kubernetes)

Jobgether

España

Presencial

EUR 90.000 - 150.000

Jornada completa

Hace 8 días
Generador de candidaturas

Transforma esta oferta en una entrevista — un currículum y una carta de presentación creados pensando en lo que quiere el empleador.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Competitive compensation
Learning opportunities
Ownership in work
International team
High-impact AI infrastructure
GPU infrastructure

Descripción de la vacante

Jobgether in Spain is seeking a Senior Site Reliability Engineer for the Token Factory (Inference Platform). You will own the reliability, performance, and observability of a large-scale AI inference system, focusing on Kubernetes, IaC, and incident management.

You will design telemetry, build runbooks, and partner with software teams to enable self-healing, cost-aware, high-availability services running GPU-heavy workloads with vLLM, Triton, Ray.

Formación

  • Significant experience in Site Reliability Engineering, Production Engineering, DevOps, or related infrastructure discipline.
  • Deep practical knowledge of Kubernetes in production environments.
  • Strong experience with Prometheus and Grafana for monitoring and observability.
  • Advanced experience with Terraform and infrastructure-as-code practices.
  • Strong scripting and automation skills using Python and/or Bash.
  • Solid understanding of distributed systems and failure modes in production backends.

Responsabilidades

  • Own the reliability, performance, and observability of the inference platform and its supporting infrastructure.
  • Design, implement, and improve telemetry pipelines covering metrics, logs, and traces.
  • Build monitoring and observability solutions capable of processing large production signals into actionable insights.
  • Configure and optimize Kubernetes infrastructure for high availability and efficient resource use.
  • Tune Kubernetes autoscaling to optimize GPU resource utilization.
  • Develop and maintain Terraform modules and IaC patterns embedding resilience and reliability.

Conocimientos

Site Reliability Engineering
Kubernetes
Prometheus & Grafana
Terraform
Python/Bash scripting
Incident management
Observability
Automation
GPU-heavy workloads

Herramientas

Terraform
Kubernetes
Prometheus
Grafana
Python
Bash
vLLM/Triton/Ray

Descripción del empleo

Jobgether in Spain is seeking a Senior Site Reliability Engineer for the Token Factory (Inference Platform). You will own the reliability, performance, and observability of a large-scale AI inference system, focusing on Kubernetes, IaC, and incident management.

You will design telemetry, build runbooks, and partner with software teams to enable self-healing, cost-aware, high-availability services running GPU-heavy workloads with vLLM, Triton, Ray.

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Site Reliability Engineer — Token Factory (Inference Platform)
Senior Site Reliability Engineer — Token Factory (Inference Platform)

Jobgether • España

Presencial
EUR 90.000 - 150.000
Competitive compensation
Learning opportunities
Ownership in work
+3
Senior SRE - AI GPU Infra & HPC Orchestrator (Remote EU)
Senior SRE - AI GPU Infra & HPC Orchestrator (Remote EU)

Hamilton Barnes ? • España

Presencial
EUR 120.000 - 190.000
Senior SRE - Compute Node & Linux Virtualization
Senior SRE - Compute Node & Linux Virtualization

Jobgether • España

Presencial
EUR 85.000 - 125.000
Competitive compensation
Career growth
Ownership and autonomy
+1
Senior AI Infra & Platform SRE — Remote EU
Senior AI Infra & Platform SRE — Remote EU

Front Door Defense • Barcelona

A distancia
EUR 90.000 - 130.000
Remote work flexibility
Competitive compensation
Career growth opportunities
Senior SRE - AI Cloud Infra, Kubernetes & GPUs
Senior SRE - AI Cloud Infra, Kubernetes & GPUs

Mirantis • Bellprat

Presencial
EUR 90.000 - 120.000
Competitive compensation package
Professional development and training
Conference attendance and tech talks
+1
Senior Reliability & Platform Engineer — Remote
Senior Reliability & Platform Engineer — Remote

Jobgether • España

A distancia
EUR 117.000 - 261.000
Equity participation
Fully remote
Health, dental, and vision
+1
Senior SRE — Hybrid, Barcelona IoT Platform Reliability
Senior SRE — Hybrid, Barcelona IoT Platform Reliability

Tamarind Intelligence • Barcelona

Híbrido
EUR 55.000 - 70.000
Hybrid work model
Salary 55k-70k€ annually
Comprehensive health insurance
Platform SRE: AI-Driven Kubernetes & GitOps
Platform SRE: AI-Driven Kubernetes & GitOps

T-Systems Iberia • Valencia

Híbrido
EUR 70.000 - 97.000
Health insurance
Meal vouchers
Childcare
+4
SRE — AI Platform (Hybrid/Remote-Option)
SRE — AI Platform (Hybrid/Remote-Option)

Akamai Career Site • España

Híbrido
EUR 60.000 - 90.000
FlexBase program
Benefits package
Hybrid work options
Senior Platform Engineer — AI-First Infra & Scale
Senior Platform Engineer — AI-First Infra & Scale

Haddock • Barcelona

Híbrido
EUR 50.000 - 60.000
Hybrid work
Office in Barcelona Poblenou
AI tooling access
+2