Remote-First Senior SRE: Scalable AI Infra & Reliability

Runware

Italia

In loco

EUR 90.000 - 140.000

Tempo pieno

14 giorni+
Generatore di candidature

Fatti notare per questo impiego — genera un curriculum e una lettera di presentazione personalizzati in circa un minuto.

Supera i filtri ATS

Vantaggi offerti da questo lavoro

Generous paid time off
Meaningful stock options
Remote-first setup
Flexible hours
Family leave
Company retreats

Descrizione del lavoro

Runware is seeking a Site Reliability Engineer to keep our production services reliable, scalable and efficient as we grow across image, video and AI workloads. This is a highly technical, hands-on role spanning software, infrastructure and production operations.

You will own reliability practices, lead incidents, and partner with Engineering and DevOps to improve observability, capacity planning and system resilience in a remote-first, Europe-friendly environment.

Competenze

  • Experience operating production systems at scale in an SRE/Production/Platform role.
  • Strong understanding of distributed systems and debugging across applications, databases, queues and networks.
  • Experience designing and operating observability using metrics, logs and distributed tracing.

Mansioni

  • Own and improve the reliability, availability and performance of critical production services across the Runware platform.
  • Define and evolve reliability practices, including SLIs, SLOs, alerting, observability and production-readiness standards.
  • Investigate complex production issues across distributed systems, APIs, networking, queues, databases and GPU-backed workloads, participating in our engineering on-call rotation.
  • Lead and contribute to incident reviews and RCAs, turning recurring failure modes into lasting engineering improvements.
  • Reduce operational toil through automation, automated remediation and improvements to deployment safety, recovery and system resilience.
  • Work closely with Engineering and DevOps teams on capacity planning, performance, scaling and architectural improvements as the platform grows.

Conoscenze

SRE fundamentals
Distributed systems
Observability
On-call rotation
Incident management
Automation scripting
Kubernetes

Strumenti

Kubernetes
IaC
Docker
CI/CD

Descrizione del lavoro

Runware is seeking a Site Reliability Engineer to keep our production services reliable, scalable and efficient as we grow across image, video and AI workloads. This is a highly technical, hands-on role spanning software, infrastructure and production operations.

You will own reliability practices, lead incidents, and partner with Engineering and DevOps to improve observability, capacity planning and system resilience in a remote-first, Europe-friendly environment.

Ottieni la revisione del curriculum gratis e riservata.
o trascina qui il file.
Similar jobs

Offerte di lavoro simili che vale la pena confrontare

Remote SRE - Cloud Reliability Engineer (Mid-Level)
Remote SRE - Cloud Reliability Engineer (Mid-Level)

Improove • Bologna

In loco
EUR 35.000 - 50.000
ESOP
Remote-friendly environment
Flexible schedules
Remote Storage SRE: Build Resilient, Self-Serve Infra
Remote Storage SRE: Build Resilient, Self-Serve Infra

Qonto • Milano

In loco
EUR 70.000 - 110.000
Fully remote team across Europe
Ownership from day one
Team growth opportunities
Senior SRE/DevOps Engineer — Remote, Cloud & Edge Infra
Senior SRE/DevOps Engineer — Remote, Cloud & Edge Infra

E80 Group • Reggio Emilia

Remoto
EUR 70.000 - 110.000
Permanent contract
Full time
Full remote
+1
Senior Site Reliability Engineer - Remote & Scalable Infra
Senior Site Reliability Engineer - Remote & Scalable Infra

Prima • Turbigo

Ibrido
EUR 45.000 - 70.000
Senior SRE & DevOps Engineer - Remote
Senior SRE & DevOps Engineer - Remote

E80 Group • Parma

In loco
EUR 40.000 - 60.000
Permanent contract
Full remote
Senior SRE: AI-Scale Infra, Observability & Resilience
Senior SRE: AI-Scale Infra, Observability & Resilience

Domyn • Milano

In loco
EUR 50.000 - 70.000
Learning Friday
Smart Working
Stock options
SRE: Build Scalable, Reliable Cloud Infra
SRE: Build Scalable, Reliable Cloud Infra

Cacheflow • Milano

In loco
EUR 90.000 - 130.000
Senior IT/Ops Lead (SRE) — Global SaaS Reliability
Senior IT/Ops Lead (SRE) — Global SaaS Reliability

ToolsGroup • Milano

In loco
EUR 55.000 - 68.000
Bonus up to 10%
Senior SRE - AI-Powered Identity Infra on EKS
Senior SRE - AI-Powered Identity Infra on EKS

SecuredTouch (acquired by Ping Identity) • Milano

In loco
EUR 90.000 - 140.000
Generous PTO & Holiday Schedule
Parental Leave
Progressive Healthcare Options
+3
Senior ML SRE: Platform Reliability & Growth
Senior ML SRE: Platform Reliability & Growth

Helloprima • Milano

In loco
EUR 90.000 - 130.000
Private healthcare
Gym discounts
Wellbeing programs
+1