Senior Site Reliability Engineer for AI Supercomputer (Remote)

Domyn

Milano

In loco

EUR 50.000 - 70.000

Tempo pieno

14 giorni+

Ricevi più risposte dai datori di lavoro

Invia un CV specifico per questa offerta in pochi minuti.

Vantaggi offerti da questo lavoro

Learning Friday
Smart Working
Equity / stock options

Descrizione del lavoro

Domyn in Milan is seeking an experienced Site Reliability Engineer to join our team. You will design observability and control mechanisms to extract operational data from infrastructure and feed it into automated systems to optimize power, cooling and service budgets.

You will guard and maintain these budgets as part of daily reliability and performance management, contribute to blameless post-mortem analysis, and collaborate with Platform Engineering in a cybersecurity-focused model to ensure

Competenze

  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
  • At least 6 years of experience as a Site Reliability Engineer or in similar roles.
  • Proficient with observability/monitoring tools: Prometheus, Thanos, Grafana, OpenTelemetry.
  • Experience with low-level instrumentation using eBPF.
  • Experience with security monitoring tools like Zeek or Wazuh.
  • Strong Kubernetes experience in cloud-native environments.
  • Strong Python development for automation and tooling.
  • Experience integrating heterogeneous infrastructure across multiple vendors.
  • Familiarity with MCP/A2A or similar agent-based infra frameworks.
  • Exposure to NVIDIA Omniverse or simulation platforms.

Mansioni

  • Design and implement observability and control mechanisms that extract operational data from infrastructure and feed it into automated systems.
  • Guard and maintain operational budgets such as power, cooling and service level objectives.
  • Contribute to blameless post-mortems and structured incident learning.
  • Collaborate with Platform Engineering in a security-focused, shared cybersecurity model.
  • Design, build, and maintain software-driven infrastructure solutions in large-scale environments.

Conoscenze

Python
Observability
Systems thinking
Collaboration
English fluency

Formazione

Bachelor's or Master's in CS/CE/EE

Strumenti

Prometheus
Thanos
Grafana
OpenTelemetry
eBPF
Zeek
Wazuh
Kubernetes

Descrizione del lavoro

Domyn in Milan is seeking an experienced Site Reliability Engineer to join our team. You will design observability and control mechanisms to extract operational data from infrastructure and feed it into automated systems to optimize power, cooling and service budgets.

You will guard and maintain these budgets as part of daily reliability and performance management, contribute to blameless post-mortem analysis, and collaborate with Platform Engineering in a cybersecurity-focused model to ensure

Ottieni la revisione del curriculum gratis e riservata.
o trascina qui il file.
Similar jobs

Offerte di lavoro simili che vale la pena confrontare

Senior SRE for AI Supercomputer — Remote & Equity
Senior SRE for AI Supercomputer — Remote & Equity

PLP Group • Milano

In loco
EUR 50.000 - 70.000
Learning Friday
Smart Working
Equity opportunity
Senior Data Center Ops Engineer Remote AI Supercomputers
Senior Data Center Ops Engineer Remote AI Supercomputers

PLP Group • Milano

In loco
EUR 50.000 - 70.000
Learning Friday
Smart Working
Equity
Senior AI Cloud Engineer — Remote HPC & AI Platform
Senior AI Cloud Engineer — Remote HPC & AI Platform

PLP Group • Milano

In loco
EUR 55.000 - 75.000
Learning Friday
Smart Working
Equity
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Domyn • Milano

In loco
EUR 50.000 - 70.000
Learning Friday
Smart Working
Equity / stock options
Senior Data Center Engineer — AI Compute Backbone
Senior Data Center Engineer — AI Compute Backbone

Domyn • Milano

In loco
EUR 50.000 - 70.000
Learning budget
Smart Working
Stock options
Senior AI Cloud Engineer - HPC & AI Workloads
Senior AI Cloud Engineer - HPC & AI Workloads

Domyn • Milano

In loco
EUR 55.000 - 75.000
Training budget
Smart Working
Stock options
Senior Data Center Operations Engineer
Senior Data Center Operations Engineer

Domyn • Milano

In loco
EUR 50.000 - 70.000
Learning budget
Smart Working
Stock options
Senior Platform Engineer: AI Infra & Multi-Tenancy Orchestrator
Senior Platform Engineer: AI Infra & Multi-Tenancy Orchestrator

PLP Group • Milano

Ibrido
EUR 50.000 - 70.000
Learning Friday
Smart Working
Equity
Senior Site Reliability Engineer
Senior Site Reliability Engineer

PLP Group • Milano

In loco
EUR 50.000 - 70.000
Learning Friday
Smart Working
Equity opportunity
AI HPC Infrastructure Project Manager
AI HPC Infrastructure Project Manager

PLP Group • Milano

In loco
EUR 30.000 - 60.000
Learning Friday
Smart Working
Stock options