Senior Site Reliability Engineer

Domyn

Milano

In loco

EUR 50.000 - 70.000

Tempo pieno

14 giorni+

Ricevi più risposte dai datori di lavoro

Invia un CV specifico per questa offerta in pochi minuti.

Vantaggi offerti da questo lavoro

Learning Friday
Smart Working
Equity / stock options

Descrizione del lavoro

Domyn in Milan is seeking an experienced Site Reliability Engineer to join our team. You will design observability and control mechanisms to extract operational data from infrastructure and feed it into automated systems to optimize power, cooling and service budgets.

You will guard and maintain these budgets as part of daily reliability and performance management, contribute to blameless post-mortem analysis, and collaborate with Platform Engineering in a cybersecurity-focused model to ensure

Competenze

  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
  • At least 6 years of experience as a Site Reliability Engineer or in similar roles.
  • Proficient with observability/monitoring tools: Prometheus, Thanos, Grafana, OpenTelemetry.
  • Experience with low-level instrumentation using eBPF.
  • Experience with security monitoring tools like Zeek or Wazuh.
  • Strong Kubernetes experience in cloud-native environments.
  • Strong Python development for automation and tooling.
  • Experience integrating heterogeneous infrastructure across multiple vendors.
  • Familiarity with MCP/A2A or similar agent-based infra frameworks.
  • Exposure to NVIDIA Omniverse or simulation platforms.

Mansioni

  • Design and implement observability and control mechanisms that extract operational data from infrastructure and feed it into automated systems.
  • Guard and maintain operational budgets such as power, cooling and service level objectives.
  • Contribute to blameless post-mortems and structured incident learning.
  • Collaborate with Platform Engineering in a security-focused, shared cybersecurity model.
  • Design, build, and maintain software-driven infrastructure solutions in large-scale environments.

Conoscenze

Python
Observability
Systems thinking
Collaboration
English fluency

Formazione

Bachelor's or Master's in CS/CE/EE

Strumenti

Prometheus
Thanos
Grafana
OpenTelemetry
eBPF
Zeek
Wazuh
Kubernetes

Descrizione del lavoro

We are looking for an experienced Site Reliability Engineer to join our growing team in Milan and help shape the future of our flagship project, Colosseum, one of Europe’s most powerful AI supercomputers, currently in development.

Designed to run our proprietary AI models at scale, it forms the compute backbone behind the intelligence we deliver to the world’s most demanding industries.

In this role, you will design and implement observability and control mechanisms that extract operational data from infrastructure and feed it into automated systems to enable continuous optimization, including key system budgets such as power, cooling and service level, security-level objectives.

You will be responsible for actively guarding and maintaining these operational budgets as part of day-to-day system reliability and performance management.

You will also contribute to operational excellence through blameless post-mortem analysis and structured incident learning, ensuring continuous improvement of system behavior and resilience.

As a part of the team, you will work closely with Platform Engineering in a shared cybersecurity model, where SRE focuses on detection and monitoring, while Platform Engineering ensures the secure design and operation of the underlying infrastructure.

What You Have
  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
  • At least 6 years of experience as a Site Reliability Engineer or in similar roles.
  • Strong experience with observability and monitoring systems such as Prometheus, Thanos, Grafana, and OpenTelemetry
  • Experience with low-level system instrumentation and performance visibility using technologies such as eBPF
  • Experience with security monitoring and threat detection tools such as Zeek, Wazuh, or equivalent SIEM / security observability platforms
  • Strong experience with containerized and cloud-native environments, particularly Kubernetes
  • Strong software development skills, particularly in Python, with the ability to build automation, integrations, and custom tooling
  • Experience integrating heterogeneous infrastructure systems across multiple vendors, APIs, and evolving tool ecosystems
  • Familiarity with modern infrastructure automation and emerging agent-based frameworks such as MCP / A2A (or equivalent technologies)
  • Exposure to digital twin technologies and simulation platforms such as NVIDIA Omniverse or equivalent
  • Strong ability to design, build, and maintain software-driven infrastructure solutions in complex, large-scale environments
Who You Are
  • A versatile engineer, comfortable operating in complex and fast-paced environments.
  • Driven and fearless, you proactively tackle challenges and overcome obstacles with determination.
  • A systems thinker, capable of understanding the broader architecture and identifying dependencies across platforms and technologies.
  • A collaborative team player who is enthusiastic, curious, and passionate about problem-solving, thriving both independently and within cross-functional teams.
  • An effective communicator with strong interpersonal skills, able to engage with diverse stakeholders and foster collaboration.
  • Fluent in English and eager to contribute in a multicultural and international environment.
Benefits
Perks
  • Learning Friday. If our team members know more, so do we. That’s why we give everyone a training budget that they can spend on books, online courses or other training materials.
  • Smart Working. Trains can be a drag, you can save some commuting time by working from home.
  • Salary is based on experience and topped up with other bonuses.

We offer a competitive salary, as well as an opportunity to receive company equity. The typical salary for this role ranges between € 50.000 and € 70.000. As you gain experience and make more significant contributions to the business, your compensation will be reviewed to match your impact. Additionally, depending on your seniority and your performance, you'll have the opportunity to receive stock options, with a variable value calculated from your base salary, giving you the chance to directly participate in the company’s success.

About Domyn

Domyn is a company specializing in the research and development of Responsible AI for regulated industries, including financial services, government, and heavy industry. It supports enterprises with proprietary, fully governable solutions based on a composable AI architecture — including LLMs, AI agents, and one of the world’s largest supercomputers. At the core of Domyn's product offer is a chip-to-frontend architecture that allows organizations to control the entire AI stack — from hardware to application — ensuring isolation, security, and governance throughout the AI lifecycle. Its foundational LLMs, Domyn Large and Domyn Small, are designed for advanced reasoning and optimized to understand each business's specific language, logic, and context. Provided under an open-enterprise license, these models can be fully transferred and owned by clients. Once deployed, they enable customizable agents that operate on proprietary data to solve complex, domain‑specific problems. All solutions are managed via a unified platform with native tools for access management, traceability, and security. Powering it all, Colosseum — a supercomputer in development using NVIDIA Grace Blackwell Superchips — will train next-gen models exceeding 1T parameters. Domyn partners with Microsoft, NVIDIA, and G42. Clients include Allianz, Intesa Sanpaolo, and Fincantieri.

Please review our Privacy Policy here https://bit.ly/4tndszN .

Ottieni la revisione del curriculum gratis e riservata.
o trascina qui il file.
Similar jobs

Offerte di lavoro simili che vale la pena confrontare

Senior AI Cloud Engineer
Senior AI Cloud Engineer

Domyn • Milano

In loco
EUR 55.000 - 75.000
Training budget
Smart Working
Stock options
Senior Data Center Operations Engineer
Senior Data Center Operations Engineer

Domyn • Milano

In loco
EUR 50.000 - 70.000
Learning budget
Smart Working
Stock options
Junior AI Research Engineer
Junior AI Research Engineer

Domyn • Milano

In loco
EUR 30.000 - 50.000
Learning Friday
Smart Working
Stock options
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

PLP Group • Milano

In loco
EUR 50.000 - 70.000
Learning Friday
Smart Working
Equity opportunity
AI Research Engineer
AI Research Engineer

Domyn • Milano

Ibrido
EUR 50.000 - 80.000
Learning Budget
Smart Working
Stock options
Join Our Team - Submit Your Spontaneous Application
Join Our Team - Submit Your Spontaneous Application

Domyn • Milano

In loco
Technical Project Manager (HPC infrastructure)
Technical Project Manager (HPC infrastructure)

Domyn • Milano

In loco
EUR 30.000 - 60.000
Learning Friday
Smart Working
Competitive salary + equity
Senior Platform Engineer
Senior Platform Engineer

PLP Group • Milano

Ibrido
EUR 50.000 - 70.000
Learning Friday
Smart Working
Equity
Corporate Development Manager
Corporate Development Manager

Domyn • Turbigo

In loco
EUR 60.000 - 70.000
Competitive salary
Opportunity to receive company equity
Training budget for professional development
+1
Open Call for Talent: Submit Your Spontaneous Application
Open Call for Talent: Submit Your Spontaneous Application

Domyn • Milano

In loco