Senior Site Reliability Engineer

Domyn

Milano

In loco

EUR 50.000 - 70.000

Tempo pieno

6 giorni fa
Candidati tra i primi

Ricevi più risposte dai datori di lavoro

Invia un CV specifico per questa offerta in pochi minuti.

Vantaggi offerti da questo lavoro

Learning Friday
Smart Working
Stock options

Descrizione del lavoro

Domyn in Milan is seeking an experienced Site Reliability Engineer to help shape the future of Colosseum, our flagship AI supercomputer in development. You will design observability and control mechanisms, extract operational data, and feed it into automated systems to optimize power, cooling, and service levels.

You will guard budgets, improve system resilience, and participate in blameless post-mortems. Join a cross-functional team with Platform Engineering to ensure secure, scalable

Competenze

  • Bachelor’s or Master’s degree in CS/CE/EE or related field.
  • 6+ years as an SRE or similar role.
  • Strong experience with Prometheus, Thanos, Grafana, OpenTelemetry.
  • Experience with low-level instrumentation eBPF.
  • Security monitoring tools such as Zeek or Wazuh.
  • Kubernetes and cloud-native environments; Python automation.

Mansioni

  • Design and implement observability and control mechanisms for infrastructure.
  • Extract operational data and feed it into automated systems for optimization.
  • Guard and maintain operational budgets (power, cooling, SLOs).
  • Contribute to blameless post-mortems and continuous improvement.

Conoscenze

Observability systems
Kubernetes
Prometheus
OpenTelemetry
Python automation
eBPF instrumentation
Security monitoring
Cloud-native

Formazione

Bachelor’s or Master’s degree in Computer Science/Engineering

Strumenti

Prometheus
Thanos
Grafana
OpenTelemetry
NVIDIA Omniverse

Descrizione del lavoro

We are looking for an experienced Site Reliability Engineer to join our growing team in Milan and help shape the future of our flagship project, Colosseum, one of Europe’s most powerful AI supercomputers, currently in development.

Designed to run our proprietary AI models at scale, it forms the compute backbone behind the intelligence we deliver to the world’s most demanding industries.

In this role, you will design and implement observability and control mechanisms that extract operational data from infrastructure and feed it into automated systems to enable continuous optimization, including key system budgets such as power, cooling and service level, security-level objectives.

You will be responsible for actively guarding and maintaining these operational budgets as part of day-to-day system reliability and performance management.

You will also contribute to operational excellence through blameless post-mortem analysis and structured incident learning, ensuring continuous improvement of system behavior and resilience.

As a part of the team, you will work closely with Platform Engineering in a shared cybersecurity model, where SRE focuses on detection and monitoring, while Platform Engineering ensures the secure design and operation of the underlying infrastructure.

What You Have
  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
  • At least 6 years of experience as a Site Reliability Engineer or in similar roles.
  • Strong experience with observability and monitoring systems such as Prometheus, Thanos, Grafana, and OpenTelemetry
  • Experience with low-level system instrumentation and performance visibility using technologies such as eBPF
  • Experience with security monitoring and threat detection tools such as Zeek, Wazuh, or equivalent SIEM / security observability platforms
  • Strong experience with containerized and cloud-native environments, particularly Kubernetes
  • Strong software development skills, particularly in Python, with the ability to build automation, integrations, and custom tooling
  • Experience integrating heterogeneous infrastructure systems across multiple vendors, APIs, and evolving tool ecosystems
  • Familiarity with modern infrastructure automation and emerging agent-based frameworks such as MCP / A2A (or equivalent technologies)
  • Exposure to digital twin technologies and simulation platforms such as NVIDIA Omniverse or equivalent
  • Strong ability to design, build, and maintain software-driven infrastructure solutions in complex, large-scale environments
Who You Are
  • A versatile engineer, comfortable operating in complex and fast-paced environments.
  • Driven and fearless, you proactively tackle challenges and overcome obstacles with determination.
  • A systems thinker, capable of understanding the broader architecture and identifying dependencies across platforms and technologies.
  • A collaborative team player who is enthusiastic, curious, and passionate about problem-solving, thriving both independently and within cross-functional teams.
  • An effective communicator with strong interpersonal skills, able to engage with diverse stakeholders and foster collaboration.
  • Fluent in English and eager to contribute in a multicultural and international environment.
Benefits
Perks
  • Learning Friday. If our team members know more, so do we. That’s why we give everyone a training budget that they can spend on books, online courses or other training materials.
  • Smart Working. Trains can be a drag, you can save some commuting time by working from home.
  • Salary is based on experience and topped up with other bonuses.

We offer a competitive salary, as well as an opportunity to receive company equity. The typical salary for this role ranges between € 50.000 and € 70.000. As you gain experience and make more significant contributions to the business, your compensation will be reviewed to match your impact. Additionally, depending on your seniority and your performance, you'll have the opportunity to receive stock options, with a variable value calculated from your base salary, giving you the chance to directly participate in the company’s success.

Employment terms are governed by the CCNL Commercio, the Italian National Collective Bargaining Agreement for the Commerce, Distribution and Services sector.

About Domyn

Domyn is a company specializing in the research and development of Responsible AI for regulated industries, including financial services, government, and heavy industry. It supports enterprises with proprietary, fully governable solutions based on a composable AI architecture — including LLMs, AI agents, and one of the world's largest supercomputers. At the core of Domyn's product offer is a chip-to-frontend architecture that allows organizations to control the entire AI stack — from hardware to application — ensuring isolation, security, and governance throughout the AI lifecycle. Its foundational LLMs, Domyn Large and Domyn Small, are designed for advanced reasoning and optimized to understand each business's specific language, logic, and context. Provided under an open-enterprise license, these models can be fully transferred and owned by clients. Once deployed, they enable customizable agents that operate on proprietary data to solve complex, domain‑specific problems. All solutions are managed via a unified platform with native tools for access management, traceability, and security. Powering it all, Colosseum — a supercomputer in development using NVIDIA Grace Blackwell Superchips — will train next‑gen models exceeding 1T parameters. Domyn partners with Microsoft, NVIDIA, and G42. Clients include Allianz, Intesa Sanpaolo, and Fincantieri. Please review our Privacy Policy here https://bit.ly/4tndszN .

Ottieni la revisione del curriculum gratis e riservata.
o trascina qui il file.
Similar jobs

Offerte di lavoro simili che vale la pena confrontare

Senior AI Cloud Engineer
Senior AI Cloud Engineer

Domyn • Milano

In loco
EUR 55.000 - 75.000
Training budget
Smart Working
Stock options
Senior Platform Engineer
Senior Platform Engineer

PLP Group • Milano

Ibrido
EUR 50.000 - 70.000
Learning Friday
Smart Working
Equity
Technical Project Manager (HPC Infrastructure)
Technical Project Manager (HPC Infrastructure)

PLP Group • Milano

In loco
EUR 30.000 - 60.000
Learning Friday
Smart Working
Stock options
Senior Data Center Operations Engineer
Senior Data Center Operations Engineer

PLP Group • Milano

In loco
EUR 50.000 - 70.000
Learning Friday
Smart Working
Equity
Senior AI Cloud Engineer
Senior AI Cloud Engineer

PLP Group • Milano

In loco
EUR 55.000 - 75.000
Learning Friday
Smart Working
Equity
Corporate Development Manager
Corporate Development Manager

Domyn • Turbigo

In loco
EUR 60.000 - 70.000
Competitive salary
Opportunity to receive company equity
Training budget for professional development
+1
Corporate Development Manager
Corporate Development Manager

Domyn • Milano

In loco
EUR 60.000 - 70.000
Learning Friday
Smart Working
Bonus potential
Business Analyst
Business Analyst

Domyn • Milano

In loco
EUR 30.000 - 60.000
Training budget for courses and materials
Work from home flexibility
Employee equity options
Senior Legal Counsel, Corporate
Senior Legal Counsel, Corporate

PLP Group • Milano

Ibrido
EUR 60.000 - 85.000
Smart Working
Training budget
Company equity
+1
Legal Counsel, Privacy
Legal Counsel, Privacy

PLP Group • Milano

Ibrido
EUR 45.000 - 70.000
Learning Friday
Smart Working
Equity opportunity