Senior Platform Engineer

Domyn

Milano

In loco

EUR 50.000 - 70.000

Tempo pieno

5 giorni fa
Candidati tra i primi

Ricevi più risposte dai datori di lavoro

Invia un CV specifico per questa offerta in pochi minuti.

Vantaggi offerti da questo lavoro

Learning Friday
Smart Working
Salary bonuses

Descrizione del lavoro

Domyn is seeking an experienced Platform Engineer in Milan to help shape the Colosseum AI supercomputer's platform. You will design and operate scalable infrastructure, orchestrate workloads, and ensure secure, efficient resource management across tenants.

The role emphasizes Kubernetes-based multitenancy, integration of performance metrics, and collaboration with SRE and AI cloud engineers to continuously improve the platform. Fluency in English is required; remote work is not indicated.

Competenze

  • Bachelor's or Master's degree in Computer Science, Software Engineering, Data Science, or a related field.
  • 6+ years of experience as a Platform Engineer or in similar roles.
  • Knowledge of NVIDIA infrastructure technologies (NicO, DSX, DCGM).
  • Expertise in leak detection technologies and monitoring systems.
  • Strong experience with container orchestration and virtualization (Kubernetes) for multitenancy, bare metal provisioning, and system network configuration.
  • Knowledge of NVIDIA Omniverse and digital twins.
  • Solid understanding of MCP/A2A or equivalent.
  • Proficiency in Python and software development best practices (version control, testing, CI/CD).
  • Experience with Run:AI or similar AI workload orchestration and job scheduling frameworks.
  • Expertise in parallel file systems (Weka, NetApp, Hammerspace).
  • Strong knowledge of HPC networking (Ethernet, InfiniBand, NVLink, NVIDIA BlueField).
  • Experience with workload optimization and resource management frameworks.

Mansioni

  • Build and operate the foundational platform capabilities to orchestrate workloads and resources with precision and efficiency.
  • Scale and dynamically redistribute resources across tenants while integrating infrastructure metrics for intelligent scheduling.
  • Collaborate with Site Reliability Engineers and AI Cloud Engineers to design, implement, and improve the platform.
  • Embed security controls into the platform by implementing firewalls, enforcing least-privilege access, and applying network hardening.

Conoscenze

Kubernetes
Python
CI/CD pipelines
Run:AI
NVIDIA infrastructure
Weka
NetApp
Hammerspace
InfiniBand
NVIDIA Omniverse

Formazione

Bachelor's or Master's degree in Computer Science, Software Engineering, Data Science, or related field

Strumenti

NVIDIA NicO
NVIDIA DSX
NVIDIA DCGM
NVIDIA Omniverse
Run:AI
Bare metal provisioning
Networking configuration

Descrizione del lavoro

We are looking for an experienced Platform Engineer to join our growing team in Milan and help shape the future of our flagship project, Colosseum, one of Europe’s most powerful AI supercomputers, currently in development.

Designed to run our proprietary AI models at scale, it forms the compute backbone behind the intelligence we deliver to the world’s most demanding industries.

In this role, you will play a key part in ensuring optimal resource utilization across our AI infrastructure. You will build and operate the foundational platform capabilities required to orchestrate workloads and resources with precision and efficiency.

Furthermore, you will be responsible for scaling and dynamically redistributing resources across tenants while integrating critical infrastructure metrics into the resource management framework to enable intelligent scheduling and optimization. You will also work closely with Site Reliability Engineers and AI Cloud Engineers to design, implement, and continuously improve the platform.

You will embed security controls directly into the platform foundation by implementing firewalls, enforcing least-privilege access policies, and applying network hardening practices, ensuring that security is built in from the ground up—before code is even deployed.

What You Have
  • Bachelor's or Master's degree in Computer Science, Software Engineering, Data Science, or a related field.
  • At least 6 years of experience as a Platform Engineer or in similar roles.
  • Knowledge of NVIDIA infrastructure technologies, including NicO, DSX, and DCGM.
  • Expertise in leak detection technologies and monitoring systems.
  • Strong experience with container orchestration and virtualization technologies, including Kubernetes implementations for strong multitenancy, bare metal provisioning and system network configuration.
  • Knowledge of digital twin technologies and platforms such as NVIDIA Omniverse.
  • Solid understanding of MCP/A2A or equivalent
  • Proficiency in Python and experience with software development best practices, including version control, testing, and CI/CD pipelines.
  • Experience with Run:ai or similar AI workload orchestration and job scheduling frameworks.
  • Expertise in parallel file systems, including Weka, NetApp, and Hammerspace.
  • Strong knowledge of HPC networking technologies, including Ethernet, InfiniBand, NVLink, and NVIDIA BlueField.
  • Experience with workload optimization and resource management frameworks.
Who you are
  • A versatile engineer, comfortable operating in complex and fast-paced environments.
  • Driven and fearless, you proactively tackle challenges and overcome obstacles with determination.
  • A systems thinker, capable of understanding the broader architecture and identifying dependencies across platforms and technologies.
  • A collaborative team player who is enthusiastic, curious, and passionate about problem-solving, thriving both independently and within cross-functional teams.
  • An effective communicator with strong interpersonal skills, able to engage with diverse stakeholders and foster collaboration.
  • Fluent in English and eager to contribute in a multicultural and international environment.
Benefits
Perks
  • Learning Friday. If our team members know more, so do we. That’s why we give everyone a training budget that they can spend on books, online courses or other training materials.
  • Smart Working. Trains can be a drag, you can save some commuting time by working from home.
  • Salary is based on experience and topped up with other bonuses.

We offer a competitive salary, as well as an opportunity to receive company equity. The typical salary for this role ranges between € 50.000 and € 70.000. As you gain experience and make more significant contributions to the business, your compensation will be reviewed to match your impact. Additionally, depending on your seniority and your performance, you'll have the opportunity to receive stock options, with a variable value calculated from your base salary, giving you the chance to directly participate in the company’s success.

Employment terms are governed by the CCNL Commercio, the Italian National Collective Bargaining Agreement for the Commerce, Distribution and Services sector.

About Domyn

Domyn is a company specializing in the research and development of Responsible AI for regulated industries, including financial services, government, and heavy industry. It supports enterprises with proprietary, fully governable solutions based on a composable AI architecture — including LLMs, AI agents, and one of the world's largest supercomputers. At the core of Domyn's product offer is a chip-to-frontend architecture that allows organizations to control the entire AI stack — from hardware to application — ensuring isolation, security, and governance throughout the AI lifecycle. Its foundational LLMs, Domyn Large and Domyn Small, are designed for advanced reasoning and optimized to understand each business's specific language, logic, and context. Provided under an open-enterprise license, these models can be fully transferred and owned by clients. Once deployed, they enable customizable agents that operate on proprietary data to solve complex, domain-specific problems. All solutions are managed via a unified platform with native tools for access management, traceability, and security. Powering it all, Colosseum — a supercomputer in development using NVIDIA Grace Blackwell Superchips — will train next-gen models exceeding 1T parameters. Domyn partners with Microsoft, NVIDIA, and G42. Clients include Allianz, Intesa Sanpaolo, and Fincantieri. Please review our Privacy Policy here https://bit.ly/4tndszN .

Ottieni la revisione del curriculum gratis e riservata.
o trascina qui il file.
Similar jobs

Offerte di lavoro simili che vale la pena confrontare

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Domyn • Milano

In loco
EUR 50.000 - 70.000
Learning Friday
Smart Working
Stock options
Senior AI Cloud Engineer
Senior AI Cloud Engineer

Domyn • Milano

In loco
EUR 55.000 - 75.000
Training budget
Smart Working
Stock options
Senior Platform Engineer
Senior Platform Engineer

PLP Group • Milano

Ibrido
EUR 50.000 - 70.000
Learning Friday
Smart Working
Equity
Senior AI Cloud Engineer
Senior AI Cloud Engineer

PLP Group • Milano

In loco
EUR 55.000 - 75.000
Learning Friday
Smart Working
Equity
Technical Project Manager (HPC Infrastructure)
Technical Project Manager (HPC Infrastructure)

PLP Group • Milano

In loco
EUR 30.000 - 60.000
Learning Friday
Smart Working
Stock options
Senior Data Center Operations Engineer
Senior Data Center Operations Engineer

PLP Group • Milano

In loco
EUR 50.000 - 70.000
Learning Friday
Smart Working
Equity
Corporate Development Manager
Corporate Development Manager

Domyn • Turbigo

In loco
EUR 60.000 - 70.000
Competitive salary
Opportunity to receive company equity
Training budget for professional development
+1
Senior Platform Engineer: AI Infra & Multi-Tenancy Orchestrator
Senior Platform Engineer: AI Infra & Multi-Tenancy Orchestrator

PLP Group • Milano

Ibrido
EUR 50.000 - 70.000
Learning Friday
Smart Working
Equity
Corporate Development Manager
Corporate Development Manager

Domyn • Milano

In loco
EUR 60.000 - 70.000
Learning Friday
Smart Working
Bonus potential
Senior Legal Counsel, Corporate
Senior Legal Counsel, Corporate

PLP Group • Milano

Ibrido
EUR 60.000 - 85.000
Smart Working
Training budget
Company equity
+1