We are looking for a Service Delivery Manager to support strategic hyperscale customers using bare metal and GPU cloud AI infrastructure in Nebius data centers.
This role acts as the technical and operational interface between the customer/platform and Nebius infrastructure teams, ensuring reliable service delivery, SLA compliance, and smooth operations of large-scale GPU clusters and bare metal environments.
You will coordinate across data center operations, network, hardware lifecycle, and infrastructure engineering teams to deliver world‑class infrastructure services for large AI workloads.
Customer Infrastructure Ownership
- Serve as the primary technical point of contact for the customer.
- Manage operational relationship with hyperscale customers.
- Coordinate infrastructure lifecycle including provisioning, maintenance, and incident management.
Service Delivery & SLA Management
- Ensure SLA and SLO compliance for infrastructure services.
- Drive incident management and root cause analysis.
- Track service performance and operational metrics.
Data Center & Infrastructure Coordination
- Work closely with data center operations teams to manage hardware support and infrastructure maintenance.
- Coordinate bare metal deployments, replacements, and capacity expansion.
- Align infrastructure operations with customer workload requirements.
Operational Excellence
- Establish operational processes for hyperscale infrastructure environments.
- Lead service reviews and operational planning with internal and customer teams.
- Improve reliability, response times, and operational workflows.
Cross-Functional Collaboration
- Partner with engineering, networking, and hardware lifecycle teams.
- Support infrastructure scaling and new cluster deployments.
- Participate in planning for large-scale AI compute infrastructure.
We expect you to have:
- 5+ years in technical account management, service delivery, or infrastructure operations
- Experience working with hyperscale customers or large enterprise clients
- Background in cloud, AI infrastructure, or data center operations
- Understanding of bare metal infrastructure
- Experience with data center environments
- Familiarity with GPU clusters / AI workloads (preferred)
- Knowledge of networking, hardware lifecycle, and infrastructure monitoring
- Strong stakeholder management
- Ability to operate in high-scale infrastructure environments
- Excellent communication between technical and business teams
It will be an added bonus of you have:
- Experience working with hyperscale companies, Experience supporting AI / ML infrastructure, Experience with GPU clusters 5+ years in technical account management, service delivery, or infrastructure operations, Experience working with hyperscale customers or large enterprise clients, Background in cloud, AI infrastructure, or data center operations, Understanding of bare metal infrastructure, Experience with data center environments, Knowledge of networking, hardware lifecycle, and infrastructure monitoring, Strong stakeholder management, Ability to operate in high-scale infrastructure environments, Excellent communication between technical and business teams