About The Role
We are seeking an Infrastructure Operations Lead to oversee operational performance across our high-density GPU data center footprint.
About The Role
We are seeking an Infrastructure Operations Lead to oversee operational performance across our high-density GPU data center footprint.
This is an on-site vendor governance and technical escalation role. You will act as our primary operational anchor—managing vendor performance and holding our Managed Service Provider (MSP), Co-Location Facility Provider, and Enterprise Customers accountable to their operational standards, service contracts, and SLAs.
Key Responsibilities
- 360° Vendor & Service Governance: Oversee third-party service delivery to enforce hardware repair SLAs, ticket response times, and spare parts/RMA workflows. Ensure the co-location provider meets power and high-density cooling guarantees, while keeping enterprise customers within agreed operating boundaries.
- Crisis Management & Incident Command: Act as On-Site Incident Commander during high-severity outages or infrastructure degradation. Lead recovery efforts by coordinating across vendor engineering teams and internal stakeholders.
- Incident & Executive Communication: Provide concise, real-time updates to executive leadership and enterprise clients during major outages, translating complex technical failures into clear operational impact.
- Root Cause Analysis (RCA) & Post-Mortems: Lead technical investigations following major incidents. Audit and challenge technical RCAs provided by vendors (MSP/Colo) to identify systemic hardware, environmental, or workflow issues, ensuring permanent corrective actions are executed.
- Technical Escalation & Standards: Serve as the Subject Matter Expert (SME) for complex GPU/HPC hardware escalations that exceed standard vendor runbooks, and maintain ownership of operational standards.
Requirements
- 5+ years in data center infrastructure operations, with a strong focus on third-party vendor management, SLA enforcement, and service delivery for high-density GPU/HPC platforms.
- Crisis management: Proven ability to drive incident recovery and direct third-party vendor teams during critical data center outages under high‑pressure conditions.
- Stakeholder Communication: Outstanding verbal and written communication skills to bridge technical vendor teams, internal stakeholders, and enterprise client representatives.
- Root Cause Analysis (RCA) Expertise: Demonstrated capability in methodical troubleshooting, post-incident investigations, and auditing vendor‑supplied RCAs (using frameworks like 5‑Whys or Fishbone) to drive long‑term infrastructure reliability.
- Technical & Facility Knowledge: Solid understanding of enterprise server platforms (NVIDIA Blackwell GPU architectures, high‑speed networking RoCE,infiniband) and data center facility constraints (high‑density power distribution, liquid cooling).
- Education: Bachelor’s degree in Computer Science, Engineering, or equivalent practical experience.
Skills: oems,infrastructure,commissioning,data,operations,field execution,skills,critical infrastructure,maintenance,contractors