AI Infra Ops Leader: Scale Mission-Critical Data Centers

Designworks Talent

Bellevue (WA)

Hybrid

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
Vision coverage
401(k) plan

Job summary

Designworks Talent seeks a Data Center Operations and Maintenance Engineering Leader to build and scale the Operations & Maintenance function for next‑generation AI infrastructure. This leadership role guides regional teams and partners with Engineering, Infrastructure, Networking, Hardware, and Customer Operations to deliver reliability at scale.

You will define the operating model, establish SLAs and KPIs, mentor high‑performing teams, and drive proactive monitoring, automation, and

Qualifications

  • Proven experience leading large-scale operations in cloud or AI infrastructure.
  • Strong incident management and service reliability track record.
  • Executive communication with engineering and business leadership.

Responsibilities

  • Build, lead, mentor, and grow high-performing Operations & Maintenance teams.
  • Develop the operational strategy, organizational structure, and execution model supporting large-scale AI infrastructure.
  • Establish world-class operational processes, including incident response, change management, problem management, and service reliability.
  • Lead major incident management efforts and executive communications during production events.
  • Drive operational excellence through proactive monitoring, observability, automation, and continuous improvement initiatives.
  • Partner with Engineering to ensure operational readiness for new infrastructure deployments and platform launches.
  • Define service level objectives (SLOs), operational KPIs, and reliability metrics across the infrastructure portfolio.
  • Build scalable on-call programs, escalation models, runbooks, and operational governance.
  • Champion root cause analysis and long-term corrective actions to improve platform resilience.
  • Influence infrastructure architecture and operational tooling to improve availability, efficiency, and customer experience.
  • Help shape the long-term operations organization as the company expands globally.

Skills

Operations leadership
Incident management
Site Reliability
Cloud ops leadership
Executive communication
Cross-functional collaboration

Tools

Grafana
Prometheus
Datadog
PagerDuty
Opsgenie

Job description

Designworks Talent seeks a Data Center Operations and Maintenance Engineering Leader to build and scale the Operations & Maintenance function for next‑generation AI infrastructure. This leadership role guides regional teams and partners with Engineering, Infrastructure, Networking, Hardware, and Customer Operations to deliver reliability at scale.

You will define the operating model, establish SLAs and KPIs, mentor high‑performing teams, and drive proactive monitoring, automation, and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Data Center Operations & Reliability Engineer
AI Data Center Operations & Reliability Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 130,000 - 190,000
Medical insurance
Dental and vision insurance
401(k) with company match
+1
AI Data Center Facilities Operations Lead
AI Data Center Facilities Operations Lead

OpenAI • San Francisco (CA)

On-site
USD 150,000 - 190,000
Data Center Operations and Maintenance Engineering Leader
Data Center Operations and Maintenance Engineering Leader

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Health insurance
Vision coverage
401(k) plan
AI-Driven Data Center Operations Lead
AI-Driven Data Center Operations Lead

Milestone Technologies, Inc. • Hubbard (TX)

On-site
USD 90,000 - 120,000
Senior Data Center Operations & Commissioning Lead
Senior Data Center Operations & Commissioning Lead

OpenAI • United States

On-site
USD 159,000 - 285,000
AI Data Center Ops Lead — GPU/HPC Infrastructure
AI Data Center Ops Lead — GPU/HPC Infrastructure

Nscale • Town of Norway (WI)

On-site
USD 120,000 - 170,000
Data Center Delivery Program Lead for AI Infra
Data Center Delivery Program Lead for AI Infra

Together AI • United States

Remote
USD 190,000 - 240,000
Equity
Remote work flexibility
Global Data Center Ops SVP — AI Infra Leader
Global Data Center Ops SVP — AI Infra Leader

Nscale • Seattle (WA)

On-site
USD 250,000 - 350,000
Equity
Flexible PTO
Medical insurance
+4
Mechanical Engineer - AI Data Center Operations
Mechanical Engineer - AI Data Center Operations

Intelletec Energy • San Francisco (CA)

On-site
USD 190,000 - 260,000
Global Data Center Engineer — Lead HPC Deployments & DCIM
Global Data Center Engineer — Lead HPC Deployments & DCIM

EngineersOfAI • San Francisco (CA)

Hybrid
USD 90,000 - 120,000