Director, AI Accelerator Infrastructure & SRE

d-Matrix

Santa Clara (CA)

On-site

USD 250,000 - 350,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

d-Matrix is seeking a Director of Site Reliability Engineering to architect and lead our AI accelerator infrastructure. You will own the end‑to‑end reliability of colocation, on‑premise lab clusters, and cloud environments (AWS/Azure/GCP), while shaping SRE as a discipline.

Lead 3–5 engineers, drive SLOs, RCA, and observability from day one. The role requires deep Linux, storage, multi‑cloud ops, and executive communication to align stakeholders with strategy and delivery.

Qualifications

  • 15+ years in SRE, infrastructure engineering, or production engineering.
  • Experience building or rebuilding an SRE function with measurable impact.
  • Strong Linux, networking, and storage fundamentals at scale.
  • Proven ability to define SLOs, on-call models, and observability from scratch.
  • Executive communication to translate risk into clear narratives.

Responsibilities

  • Build and lead SRE from the ground up, owning global infrastructure.
  • Define charter, roadmap, and on-call model; hire 3–5 SRE engineers.
  • Own end-to-end reliability across cloud, colocation, and on‑premises.
  • Establish incident management, RCA, and continuous improvement.
  • Own cost visibility, capacity planning, and FinOps across tiers.
  • Drive IaC discipline and platform automation at scale.
  • Partner with DevOps to align pipelines with reliability goals.

Skills

SRE leadership
Linux systems
IaC (Terraform, Ansible)
Kubernetes
Observability (Prometheus, Grafana, or

Education

Bachelor’s or Master’s in CS/EE or related field

Tools

Terraform
Ansible
Kubernetes
Prometheus
Grafana
Datadog

Job description

d-Matrix is seeking a Director of Site Reliability Engineering to architect and lead our AI accelerator infrastructure. You will own the end‑to‑end reliability of colocation, on‑premise lab clusters, and cloud environments (AWS/Azure/GCP), while shaping SRE as a discipline.

Lead 3–5 engineers, drive SLOs, RCA, and observability from day one. The role requires deep Linux, storage, multi‑cloud ops, and executive communication to align stakeholders with strategy and delivery.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Director, SRE for AI Accelerator Infra
Director, SRE for AI Accelerator Infra

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 270,000
SRE for AI Accelerator Infra: Cloud & On-Prem Automation
SRE for AI Accelerator Infra: Cloud & On-Prem Automation

Entrada Ventures • Santa Clara (CA)

On-site
USD 120,000 - 180,000
Director, Site Reliability Engineering - AI Accelerator Infrastructure - Contract
Director, Site Reliability Engineering - AI Accelerator Infrastructure - Contract

d-Matrix • Santa Clara (CA)

On-site
USD 250,000 - 350,000
SRE: AI Accelerator Infrastructure (6-Month Contract)
SRE: AI Accelerator Infrastructure (6-Month Contract)

d-Matrix • Santa Clara (CA)

On-site
USD 150,000 - 210,000
Director, Site Reliability Engineering - AI Accelerator Infrastructure - Contract
Director, Site Reliability Engineering - AI Accelerator Infrastructure - Contract

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 270,000
Remote SRE Manager: Lead AI-Driven Reliability & Cloud Ops
Remote SRE Manager: Lead AI-Driven Reliability & Cloud Ops

Arcoro Holdings Corp • Phoenix (AZ), Northern (KY)

Hybrid
USD 200,000 - 220,000
Remote Work
401(k) with Company match
Flexible PTO and Company-paid holidays
Site Reliability Engineer - AI Accelerator Infrastructure - Contract
Site Reliability Engineer - AI Accelerator Infrastructure - Contract

d-Matrix • Santa Clara (CA)

On-site
USD 150,000 - 210,000
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Remote AI Infrastructure SRE — Kubernetes & Reliability
Remote AI Infrastructure SRE — Kubernetes & Reliability

Andromeda • San Francisco (CA)

On-site
USD 120,000 - 160,000
Site Reliability Engineer - AI Accelerator Infrastructure - Contract
Site Reliability Engineer - AI Accelerator Infrastructure - Contract

Entrada Ventures • Santa Clara (CA)

On-site
USD 120,000 - 180,000