SRE: AI Accelerator Infrastructure (6-Month Contract)

d-Matrix

Santa Clara (CA)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

d-Matrix is seeking a Site Reliability Engineer for AI Accelerator Infrastructure on a 6‑month contract with potential full‑time conversion. You will own end‑to‑end reliability, automation, and observability across colo, on‑prem lab clusters, and cloud environments, partnering with the DevOps lead team to keep systems running and scalable.

You will implement IaC with Terraform and Ansible, build dashboards in Prometheus/Grafana or DataDog, and support customer‑facing platforms while ensuring QoS

Qualifications

  • Bachelor's or Master's in Computer Science, Electrical Engineering, or related field.
  • 5+ years in SRE, infrastructure engineering, or systems administration.
  • Strong Linux knowledge: networking, storage, systemd, kernel parameters.
  • Hands-on experience with colocation or on-prem server infra & bare-metal provisioning.
  • IaC experience with Terraform and/or Ansible; production configurations.
  • Kubernetes operational experience: troubleshooting, storage, networking.

Responsibilities

  • Own reliability and availability of assigned infrastructure domains (colo, on-prem lab, cloud).
  • Perform hands-on provisioning, OS config, network setup, storage management, hardware troubleshooting.
  • Build and maintain monitoring dashboards and alerting (Prometheus/Grafana, DataDog).
  • Lead incident response and RCA reporting; drive improvements to prevent recurrence.
  • Collaborate with DevOps and customer-facing environments to ensure QoS and uptime.

Skills

Linux administration
Networking
Storage management
Kernel tuning
IaC (Terraform/Ansible)
Kubernetes operations
Monitoring (Prometheus/Grafana or Data

Education

Bachelor's or Master's in CS/EE or related field

Tools

Terraform
Ansible
Kubernetes
Prometheus
Grafana
DataDog
Python
Bash

Job description

d-Matrix is seeking a Site Reliability Engineer for AI Accelerator Infrastructure on a 6‑month contract with potential full‑time conversion. You will own end‑to‑end reliability, automation, and observability across colo, on‑prem lab clusters, and cloud environments, partnering with the DevOps lead team to keep systems running and scalable.

You will implement IaC with Terraform and Ansible, build dashboards in Prometheus/Grafana or DataDog, and support customer‑facing platforms while ensuring QoS

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE for AI Accelerator Infra: Cloud & On-Prem Automation
SRE for AI Accelerator Infra: Cloud & On-Prem Automation

Entrada Ventures • Santa Clara (CA)

On-site
USD 120,000 - 180,000
Site Reliability Engineer - AI Accelerator Infrastructure - Contract
Site Reliability Engineer - AI Accelerator Infrastructure - Contract

d-Matrix • Santa Clara (CA)

On-site
USD 150,000 - 210,000
SRE Director — AI Accelerator Infrastructure
SRE Director — AI Accelerator Infrastructure

d-Matrix • United States

On-site
USD 180,000 - 240,000
Director, AI Accelerator Infrastructure & SRE
Director, AI Accelerator Infrastructure & SRE

d-Matrix • Santa Clara (CA)

On-site
USD 250,000 - 350,000
Director, SRE for AI Accelerator Infra
Director, SRE for AI Accelerator Infra

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 270,000
Site Reliability Engineer - AI Accelerator Infrastructure - Contract
Site Reliability Engineer - AI Accelerator Infrastructure - Contract

Entrada Ventures • Santa Clara (CA)

On-site
USD 120,000 - 180,000
Director, Site Reliability Engineering - AI Accelerator Infrastructure - Contract
Director, Site Reliability Engineering - AI Accelerator Infrastructure - Contract

d-Matrix • United States

On-site
USD 180,000 - 240,000
Director, Site Reliability Engineering - AI Accelerator Infrastructure - Contract
Director, Site Reliability Engineering - AI Accelerator Infrastructure - Contract

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 270,000
AI Hardware Lab & Data Center Technician (1-Year Contract)
AI Hardware Lab & Data Center Technician (1-Year Contract)

Entrada Ventures • Santa Clara (CA)

On-site
USD 80,000 - 110,000
AI Reliability Engineer (AI SRE)
AI Reliability Engineer (AI SRE)

DeWinter Group • Campbell (CA)

Remote