SRE for AI Accelerator Infra: Cloud & On-Prem Automation

Entrada Ventures

Santa Clara (CA)

On-site

USD 120,000 - 180,000

Part time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

d-Matrix is seeking a senior SRE to own reliability for colo/server fleets, on-prem lab clusters, and cloud environments. You will write IaC using Terraform/Ansible, automate workflows, and manage monitoring.

The role emphasizes hands-on hardware, network, and performance tuning to keep silicon development and customer deployments running smoothly. The position requires 5+ years in SRE or infra, strong Linux skills, and experience with cloud/on-prem and monitoring tools.

Qualifications

  • Bachelor's or master's in CS/EE or equivalent experience; 5+ years in SRE or infra.
  • Strong Linux skills: networking, storage, systemd, performance diagnostics.
  • Hands-on with colocation or on-prem hardware and bare-metal provisioning.
  • IaC experience with Terraform and Ansible; production configurations.

Responsibilities

  • Own reliability and availability of assigned infrastructure domains: colo server fleets, on-premises lab clusters, cloud environments (AWS, Azure, GCP), and customer-facing platform services.
  • Perform hands-on infrastructure work: server provisioning, OS configuration, network setup, storage management, and hardware troubleshooting from bare metal up.
  • Support and operate high-speed interconnect environments — InfiniBand, RoCE, or high-speed Ethernet — in lab and colo settings.
  • Conduct capacity planning and hardware lifecycle management for assigned infrastructure domains.

Skills

Linux administration
IaC (Terraform/Ansible)
Kubernetes operations
Monitoring (Prometheus/Grafana or Data
Scripting (Python/Bash)
Incident response
Cloud & on-prem
Go (optional)

Education

Bachelor's/Master's in CS or related field

Tools

Terraform
Ansible
Kubernetes
Prometheus
Grafana
DataDog
Python

Job description

d-Matrix is seeking a senior SRE to own reliability for colo/server fleets, on-prem lab clusters, and cloud environments. You will write IaC using Terraform/Ansible, automate workflows, and manage monitoring.

The role emphasizes hands-on hardware, network, and performance tuning to keep silicon development and customer deployments running smoothly. The position requires 5+ years in SRE or infra, strong Linux skills, and experience with cloud/on-prem and monitoring tools.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE: AI Accelerator Infrastructure (6-Month Contract)
SRE: AI Accelerator Infrastructure (6-Month Contract)

d-Matrix • Santa Clara (CA)

On-site
USD 150,000 - 210,000
Director, SRE for AI Accelerator Infra
Director, SRE for AI Accelerator Infra

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 270,000
Director, AI Accelerator Infrastructure & SRE
Director, AI Accelerator Infrastructure & SRE

d-Matrix • Santa Clara (CA)

On-site
USD 250,000 - 350,000
SRE / DevOps Engineer: AI-Driven Infra & Cloud Automation
SRE / DevOps Engineer: AI-Driven Infra & Cloud Automation

Lineate • Georgia

On-site
USD 90,000 - 150,000
Freedom to develop
Career path
Benefits package
+6
Senior SRE: Automation, AI-Driven Reliability (Enterprise)
Senior SRE: Automation, AI-Driven Reliability (Enterprise)

ManpowerGroup Global, Inc. • Austin (TX)

On-site
USD 66,000 - 90,000
Senior SRE: AI Cloud Infra, Kubernetes & Terraform (Remote)
Senior SRE: AI Cloud Infra, Kubernetes & Terraform (Remote)

Motion Recruitment • Mount Laurel Township (NJ)

Remote
USD 140,000 - 190,000
Remote equipment stipend
Annual learning and development budget
Equity / Stock Options
+1
SRE (Site Realiability Engineer)
SRE (Site Realiability Engineer)

STRATIS Cloud Tech Solutions INC • Arkansas

On-site
USD 110,000 - 150,000
Competitive salary
Growth and learning opportunities
Friendly, collaborative team
Senior SRE: AI-Driven Reliability & Automation (Hybrid)
Senior SRE: AI-Driven Reliability & Automation (Hybrid)

Namely • United States

Hybrid
USD 120,000 - 150,000
Site Reliability Engineer - AI Accelerator Infrastructure - Contract
Site Reliability Engineer - AI Accelerator Infrastructure - Contract

Entrada Ventures • Santa Clara (CA)

On-site
USD 120,000 - 180,000
Infra Engineer (SRE) for AI & GPU Compute
Infra Engineer (SRE) for AI & GPU Compute

Blackwing • Detroit (MI), Northern (KY)

Hybrid
USD 140,000 - 190,000
Competitive salary
Equity package
Medical/dental/vision coverage
+5