Site Reliability Engineer - Research Compute for AI

Astera Institute

United States

On-site

USD 120,000 - 190,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

A leading research foundation in the United States seeks a Site Reliability Engineer to manage its digital infrastructure, ensuring reliable access to compute resources for research. The role involves automating processes, enhancing resource visibility, and managing access efficiently. Candidates should have strong experience with technologies like Kubernetes and Docker, and a keen understanding of systems operations. This position is on-site in Emeryville, with potential visa sponsorship for qualified applicants.

Qualifications

  • Comfortable being accountable for cluster health and capacity.
  • Understanding of schedulers, containers, and networking interactions.
  • Values observability, reproducibility, and clear operational boundaries.
  • Can support experimental workloads and adapt to project needs.

Responsibilities

  • Ensure easy access to compute resources for researchers.
  • Provide visibility into resource utilization and cluster health.
  • Enable automatic scaling of compute resources.
  • Manage access to resources effectively.
  • Drive towards reproducible research environments.
  • Automate operational processes to increase efficiency.

Skills

Ansible
Kubernetes
Docker
Python
Grafana
Prometheus
Tailscale
Talos Linux

Job description

A leading research foundation in the United States seeks a Site Reliability Engineer to manage its digital infrastructure, ensuring reliable access to compute resources for research. The role involves automating processes, enhancing resource visibility, and managing access efficiently. Candidates should have strong experience with technologies like Kubernetes and Docker, and a keen understanding of systems operations. This position is on-site in Emeryville, with potential visa sponsorship for qualified applicants.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer — ML Infra, Scale & Equity
Site Reliability Engineer — ML Infra, Scale & Equity

Baseten • New York (NY)

On-site
USD 165,000 - 330,000
Global Remote SRE for AI Infrastructure & Kubernetes
Global Remote SRE for AI Infrastructure & Kubernetes

Andromeda Cluster • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Site Reliability Engineer — Cloud-Scale & Automation
Site Reliability Engineer — Cloud-Scale & Automation

ByteDance • San Jose (CA)

On-site
USD 136,000 - 360,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+3
Site Reliability Engineer, Compute — Scale & Automate Systems
Site Reliability Engineer, Compute — Scale & Automate Systems

TikTok • Seattle (WA)

On-site
USD 100,000 - 130,000
Site Reliability Engineer — Scale & Resilience for AI Ops
Site Reliability Engineer — Scale & Resilience for AI Ops

HappyRobot • San Francisco (CA)

On-site
USD 120,000 - 160,000
Site Reliability Engineer — Build Scalable AI Infra
Site Reliability Engineer — Build Scalable AI Infra

Future Secure AI Pty • Austin (TX)

On-site
USD 100,000 - 140,000
Flexible work environment
Competitive salary
Growth trajectory
Senior SRE — AI-Driven Cloud Reliability
Senior SRE — AI-Driven Cloud Reliability

BetterUp • New York (NY)

Hybrid
USD 164,000 - 205,000
Site Reliability Engineer - Kubernetes & Cloud
Site Reliability Engineer - Kubernetes & Cloud

Hydrolix • United States

On-site
USD 110,000 - 150,000
Senior SRE: AI Cloud Reliability & Observability (Remote)
Senior SRE: AI Cloud Reliability & Observability (Remote)

SDI International • Chicago (IL)

Hybrid
USD 130,000 - 180,000
Site Reliability Engineer — Automate, Scale & On-Call
Site Reliability Engineer — Automate, Scale & On-Call

TikTok • San Jose (CA)

Hybrid
USD 187,000 - 360,000