Site Reliability Engineer, AI Cloud Infrastructure

Anyscale

San Francisco, Northern (CA, KY)

Hybrid

USD 140,000 - 210,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Anyscale in San Francisco is seeking a Site Reliability Engineer to join the Infrastructure team. You will help build the scalable, secure backbone powering distributed AI workloads in the cloud, spanning control plane and data plane.

The role emphasizes Kubernetes, container orchestration, and cloud-native infrastructure, with opportunities to work on Ray integrations and open‑source contributions. You will collaborate across teams to deliver high‑impact infrastructure features.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or equivalent practical experience.
  • 3+ years of experience writing high-quality production code.
  • Hands-on experience in building and maintaining highly available, scalable distributed systems.
  • Expertise in cloud-native technologies (AWS, Azure, GCP) and Kubernetes-based deployments.
  • Deep understanding of networking, security, and authentication mechanisms in cloud environments.
  • Familiarity with observability stacks (Prometheus, Grafana).
  • Proficiency in Go and Python.
  • Knowledge of low-level OS foundations (Linux kernel, filesystems, containers).

Responsibilities

  • Design, build, and scale services that orchestrate Ray clusters across cloud and on‑prem environments.
  • Optimize control plane components for large‑scale distributed AI/ML workloads.
  • Build intelligent scheduling and resource management systems for heterogeneous compute clusters.
  • Develop features to enhance reliability, performance, scalability, and observability of Ray workloads.
  • Support and optimize accelerator integration (e.g., GPUs).
  • Handle container image management and dependency resolution for distributed workloads.
  • Participate in code reviews, design and architecture discussions.
  • Provide on‑call support, collaborating with customer and field teams to troubleshoot infrastructure issues.

Skills

Go
Python
Distributed systems
Observability
Networking & security
On-call experience

Education

Bachelor's degree in CS/Engineering or equivalent

Tools

Kubernetes
AWS
Azure
GCP
Linux

Job description

Anyscale in San Francisco is seeking a Site Reliability Engineer to join the Infrastructure team. You will help build the scalable, secure backbone powering distributed AI workloads in the cloud, spanning control plane and data plane.

The role emphasizes Kubernetes, container orchestration, and cloud-native infrastructure, with opportunities to work on Ray integrations and open‑source contributions. You will collaborate across teams to deliver high‑impact infrastructure features.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Platform Infrastructure Engineer – Scalable AI Cloud
Staff Platform Infrastructure Engineer – Scalable AI Cloud

Anyscale • San Francisco (CA)

On-site
USD 260,000 - 380,000
Competitive salary and equity
Health/dental/vision coverage
Flexible time off
+2
Site Reliability Engineer — AI Infra & Equity
Site Reliability Engineer — AI Infra & Equity

Zof AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 200,000
Competitive salary
Meaningful equity
Platform Infrastructure Engineering Leader
Platform Infrastructure Engineering Leader

Anyscale • San Francisco (CA)

On-site
USD 260,000 - 340,000
Site Reliability Engineer, Platform Infrastructure (Foundations)
Site Reliability Engineer, Platform Infrastructure (Foundations)

Anyscale • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineer — Scale an AI‑Powered SaaS Platform
Site Reliability Engineer — Scale an AI‑Powered SaaS Platform

Instrumental Inc. • Palo Alto (CA)

On-site
USD 140,000 - 165,000
Health insurance
Vision insurance
Dental plan
+2
Site Reliability Engineer — ML Infra, Scale & Equity
Site Reliability Engineer — ML Infra, Scale & Equity

Baseten • New York (NY)

On-site
USD 165,000 - 330,000
Competitive compensation
100% coverage of medical, dental, and vision insurance
Generous PTO policy
+3
Senior Site Reliability Engineer – Scalable AI Infra
Senior Site Reliability Engineer – Scalable AI Infra

Tavily Inc. • New York (NY)

Hybrid
USD 156,000 - 262,000
100% company-paid medical, dental, and vision coverage
Up to 4% company match 401(k) plan
20 weeks paid parental leave for primary caregivers
+2
Senior Site Reliability Engineer - AI-Driven Cloud Infra
Senior Site Reliability Engineer - AI-Driven Cloud Infra

Zscaler, Inc. • San Jose (CA), Northern (KY)

Hybrid
USD 123,000 - 175,000
Time off plans
Parental leave options
Retirement options
+1
Senior AI Platform SRE: Scale Cloud Infra & Kubernetes
Senior AI Platform SRE: Scale Cloud Infra & Kubernetes

GCS Recruitment • Mount Laurel Township (NJ)

On-site
USD 110,000 - 170,000