Site Reliability Engineer, Cloud Infra for AI Platform

Anyscale

San Francisco (CA)

On-site

USD 150,000 - 210,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Anyscale in San Francisco, CA is seeking a Site Reliability Engineer to join the Infrastructure team. You will design and build the scalable control and data plane that powers distributed AI workloads across cloud and on‑prem environments.

The role emphasizes Kubernetes, cloud-native infrastructure, Go and Python, with on‑call responsibilities and collaboration with field teams to deliver high‑availability systems for Ray workloads.

Qualifications

  • Bachelor's degree in CS or equivalent.
  • 3+ years of production code experience.
  • Experience building scalable distributed systems.
  • Proficiency with Go and Python.
  • Deep understanding of cloud-native infrastructures and Kubernetes.
  • Knowledge of networking, security, and auth in cloud environments.
  • Experience with observability stacks (Prometheus, Grafana).

Responsibilities

  • Design and build scalable control and data plane components for Ray-based workloads.
  • Scale services across cloud and on-prem environments (VM and Kubernetes).
  • Develop intelligent scheduling and resource management.
  • Improve reliability, performance, and observability.
  • Provide on-call support and collaborate with customers.

Skills

Go
Python
Distributed systems
Cloud-native
Networking
Observability stacks

Education

Bachelor's degree in Computer Science or equivalent

Tools

Kubernetes
AWS
Azure
GCP
Linux
Prometheus
Grafana

Job description

Anyscale in San Francisco, CA is seeking a Site Reliability Engineer to join the Infrastructure team. You will design and build the scalable control and data plane that powers distributed AI workloads across cloud and on‑prem environments.

The role emphasizes Kubernetes, cloud-native infrastructure, Go and Python, with on‑call responsibilities and collaboration with field teams to deliver high‑availability systems for Ray workloads.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineer for Scalable ML Platform
Site Reliability Engineer for Scalable ML Platform

anyscale • San Francisco (CA)

On-site
USD 140,000 - 200,000
Stock options
Healthcare coverage
401k
+6
Site Reliability Engineer, Platform Infrastructure (Foundations)
Site Reliability Engineer, Platform Infrastructure (Foundations)

Anyscale • San Francisco (CA)

On-site
USD 150,000 - 210,000
Site Reliability Engineer — ML Infra, Scale & Equity
Site Reliability Engineer — ML Infra, Scale & Equity

Baseten • New York (NY)

On-site
USD 165,000 - 330,000
Platform Infrastructure Engineering Leader
Platform Infrastructure Engineering Leader

Anyscale • San Francisco (CA)

On-site
USD 260,000 - 340,000
Senior Site Reliability Engineer – Scalable AI Infra
Senior Site Reliability Engineer – Scalable AI Infra

Tavily Inc. • New York (NY)

Hybrid
USD 156,000 - 262,000
100% company-paid medical, dental, and vision coverage
Up to 4% company match 401(k) plan
20 weeks paid parental leave for primary caregivers
+2
Site Reliability Engineer, AI Infra & Observability
Site Reliability Engineer, AI Infra & Observability

Sierra • California (MO)

On-site
USD 150,000 - 210,000
Unlimited PTO
Medical, dental, vision
Retirement plan
+4
Site Reliability Engineer — Scale & Resilience for AI Ops
Site Reliability Engineer — Scale & Resilience for AI Ops

HappyRobot • San Francisco (CA)

On-site
USD 120,000 - 160,000
AI SRE: Scale, Resilience & GPU Inference
AI SRE: Scale, Resilience & GPU Inference

Seekr • Austin (TX)

Hybrid
USD 140,000 - 200,000
Equity Ownership – RSUs
Unlimited PTO
14 paid company holidays
+4
Site Reliability Engineer — Scale AI Infra with Ownership
Site Reliability Engineer — Scale AI Infra with Ownership

Happyrobot Inc. • San Francisco (CA)

On-site
USD 100,000 - 140,000
Competitive salary + equity
Ownership & autonomy in projects
Opportunity to work with top-tier engineers