Site Reliability Engineer, AI Infra & Observability

Sierra

California (MO)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Unlimited PTO
Medical, dental, vision
Retirement plan
Parental leave
Equity plans
Snacks and coffee
Discretionary stipend

Job summary

Sierra is hiring a Software Engineer for the Site Reliability team to build the foundation of reliability, observability, and scalability across our AI-driven platform. You will work with core engineering and product teams to ensure high availability, efficiency, and growth readiness.

In a fast-moving, in-person setting headquartered in San Francisco, you will own the observability stack, design scalable infrastructure with Terraform on AWS, and advance deployment pipelines and incident response

Qualifications

  • 5+ years of hands-on experience in Site Reliability or Infrastructure engineering.
  • Experience designing for availability, scalability and reliability.
  • Strong background with Terraform, AWS services, container orchestration and cloud networking (IAM, VPC).

Responsibilities

  • Own Sierra's observability stack: monitoring, alerting, logging and tracing.
  • Design scalable, reliable cloud infrastructure (AWS) using Terraform.
  • Improve deployment pipelines, CI/CD tooling, and incident management processes.
  • Lead improvements to SRE culture and best practices across the engineering org.
  • Partner with product and platform engineers to ensure reliability from day one.
  • Ensure LLM deployments are robust, performant, and cost-efficient.

Skills

SRE
Terraform
AWS
Kubernetes
Observability

Education

Degree in Computer Science

Tools

Terraform
AWS
Kubernetes
Datadog

Job description

Sierra is hiring a Software Engineer for the Site Reliability team to build the foundation of reliability, observability, and scalability across our AI-driven platform. You will work with core engineering and product teams to ensure high availability, efficiency, and growth readiness.

In a fast-moving, in-person setting headquartered in San Francisco, you will own the observability stack, design scalable infrastructure with Terraform on AWS, and advance deployment pipelines and incident response

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Infrastructure Software Engineer for Scalable AI Platform
Infrastructure Software Engineer for Scalable AI Platform

Sierra • California (MO)

On-site
USD 180,000 - 250,000
Unlimited PTO
Medical, dental, and vision benefits
Life insurance and disability benefits
+6
Site Reliability Engineer — ML Infra & Observability
Site Reliability Engineer — ML Infra & Observability

Baseten • San Francisco (CA)

On-site
USD 135,000 - 285,000
Competitive compensation including equity
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
+3
Senior Site Reliability Engineer — AI-First Observability
Senior Site Reliability Engineer — AI-First Observability

Qualitest • Riverwoods (IL)

On-site
USD 110,000 - 130,000
Diversity and inclusion
Internal rotation
Career progression
+3
Remote Site Reliability Engineer — AI-Driven Observability
Remote Site Reliability Engineer — AI-Driven Observability

OhioX • Northern (KY)

Hybrid
USD 142,000 - 197,000
AI-Driven Site Reliability Engineer – Remote
AI-Driven Site Reliability Engineer – Remote

Upstart • Austin (TX), San Francisco (CA), New York (NY)

Hybrid
USD 142,000 - 197,000
Competitive pay
Annual equity grants
401(k) matching
+2
Remote Site Reliability Engineer: AI-Driven Ops
Remote Site Reliability Engineer: AI-Driven Ops

Upstart • United States

On-site
USD 142,000 - 197,000
401k
ESPP
Health coverage
+3
Site Reliability Engineer, Cloud Infra for AI Platform
Site Reliability Engineer, Cloud Infra for AI Platform

Anyscale • San Francisco (CA)

On-site
USD 150,000 - 210,000
Software Engineer, Site Reliability (SRE)
Software Engineer, Site Reliability (SRE)

Sierra • California (MO)

On-site
USD 150,000 - 210,000
Unlimited PTO
Medical, dental, vision
Retirement plan
+4
Site Reliability Engineer — Scale & Resilience for AI Ops
Site Reliability Engineer — Scale & Resilience for AI Ops

HappyRobot • San Francisco (CA)

On-site
USD 120,000 - 160,000
SF Platform Engineer — AI Infra, Scale & Reliability
SF Platform Engineer — AI Infra, Scale & Reliability

Harper • San Francisco (CA)

On-site
USD 140,000 - 280,000
Uber commuter benefits
Meals provided (breakfast, lunch, and/
Snacks, drinks and coffee daily
+2