Senior SRE: AI-Driven Cloud Reliability

SupportFinity™

San Francisco (CA)

Hybrid

USD 164,000 - 205,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

BetterUp coaching
Competitive pay
Medical, dental, and vision insurance
Flexible paid time off
Holiday observances
Inner Workdays
Volunteer Days
L&D stipend
Summer & Winter breaks
401(k)

Job summary

BetterUp Inc. is seeking a Site Reliability Engineer to build and operate scalable cloud infrastructure and observability for our AI-forward platform.

You will deploy on AWS with Terraform, manage Kubernetes clusters, and design resilient, automated systems with modern observability stacks. The role emphasizes collaboration, automation, and a maker mindset within a hybrid work environment.

Qualifications

  • 4+ years of experience in SRE or infrastructure roles.
  • Genuine excitement about AI tooling and using copilots/LLMs.
  • Deep experience with AWS.
  • Hands-on Kubernetes experience deploying, scaling, debugging, and securing clusters.
  • Strong Terraform skills for multi-environment infra.
  • Familiarity with modern observability stacks (Datadog, Prometheus, OpenTelemetry).
  • Strong debugging instincts and ability to explain incidents clearly.
  • Builder mindset to automate manual processes.

Responsibilities

  • Leverage AI-powered tools and automation to transform monitoring and maintenance of production systems.
  • Build and operate cloud infrastructure on AWS using Terraform.
  • Manage and scale Kubernetes clusters for high availability and performance.
  • Design alerting and observability systems.
  • Collaborate with engineers to embed reliability in development lifecycle.
  • Automate incident response workflows and build self-healing infra.
  • Experiment with AI tools for log analysis and anomaly detection.
  • Drive continuous improvement with data-driven retrospectives and reliability metrics.

Skills

4+ years SRE/infra experience
AI tooling enthusiasm
Strong debugging instincts
Clear communication
Builder's mindset

Tools

AWS
Kubernetes
Terraform
Datadog
Prometheus
OpenTelemetry

Job description

BetterUp Inc. is seeking a Site Reliability Engineer to build and operate scalable cloud infrastructure and observability for our AI-forward platform.

You will deploy on AWS with Terraform, manage Kubernetes clusters, and design resilient, automated systems with modern observability stacks. The role emphasizes collaboration, automation, and a maker mindset within a hybrid work environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI-Driven Site Reliability Engineer – Remote
AI-Driven Site Reliability Engineer – Remote

Upstart • Austin (TX), San Francisco (CA), New York (NY)

Hybrid
USD 142,000 - 197,000
Competitive pay
Annual equity grants
401(k) matching
+2
Senior SRE: AI-Driven Reliability & Cloud Automation
Senior SRE: AI-Driven Reliability & Cloud Automation

Quality Ai • Northern (KY)

Hybrid
USD 110,000 - 130,000
Competitive pay
Global opportunities
Technical training & certification
Senior SRE: AI-Driven Cloud Reliability & Automation
Senior SRE: AI-Driven Cloud Reliability & Automation

Sight Machine • United States

Hybrid
USD 170,000 - 250,000
Hybrid work flexibility
Catered Lunches, Snacks and Beverages
Commuter Savings Program
+2
Senior SRE: AI-Driven Infra & Reliability
Senior SRE: AI-Driven Infra & Reliability

Jobless • Ann Arbor (MI)

Hybrid
USD 180,000 - 240,000
Health Care Coverage
Life Insurance
Health Savings Account
+3
Senior SRE: AI Cloud Platform & Kubernetes Expert
Senior SRE: AI Cloud Platform & Kubernetes Expert

Lambda • Bellevue (WA)

On-site
USD 180,000 - 260,000
Health insurance
Dental insurance
Vision insurance
+3
Senior SRE: AI-Driven Reliability & Automation (Hybrid)
Senior SRE: AI-Driven Reliability & Automation (Hybrid)

Namely • United States

Hybrid
USD 120,000 - 150,000
Senior SRE: Deployments, AI-Driven Infra on AWS
Senior SRE: Deployments, AI-Driven Infra on AWS

Supio • San Francisco (CA)

On-site
USD 170,000 - 220,000
Remote Site Reliability Engineer: AI-Driven Ops
Remote Site Reliability Engineer: AI-Driven Ops

Upstart • United States

On-site
USD 142,000 - 197,000
401k
ESPP
Health coverage
+3
Senior SRE — AI-Driven Cloud Reliability
Senior SRE — AI-Driven Cloud Reliability

BetterUp • New York (NY)

Hybrid
USD 164,000 - 205,000
Senior SRE: AI-Powered Reliability on AWS & Kubernetes
Senior SRE: AI-Powered Reliability on AWS & Kubernetes

BetterUp • Austin (TX)

Hybrid
USD 147,000 - 185,000