Remote SRE: AI-Driven Reliability & Observability

Baseten

United States

On-site

USD 140,000 - 180,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Remote‑first
In-person team summits
Unlimited PTO
Healthcare coverage
Parental leave

Job summary

Baseten is seeking an experienced Site Reliability Engineer to define day-2 operations for our ML infrastructure. You will build robust systems, automations, and observability tools to keep Baseten reliable at scale, collaborating with engineering, product, and forward‑deployed teams to turn failure patterns into automated mitigations.

You will own multi-cloud Kubernetes reliability, develop runbooks, and instrument SLOs/SLIs with a focus on low-context, safe execution.

Qualifications

  • Strong foundation in observability tooling (metrics, logging, dashboards)
  • Kubernetes multi-cloud experience (EKS, GKE, or similar)
  • Experience with Terraform/Helm and GitOps workflows (Flux CD, ArgoCD)
  • Incident management experience and post-mortem analysis
  • Experience writing runbooks and leading incident response
  • Familiarity with infrastructure-as-code and scalable infra design
  • Interest in ML model deployment and serving at scale

Responsibilities

  • Define and codify the gold standards of day 2 operations for our ML infrastructure platform
  • Build robust systems, processes, automations, and observability tooling for reliability at scale
  • Collaborate with engineering, forward‑deployed and product teams to convert failures into automated mitigations
  • Improve SRE practices by instrumenting SLOs and SLIs, and improving alerting and observability
  • Develop AI-assisted tooling for incident triage and response
  • Own the reliability of multi-cloud Kubernetes infrastructure, including post‑mortems and remediation tracking
  • Create and maintain observability infrastructure—metrics, logging, dashboards, and alerting—as code
  • Author and improve runbooks for recurring failures for low-context, safe execution
  • Identify high-frequency failure patterns and convert them into automated mitigations or self‑healing automations
  • Diagnose runtime issues related to latency, memory, GPU utilization, concurrency, and model lifecycle management
  • Define SLOs/SLIs across customer workloads and internal services
  • Navigate ambiguity and balance tradeoffs to avoid unnecessary complexity

Skills

Observability tooling
Runbooks & incident response
Kubernetes multi-cloud
Terraform & Helm
GitOps (Flux/ArgoCD)
Incident management

Tools

VictoriaMetrics
Prometheus
Loki
ELK
Grafana
Flux CD
ArgoCD
Terraform
Helm
incident.io

Job description

Baseten is seeking an experienced Site Reliability Engineer to define day-2 operations for our ML infrastructure. You will build robust systems, automations, and observability tools to keep Baseten reliable at scale, collaborating with engineering, product, and forward‑deployed teams to turn failure patterns into automated mitigations.

You will own multi-cloud Kubernetes reliability, develop runbooks, and instrument SLOs/SLIs with a focus on low-context, safe execution.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE
SRE

The Consensus • New York (NY)

On-site
USD 110,000 - 140,000
Competitive compensation
100% insurance coverage
Flexible PTO policy
+3
Remote SRE: AI Platform Reliability & Automation
Remote SRE: AI Platform Reliability & Automation

Runpod • United States

On-site
USD 150,000 - 200,000
Remote work first
Competitive base salary
Stock options equity
+2
Site Reliability Engineer
Site Reliability Engineer

Baseten • San Francisco (CA)

On-site
USD 135,000 - 285,000
Competitive compensation including equity
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
+3
Senior SRE: AI-Driven Cloud Reliability
Senior SRE: AI-Driven Cloud Reliability

SupportFinity™ • San Francisco (CA)

Hybrid
USD 164,000 - 205,000
BetterUp coaching
Competitive pay
Medical, dental, and vision insurance
+7
Remote Site Reliability Engineer: AI-Driven Ops
Remote Site Reliability Engineer: AI-Driven Ops

Upstart • United States

On-site
USD 142,000 - 197,000
401k
ESPP
Health coverage
+3
Observability Engineer: Scale Telemetry & Reliability
Observability Engineer: Scale Telemetry & Reliability

The Consensus • New York (NY)

On-site
USD 150,000 - 210,000
Competitive compensation
Equity
Medical, dental, vision coverage
+4
SRE for AI Platform: Reliability at Scale
SRE for AI Platform: Reliability at Scale

Mosaic.tech • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Remote SRE Technical Lead — Automate & Scale Reliability
Remote SRE Technical Lead — Automate & Scale Reliability

Bright Vision Technologies • New Albany (IN), City of Albany (NY)

On-site
USD 100,000 - 150,000
Remote SRE Lead - Incident, Reliability & Observability
Remote SRE Lead - Incident, Reliability & Observability

NightDragon Acquisition Corp. • United States

On-site
USD 260,000 - 280,000
Hybrid & Remote Work
Competitive Compensation
Equity package