Remote SRE: AI Platform Reliability & Automation

Runpod

United States

On-site

USD 150,000 - 200,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Remote work first
Competitive base salary
Stock options equity
Flexible PTO
Medical, dental & vision plans

Job summary

Runpod, a remote-first AI developer cloud, seeks a Site Reliability Engineer to own reliability, observability, and operations for a distributed platform. You will design SLIs/SLOs, drive incident response, and automate deployments while strengthening production readiness.

You’ll work with cross-functional teams to reduce toil, improve MTTR, and ensure scalable performance in a fast-growing environment. Strong scripting and GPU awareness are valued.

Qualifications

  • 5+ years in SRE, Reliability Engineering, or Production Engineering.
  • Strong Linux systems and Networking expertise.
  • Experience managing containerized production systems.
  • Strong understanding of distributed systems and failure modes.
  • Experience defining and managing SLIs/SLOs.
  • Proven incident response and postmortem leadership experience.
  • Strong scripting or programming skills.
  • Experience with monitoring and alerting systems.
  • Excellent written communication skills.
  • Background check completed.

Responsibilities

  • Define and implement SLIs/SLOs for critical services.
  • Lead incident response and coordinate cross-team mitigation efforts.
  • Conduct blameless postmortems and ensure corrective actions are completed.
  • Perform production readiness reviews for new services and features.
  • Identify systemic risks and drive preventative improvements.
  • Partner with engineering to improve system resilience and fault tolerance.
  • Contribute to architectural discussions with a reliability-first mindset.

Skills

Linux
Networking
SRE
Incident response
Scripting

Tools

Prometheus
Grafana
CI/CD
Kubernetes

Job description

Runpod, a remote-first AI developer cloud, seeks a Site Reliability Engineer to own reliability, observability, and operations for a distributed platform. You will design SLIs/SLOs, drive incident response, and automate deployments while strengthening production readiness.

You’ll work with cross-functional teams to reduce toil, improve MTTR, and ensure scalable performance in a fast-growing environment. Strong scripting and GPU awareness are valued.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote Senior SRE: Build Reliable, Scalable AI Infra
Remote Senior SRE: Build Reliable, Scalable AI Infra

Runware • Town of Sweden (NY)

On-site
USD 140,000 - 190,000
Generous paid time off
Meaningful stock options
Remote-first setup
+3
Remote SRE Manager: Lead AI-Driven Reliability & Cloud Ops
Remote SRE Manager: Lead AI-Driven Reliability & Cloud Ops

Arcoro Holdings Corp • Phoenix (AZ), Northern (KY)

Hybrid
USD 200,000 - 220,000
Remote Work
401(k) with Company match
Flexible PTO and Company-paid holidays
SRE — AI Platform Infra & Observability
SRE — AI Platform Infra & Observability

Runloop • San Francisco (CA)

Hybrid
USD 150,000 - 180,000
Health insurance
Daily catered lunch
Equity & competitive salary
Head of Scalable AI Infrastructure & SRE
Head of Scalable AI Infrastructure & SRE

Runpod • United States

On-site
USD 225,000 - 325,000
Stock options
Remote-first culture
Competitive salary
+2
Real-Time AI SRE — Remote, High-Impact Reliability
Real-Time AI SRE — Remote, High-Impact Reliability

Granite State Trade School LLC • New York (NY)

Remote
USD 120,000 - 150,000
Equity ownership
Remote-first flexibility
Impact on user platform
Remote Site Reliability Engineer: AI-Driven Ops
Remote Site Reliability Engineer: AI-Driven Ops

Upstart • United States

On-site
USD 142,000 - 197,000
401k
ESPP
Health coverage
+3
Remote SRE: AI-Driven Reliability & Observability
Remote SRE: AI-Driven Reliability & Observability

Baseten • United States

On-site
USD 140,000 - 180,000
Remote‑first
In-person team summits
Unlimited PTO
+2
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Remote SRE – Scale, Automation & Ownership (US CT/ET)
Remote SRE – Scale, Automation & Ownership (US CT/ET)

PostHog • United States

On-site
USD 140,000 - 210,000
Remote Full-Stack Engineer for AI Platform Open-Source
Remote Full-Stack Engineer for AI Platform Open-Source

Runpod • United States

On-site
USD 130,000 - 200,000
Equity
Health plans
Home Office stipend
+1