SRE — AI Platform Infra & Observability

Runloop

San Francisco (CA)

Hybrid

USD 150,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
Daily catered lunch
Equity & competitive salary

Job summary

Runloop is seeking a Site Reliability Engineer to own the reliability, observability, performance, and security of our code sandbox platform. You will work with the engineering team to build and maintain distributed systems ensuring a seamless developer experience.

The role requires strong CS fundamentals, 5+ years in software engineering with 3+ years in SRE/DevOps, and expertise in Python/Go, Docker, Kubernetes, and cloud tooling. Hybrid onsite in San Francisco with on-call responsibilities.

Qualifications

  • Strong CS fundamentals backed by a degree or equivalent experience.
  • 5+ years of software engineering with 3+ years in site reliability, DevOps, or infra ops.
  • Proficient in Python or Go.
  • Expertise with Docker and Kubernetes.
  • Experience with Terraform and/or Pulumi.
  • Familiarity with monitoring/alerting tools like Prometheus, Grafana, or Datadog.
  • Solid understanding of networking, security, and Linux systems administration.
  • Experience designing, scaling, and maintaining distributed systems.
  • Ability to implement observability and balance reliability with developer velocity.
  • Hands-on incident management and blameless post-mortems.

Responsibilities

  • Design and maintain production infrastructure on cloud platforms (AWS, GCP, Azure).
  • Monitor alerts and incidents to ensure high availability and security.
  • Collaborate with engineers to ensure scalable, reliable features.
  • Troubleshoot complex infra issues across networks and sandbox environments.
  • Participate in on-call rotation for production support.
  • Define SLIs/SLOs and manage error budgets.
  • Automate deployments, scaling, provisioning, and recovery tasks.
  • Lead incident response and conduct root-cause analyses.
  • Mentor engineers and influence reliability practices across teams.
  • Plan for capacity growth and safe release/change management.

Skills

Strong problem solving
System design
Leadership
Mentoring engineers

Education

Bachelor's degree in CS/EE

Tools

Python
Go
Docker
Kubernetes
Terraform
Pulumi
Prometheus
Grafana
Datadog
Linux
Networking
Security

Job description

Runloop is seeking a Site Reliability Engineer to own the reliability, observability, performance, and security of our code sandbox platform. You will work with the engineering team to build and maintain distributed systems ensuring a seamless developer experience.

The role requires strong CS fundamentals, 5+ years in software engineering with 3+ years in SRE/DevOps, and expertise in Python/Go, Docker, Kubernetes, and cloud tooling. Hybrid onsite in San Francisco with on-call responsibilities.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Runloop • San Francisco (CA)

Hybrid
USD 150,000 - 180,000
Health insurance
Daily catered lunch
Equity & competitive salary
Remote SRE: AI Platform Reliability & Automation
Remote SRE: AI Platform Reliability & Automation

Runpod • United States

On-site
USD 150,000 - 200,000
Remote work first
Competitive base salary
Stock options equity
+2
Senior SRE — Flexible, AI-Driven Reliability
Senior SRE — Flexible, AI-Driven Reliability

Salesforce, Inc. • San Francisco (CA)

Hybrid
USD 148,000 - 224,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Methodic • San Francisco (CA)

On-site
USD 140,000 - 210,000
Senior SRE — AI-Driven Reliability & Oncall Leadership
Senior SRE — AI-Driven Reliability & Oncall Leadership

Block • San Francisco (CA)

On-site
USD 160,700 - 283,600
Healthcare coverage
Health Savings Account
Retirement Plans
+5
Hybrid SRE: Infra & Assurance Services
Hybrid SRE: Infra & Assurance Services

TikTok • Seattle (WA)

Hybrid
USD 112,000 - 178,000
Senior SRE: Scale Reliability, Observability & CI/CD
Senior SRE: Scale Reliability, Observability & CI/CD

Breakout Tools • San Francisco (CA)

On-site
USD 120,000 - 160,000
SRE – SaaS Platform Reliability (Kubernetes, CI/CD)
SRE – SaaS Platform Reliability (Kubernetes, CI/CD)

Obsidian • Palo Alto (CA)

On-site
USD 165,000 - 190,000
Equity + 401k
Healthcare coverage
Flexible PTO
+2
Site Reliability Engineer
Site Reliability Engineer

Latent • San Francisco (CA)

On-site
USD 140,000 - 200,000
SRE, Cloud Platform – AI-Driven Reliability & Observability
SRE, Cloud Platform – AI-Driven Reliability & Observability

Visa • Austin (TX)

On-site
USD 88,000 - 137,000