Site Reliability Engineer

Runloop

San Francisco (CA)

Hybrid

USD 150,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
Daily catered lunch
Equity & competitive salary

Job summary

Runloop is seeking a Site Reliability Engineer to own the reliability, observability, performance, and security of our code sandbox platform. You will work with the engineering team to build and maintain distributed systems ensuring a seamless developer experience.

The role requires strong CS fundamentals, 5+ years in software engineering with 3+ years in SRE/DevOps, and expertise in Python/Go, Docker, Kubernetes, and cloud tooling. Hybrid onsite in San Francisco with on-call responsibilities.

Qualifications

  • Strong CS fundamentals backed by a degree or equivalent experience.
  • 5+ years of software engineering with 3+ years in site reliability, DevOps, or infra ops.
  • Proficient in Python or Go.
  • Expertise with Docker and Kubernetes.
  • Experience with Terraform and/or Pulumi.
  • Familiarity with monitoring/alerting tools like Prometheus, Grafana, or Datadog.
  • Solid understanding of networking, security, and Linux systems administration.
  • Experience designing, scaling, and maintaining distributed systems.
  • Ability to implement observability and balance reliability with developer velocity.
  • Hands-on incident management and blameless post-mortems.

Responsibilities

  • Design and maintain production infrastructure on cloud platforms (AWS, GCP, Azure).
  • Monitor alerts and incidents to ensure high availability and security.
  • Collaborate with engineers to ensure scalable, reliable features.
  • Troubleshoot complex infra issues across networks and sandbox environments.
  • Participate in on-call rotation for production support.
  • Define SLIs/SLOs and manage error budgets.
  • Automate deployments, scaling, provisioning, and recovery tasks.
  • Lead incident response and conduct root-cause analyses.
  • Mentor engineers and influence reliability practices across teams.
  • Plan for capacity growth and safe release/change management.

Skills

Strong problem solving
System design
Leadership
Mentoring engineers

Education

Bachelor's degree in CS/EE

Tools

Python
Go
Docker
Kubernetes
Terraform
Pulumi
Prometheus
Grafana
Datadog
Linux
Networking
Security

Job description

About Runloop

Runloop.ai is pioneering the next generation of infrastructure and orchestration to power the Agentic Web/age of AI Agents. Our platform empowers developers to deploy agents that write code, browse the web, and use computers the way a human would. We're a small team of former Google and Stripe engineers, including the founding team of Google Wallet, dedicated to solving the complex challenges of productionizing AI for software engineering at scale.

The Role

We're looking for a skilled and passionate Site Reliability Engineer to join our team. As a SRE, you'll be responsible for the reliability, observability, performance, and security of our core platform. You'll work closely with our engineering team to develop and maintain the systems that power our code sandboxes, ensuring a seamless and stable experience for our customers. This is a critical role that blends a deep understanding of distributed systems with a software engineering mindset.

Responsibilities
  • Design and maintain our production infrastructure on cloud platforms like AWS, GCP, Azure, and emergent Neo-Clouds
  • Monitor and respond to system alerts and incidents using Grafana and Prometheus, ensuring high availability and a secure environment for our users
  • Collaborate with developers to ensure new features and services are designed with scalability and reliability in mind
  • Troubleshoot and resolve complex issues related to our infrastructure, networking, and the sandbox environment
  • Participate in an on-call rotation to support our production systems
  • Define and track SLIs/SLOs, manage error budgets, and proactively monitor distributed systems with logging and tracing
  • Automate deployments, scaling, provisioning, and recovery tasks to reduce toil and build self-healing systems
  • Lead incident response, conduct root‑cause analysis, and facilitate blameless post‑mortems to drive continual improvement
  • Collaborate cross‑functionally with product, engineering, and developer relations to ensure reliable releases and an outstanding developer experience
  • Plan for capacity growth, forecast system usage, and contribute to safe release and change management processes
Qualifications
  • Strong computer science fundamentals, backed by a degree from a top‑tier CS/EE program, or equivalent experience
  • 5+ years of experience in software engineering, with at least 3 years focused explicitly on site reliability, DevOps, or infrastructure operations
  • Strong programming skills in languages like Python or Go
  • Deep expertise in containerization technologies such as Docker and Kubernetes
  • Experience with cloud infrastructure and tools like Terraform and/or Pulumi
  • Familiarity with monitoring and alerting tools like Prometheus, Grafana, or Datadog
  • A solid understanding of networking, security, and Linux systems administration
  • Experience designing, scaling, and maintaining distributed systems (backend platforms, APIs, or front‑end infrastructure)
  • Proficiency in implementing observability frameworks (metrics, logging, tracing) and aligning reliability goals with developer velocity
  • Hands‑on experience managing incidents, running on‑call operations, and producing actionable post‑mortems
  • Ability to mentor engineers and influence reliability practices across teams, especially for front‑end infrastructure and performance
Bonus Points
  • Experience with chaos engineering techniques, front‑end observability tools (e.g., Sentry, RUM, synthetic monitoring), or building CI/CD pipelines for front‑end delivery
Benefits
  • Competitive salary and equity
  • Comprehensive health, dental, and vision insurance for employee and dependents
  • Opportunity to work on cutting‑edge technology and make a real impact on the future of software engineering
  • Daily catered lunch for all employees and a fridge full of your favorite snacks and drinks
Location
  • Onsite 4 days a week in San Francisco; Optional 1 day a week remote

Join Us! If you're excited about shaping the future of AI‑driven software engineering and empowering developers to build the next generation of AI powered coding tools, we want to hear from you. Join the Runloop team and be at the forefront of the AI revolution in software development.

Runloop AI is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, disability status, protected veteran status, sexual orientation, gender identity, or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Infrastructure
Software Engineer, Infrastructure

Runloop • San Francisco (CA)

Hybrid
USD 120,000 - 150,000
Competitive salary and equity
Health, dental, and vision insurance
Free lunch and snacks
+1
Software Engineer, Full Stack
Software Engineer, Full Stack

Runloop • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Competitive salary and equity
Comprehensive health, dental, and vision insurance
Opportunity to work on cutting-edge technology
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Replit • Town of Italy (NY)

On-site
USD 140,000 - 190,000
Competitive Salary & Equity
401(k) with 4% match (US Only)
Health, Dental, Vision, Life Insurance
+8
Senior Site Reliability Engineer: Scale, Automate, Observe
Senior Site Reliability Engineer: Scale, Automate, Observe

Replit • Town of Italy (NY)

On-site
USD 140,000 - 190,000
SRE — AI Platform Infra & Observability
SRE — AI Platform Infra & Observability

Runloop • San Francisco (CA)

Hybrid
USD 150,000 - 180,000
Health insurance
Daily catered lunch
Equity & competitive salary
Senior Site Reliability Engineer
Senior Site Reliability Engineer

hardrockdigital • United States

Hybrid
USD 120,000 - 160,000
Competitive pay and benefits
Flexible vacation allowance
Startup culture with global brand support
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Supio • San Francisco (CA)

On-site
USD 170,000 - 220,000
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

Hybrid
USD 146,000 - 163,000
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1
Founding Engineer - Site Reliability
Founding Engineer - Site Reliability

uRun • San Francisco (CA)

On-site
USD 120,000 - 140,000
Health, dental, and vision – full coverage
401(k) – company-supported retirement savings
FSA/HSA – flexible spending accounts
+3
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health benefits
Unlimited PTO
Paid parental leave
+1