Staff Site Reliability Engineer: Scale & Reliability Lead

Reddit, Inc.

New York (NY)

On-site

USD 217,000 - 304,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Global Benefit programs
Family Planning Support
Gender-Affirming Care
Mental Health & Coaching Benefits
Private Medical, Dental, and Vision

Job summary

Reddit, Inc. is seeking a Staff Site Reliability Engineer to lead reliability engineering for critical user-facing systems at internet scale.

You will partner with product and infrastructure teams to improve availability, latency, scalability, and operational excellence across APIs, content delivery, feed generation, search, messaging, and real-time experiences. This is a highly technical leadership role for someone who thrives in large-scale distributed systems and can influence engineering

Qualifications

  • 8+ years in Site Reliability Engineering, Infrastructure Engineering, or related roles operating large scale distributed systems.
  • Strong collaboration and communication skills with the ability to influence technical direction across teams.
  • Strong experience supporting high traffic, user facing production environments.
  • Deep understanding of distributed systems, networking, Linux systems, cloud native architectures.
  • Experience designing highly available systems with strong operational and reliability practices.
  • Strong programming skills in Go, Python, or similar.
  • Strong understanding of observability systems including metrics, logging, tracing, and alerting.
  • Experience improving reliability through SLOs, automation, incident management, and performance optimization.

Responsibilities

  • Lead Reliability Engineering for User Experience.
  • Architect for Scale and capacity planning with partner teams.
  • Reduce Operational Risk by identifying bottlenecks and driving mitigations.
  • Drive Automation to eliminate repetitive toil and improve deployment safety.
  • Incident Management including blameless postmortems and long-term fixes.
  • Influence Engineering Standards around SLIs/SLOs, release engineering, and operational maturity.
  • Mentor and multiply impact across SRE and software engineering teams.

Skills

SRE experience
Distributed systems
Go/Python
Observability
Incident management
Automation
Leadership
Communication

Tools

Kubernetes
Prometheus
Grafana
OpenTelemetry
Envoy
Kafka
Redis

Job description

Reddit, Inc. is seeking a Staff Site Reliability Engineer to lead reliability engineering for critical user-facing systems at internet scale.

You will partner with product and infrastructure teams to improve availability, latency, scalability, and operational excellence across APIs, content delivery, feed generation, search, messaging, and real-time experiences. This is a highly technical leadership role for someone who thrives in large-scale distributed systems and can influence engineering

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Site Reliability Engineer: Lead Global Reliability
Staff Site Reliability Engineer: Lead Global Reliability

Reddit, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 217,000 - 304,000
Medical, Dental, Vision
401(k) with employer match
Generous time off
+1
Senior Site Reliability Engineer — Ads Platform
Senior Site Reliability Engineer — Ads Platform

United States Digital Space LLC • United States

Remote
USD 217,000 - 304,000
Health benefits
401k Matching
Home office workspace
+5
Senior Site Reliability Lead: Scale & Reliability Champion
Senior Site Reliability Lead: Scale & Reliability Champion

Twitter • San Francisco (CA)

On-site
USD 130,000 - 160,000
Staff Site Reliability Engineer – Ads (Remote)
Staff Site Reliability Engineer – Ads (Remote)

Reddit, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 217,000 - 304,000
401k Matching
Workspace benefits
Family Planning Support
+2
Staff SRE - Ads Reliability & Platform
Staff SRE - Ads Reliability & Platform

Jackalope Digital LLC • San Francisco (CA), Northern (KY)

Hybrid
USD 217,000 - 304,000
Comprehensive Health benefits
401k Matching
Workspace benefits for home office
+3
Senior Site Reliability Engineer, Ads Platform
Senior Site Reliability Engineer, Ads Platform

Reddit, Inc. • San Francisco (CA)

On-site
USD 191,000 - 267,000
Comprehensive Health benefits
401k Matching
Workspace benefits for home office
+5
Staff Site Reliability Engineer, Ads
Staff Site Reliability Engineer, Ads

Reddit, Inc. • New York (NY)

On-site
USD 217,000 - 304,000
Global Benefit programs
Family Planning Support
Gender-Affirming Care
+2
Senior Site Reliability Engineer – Scale & Reliability
Senior Site Reliability Engineer – Scale & Reliability

Google • Seattle (WA)

On-site
USD 174,000 - 252,000
Health insurance
Dental insurance
Vision insurance
+6
Lead Site Reliability Engineer - Scale & Reliability
Lead Site Reliability Engineer - Scale & Reliability

TQC Ltd • United States

Remote
USD 170,000 - 230,000
Staff Site Reliability Engineer: Scale, Automate & Improve
Staff Site Reliability Engineer: Scale, Automate & Improve

Google • San Jose (CA)

On-site
USD 207,000 - 300,000
Equity
Benefits