Senior Site Reliability Engineer

Brez Technology Inc.

San Francisco (CA)

Hybrid

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Private Medical, Dental and Vision Benefits
Retirement Savings plan with matching contributions
Workspace benefits for your home office
Personal & Professional development funds
Family Planning Support
Commuter Benefits
Flexible Vacation & Reddit Global Days Off

Job summary

A leading technology company in San Francisco is seeking a Senior Site Reliability Engineer to enhance the performance and reliability of systems. The ideal candidate will have over 5 years of experience, proficiency in programming languages like Go and Python, and familiarity with Kubernetes. Responsibilities include advising engineering teams, automating tasks, and troubleshooting issues. Competitive benefits include private medical, dental, and vision, as well as a flexible vacation policy.

Qualifications

  • 5+ years of experience in Software Engineering, Site Reliability Engineering, or DevOps.
  • Proficiency in Go and Python programming languages.
  • Experience with Kubernetes and cloud systems.
  • Strong debugging and troubleshooting skills.

Responsibilities

  • Advise engineering teams on resilient system design.
  • Identify and build capabilities into foundational services.
  • Deliver software to improve observability components.
  • Automate repetitive tasks to enhance efficiency.
  • Diagnose and fix service-level issues.

Skills

Software Engineering
Site Reliability Engineering
Kubernetes
Go
Python
Distributed systems
Linux

Tools

Prometheus
Grafana
Thanos
Vector
ClickHouse

Job description

Reddit is a community of communities. It’s built on shared interests, passion, and trust and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximately 101M+ daily active unique visitors, Reddit is one of the internet’s largest sources of information. For more information, visit redditinc.com .

Reddit SRE is rapidly innovating and our teams are working to meet the needs of infrastructure and development teams as they evolve our product faster than ever before. This is a unique opportunity to leave your mark on one of the most influential and trafficked corners of the internet.

As a Senior Site Reliability Engineer on Reddit’s Infrastructure SRE team, you’ll use your knowledge of distributed systems and architecture to improve the reliability and performance of Reddit’s engineering platforms and services. We are looking for someone who thrives at the intersection of infrastructure and software development. This team will work very closely with the Compute, Traffic, and Observability infrastructure teams. They will own a suite of tools for allowing engineers to understand their creations, based primarily on open‑source solutions at scale. We’re active users of and contributors to Prometheus, Thanos, Grafana, Vector and more.

In this role, you will also take ownership of risk management, ensuring the reliability and performance of our systems. You will collaborate with cross‑functional teams to identify, assess, and mitigate risks, implementing best practices to enhance system resilience. Your expertise will drive proactive measures to maintain uptime and optimize service delivery, making a significant impact on our operational excellence.

Join us and help build the future of Reddit!

Responsibilities
  • Advise: Work closely with engineering teams in designing and developing systems that are resilient and highly performant at a tremendous scale, and maintaining the foundational platform for running Reddit’s infrastructure.
  • Amplify: Identify and build capabilities into our foundational Infrastructure and Platform services, which are used by Reddit engineering teams to build, deploy, and operate Reddit.
  • Deliver software to improve the availability, scalability, latency, and efficiency of observability components.
  • Identify and engineer away risk across Reddit’s systems.
  • Automate: Take repetitive, manual, or risky tasks and automate them out of existence. Build tools and integrate systems to support Reddit’s evolution.
  • Automate critical aspects of the event driven development process.
  • Diagnose: Draw on your knowledge of distributed systems to identify and fix network, system, and service‑level issues. Practice sustainable incident response, and drive structural improvement with blameless post‑mortem.
  • Share on‑call responsibilities.
  • Optimize: Observe and improve performance, reduce cost, and improve the experience for millions of users.
  • Contribute upstream changes to the open source projects we use.
Qualifications
  • 5+ years of experience in Software Engineering, Site Reliability Engineering, or a development‑focused DevOps role.
  • Proficiency in one or more programming languages. We’re predominantly writing code in Go and Python.
  • Experience with Kubernetes and Cloud systems.
  • Familiarity with distributed systems development, bonus if familiar with any of the specific tools (Prometheus, Thanos, Grafana, Vector, ClickHouse, Otel, Loki).
  • Experience with the development and operation of high‑traffic backend systems.
  • A demonstrated ability to debug, fix, and optimize code.
  • Troubleshooting skills that span applications, networking (TCP/IP), and systems.
  • Strong working knowledge of Linux and containers.
  • Excellent communication and collaborative skills.
Benefits
  • Private Medical, Dental and Vision Benefits
  • Retirement Savings plan with matching contributions
  • Workspace benefits for your home office
  • Personal & Professional development funds
  • Family Planning Support
  • Commuter Benefits
  • Flexible Vacation & Reddit Global Days Off

Reddit is proud to be an equal opportunity employer, and is committed to building a workforce representative of the diverse communities we serve. Reddit is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. If, due to a disability, you need an accommodation during the interview process, please let your recruiter know.

How to Apply

Interested in this position? Please submit your resume and cover letter through the application portal.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Alien Blue • Chicago (IL)

On-site
USD 190,000 - 268,000
Comprehensive Healthcare Benefits
401k Matching
Flexible Vacation
+2
Software Engineer, Content Platform
Software Engineer, Content Platform

Reddit, Inc. • Northern (KY)

Hybrid
USD 164,000 - 230,000
Healthcare benefits
401k match
Family planning support
+3
Software Engineer, Content Platform
Software Engineer, Content Platform

EngineersOfAI • Northern (KY)

Hybrid
USD 120,000 - 180,000
Healthcare benefits
401k match
Family planning support
+1
Senior Software Engineer, Home Experience
Senior Software Engineer, Home Experience

Jackalope Digital LLC • Northern (KY)

Hybrid
USD 190,000 - 267,000
Healthcare Benefits
401k with Employer Match
Global Benefits Program
+1
Senior Software Engineer - DevX
Senior Software Engineer - DevX

Reddit, Inc. • San Francisco (CA)

On-site
USD 190,000 - 268,000
Comprehensive Healthcare Benefits and
401k with Employer Match
Global Benefit programs; professional,
+5
Senior Backend Engineer, Compliance Engineering
Senior Backend Engineer, Compliance Engineering

Reddit, Inc. • San Francisco (CA)

On-site
USD 191,000 - 267,000
Comprehensive Healthcare Benefits and
401k with Employer Match
Global Benefit programs
+5
Engineering Manager, Search Storage
Engineering Manager, Search Storage

Reddit, Inc. • San Francisco (CA)

On-site
USD 217,000 - 303,000
Healthcare benefits
401k with employer match
Global benefits program
+5
Staff Software Engineer , Observability New Remote - United States
Staff Software Engineer , Observability New Remote - United States

Reddit, Inc. • Northern (KY)

Hybrid
USD 217,000 - 304,000
Healthcare benefits
401k with employer match
Global benefit programs
+2
Senior Software Engineer, Core Platform
Senior Software Engineer, Core Platform

Reddit, Inc. • Chicago (IL)

On-site
USD 190,000 - 268,000
Comprehensive Healthcare Benefits and
Income Replacement Programs
401k with Employer Match
+6
Senior Software Engineer, Storage
Senior Software Engineer, Storage

Reddit, Inc. • San Francisco (CA)

On-site
USD 217,000 - 304,000
Restricted stock units
401(k) with employer match
Medical insurance
+4