Staff Software Engineer , Observability

EngineersOfAI

Northern (KY)

On-site

USD 140,000 - 210,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

&

Job summary

Reddit's Observability (OBS) team is hiring an Engineer who thrives at the intersection of infrastructure and software development. You will work on a stack built on open-source tools at scale, including Prometheus, Thanos, and more, to help engineers understand their systems.

You will improve availability, latency, and efficiency of observability components, contribute to upstream OSS, and share on-call responsibilities while collaborating with a service-oriented team across the globe.

Qualifications

  • 7+ years of experience developing internet-scale software, preferably in infrastructure.
  • Experience with distributed systems development and the listed tools (Prometheus, Thanos, Grafana, Vector, Clickhouse, OTel, Loki).
  • Experience developing on top of Kubernetes or similar distributed systems.
  • Strong troubleshooting capabilities surrounding both systems and software.
  • Excellent communication skills to collaborate with a service-oriented team.

Responsibilities

  • Collaborate with software engineers to create and maintain the foundational platform for running Reddit's infrastructure.
  • Deliver software to improve availability, scalability, latency, and efficiency of observability components.
  • Contribute feedback to the technical and strategic direction of eventing at Reddit.
  • Automate critical aspects of the event-driven development process.
  • Share on-call responsibilities.
  • Contribute upstream changes to the open source projects we use.

Skills

Distributed systems
Troubleshooting
Communication
Self-starting

Tools

Prometheus
Thanos
Grafana
Vector
Clickhouse
OTel
Loki
Kubernetes

Job description

This role is fully remote. Reddit has a flexible first workforce.

The Observability (OBS) team is looking to hire an Engineer that thrives at the intersection of infrastructure and software development. This team owns a suite of tools for allowing engineers to understand their creations, based primarily on open-source solutions at scale. Were active users of and contributors to our open source tools, including Prometheus, Thanos, and more.

Monitoring

Our monitoring stack at Reddit processes billions of data-points a minute, delivering insights at scale. Reddit has one of the larger deployments of Prometheus and Thanos in the world , and with this come unique challenges of scale for these systems. Fun problems include performance engineering on a distributed query system and product thinking around new features to remove the user pain from this stack.

Logging

We operate a platform for logging based on Vector and Clickhouse that processes millions of logs a day for Reddit. Thousands of people at Reddit use this system every day as they maintain, deploy, and debug their code. Join us as we invest in this system to help users better understand the data they have.

Distributed Tracing

Tracing at Reddit is based on OTEL, Clickhouse, and Grafana. Its light on features, but currently processing tens of millions of events, and we will be building new ways for users to leverage and understand this data in the coming year.

As a member of the Observability team, your work will span these domains, which are rich with challenging infrastructure and software engineering problems. Your work will directly impact hundreds of millions of users around the world. Join us and help build the future of Reddit!

In your day-to-day, you can expect to:

  • Work collaboratively with a team of software engineers to create and maintain the foundational platform for running Reddits infrastructure.
  • Deliver software to improve the availability, scalability, latency, and efficiency of observability components.
  • Contribute feedback to the technical and strategic direction of eventing at Reddit.
  • Automate critical aspects of the event driven development process
  • Share on-call responsibilities.
  • Contribute upstream changes to the open source projects we use

You have:

  • 7+ years of experience developing internet-scale software, preferably in the context of infrastructure.
  • Familiarity with distributed systems development, bonus if familiar with any of the specific tools (Prometheus, Thanos, Grafana, Vector, Clickhouse, Otel, Loki)
  • Experience developing on top of Kubernetes or similar distributed systems.
    • Kubernetes controller or operator development experience is a huge plus.
  • Strong troubleshooting capabilities surrounding both systems and software.
  • Experience engineering large systems, tracking work, and being a self-starter on projects.
  • Excellent communication skills to collaborate with a service-oriented team and company.

Benefits:

  • &
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Software Engineer, Observability Platform (Remote)
Staff Software Engineer, Observability Platform (Remote)

Reddit, Inc. • Northern (KY)

Hybrid
USD 217,000 - 304,000
Healthcare benefits
401k with employer match
Global benefit programs
+2
Staff Software Engineer, Observability at Scale (Remote)
Staff Software Engineer, Observability at Scale (Remote)

EngineersOfAI • Northern (KY)

Hybrid
USD 140,000 - 210,000
&
Senior Observability & Infra Engineer - Remote
Senior Observability & Infra Engineer - Remote

Reddit, Inc. • San Francisco (CA)

On-site
USD 217,000 - 304,000
Healthcare benefits
401k with employer match
Paid parental leave
+2
Staff Software Engineer , Observability New Remote - United States
Staff Software Engineer , Observability New Remote - United States

Reddit, Inc. • Northern (KY)

Remote
USD 217,000 - 304,000
Healthcare benefits
401k with employer match
Global benefit programs
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Brez Technology Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Private Medical, Dental and Vision Benefits
Retirement Savings plan with matching contributions
Workspace benefits for your home office
+4
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Alien Blue • Chicago (IL)

On-site
USD 190,800 - 267,100
Comprehensive Healthcare Benefits
401k Matching
Flexible Vacation
+2
Software Engineer, Content Platform
Software Engineer, Content Platform

EngineersOfAI • Northern (KY)

On-site
USD 120,000 - 180,000
Healthcare benefits
401k match
Family planning support
+1
Senior Software Engineer - Full Stack Internal Tooling (Build and Deployment Platform)
Senior Software Engineer - Full Stack Internal Tooling (Build and Deployment Platform)

EngineersOfAI • United States

On-site
USD 110,000 - 150,000
Staff Site Reliability Engineer, Ads
Staff Site Reliability Engineer, Ads

Reddit, Inc. • New York (NY)

On-site
USD 217,000 - 304,000
Global Benefit programs
Family Planning Support
Gender-Affirming Care
+2
Staff Software Engineer - Ingestion Platform New Remote - United States
Staff Software Engineer - Ingestion Platform New Remote - United States

Reddit, Inc. • Northern (KY)

Remote
USD 217,000 - 304,000
Equity
Health benefits (medical, dental, and視
401(k) match
+2