Staff Software Engineer , Observability

EngineersOfAI

Northern (KY)

On-site

USD 140,000 - 210,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

&

Job summary

Reddit's Observability (OBS) team is hiring an Engineer who thrives at the intersection of infrastructure and software development. You will work on a stack built on open-source tools at scale, including Prometheus, Thanos, and more, to help engineers understand their systems.

You will improve availability, latency, and efficiency of observability components, contribute to upstream OSS, and share on-call responsibilities while collaborating with a service-oriented team across the globe.

Qualifications

  • 7+ years of experience developing internet-scale software, preferably in infrastructure.
  • Experience with distributed systems development and the listed tools (Prometheus, Thanos, Grafana, Vector, Clickhouse, OTel, Loki).
  • Experience developing on top of Kubernetes or similar distributed systems.
  • Strong troubleshooting capabilities surrounding both systems and software.
  • Excellent communication skills to collaborate with a service-oriented team.

Responsibilities

  • Collaborate with software engineers to create and maintain the foundational platform for running Reddit's infrastructure.
  • Deliver software to improve availability, scalability, latency, and efficiency of observability components.
  • Contribute feedback to the technical and strategic direction of eventing at Reddit.
  • Automate critical aspects of the event-driven development process.
  • Share on-call responsibilities.
  • Contribute upstream changes to the open source projects we use.

Skills

Distributed systems
Troubleshooting
Communication
Self-starting

Tools

Prometheus
Thanos
Grafana
Vector
Clickhouse
OTel
Loki
Kubernetes

Job description

This role is fully remote. Reddit has a flexible first workforce.

The Observability (OBS) team is looking to hire an Engineer that thrives at the intersection of infrastructure and software development. This team owns a suite of tools for allowing engineers to understand their creations, based primarily on open-source solutions at scale. Were active users of and contributors to our open source tools, including Prometheus, Thanos, and more.

Monitoring

Our monitoring stack at Reddit processes billions of data-points a minute, delivering insights at scale. Reddit has one of the larger deployments of Prometheus and Thanos in the world , and with this come unique challenges of scale for these systems. Fun problems include performance engineering on a distributed query system and product thinking around new features to remove the user pain from this stack.

Logging

We operate a platform for logging based on Vector and Clickhouse that processes millions of logs a day for Reddit. Thousands of people at Reddit use this system every day as they maintain, deploy, and debug their code. Join us as we invest in this system to help users better understand the data they have.

Distributed Tracing

Tracing at Reddit is based on OTEL, Clickhouse, and Grafana. Its light on features, but currently processing tens of millions of events, and we will be building new ways for users to leverage and understand this data in the coming year.

As a member of the Observability team, your work will span these domains, which are rich with challenging infrastructure and software engineering problems. Your work will directly impact hundreds of millions of users around the world. Join us and help build the future of Reddit!

In your day-to-day, you can expect to:

  • Work collaboratively with a team of software engineers to create and maintain the foundational platform for running Reddits infrastructure.
  • Deliver software to improve the availability, scalability, latency, and efficiency of observability components.
  • Contribute feedback to the technical and strategic direction of eventing at Reddit.
  • Automate critical aspects of the event driven development process
  • Share on-call responsibilities.
  • Contribute upstream changes to the open source projects we use

You have:

  • 7+ years of experience developing internet-scale software, preferably in the context of infrastructure.
  • Familiarity with distributed systems development, bonus if familiar with any of the specific tools (Prometheus, Thanos, Grafana, Vector, Clickhouse, Otel, Loki)
  • Experience developing on top of Kubernetes or similar distributed systems.
    • Kubernetes controller or operator development experience is a huge plus.
  • Strong troubleshooting capabilities surrounding both systems and software.
  • Experience engineering large systems, tracking work, and being a self-starter on projects.
  • Excellent communication skills to collaborate with a service-oriented team and company.

Benefits:

  • &
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Software Engineer, Observability Platform (Remote)
Staff Software Engineer, Observability Platform (Remote)

Reddit, Inc. • Northern (KY)

Hybrid
USD 217,000 - 304,000
Healthcare benefits
401k with employer match
Global benefit programs
+2
Staff Software Engineer, Observability at Scale (Remote)
Staff Software Engineer, Observability at Scale (Remote)

EngineersOfAI • Northern (KY)

Hybrid
USD 140,000 - 210,000
&
Senior Observability & Infra Engineer - Remote
Senior Observability & Infra Engineer - Remote

Reddit, Inc. • San Francisco (CA)

On-site
USD 217,000 - 304,000
Healthcare benefits
401k with employer match
Paid parental leave
+2
Staff Software Engineer , Observability New Remote - United States
Staff Software Engineer , Observability New Remote - United States

Reddit, Inc. • Northern (KY)

Hybrid
USD 217,000 - 304,000
Healthcare benefits
401k with employer match
Global benefit programs
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Brez Technology Inc. • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Private Medical, Dental and Vision Benefits
Retirement Savings plan with matching contributions
Workspace benefits for your home office
+4
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Alien Blue • Chicago (IL)

On-site
USD 190,000 - 268,000
Comprehensive Healthcare Benefits
401k Matching
Flexible Vacation
+2
Senior Software Engineer - Full Stack Internal Tooling (Build and Deployment Platform)
Senior Software Engineer - Full Stack Internal Tooling (Build and Deployment Platform)

EngineersOfAI • United States

Hybrid
USD 110,000 - 150,000
Software Engineer, Content Platform
Software Engineer, Content Platform

Reddit, Inc. • Northern (KY)

Hybrid
USD 164,000 - 230,000
Healthcare benefits
401k match
Family planning support
+3
Fullstack Software Engineer, Notifications Lifecycle
Fullstack Software Engineer, Notifications Lifecycle

Reddit • United States

Hybrid
USD 164,000 - 229,000
Comprehensive Healthcare Benefits
401k with Employer Match
Flexible Vacation & Paid Volunteer Time Off
+1
Staff Software Engineer, Observability
Staff Software Engineer, Observability

United States Digital Space LLC • Menlo Park (CA)

On-site
USD 180,000 - 250,000
Health insurance
Equity ownership
401(k) matching
+1