Senior Site Reliability Engineer

Programming.com

United States

Hybrid

USD 180,000 - 240,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Programming.com seeks a deeply hands-on Staff Observability Platform Engineer to design, build, and operate large-scale metrics, logging, and telemetry infrastructure in a hybrid Kubernetes environment. You will own the observability backend across multiple backends and drive decisions on cardin ality, retention, and cost, collaborating with platform, infra, security, and application teams.

The role demands engineering telemetry at scale, performance-minded storage and query design, and a track

Qualifications

  • Operate a metrics/logs backend at production scale and quantify scale.
  • Experience with Prometheus-based architectures and remote write.
  • Ability to explain architectural decisions around cardinality, retention, and cost.
  • Hands-on Kubernetes experience at scale and strong code-reading skills.

Responsibilities

  • Design, build, operate, and scale production metrics, logging, tracing, and telemetry platforms.
  • Own observability backend architecture supporting large distributed Kubernetes environments.
  • Operate and scale backends like Mimir, Thanos, VictoriaMetrics, Cortex, Loki, Elasticsearch.
  • Design Prometheus-based architectures including remote write, ingestion pipelines, high availability, retention, and global querying.
  • Engineer telemetry systems for growth in data volumes.
  • Identify and remediate high-cardinality metrics.
  • Make trade-offs across cardinality, retention, storage, ingestion, query performance, reliability, and cost.
  • Build and operate OpenTelemetry Collector pipelines.
  • Establish observability standards across engineering teams.
  • Troubleshoot production issues spanning metrics, logs, traces, Kubernetes, networking, and storage.
  • Write and review production-quality code with security and reliability.
  • Partner with platform, infrastructure, security, and application engineering teams.

Skills

Hands-on SRE experience
Observability design
Distributed systems problem-solving
Security and reliability focus

Tools

Go
Python
Kubernetes
OpenTelemetry
Prometheus
Loki
Thanos
Mimir
VictoriaMetrics
Cortex
Elasticsearch

Job description

Staff Observability Platform Engineer (SRE)

Locations: Seattle, WA (Hybrid), Houston, TX (Hybrid), New York, NY (Hybrid)

Role Overview

We are seeking a deeply hands-on Staff Observability Platform Engineer with experience designing, building, and operating large-scale metrics, logging, and telemetry infrastructure. The environment includes distributed Kubernetes and AI/GPU infrastructure where observability must remain reliable, performant, and cost-effective as telemetry volume grows. The ideal candidate has personally owned or operated the underlying observability backend at meaningful production scale and can quantify that scale. They should be able to explain real architectural decisions involving cardinality, ingestion, retention, storage, reliability, query performance, and cost.

Key Responsibilities
  • Design, build, operate, and scale production metrics, logging, tracing, and telemetry platforms.
  • Own observability backend architecture supporting large distributed Kubernetes environments.
  • Operate and scale platforms such as Mimir, Thanos, VictoriaMetrics, Cortex, Loki, Elasticsearch, or comparable distributed telemetry backends.
  • Design Prometheus-based architectures including remote write, ingestion pipelines, high availability, long-term retention, and global querying.
  • Engineer telemetry systems for growth in active series, samples per second, logs/events per second, and overall data volume.
  • Identify and remediate high-cardinality metrics and inefficient telemetry models.
  • Make informed trade-offs across cardinality, retention, storage tiers, ingestion volume, query performance, reliability, and infrastructure cost.
  • Build and operate OpenTelemetry Collector pipelines, including receivers, processors, exporters, routing, filtering, and sampling.
  • Establish observability standards and telemetry practices across engineering teams.
  • Troubleshoot complex production issues spanning metrics, logs, traces, Kubernetes, networking, storage, and distributed systems.
  • Write and review production-quality code with strong attention to security, scalability, reliability, and failure modes.
  • Partner with platform, infrastructure, security, and application engineering teams.
Required Experience

Large-Scale Observability Backend Ownership

Candidates should have personally operated a metrics or logs backend at meaningful production scale. Strong examples include Mimir, Thanos, VictoriaMetrics, Cortex, large-scale Loki, large-scale Elasticsearch, or comparable distributed telemetry platforms.

Experience limited primarily to dashboards, alerts, PromQL queries, or consuming an observability platform operated by another team is not sufficient for this role.

Demonstrated Production Scale

Candidates should be able to quantify the systems they have personally operated using one or more meaningful measures, such as:

  • Active time series or samples ingested per second
  • Metrics, events, or logs ingestion rate
  • Number of Kubernetes clusters, nodes, workloads, services, or tenants
  • Storage footprint and retention period. The emphasis is on the underlying size and complexity of the platform, not percentage-based improvement claims without production-scale context.
Cardinality, Retention, and Storage Trade-offs

Candidates should have personally made and be able to explain production engineering decisions involving one or more of the following:

  • Metrics cardinality and label design
  • Retention and storage architecture
  • Telemetry storage cost and capacity
  • Ingestion volume, aggregation, filtering, or sampling
  • Storage tiering or query-performance trade-offs Strong candidates can clearly explain the problem, why it mattered, the decision they made, the trade-offs involved, and the resulting operational impact
Kubernetes and Distributed Systems
  • Strong hands-on experience operating production Kubernetes environments.
  • Understanding of multi-cluster architectures, service discovery, networking, storage, autoscaling, reliability, capacity management, and failure domains.
  • Ability to approach observability as a distributed-systems problem rather than simply a collection of monitoring tools.
  • Strong production code-reading and code-review ability.
  • Ability to identify security, reliability, scalability, concurrency, and resource-exhaustion risks.
  • Understanding of retries, timeouts, backpressure, buffering, circuit breaking, graceful degradation, and downstream failure handling.
  • Strong Go and/or Python experience preferred.
Highly Valuable Experience

The following experience is particularly relevant but not required:

  • Bare-metal infrastructure and large GPU fleets
  • NVIDIA DCGM and GPU telemetry
  • Slurm or other HPC scheduling environments
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Observability Engineer
Senior Observability Engineer

Veriipro • Boston (MA)

On-site
USD 140,000 - 190,000
Principal Observability Platform Engineer
Principal Observability Platform Engineer

Programming.com • San Francisco (CA)

On-site
USD 180,000 - 240,000
SRE/Observability Engineer
SRE/Observability Engineer

BlueSky Resource Solutions • United States

Remote
USD 100,000 - 130,000
Staff Observability Platform Engineer: Scale Metrics & Telemetry
Staff Observability Platform Engineer: Scale Metrics & Telemetry

Programming.com • United States

Hybrid
USD 180,000 - 240,000
Observability Backend Engineer - Distributed Systems (Mandarin required)
Observability Backend Engineer - Distributed Systems (Mandarin required)

Applied Intelligence Consulting (Singapore) • San Francisco (CA)

On-site
USD 120,000 - 170,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobtailor • Arlington (VA)

On-site
USD 140,000 - 200,000
Back fill Engineer
Back fill Engineer

JPC TECHNO INC • Phoenix (AZ)

On-site
USD 120,000 - 180,000
Observability Operations Engineer
Observability Operations Engineer

IntraEdge • Phoenix (AZ)

On-site
USD 140,000 - 190,000
Staff Observability Platform Engineer
Staff Observability Platform Engineer

Nscale • Seattle (WA)

On-site
USD 180,000 - 240,000
Observability Architect
Observability Architect

TechDigital Group • Atlanta (GA)

On-site
USD 120,000 - 150,000