Staff Observability Platform Engineer (SRE)
Locations: Seattle, WA (Hybrid), Houston, TX (Hybrid), New York, NY (Hybrid)
Role Overview
We are seeking a deeply hands-on Staff Observability Platform Engineer with experience designing, building, and operating large-scale metrics, logging, and telemetry infrastructure. The environment includes distributed Kubernetes and AI/GPU infrastructure where observability must remain reliable, performant, and cost-effective as telemetry volume grows. The ideal candidate has personally owned or operated the underlying observability backend at meaningful production scale and can quantify that scale. They should be able to explain real architectural decisions involving cardinality, ingestion, retention, storage, reliability, query performance, and cost.
Key Responsibilities
- Design, build, operate, and scale production metrics, logging, tracing, and telemetry platforms.
- Own observability backend architecture supporting large distributed Kubernetes environments.
- Operate and scale platforms such as Mimir, Thanos, VictoriaMetrics, Cortex, Loki, Elasticsearch, or comparable distributed telemetry backends.
- Design Prometheus-based architectures including remote write, ingestion pipelines, high availability, long-term retention, and global querying.
- Engineer telemetry systems for growth in active series, samples per second, logs/events per second, and overall data volume.
- Identify and remediate high-cardinality metrics and inefficient telemetry models.
- Make informed trade-offs across cardinality, retention, storage tiers, ingestion volume, query performance, reliability, and infrastructure cost.
- Build and operate OpenTelemetry Collector pipelines, including receivers, processors, exporters, routing, filtering, and sampling.
- Establish observability standards and telemetry practices across engineering teams.
- Troubleshoot complex production issues spanning metrics, logs, traces, Kubernetes, networking, storage, and distributed systems.
- Write and review production-quality code with strong attention to security, scalability, reliability, and failure modes.
- Partner with platform, infrastructure, security, and application engineering teams.
Required Experience
Large-Scale Observability Backend Ownership
Candidates should have personally operated a metrics or logs backend at meaningful production scale. Strong examples include Mimir, Thanos, VictoriaMetrics, Cortex, large-scale Loki, large-scale Elasticsearch, or comparable distributed telemetry platforms.
Experience limited primarily to dashboards, alerts, PromQL queries, or consuming an observability platform operated by another team is not sufficient for this role.
Demonstrated Production Scale
Candidates should be able to quantify the systems they have personally operated using one or more meaningful measures, such as:
- Active time series or samples ingested per second
- Metrics, events, or logs ingestion rate
- Number of Kubernetes clusters, nodes, workloads, services, or tenants
- Storage footprint and retention period. The emphasis is on the underlying size and complexity of the platform, not percentage-based improvement claims without production-scale context.
Cardinality, Retention, and Storage Trade-offs
Candidates should have personally made and be able to explain production engineering decisions involving one or more of the following:
- Metrics cardinality and label design
- Retention and storage architecture
- Telemetry storage cost and capacity
- Ingestion volume, aggregation, filtering, or sampling
- Storage tiering or query-performance trade-offs Strong candidates can clearly explain the problem, why it mattered, the decision they made, the trade-offs involved, and the resulting operational impact
Kubernetes and Distributed Systems
- Strong hands-on experience operating production Kubernetes environments.
- Understanding of multi-cluster architectures, service discovery, networking, storage, autoscaling, reliability, capacity management, and failure domains.
- Ability to approach observability as a distributed-systems problem rather than simply a collection of monitoring tools.
- Strong production code-reading and code-review ability.
- Ability to identify security, reliability, scalability, concurrency, and resource-exhaustion risks.
- Understanding of retries, timeouts, backpressure, buffering, circuit breaking, graceful degradation, and downstream failure handling.
- Strong Go and/or Python experience preferred.
Highly Valuable Experience
The following experience is particularly relevant but not required:
- Bare-metal infrastructure and large GPU fleets
- NVIDIA DCGM and GPU telemetry
- Slurm or other HPC scheduling environments