Observability Platform Engineer — Scale & Reliability

SpaceXAI

Palo Alto (CA)

On-site

USD 180,000 - 440,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Medical, Vision & Dental insurance
401(k) retirement plan
Disability insurance
Life insurance
Discounts and perks

Job summary

SpaceXAI is seeking a Member of Technical Staff for Observability to design and operate the core observability platform at scale. You will own metrics, logs, tracing, and alerting capabilities that empower engineers to monitor and optimize services across the fleet.

You will work on high-impact telemetry pipelines and APIs, partnering with infra and product teams to embed observability into internal platforms and ensure reliability and performance under massive load.

Qualifications

  • Proficiency in Go, Rust or Scala and building scalable systems.
  • Strong knowledge of distributed systems and telemetry architecture.
  • Experience building and operating infrastructure at scale.
  • Familiarity with Prometheus, Grafana, OpenTelemetry, VictoriaMetrics or ClickHouse.
  • Experience with Kafka, Redis or large-scale time-series databases.
  • Experience operating observability pipelines in Kubernetes or similar environments.

Responsibilities

  • Design and implement scalable observability infrastructure for metrics, logging, and tracing.
  • Build high-performance telemetry pipelines that handle massive ingestion volumes.
  • Develop APIs, query engines, and UIs for real-time service insights.
  • Define and enforce best practices for instrumentation, alerting, and reliability.
  • Partner with infra and product teams to integrate observability into platforms.
  • Own the reliability, scalability, and performance of the observability stack end-to-end.

Skills

Go
Rust
Scala
Distributed systems
Telemetry architecture
Kubernetes
Prometheus
Grafana
OpenTelemetry
VictoriaMetrics
ClickHouse

Tools

Prometheus
Grafana
OpenTelemetry
VictoriaMetrics
ClickHouse
Kafka
Redis
Kubernetes

Job description

SpaceXAI is seeking a Member of Technical Staff for Observability to design and operate the core observability platform at scale. You will own metrics, logs, tracing, and alerting capabilities that empower engineers to monitor and optimize services across the fleet.

You will work on high-impact telemetry pipelines and APIs, partnering with infra and product teams to embed observability into internal platforms and ensure reliability and performance under massive load.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Observability Engineer - Scale, AI-Driven Systems
Staff Observability Engineer - Scale, AI-Driven Systems

United States Digital Space LLC • Menlo Park (CA)

On-site
USD 180,000 - 250,000
Health insurance
Equity ownership
401(k) matching
+1
Software Engineer, Observability
Software Engineer, Observability

SpaceXAI • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical, Vision & Dental insurance
401(k) retirement plan
+3
Full-Stack Observability Engineer — Real-Time Telemetry
Full-Stack Observability Engineer — Real-Time Telemetry

SPACE EXPLORATION TECHNOLOGIES CORP • Hawthorne (CA)

On-site
USD 125,000 - 175,000
Stock options and long-term incentives
Discretionary bonuses
Comprehensive medical, vision, dental
+4
Software Engineer - Observability
Software Engineer - Observability

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Staff Software Engineer, Observability Platform
Staff Software Engineer, Observability Platform

Astronomer • United States

Hybrid
USD 275,000 - 377,000
Equity component
Comprehensive benefits package
Hybrid work model
Datacenter Software Engineer — Scale Ops & Automation
Datacenter Software Engineer — Scale Ops & Automation

Worky • Southaven (MS)

On-site
USD 120,000 - 160,000
SiteOps Platform Engineer for Data-Center AI
SiteOps Platform Engineer for Data-Center AI

Socket.dev • Memphis (TN)

On-site
USD 100,000 - 150,000
Remote Observability & Reliability Monitoring Engineer
Remote Observability & Reliability Monitoring Engineer

United States Digital Space LLC • United States

Remote
USD 100,000 - 150,000
Observability Engineer: Scale Telemetry & AI Reliability
Observability Engineer: Scale Telemetry & AI Reliability

Baseten • New York (NY)

On-site
USD 165,000 - 330,000
Competitive compensation including 1x/
Full medical/dental/vision insurance
Flexible PTO with Winter Break
+3
Remote Observability Platform Engineer — Scale Telemetry
Remote Observability Platform Engineer — Scale Telemetry

BairesDev • Peru (IL)

On-site
USD 120,000 - 170,000
100% remote work
USD or local currency pay
Home office setup
+3