Principal Observability Platform Engineer

Programming.com

San Francisco (CA)

On-site

USD 180,000 - 240,000

Part time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Programming.com is seeking a Principal/Staff Observability Platform Engineer to own the technical direction of our observability platform. You will shape architecture for metrics, logs, traces, and alerts that scale with GPU clusters, AI workloads, and infrastructure.

You will lead cross-team initiatives, simplify complex systems, and mentor the observability group while partnering with SRE, infra, and AI/ML teams to bake observability into builds and operations.

Qualifications

  • 8+ years in SRE, infra, platform engineering, or observability roles.
  • Led observability at scale and designed self-operable platforms.
  • Strong bias toward simple, scalable solutions.

Responsibilities

  • Own strategy and architecture for observability across metrics, logs, traces, and alerts at scale.
  • Drive long-term platform decisions: data models, ingestion patterns, retention, cardinality.
  • Identify gaps before incidents; design platforms that are easy to diagnose.
  • Partner with SRE, infra, and AI/ML teams to embed observability by design.
  • Define standards that engineers adopt through clear benefits.
  • Mentor the observability team and raise engineering velocity.
  • Lead postmortems to drive durable platform improvements.
  • Evaluate and retire tools that don’t improve signal quality or scalability.

Skills

SRE / Infra
Observability
Python/Go
Kubernetes
IaC (Terraform/Ansible)
Incident response

Tools

Prometheus
Thanos
VictoriaMetrics
Grafana
Loki
Tempo
OpenTelemetry
ClickHouse
Elastic
Kafka
Vector
Fluent Bit

Job description

Role: Principal Observability Platform Engineer

Duration: Long Term Contract Role

About The Role

As a Principal/Staff Observability Platform Engineer, you'll own the technical direction of observability platform: the systems that give us deep visibility into GPU clusters, AI workloads, and the infrastructure running them. You treat observability as a product and a discipline, not a tooling exercise. You'll set the architectural roadmap, raise the engineering bar across teams, and ensure our platform scales ahead of the business, not behind it.

You understand that complexity is a cost. Solutions that require constant babysitting don't scale, and neither does operational burden. The platforms you build should be simple to operate, easy to understand, and self-evidently correct when something goes wrong.

This isn't a "maintain and operate" role. It's a "define, build, and lead" role.

What You'll Do

  • Own the technical strategy and architecture for observability across metrics, logs, traces, and alerting at scale.
  • Drive platform decisions that have multi-year impact: tooling, data models, ingestion patterns, retention, cardinality management.
  • Identify systemic gaps before they become incidents; design platforms that make failure visible and fast to diagnose.
  • Partner with SRE, infrastructure, and AI/ML teams to embed observability natively into how builds and operates.
  • Define standards and patterns that other engineers adopt, not by mandate, but because they're clearly better.
  • Mentor and technically grow the observability team; raise the ceiling on what the team can build and own.
  • Lead incident postmortems and use them to drive durable platform improvements.
  • Evaluate and introduce tooling that meaningfully improves signal quality, operational efficiency, or scalability, and retire what doesn't.

About You

  • 8+ years in SRE, infrastructure engineering, platform engineering, or observability-focused roles.
  • You've operated observability infrastructure at serious scale. You know what breaks at 10x and you design for it.
  • You have a strong bias toward simplicity. You've seen over-engineered observability stacks collapse under their own weight and you build accordingly.
  • Deep hands-on experience with a significant subset of: Prometheus, Thanos, VictoriaMetrics, Grafana, Loki, Tempo, OpenTelemetry, ClickHouse, Elastic.
  • Strong engineering fundamentals, proficient in Python, Go, or similar; comfortable owning complex systems end to end.
  • Experience with Kubernetes at scale; familiarity with GPU infrastructure or HPC environments (Slurm) is a strong plus.
  • You can architect systems, write the code, review others' work, and explain the tradeoffs clearly, all in the same week.
  • Infrastructure-as-Code is default, not optional (Terraform, Ansible, or equivalent).
  • You influence without authority. Teams want your opinion because it makes their work better.

Preferred

  • Experience with high-volume streaming pipelines for observability data (Kafka, Vector, Fluent Bit, etc.).
  • Background in AI/ML infrastructure observability: GPU utilisation, training job visibility, inference latency.
  • Prior experience defining observability strategy at an organisation level.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Infrastructure Engineer, Observability
Infrastructure Engineer, Observability

Jobtailor • California (MO)

On-site
USD 140,000 - 190,000
Principal Observability Platform Engineer
Principal Observability Platform Engineer

Nscale • San Francisco (CA)

On-site
USD 150,000 - 215,000
Medical benefits
Flexible paid time off
Parental leave
Principal Observability Platform Engineer
Principal Observability Platform Engineer

Nscale • Seattle (WA)

On-site
USD 150,000 - 215,000
Medical, dental, vision benefits
Flexible paid time off
Parental leave
+1
Staff Observability Platform Engineer
Staff Observability Platform Engineer

Nscale • Seattle (WA)

On-site
USD 180,000 - 240,000
Staff Observability Platform Engineer
Staff Observability Platform Engineer

Nscale • San Francisco (CA)

On-site
USD 130,000 - 160,000
Staff Observability Platform Engineer
Staff Observability Platform Engineer

Nscale • New York (NY)

On-site
USD 120,000 - 150,000
Principal Full Stack Software Engineer (Observability)
Principal Full Stack Software Engineer (Observability)

Salesforce • San Francisco (CA)

On-site
USD 180,000 - 260,000
Medical Care
Life Insurance
Retirement Savings
+2
Observability Engineer: OpenTelemetry & Open-Source
Observability Engineer: OpenTelemetry & Open-Source

Softility Tech Pvt. Ltd. • United States

Hybrid
Staff Software Engineer, Observability
Staff Software Engineer, Observability

United States Digital Space LLC • Menlo Park (CA)

On-site
USD 180,000 - 250,000
Health insurance
Equity ownership
401(k) matching
+1
Solution Architect / Team Lead - Observability
Solution Architect / Team Lead - Observability

VOLTO Consulting • Irvine (CA)

On-site
USD 120,000 - 160,000