Principal Observability Platform Engineer — GPU AI Scale

Nscale

Seattle (WA)

On-site

USD 150,000 - 215,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, vision benefits
Flexible paid time off
Parental leave
Retirement plan participation

Job summary

Nscale in Seattle is seeking a Principal Observability Platform Engineer to own the technical direction of its observability platform. You will lead the design and architecture for the systems ensuring deep visibility into GPU clusters and AI workloads.

The ideal candidate will possess over 8 years in relevant roles with deep hands-on experience in tools like Prometheus and Grafana, emphasizing simplicity and scalability. This role offers a competitive salary, potential bonuses, and a comprehensive benefits package.

Qualifications

  • 8+ years in SRE, infrastructure engineering, platform engineering, or observability-focused roles.
  • Operated observability infrastructure at serious scale.
  • Deep hands-on experience with Prometheus, Thanos, and Grafana.

Responsibilities

  • Own the technical strategy and architecture for observability across metrics, logs, traces, and alerting.
  • Drive platform decisions that have multi-year impact.
  • Mentor and technically grow the observability team.

Skills

SRE
Infrastructure engineering
Platform engineering
Observability-focused roles
Python
Go
Kubernetes
Terraform

Tools

Prometheus
Grafana
OpenTelemetry
ClickHouse

Job description

Nscale in Seattle is seeking a Principal Observability Platform Engineer to own the technical direction of its observability platform. You will lead the design and architecture for the systems ensuring deep visibility into GPU clusters and AI workloads.

The ideal candidate will possess over 8 years in relevant roles with deep hands-on experience in tools like Prometheus and Grafana, emphasizing simplicity and scalability. This role offers a competitive salary, potential bonuses, and a comprehensive benefits package.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Observability Platform Engineer — GPU AI Scale
Principal Observability Platform Engineer — GPU AI Scale

Nscale • San Francisco (CA)

On-site
USD 150,000 - 215,000
Medical benefits
Flexible paid time off
Parental leave
Senior Observability Platform Engineer for AI/GPU Infra
Senior Observability Platform Engineer for AI/GPU Infra

Nscale • Seattle (WA)

On-site
USD 180,000 - 240,000
Senior Observability Platform Engineer - GPU & AI
Senior Observability Platform Engineer - GPU & AI

Nscale • New York (NY)

On-site
USD 120,000 - 150,000
Senior Observability Platform Engineer – GPU/AI Infra
Senior Observability Platform Engineer – GPU/AI Infra

Nscale • San Francisco (CA)

On-site
USD 130,000 - 160,000
Architect of AI/GPU Observability Platform
Architect of AI/GPU Observability Platform

Programming.com • San Francisco (CA)

On-site
USD 180,000 - 240,000
Principal Observability Platform Engineer
Principal Observability Platform Engineer

Nscale • Seattle (WA)

On-site
USD 150,000 - 215,000
Medical, dental, vision benefits
Flexible paid time off
Parental leave
+1
Principal Observability Platform Engineer
Principal Observability Platform Engineer

Nscale • San Francisco (CA)

On-site
USD 150,000 - 215,000
Medical benefits
Flexible paid time off
Parental leave
Staff Observability Platform Engineer
Staff Observability Platform Engineer

Nscale • New York (NY)

On-site
USD 120,000 - 150,000
Staff Observability Platform Engineer
Staff Observability Platform Engineer

Nscale • San Francisco (CA)

On-site
USD 130,000 - 160,000
Staff Observability Platform Engineer
Staff Observability Platform Engineer

Nscale • Seattle (WA)

On-site
USD 180,000 - 240,000