Senior Observability Platform Engineer for AI/GPU Infra

Nscale

Seattle (WA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Nscale is seeking a Staff Observability Platform Engineer to build and evolve our observability stack, giving deep visibility into GPU clusters, AI workloads, and infrastructure. You will treat observability as a product, define scalable solutions, mentor engineers, and collaborate with SRE, infra, platform, and AI/ML teams to improve reliability and developer experience.

This role emphasizes hands-on engineering, architectural input, and guiding platform improvements while balancing innovation

Qualifications

  • 6+ years of experience in SRE, platform or observability engineering.

Responsibilities

  • Design and evolve observability platforms across metrics, logs, traces, and telemetry.
  • Lead scalable observability solutions for Nscale’s GPU and AI infrastructure.
  • Collaborate with SRE, infrastructure, platform, and AI/ML teams to embed observability throughout the lifecycle.
  • Drive improvements in monitoring coverage, alert quality, service health visibility, and incident response.
  • Develop standards, frameworks, and reusable patterns for observability adoption across teams.
  • Identify reliability risks and operational blind spots, helping teams address them before they impact customers.
  • Contribute to architectural decisions around telemetry collection, storage, retention, and performance.
  • Lead initiatives that improve platform scalability, reliability, and efficiency.
  • Mentor engineers and provide guidance through reviews and knowledge sharing.
  • Participate in incident investigations and postmortems to translate learnings into durable improvements.
  • Evaluate new observability technologies balancing innovation with maintainability.

Skills

Observability
Go
Python
Kubernetes
Terraform
Cloud native
Leadership
Communication

Tools

Prometheus
Grafana
OpenTelemetry
ClickHouse
Elastic
Loki
Tempo
Thanos
VictoriaMetrics

Job description

Nscale is seeking a Staff Observability Platform Engineer to build and evolve our observability stack, giving deep visibility into GPU clusters, AI workloads, and infrastructure. You will treat observability as a product, define scalable solutions, mentor engineers, and collaborate with SRE, infra, platform, and AI/ML teams to improve reliability and developer experience.

This role emphasizes hands-on engineering, architectural input, and guiding platform improvements while balancing innovation

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Observability Platform Engineer - GPU & AI
Senior Observability Platform Engineer - GPU & AI

Nscale • New York (NY)

On-site
USD 120,000 - 150,000
Senior Observability Platform Engineer – GPU/AI Infra
Senior Observability Platform Engineer – GPU/AI Infra

Nscale • San Francisco (CA)

On-site
USD 130,000 - 160,000
Principal Observability Platform Engineer — GPU AI Scale
Principal Observability Platform Engineer — GPU AI Scale

Nscale • Seattle (WA)

On-site
USD 150,000 - 215,000
Medical, dental, vision benefits
Flexible paid time off
Parental leave
+1
Staff Observability Platform Engineer
Staff Observability Platform Engineer

Nscale • New York (NY)

On-site
USD 120,000 - 150,000
Staff Observability Platform Engineer
Staff Observability Platform Engineer

Nscale • Seattle (WA)

On-site
USD 180,000 - 240,000
Staff Observability Platform Engineer
Staff Observability Platform Engineer

Nscale • San Francisco (CA)

On-site
USD 130,000 - 160,000
Architect of AI/GPU Observability Platform
Architect of AI/GPU Observability Platform

Programming.com • San Francisco (CA)

On-site
USD 180,000 - 240,000
Principal Observability Platform Engineer — GPU AI Scale
Principal Observability Platform Engineer — GPU AI Scale

Nscale • San Francisco (CA)

On-site
USD 150,000 - 215,000
Medical benefits
Flexible paid time off
Parental leave
Senior Observability Product Manager, GPU Fleet
Senior Observability Product Manager, GPU Fleet

Nscale • United States

On-site
USD 200,000 - 280,000
Competitive benefits package
Flexible paid time off
Parental leave
Principal Observability Platform Engineer
Principal Observability Platform Engineer

Nscale • Seattle (WA)

On-site
USD 150,000 - 215,000
Medical, dental, vision benefits
Flexible paid time off
Parental leave
+1