Principal Observability Platform Engineer — GPU AI Scale

Nscale

San Francisco (CA)

On-site

USD 150,000 - 215,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical benefits
Flexible paid time off
Parental leave

Job summary

Nscale is looking for a Principal Observability Platform Engineer in San Francisco, California. You will own the technical strategy for Nscale's observability platform, driving decisions that impact infrastructure and operations at scale.

The ideal candidate will have over 8 years in SRE or platform engineering with hands-on experience in tools like Prometheus and Grafana. This position offers a salary range of $150,000 to $215,000 along with a competitive benefits package.

Qualifications

  • 8+ years in SRE, infrastructure engineering, or observability-related roles.
  • Experience with Kubernetes at scale; familiarity with GPU infrastructure is a strong plus.
  • Proficient in Python, Go, or similar languages.

Responsibilities

  • Own the technical strategy and architecture for observability at scale.
  • Drive platform decisions impacting multi-year outcomes.
  • Mentor and grow the observability team.

Skills

Operating observability infrastructure at scale
Strong bias toward simplicity
Deep hands-on experience with Prometheus and Grafana
Strong engineering fundamentals
Infrastructure-as-Code (Terraform, Ansible)
Ability to influence without authority

Tools

Prometheus
Grafana
Kubernetes
Terraform

Job description

Nscale is looking for a Principal Observability Platform Engineer in San Francisco, California. You will own the technical strategy for Nscale's observability platform, driving decisions that impact infrastructure and operations at scale.

The ideal candidate will have over 8 years in SRE or platform engineering with hands-on experience in tools like Prometheus and Grafana. This position offers a salary range of $150,000 to $215,000 along with a competitive benefits package.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Observability Platform Engineer — GPU AI Scale
Principal Observability Platform Engineer — GPU AI Scale

Nscale • Seattle (WA)

On-site
USD 150,000 - 215,000
Medical, dental, vision benefits
Flexible paid time off
Parental leave
+1
Senior Observability Platform Engineer – GPU/AI Infra
Senior Observability Platform Engineer – GPU/AI Infra

Nscale • San Francisco (CA)

On-site
USD 130,000 - 160,000
Senior Observability Platform Engineer for AI/GPU Infra
Senior Observability Platform Engineer for AI/GPU Infra

Nscale • Seattle (WA)

On-site
USD 180,000 - 240,000
Principal Observability Platform Engineer
Principal Observability Platform Engineer

Nscale • San Francisco (CA)

On-site
USD 150,000 - 215,000
Medical benefits
Flexible paid time off
Parental leave
Principal Observability Platform Engineer
Principal Observability Platform Engineer

Nscale • Seattle (WA)

On-site
USD 150,000 - 215,000
Medical, dental, vision benefits
Flexible paid time off
Parental leave
+1
Senior Observability Platform Engineer - GPU & AI
Senior Observability Platform Engineer - GPU & AI

Nscale • New York (NY)

On-site
USD 120,000 - 150,000
Staff Observability Platform Engineer
Staff Observability Platform Engineer

Nscale • New York (NY)

On-site
USD 120,000 - 150,000
Staff Observability Platform Engineer
Staff Observability Platform Engineer

Nscale • San Francisco (CA)

On-site
USD 130,000 - 160,000
Staff Observability Platform Engineer
Staff Observability Platform Engineer

Nscale • Seattle (WA)

On-site
USD 180,000 - 240,000
Architect of AI/GPU Observability Platform
Architect of AI/GPU Observability Platform

Programming.com • San Francisco (CA)

On-site
USD 180,000 - 240,000