Senior Observability Platform Engineer - GPU & AI

Nscale

New York (NY)

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Nscale is seeking a Staff Observability Platform Engineer to enhance our observability platform for GPU clusters and AI workloads. This role requires a strong technical leader capable of implementing scalable solutions and improving operational efficiency.

The ideal candidate has over 6 years of experience in relevant fields, strong software engineering skills in Go or Python, and familiarity with various observability technologies. Join us in driving innovation and operational excellence.

Qualifications

  • 6+ years of experience in SRE, platform engineering, infrastructure engineering, observability engineering, or related disciplines.
  • Strong experience building and operating observability platforms in cloud-native, distributed environments.
  • Deep hands-on experience with several observability technologies.

Responsibilities

  • Design and build observability platforms across metrics, logs, traces, alerting, and telemetry pipelines.
  • Lead the implementation of scalable observability solutions.
  • Partner across teams to ensure observability is embedded throughout.

Skills

SRE
Platform engineering
Infrastructure engineering
Observability engineering
Go
Python
Prometheus
Grafana

Tools

Terraform
Kubernetes
ClickHouse

Job description

Nscale is seeking a Staff Observability Platform Engineer to enhance our observability platform for GPU clusters and AI workloads. This role requires a strong technical leader capable of implementing scalable solutions and improving operational efficiency.

The ideal candidate has over 6 years of experience in relevant fields, strong software engineering skills in Go or Python, and familiarity with various observability technologies. Join us in driving innovation and operational excellence.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Observability Platform Engineer for AI/GPU Infra
Senior Observability Platform Engineer for AI/GPU Infra

Nscale • Seattle (WA)

On-site
USD 180,000 - 240,000
Senior Observability Platform Engineer – GPU AI Infra
Senior Observability Platform Engineer – GPU AI Infra

Nscale • Northern (KY)

Hybrid
USD 160,000 - 230,000
Medical insurance
Dental insurance
Vision insurance
+3
Senior Observability Platform Engineer – AI GPU Scale
Senior Observability Platform Engineer – AI GPU Scale

Nscale • United States

On-site
USD 160,000 - 230,000
Medical, dental, vision insurance
Flexible paid time off (PTO)
Parental leave
+1
Senior Observability Platform Engineer – GPU/AI Infra
Senior Observability Platform Engineer – GPU/AI Infra

Nscale • San Francisco (CA)

On-site
USD 130,000 - 160,000
Principal Observability Platform Engineer — GPU AI Scale
Principal Observability Platform Engineer — GPU AI Scale

Nscale • Seattle (WA)

On-site
USD 150,000 - 215,000
Medical, dental, vision benefits
Flexible paid time off
Parental leave
+1
Staff Observability Platform Engineer
Staff Observability Platform Engineer

Nscale • San Francisco (CA)

On-site
USD 130,000 - 160,000
Staff Observability Platform Engineer
Staff Observability Platform Engineer

Nscale • New York (NY)

On-site
USD 120,000 - 150,000
Staff Observability Platform Engineer
Staff Observability Platform Engineer

Nscale • Seattle (WA)

On-site
USD 180,000 - 240,000
Principal Observability Platform Engineer — GPU AI Scale
Principal Observability Platform Engineer — GPU AI Scale

Nscale • San Francisco (CA)

On-site
USD 150,000 - 215,000
Medical benefits
Flexible paid time off
Parental leave
Senior Observability Product Manager, GPU Fleet
Senior Observability Product Manager, GPU Fleet

Nscale • United States

On-site
USD 200,000 - 280,000
Competitive benefits package
Flexible paid time off
Parental leave