Senior Observability Platform Engineer – AI GPU Scale

Nscale

United States

On-site

USD 160,000 - 230,000

Full time

4 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, vision insurance
Flexible paid time off (PTO)
Parental leave
Retirement plan

Job summary

Nscale is seeking a Senior Observability Platform Engineer to design, build, and scale our observability stack for GPU clusters and AI workloads. You’ll balance reliability, performance, and usability while shaping platform direction and collaborating with SRE, infrastructure, and AI/ML teams.

You’ll implement scalable data pipelines, improve signal quality, and mentor engineers. This hands-on role influences tooling, architecture, and postmortem improvements, delivering measurable reliability

Qualifications

  • 5+ years in SRE, infra, platform, or observability roles.
  • Experience operating and scaling observability in production.
  • Strong monitoring concepts: metrics, logs, traces, alerting, and SLOs.
  • Hands-on with Prometheus, Grafana, OpenTelemetry, etc.
  • Programming in Python or Go for production systems.
  • Kubernetes-based infrastructure experience.
  • Infrastructure-as-Code with Terraform/Ansible.
  • Pragmatic mindset focused on reliability and maintainability.
  • Strong collaboration across teams.

Responsibilities

  • Design, build, and operate scalable observability systems across metrics, logs, traces, and alerting
  • Contribute to architectural decisions around tooling, data pipelines, storage, and retention strategies
  • Improve signal quality by reducing noise, managing cardinality, and refining alerting practices
  • Help identify and address observability gaps before they impact reliability
  • Partner with SRE, infrastructure, and AI/ML teams to integrate observability into services and platforms
  • Develop reusable patterns, libraries, and best practices across teams
  • Participate in incident response and postmortems, driving actionable improvements
  • Evaluate and adopt tools that improve developer experience, scalability, and operational efficiency
  • Support and mentor engineers within the team through code reviews and knowledge sharing

Skills

SRE / Infra Eng
Observability
Programming: Python/Go
Kubernetes
IaC: Terraform/Ansible
Cross-team collaboration

Tools

Prometheus
Thanos
VictoriaMetrics
Grafana
Loki
Tempo
OpenTelemetry
ClickHouse
Elasticsearch/Elastic

Job description

Nscale is seeking a Senior Observability Platform Engineer to design, build, and scale our observability stack for GPU clusters and AI workloads. You’ll balance reliability, performance, and usability while shaping platform direction and collaborating with SRE, infrastructure, and AI/ML teams.

You’ll implement scalable data pipelines, improve signal quality, and mentor engineers. This hands-on role influences tooling, architecture, and postmortem improvements, delivering measurable reliability

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Observability Platform Engineer – GPU AI Infra
Senior Observability Platform Engineer – GPU AI Infra

Nscale • Northern (KY)

Hybrid
USD 160,000 - 230,000
Medical insurance
Dental insurance
Vision insurance
+3
Senior Observability Platform Engineer for AI/GPU Infra
Senior Observability Platform Engineer for AI/GPU Infra

Nscale • Seattle (WA)

On-site
USD 180,000 - 240,000
Senior Observability Platform Engineer - GPU & AI
Senior Observability Platform Engineer - GPU & AI

Nscale • New York (NY)

On-site
USD 120,000 - 150,000
Senior Observability Platform Engineer – GPU/AI Infra
Senior Observability Platform Engineer – GPU/AI Infra

Nscale • San Francisco (CA)

On-site
USD 130,000 - 160,000
Principal Observability Platform Engineer — GPU AI Scale
Principal Observability Platform Engineer — GPU AI Scale

Nscale • Seattle (WA)

On-site
USD 150,000 - 215,000
Medical, dental, vision benefits
Flexible paid time off
Parental leave
+1
Principal Observability Platform Engineer — GPU AI Scale
Principal Observability Platform Engineer — GPU AI Scale

Nscale • San Francisco (CA)

On-site
USD 150,000 - 215,000
Medical benefits
Flexible paid time off
Parental leave
Senior Observability Platform Engineer
Senior Observability Platform Engineer

Nscale • United States

On-site
USD 160,000 - 230,000
Medical, dental, vision insurance
Flexible paid time off (PTO)
Parental leave
+1
Architect of AI/GPU Observability Platform
Architect of AI/GPU Observability Platform

Programming.com • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff Observability Platform Engineer
Staff Observability Platform Engineer

Nscale • Seattle (WA)

On-site
USD 180,000 - 240,000
Staff Observability Platform Engineer
Staff Observability Platform Engineer

Nscale • San Francisco (CA)

On-site
USD 130,000 - 160,000