Senior AI Observability Platform Engineer

Socket.dev

United States

On-site

USD 160,000 - 230,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Nscale is seeking a Senior Observability Platform Engineer to design, build, and scale our observability stack for GPU clusters and AI workloads. You’ll treat observability as a product, delivering reliable metrics, logs, and traces while reducing noise and improving debugging.

You’ll partner with SRE, infrastructure, and AI/ML teams to embed observability into services, develop reusable patterns, and influence platform direction through hands-on engineering and incident postmortems.

Qualifications

  • 5+ years in SRE or observability roles.
  • Experience operating production observability systems.
  • Proficient in Python or Go.
  • Experience with Kubernetes-based infrastructure.

Responsibilities

  • Design, build, and operate scalable observability systems across metrics, logs, traces, and alerting.
  • Contribute to architectural decisions around tooling, data pipelines, storage, and retention.
  • Improve signal quality by reducing noise and refining alerting practices.
  • Identify observability gaps and address reliability impacts.
  • Collaborate with SRE, infrastructure, and AI/ML teams to embed observability into services.

Skills

SRE experience
Observability
Production systems
Python or Go

Tools

Kubernetes
Prometheus
Thanos
VictoriaMetrics
Grafana
Loki
OpenTelemetry
ClickHouse
Elastic
Terraform
Ansible

Job description

Nscale is seeking a Senior Observability Platform Engineer to design, build, and scale our observability stack for GPU clusters and AI workloads. You’ll treat observability as a product, delivering reliable metrics, logs, and traces while reducing noise and improving debugging.

You’ll partner with SRE, infrastructure, and AI/ML teams to embed observability into services, develop reusable patterns, and influence platform direction through hands-on engineering and incident postmortems.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Observability Platform Engineer – AI GPU Scale
Senior Observability Platform Engineer – AI GPU Scale

Nscale • United States

On-site
USD 160,000 - 230,000
Medical, dental, vision insurance
Flexible paid time off (PTO)
Parental leave
+1
Senior Observability Platform Engineer – GPU AI Infra
Senior Observability Platform Engineer – GPU AI Infra

Nscale • Northern (KY)

Hybrid
USD 160,000 - 230,000
Medical insurance
Dental insurance
Vision insurance
+3
Senior Observability Platform Engineer for AI/GPU Infra
Senior Observability Platform Engineer for AI/GPU Infra

Nscale • Seattle (WA)

On-site
USD 180,000 - 240,000
Senior Observability Platform Engineer - GPU & AI
Senior Observability Platform Engineer - GPU & AI

Nscale • New York (NY)

On-site
USD 120,000 - 150,000
Senior Observability Platform Engineer – GPU/AI Infra
Senior Observability Platform Engineer – GPU/AI Infra

Nscale • San Francisco (CA)

On-site
USD 130,000 - 160,000
Principal Observability Platform Engineer — GPU AI Scale
Principal Observability Platform Engineer — GPU AI Scale

Nscale • Seattle (WA)

On-site
USD 150,000 - 215,000
Medical, dental, vision benefits
Flexible paid time off
Parental leave
+1
Principal Observability Platform Engineer — GPU AI Scale
Principal Observability Platform Engineer — GPU AI Scale

Nscale • San Francisco (CA)

On-site
USD 150,000 - 215,000
Medical benefits
Flexible paid time off
Parental leave
Staff Observability Platform Engineer
Staff Observability Platform Engineer

Nscale • San Francisco (CA)

On-site
USD 130,000 - 160,000
Architect of AI/GPU Observability Platform
Architect of AI/GPU Observability Platform

Programming.com • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff Observability Platform Engineer
Staff Observability Platform Engineer

Nscale • Seattle (WA)

On-site
USD 180,000 - 240,000