Senior Observability Platform Engineer, GPU AI Cloud

nscaleoperationsukltd

United States

On-site

USD 160,000 - 230,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, vision benefits
Flexible paid time off
Parental leave
Retirement plan participation

Job summary

Nscale is seeking a Senior Observability Platform Engineer to design, build, and scale our observability platform for GPU-driven AI workloads. You will deliver reliable visibility into GPU clusters and AI infrastructure, balancing usability, scalability, and operational efficiency.

As an hands-on engineer, you will influence platform direction, implement critical systems, and collaborate closely with SRE, infra, and AI/ML teams to ensure observability is embedded in services we run.

Qualifications

  • 5+ years in SRE, infrastructure, platform engineering, or observability roles.
  • Experience operating and scaling observability systems in production.
  • Strong monitoring concepts: metrics, logs, traces, alerting, and SLOs.
  • Hands-on with Prometheus, Thanos, VictoriaMetrics, Grafana, Loki, Tempo, OpenTelemetry, ClickHouse, Elastic.
  • Solid Python/Go programming skills for production systems.
  • Experience with Kubernetes-based infrastructure.
  • Familiarity with Terraform, Ansible, or similar IaC tools.
  • Pragmatic mindset focused on simplicity, reliability, and maintainability.
  • Strong collaboration across teams.

Responsibilities

  • Design, build, and operate scalable observability systems across metrics, logs, traces, and alerting.
  • Contribute to tooling and data pipelines, storage, and retention decisions.
  • Improve signal quality by reducing noise and refining alerting practices.
  • Identify observability gaps and drive reliability improvements.
  • Collaborate with SRE, infrastructure, and AI/ML teams to embed observability.
  • Develop reusable patterns and best practices for consistency across teams.
  • Participate in incident response and postmortems with actionable follow-ups.
  • Evaluate tools that improve developer experience and scalability.
  • Mentor engineers via code reviews and knowledge sharing.

Skills

SRE expertise
Production observability
Python/Go programming

Tools

Prometheus
Thanos
VictoriaMetrics
Grafana
Loki
Tempo
OpenTelemetry
ClickHouse
Elastic
Kubernetes
Terraform
Ansible

Job description

Nscale is seeking a Senior Observability Platform Engineer to design, build, and scale our observability platform for GPU-driven AI workloads. You will deliver reliable visibility into GPU clusters and AI infrastructure, balancing usability, scalability, and operational efficiency.

As an hands-on engineer, you will influence platform direction, implement critical systems, and collaborate closely with SRE, infra, and AI/ML teams to ensure observability is embedded in services we run.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Observability Platform Engineer – GPU AI Infra
Senior Observability Platform Engineer – GPU AI Infra

Nscale • Northern (KY)

Hybrid
USD 160,000 - 230,000
Medical insurance
Dental insurance
Vision insurance
+3
Senior Observability Platform Engineer – AI GPU Scale
Senior Observability Platform Engineer – AI GPU Scale

Nscale • United States

On-site
USD 160,000 - 230,000
Medical, dental, vision insurance
Flexible paid time off (PTO)
Parental leave
+1
Senior Observability Platform Engineer for AI/GPU Infra
Senior Observability Platform Engineer for AI/GPU Infra

Nscale • Seattle (WA)

On-site
USD 180,000 - 240,000
Senior AI Observability Platform Engineer
Senior AI Observability Platform Engineer

Socket.dev • United States

On-site
USD 160,000 - 230,000
Senior Observability Platform Engineer - GPU & AI
Senior Observability Platform Engineer - GPU & AI

Nscale • New York (NY)

On-site
USD 120,000 - 150,000
Senior Observability Platform Engineer – GPU/AI Infra
Senior Observability Platform Engineer – GPU/AI Infra

Nscale • San Francisco (CA)

On-site
USD 130,000 - 160,000
Principal Observability Platform Engineer — GPU AI Scale
Principal Observability Platform Engineer — GPU AI Scale

Nscale • Seattle (WA)

On-site
USD 150,000 - 215,000
Medical, dental, vision benefits
Flexible paid time off
Parental leave
+1
Senior Observability Product Manager, GPU Fleet
Senior Observability Product Manager, GPU Fleet

Nscale • United States

On-site
USD 200,000 - 280,000
Competitive benefits package
Flexible paid time off
Parental leave
Principal Observability Platform Engineer — GPU AI Scale
Principal Observability Platform Engineer — GPU AI Scale

Nscale • San Francisco (CA)

On-site
USD 150,000 - 215,000
Medical benefits
Flexible paid time off
Parental leave
Senior Observability Platform Engineer
Senior Observability Platform Engineer

nscaleoperationsukltd • United States

On-site
USD 160,000 - 230,000
Medical, dental, vision benefits
Flexible paid time off
Parental leave
+1