Staff Observability Platform Engineer

Nscale

Seattle (WA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Nscale is seeking a Staff Observability Platform Engineer to build and evolve our observability stack, giving deep visibility into GPU clusters, AI workloads, and infrastructure. You will treat observability as a product, define scalable solutions, mentor engineers, and collaborate with SRE, infra, platform, and AI/ML teams to improve reliability and developer experience.

This role emphasizes hands-on engineering, architectural input, and guiding platform improvements while balancing innovation

Qualifications

  • 6+ years of experience in SRE, platform or observability engineering.

Responsibilities

  • Design and evolve observability platforms across metrics, logs, traces, and telemetry.
  • Lead scalable observability solutions for Nscale’s GPU and AI infrastructure.
  • Collaborate with SRE, infrastructure, platform, and AI/ML teams to embed observability throughout the lifecycle.
  • Drive improvements in monitoring coverage, alert quality, service health visibility, and incident response.
  • Develop standards, frameworks, and reusable patterns for observability adoption across teams.
  • Identify reliability risks and operational blind spots, helping teams address them before they impact customers.
  • Contribute to architectural decisions around telemetry collection, storage, retention, and performance.
  • Lead initiatives that improve platform scalability, reliability, and efficiency.
  • Mentor engineers and provide guidance through reviews and knowledge sharing.
  • Participate in incident investigations and postmortems to translate learnings into durable improvements.
  • Evaluate new observability technologies balancing innovation with maintainability.

Skills

Observability
Go
Python
Kubernetes
Terraform
Cloud native
Leadership
Communication

Tools

Prometheus
Grafana
OpenTelemetry
ClickHouse
Elastic
Loki
Tempo
Thanos
VictoriaMetrics

Job description

About Nscale

Nscale is the GPU cloud engineered for AI. We provide cost‑effective, high‑performance infrastructure for AI start‑ups and large enterprise customers. Nscale simplifies AI development while enabling superior results, supporting strategic business outcomes such as cost management, rapid innovation, and environmental responsibility. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you’ll build trust through openness and transparency while contributing to the technology that powers the future.

About The Role

As a Staff Observability Platform Engineer, you’ll play a critical role in building and evolving Nscale’s observability platform, enabling deep visibility into GPU clusters, AI workloads, and the infrastructure that powers them. You view observability as a product, not simply a collection of tools. You’ll help define and implement scalable, reliable observability solutions that empower engineering teams to understand system behavior, diagnose issues quickly, and operate complex distributed systems with confidence. You’ll combine technical leadership with hands‑on engineering, partnering across SRE, infrastructure, platform, and AI/ML teams to improve reliability, operational efficiency, and developer experience. You’ll influence architectural decisions, establish engineering best practices, and help drive the evolution of observability capabilities across the organization. This is a role for someone who enjoys solving difficult infrastructure problems, building platforms that scale, and helping engineering teams succeed through better visibility and operational insight.

Responsibilities
  • Design, build, and evolve observability platforms across metrics, logs, traces, alerting, and telemetry pipelines.
  • Lead the implementation of scalable observability solutions that support Nscale’s growing GPU and AI infrastructure.
  • Partner with SRE, infrastructure, platform, and AI/ML teams to ensure observability is embedded throughout the software and infrastructure lifecycle.
  • Drive improvements in monitoring coverage, alert quality, service health visibility, and incident response effectiveness.
  • Develop standards, frameworks, and reusable patterns that simplify observability adoption across engineering teams.
  • Identify reliability risks and operational blind spots, helping teams proactively address them before they impact customers.
  • Contribute to architectural decisions around telemetry collection, storage, retention, cardinality management, and performance optimization.
  • Lead technical initiatives and projects that improve platform scalability, reliability, and operational efficiency.
  • Mentor engineers and provide technical guidance through design reviews, code reviews, and knowledge sharing.
  • Participate in incident investigations and postmortems, translating operational learnings into durable platform improvements.
  • Evaluate new observability technologies and practices, balancing innovation with operational simplicity and long‑term maintainability.
About You
  • 6+ years of experience in SRE, platform engineering, infrastructure engineering, observability engineering, or related disciplines.
  • Strong experience building and operating observability platforms in cloud‑native, distributed environments.
  • Deep hands‑on experience with several of the following technologies: Prometheus, Thanos, VictoriaMetrics, Grafana, Loki, Tempo, OpenTelemetry, ClickHouse, Elastic, or similar platforms.
  • Strong software engineering skills with proficiency in Go, Python, or equivalent languages.
  • Experience operating and troubleshooting Kubernetes‑based platforms at scale.
  • Strong understanding of monitoring, logging, tracing, telemetry pipelines, and modern observability practices.
  • Experience designing systems with scalability, reliability, performance, and operational simplicity in mind.
  • Proficiency with Infrastructure‑as‑Code tools such as Terraform, Ansible, or equivalent.
  • Ability to lead technical initiatives and influence engineering decisions across multiple teams.
  • Excellent communication skills with the ability to explain technical tradeoffs and align stakeholders around pragmatic solutions.
Preferred
  • Experience operating observability systems in GPU, AI/ML, HPC, or large‑scale compute environments.
  • Familiarity with Slurm, Kubernetes GPU scheduling, or AI infrastructure platforms.
  • Experience with high‑volume telemetry pipelines and streaming technologies such as Kafka, Vector, or Fluent Bit.
  • Knowledge of observability challenges related to model training, inference workloads, GPU utilization, and distributed AI systems.
  • Experience mentoring engineers and helping grow technical capability across teams.
Equal Opportunities Statement

We strongly encourage applications from people of color, the LGBTQ+ community, people with disabilities, neurodivergent individuals, parents, carers, and people from lower socio‑economic backgrounds. If there’s anything we can do to accommodate your specific situation, please let us know.

Note

Responsibilities outlined are not exhaustive and may evolve as business needs change.

Privacy Notice

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Observability Platform Engineer
Staff Observability Platform Engineer

Nscale • San Francisco (CA)

On-site
USD 130,000 - 160,000
Staff Observability Platform Engineer
Staff Observability Platform Engineer

Nscale • New York (NY)

On-site
USD 120,000 - 150,000
Principal Observability Platform Engineer
Principal Observability Platform Engineer

Nscale • San Francisco (CA)

On-site
USD 150,000 - 215,000
Medical benefits
Flexible paid time off
Parental leave
Principal Observability Platform Engineer
Principal Observability Platform Engineer

Nscale • Seattle (WA)

On-site
USD 150,000 - 215,000
Medical, dental, vision benefits
Flexible paid time off
Parental leave
+1
Senior Technical Product Manager, Observability
Senior Technical Product Manager, Observability

Nscale • New York (NY)

On-site
USD 200,000 - 280,000
Competitive benefits package
Flexible paid time off
Parental leave
Infrastructure Software Engineer, Fleet & Automation New Houston; New York; San Francisco; Seattle
Infrastructure Software Engineer, Fleet & Automation New Houston; New York; San Francisco; Seattle

Nscale • New York (NY)

On-site
USD 140,000 - 210,000
Competitive package
Equity
Growth opportunities
Senior Cloud Native Platform Engineer
Senior Cloud Native Platform Engineer

Nscale • New York (NY)

On-site
USD 200,000 - 225,000
Highly competitive US compensation package
Human-first flexibility
Dynamic progression plan
Senior Observability Platform Engineer for AI/GPU Infra
Senior Observability Platform Engineer for AI/GPU Infra

Nscale • Seattle (WA)

On-site
USD 180,000 - 240,000
Infrastructure Software Engineer, Fleet & Automation
Infrastructure Software Engineer, Fleet & Automation

Nscale • Houston (TX)

On-site
USD 150,000 - 215,000
Base salary + equity
Annual reviews
Growth opportunities
Infrastructure Software Engineer, Fleet & Automation
Infrastructure Software Engineer, Fleet & Automation

Socket.dev • Houston (TX)

On-site
USD 150,000 - 200,000
Base + equity
Fast-growing startup
Growth/progression plan