Senior Observability Platform Engineer New US

Nscale

Northern (KY)

Hybrid

USD 160,000 - 230,000

Full time

6 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical insurance
Dental insurance
Vision insurance
Flexible PTO
Parental leave
Retirement plan

Job summary

Nscale is hiring a Senior Observability Platform Engineer to design, build, and scale our observability platform for GPU clusters and AI workloads. You’ll treat observability as a product, balancing usability, scalability, and operational efficiency to reduce cognitive load for engineers.

You will partner with SRE, infrastructure, and AI/ML teams to embed observability into services, mentor peers, and help drive platform direction and patterns.

Qualifications

  • 5+ years in SRE, infrastructure engineering, platform engineering, or observability-focused roles.
  • Experience operating and scaling observability systems in production environments.
  • Strong understanding of monitoring concepts: metrics, logs, traces, alerting, and SLOs.

Responsibilities

  • Design, build, and operate scalable observability systems across metrics, logs, traces, and alerting.
  • Contribute to architectural decisions around tooling, data pipelines, storage, and retention strategies.
  • Improve signal quality by reducing noise, managing cardinality, and refining alerting practices.
  • Help identify and address observability gaps before they impact reliability.
  • Partner with SRE, infrastructure, and AI/ML teams to integrate observability into services and platforms.
  • Develop reusable patterns, libraries, and best practices that improve consistency across teams.
  • Participate in incident response and postmortems, driving actionable improvements.
  • Evaluate and adopt tools that improve developer experience, scalability, and operational efficiency.
  • Support and mentor engineers within the team through code reviews and knowledge sharing.

Skills

SRE/Platform engineering
Observability
Collaboration

Tools

Prometheus
Thanos
VictoriaMetrics
Grafana
Loki
Tempo
OpenTelemetry
ClickHouse
Elastic

Job description

Senior Observability Platform Engineer – Nscale

About Nscale

Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale simplifies AI development while enabling superior results, supporting strategic business outcomes such as cost management, rapid innovation, and environmental responsibility.

We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you’ll build trust through openness and transparency while contributing to the technology that powers the future.

About the Role

As a Senior Observability Platform Engineer, you’ll play a key role in designing, building, and scaling Nscale’s observability platform. You’ll focus on delivering reliable, high-quality visibility into GPU clusters, AI workloads, and the infrastructure that powers them.

You approach observability as a product—balancing usability, scalability, and operational efficiency. You build systems that reduce cognitive load for engineers, surface meaningful signals, and enable fast, confident debugging when things go wrong.

You’ll contribute to platform direction, implement critical systems, and collaborate closely with SRE, infrastructure, and AI/ML teams to ensure observability is embedded into everything we run.

This is a hands-on engineering role with meaningful influence over platform design and evolution.

What You’ll Do
  • Design, build, and operate scalable observability systems across metrics, logs, traces, and alerting
  • Contribute to architectural decisions around tooling, data pipelines, storage, and retention strategies
  • Improve signal quality by reducing noise, managing cardinality, and refining alerting practices
  • Help identify and address observability gaps before they impact reliability
  • Partner with SRE, infrastructure, and AI/ML teams to integrate observability into services and platforms
  • Develop reusable patterns, libraries, and best practices that improve consistency across teams
  • Participate in incident response and postmortems, driving actionable improvements
  • Evaluate and adopt tools that improve developer experience, scalability, and operational efficiency
  • Support and mentor engineers within the team through code reviews and knowledge sharing
About You
  • 5+ years in SRE, infrastructure engineering, platform engineering, or observability-focused roles
  • Experience operating and scaling observability systems in production environments
  • Strong understanding of monitoring concepts: metrics, logs, traces, alerting, and SLOs
  • Hands-on experience with several of: Prometheus, Thanos, VictoriaMetrics, Grafana, Loki, Tempo, OpenTelemetry, ClickHouse, Elastic
  • Solid programming skills (Python, Go, or similar) with the ability to build and maintain production systems
  • Experience working with Kubernetes-based infrastructure
  • Familiarity with Infrastructure-as-Code (Terraform, Ansible, or similar)
  • Pragmatic mindset with a focus on simplicity, reliability, and maintainability
  • Strong collaboration skills and ability to work across teams
Preferred
  • Experience with observability data pipelines (Kafka, Vector, Fluent Bit, etc.)
  • Exposure to AI/ML infrastructure or GPU-based systems
  • Familiarity with performance monitoring for distributed systems
  • Experience improving developer experience through observability tooling

We strongly encourage applications from people of color, the LGBTQ+ community, people with disabilities, neurodivergent individuals, parents, carers, and people from lower socio-economic backgrounds.

If there’s anything we can do to accommodate your specific situation, please let us know.

Note: Responsibilities outlined are not exhaustive and may evolve as business needs change

The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.

Salary Range

$160,000 - $230,000 USD

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice:Here.

Apply for this job

*

indicates a required field

First Name *

Last Name *

Email *

Phone

Country *

Phone *

Location (City) *

Resume/CV *

Enter manually

Accepted file types: pdf, doc, docx, txt, rtf

Enter manually

Accepted file types: pdf, doc, docx, txt, rtf

LinkedIn Profile

Website

Do you have the legal right to work in the country in which this job is located, without requiring visa sponsorship? * Select...

What is your current or most recent employer? *

What is your preferred name? *

Nscale uses AI-powered tools to assist in reviewing and prioritising applications against the requirements of this role. All final hiring decisions are made by humans. To learn more about how AI is used and your rights, click "Learn more" below.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Observability Platform Engineer
Staff Observability Platform Engineer

Nscale • San Francisco (CA)

On-site
USD 130,000 - 160,000
Staff Observability Platform Engineer
Staff Observability Platform Engineer

Nscale • Seattle (WA)

On-site
USD 180,000 - 240,000
Staff Observability Platform Engineer
Staff Observability Platform Engineer

Nscale • New York (NY)

On-site
USD 120,000 - 150,000
Senior Cloud Native Platform Engineer
Senior Cloud Native Platform Engineer

Nscale • New York (NY)

Hybrid
USD 200,000 - 225,000
Bonus & equity
Flexible PTO
Medical benefits
+2
Support Desk Engineer
Support Desk Engineer

Nscale • United States

Remote
USD 50,000 - 100,000
Competitive package with reviews every 12 months
Dynamic progression plan tailored to ambitions
Human-first flexibility in a remote-first environment
Senior Software Engineering Manager - Fleet Management
Senior Software Engineering Manager - Fleet Management

Socket.dev • Seattle (WA)

On-site
USD 300,000 - 350,000
Equity
Bonus potential
Flexible work policy
+1
Principal Observability Platform Engineer
Principal Observability Platform Engineer

Nscale • Seattle (WA)

On-site
USD 150,000 - 215,000
Medical, dental, vision benefits
Flexible paid time off
Parental leave
+1
Principal Observability Platform Engineer
Principal Observability Platform Engineer

Nscale • San Francisco (CA)

On-site
USD 150,000 - 215,000
Medical benefits
Flexible paid time off
Parental leave
Infrastructure Software Engineer, Fleet & Automation Houston; New York; San Francisco; Seattle
Infrastructure Software Engineer, Fleet & Automation Houston; New York; San Francisco; Seattle

Nscale • Northern (KY)

Hybrid
USD 150,000 - 215,000
Base + equity
Equity
Growth opportunities
Infrastructure Software Engineer, Fleet & Automation
Infrastructure Software Engineer, Fleet & Automation

Socket.dev • Houston (TX)

On-site
USD 150,000 - 200,000
Base + equity
Fast-growing startup
Growth/progression plan