Operational Data & Observability Engineer

Nscale

United States

On-site

USD 145,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, vision
Flexible paid time off
Parental leave
Equity/bonus programs

Job summary

Nscale is seeking an Operational Data & Observability Engineer to build and evolve monitoring, logging, and observability capabilities for our production environments. You will partner with DevOps, SRE, and platform teams to develop modern observability solutions and enable data-driven operational excellence.

Responsibilities include designing dashboards and SLOs, building centralized logging pipelines, implementing distributed tracing, and ensuring data quality and cost-efficient storage.

Qualifications

  • 3+ years in DevOps, SRE, Operations, Platform, or Observability Engineering.
  • Hands-on experience with modern monitoring platforms (Prometheus, Grafana, Datadog, New Relic).
  • Experience with centralized logging (ELK/Elastic Stack, Splunk, CloudWatch).
  • Proficiency in Python, Go, or Bash scripting; strong observability fundamentals.

Responsibilities

  • Design and implement enterprise observability strategies across infrastructure, services, and applications.
  • Develop dashboards, alerts, and SLOs for operational visibility.
  • Build and maintain logging and log analysis pipelines; implement distributed tracing.
  • Design and manage data pipelines for monitoring and analytics; ensure data quality and retention.

Skills

Observability
DevOps
SRE
Python
Go
Shell/Bash

Tools

Prometheus
Grafana
Datadog
ELK Stack
Kubernetes

Job description

Operational Data & Observability Engineer

US

About the Role

We're looking for an Operational Data & Observability Engineer to build and evolve the monitoring, logging, and observability capabilities that power our production environments.

In this role, you'll help ensure our infrastructure and applications remain reliable, scalable, and performant by providing engineering teams with actionable operational insights.

You’ll partner closely with DevOps, Site Reliability Engineering (SRE), platform, and software engineering teams to develop modern observability solutions, improve incident response, and enable data-driven operational excellence.

What You'll Do
Design & Build Observability Solutions
  • Design and implement enterprise observability strategies across infrastructure, services, and applications.
  • Develop monitoring dashboards, alerts, and Service Level Objectives (SLOs) that provide meaningful operational visibility.
  • Build and maintain centralized logging and log analysis pipelines.
  • Implement distributed tracing to improve visibility across microservices and complex application workflows.
  • Establish performance baselines and develop anomaly detection strategies.
Operational Data Engineering
  • Deploy, configure, and maintain metrics, logs, events, and telemetry collection systems.
  • Design and manage operational data pipelines that support monitoring and analytics.
  • Develop APIs and integrations that enable operational data consumption across teams.
  • Ensure data quality, consistency, retention, and cost‑efficient storage practices.
Reliability & Operations
  • Troubleshoot production issues using monitoring, logging, and tracing data.
  • Participate in an on‑call rotation and support incident response activities.
  • Create and maintain operational documentation, runbooks, and troubleshooting guides.
  • Partner with engineering teams to improve platform reliability, scalability, and operational readiness.
  • Continuously optimize observability infrastructure for performance and resilience.
Platform & Tool Administration
  • Administer and enhance observability platforms such as Datadog, Grafana, Prometheus, ELK Stack, New Relic, or similar technologies.
  • Evaluate emerging observability tools and recommend improvements.
  • Automate monitoring deployments, instrumentation, and platform configuration.
  • Perform ongoing maintenance, upgrades, and lifecycle management of observability infrastructure.
What You’ll Bring
Required Qualifications
  • 3+ years of experience in DevOps, Site Reliability Engineering (SRE), Operations Engineering, Platform Engineering, or Observability Engineering.
  • Hands‑on experience with modern monitoring platforms such as Prometheus, Grafana, Datadog, New Relic, or equivalent.
  • Experience working with centralized logging platforms including ELK/Elastic Stack, Splunk, CloudWatch, or similar solutions.
  • Proficiency with scripting or programming languages such as Python, Go, Bash, or equivalent.
  • Strong understanding of observability fundamentals, including metrics, logging, distributed tracing, and application performance monitoring (APM).
  • Experience working with cloud platforms (AWS, Azure, or Google Cloud Platform) and Kubernetes or other container orchestration technologies.
  • Solid understanding of application, infrastructure, networking, database, and storage performance monitoring.
  • Strong analytical, troubleshooting, communication, and documentation skills with a collaborative approach to problem‑solving.
Preferred Qualifications
  • Experience supporting microservices‑based architectures.
  • Expertise across multiple observability platforms.
  • Experience with incident management, root cause analysis, and post‑incident reviews.
  • Infrastructure as Code experience using Terraform, Ansible, or similar tools.
  • Familiarity with eBPF or low‑level Linux performance monitoring.
  • Experience building custom telemetry, ETL, or operational data pipelines.
  • Understanding of security monitoring, audit logging, and compliance requirements.
What Success Looks Like
  • Improve platform visibility and operational health.
  • Reduce Mean Time to Resolution (MTTR) during incidents.
  • Increase alert quality while reducing unnecessary noise.
  • Deliver highly available, scalable observability platforms.
  • Improve engineering productivity through actionable monitoring and operational insights.
  • Optimize observability infrastructure performance and cost efficiency.
  • Participate in a rotating on‑call schedule to support production environments.
  • Support mission‑critical systems with occasional after‑hours or incident response responsibilities.
  • Hybrid or remote work arrangements available, depending on business needs.
Why Join Us?

You'll play a critical role in building the operational intelligence that keeps our platforms running at scale. If you're passionate about observability, automation, reliability, and empowering engineering teams with meaningful operational insights, we'd love to hear from you.

Salary Range

$145,000 - $180,000 USD

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.

Benefits
  • Medical, dental, vision, and flexible paid time off.
  • Parental leave and retirement plan participation.
  • Potential bonus, equity, and/or commission programs.
  • Competitive benefits package including health, dental, vision, paid time off, parental leave, and retirement plan participation.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Observability Platform Engineer
Principal Observability Platform Engineer

Nscale • San Francisco (CA)

On-site
USD 150,000 - 215,000
Medical benefits
Flexible paid time off
Parental leave
Principal Observability Platform Engineer
Principal Observability Platform Engineer

Nscale • Seattle (WA)

On-site
USD 150,000 - 215,000
Medical, dental, vision benefits
Flexible paid time off
Parental leave
+1
Staff Observability Platform Engineer
Staff Observability Platform Engineer

Nscale • Seattle (WA)

On-site
USD 180,000 - 240,000
Staff Observability Platform Engineer
Staff Observability Platform Engineer

Nscale • San Francisco (CA)

On-site
USD 130,000 - 160,000
Staff Observability Platform Engineer
Staff Observability Platform Engineer

Nscale • New York (NY)

On-site
USD 120,000 - 150,000
Senior Technical Product Manager, Observability
Senior Technical Product Manager, Observability

Nscale • New York (NY)

On-site
USD 200,000 - 280,000
Competitive benefits package
Flexible paid time off
Parental leave
Remote Observability & Reliability Engineer
Remote Observability & Reliability Engineer

Nscale • United States

Hybrid
USD 145,000 - 180,000
Medical, dental, vision
Flexible paid time off
Parental leave
+1
Senior Platform Engineer
Senior Platform Engineer

Ww • United States

On-site
USD 200,000 - 215,000
Staff Software Engineer, Observability
Staff Software Engineer, Observability

United States Digital Space LLC • Menlo Park (CA)

On-site
USD 180,000 - 250,000
Health insurance
Equity ownership
401(k) matching
+1
Infrastructure Operations Engineer Greensboro, NC
Infrastructure Operations Engineer Greensboro, NC

Nscale • Winston-Salem (NC)

Hybrid
USD 100,000 - 160,000
Competitive package (base + equity)
Flexible workplace and support for personal growth
Medical, dental, and vision benefits