Senior Operational Engineer

Greenhouse Software, Inc.

Seattle (WA)

On-site

USD 180,000 - 240,000

Full time

23 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Nscale is seeking a Senior Operational Engineer to lead cross-service improvements that enhance reliability, security, and cost awareness across our engineering services.

You will design shared operational capabilities, manage incident response, automation, and integrations with Grafana, PagerDuty, Jira and Backstage, guiding teams across the platform.

Qualifications

  • 6–10 years in operations engineering or related production-focused roles.
  • Experience designing shared monitoring, alerting, on-call, and runbooks.
  • Hands-on with Grafana, PagerDuty, Jira, Backstage and cloud integrations.
  • Experience using incident data to drive operational improvements.
  • Ability to lead cross-service work and establish practical standards.

Responsibilities

  • Own design and delivery of shared operational capabilities.
  • Establish standards for service ownership, on-call readiness, alerts and dashboards.
  • Collaborate with teams to resolve recurring operational issues.
  • Design and improve integrations across Grafana, PagerDuty, Jira, Backstage and cloud.
  • Turn incident findings into lasting engineering improvements.
  • Act as incident commander and lead technical analysis for incidents.
  • Build auditable automations to reduce manual work and improve execution.
  • Automate operational reporting for SLA, health, cost trends, and patching.
  • Mentor engineers and raise the team's technical standard.
  • Contribute to continuity testing and focused failure/recovery experiments.

Skills

years of experience
monitoring and alerting
incident response leadership
cross-service collaboration
operational reporting

Tools

Grafana
PagerDuty
Jira
Backstage
Public Cloud

Job description

Nscale is the GPU cloud built for AI. We run high-performance, cost-efficient infrastructure for AI-native startups and global enterprises, from bare metal up through the platform services teams actually build on. Our culture runs on ownership, accountability, and speed. We move with urgency, we tell each other the truth, and everyone here stays close to the infrastructure that makes AI work.

The role

We are looking for a Senior Operational Engineer to lead cross-service improvements that make Nscale's engineering services more reliable, secure, controlled, and cost-aware.

You will design and deliver shared operational capabilities—such as service-readiness standards, canary checks, runbooks, reporting, alerting workflows, and automation—and partner with service teams to address systemic operational risks. This is a hands-on technical role with significant influence across engineering.

What you'll do

Take ownership of the design and delivery of shared operational capabilities, including readiness checks, canaries, runbooks, service-health reporting, and operational automation.

Establish practical standards for service ownership, on-call readiness, alerts, dashboards, recovery procedures, and operational evidence.

Partner with service teams to identify and resolve recurring operational issues and cross-team blockers.

Design and improve integrations and workflows across Grafana, PagerDuty, Jira, Backstage, public-cloud platforms, and reporting systems.

Turn incident findings, change failures, and near misses into lasting engineering improvements.

Act as an incident commander and lead technical analysis for incidents; support corrective-action planning and delivery.

Build safe, observable, auditable automations that reduce manual work and improve the speed and quality of operational execution.

Automate operational reporting, including SLA performance, health indicators, cost trends, patching status, and action closure.

Mentor engineers, review operational designs, and raise the technical standard of the wider team.

Contribute to continuity testing and focused failure-and-recovery experiments.

What you'll bring
  • 6-10 years' experience in operations engineering, SRE, cloud infrastructure, platform engineering, or a similar production-focused engineering role.
  • Experience designing and operating shared capabilities for monitoring, alerting, on-call, runbooks, service readiness, and operational reporting.
  • Hands-on experience building integrations and automations across Grafana, PagerDuty, Jira, Backstage, and public cloud.
  • Experience using incident, change, service-health, continuity, patching, or cost data to drive operational improvement.
  • Ability to lead cross-service work, set practical standards, and support teams through incidents.

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here. Nscale does not accept unsolicited candidate submissions from recruitment agencies.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Operational Data & Observability Engineer
Operational Data & Observability Engineer

Nscale • Seattle (WA)

Hybrid
USD 145,000 - 180,000
Medical, dental, vision
Flexible PTO
Parental leave
+1
Operational Data & Observability Engineer
Operational Data & Observability Engineer

Nscale • United States

Hybrid
USD 145,000 - 180,000
Medical
Dental
Vision
+2
Site Reliability Engineer
Site Reliability Engineer

Nscale • Seattle (WA)

On-site
USD 130,000 - 200,000
Staff Observability Platform Engineer
Staff Observability Platform Engineer

Nscale • Seattle (WA)

On-site
USD 180,000 - 240,000
Staff Observability Platform Engineer
Staff Observability Platform Engineer

Nscale • San Francisco (CA)

On-site
USD 130,000 - 160,000
Senior Observability Platform Engineer
Senior Observability Platform Engineer

Socket.dev • United States

On-site
USD 160,000 - 230,000
Staff Observability Platform Engineer
Staff Observability Platform Engineer

Nscale • New York (NY)

On-site
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

Nscale • San Francisco (CA)

On-site
USD 130,000 - 200,000
Equity
Career growth
Flexible schedule
Director, DC Operations, NC
Director, DC Operations, NC

RiseMe • Charlotte (NC)

On-site
USD 200,000 - 270,000
Equity
Remote-friendly culture
Competitive package
Staff Cloud Native Software Engineer
Staff Cloud Native Software Engineer

Nscale • Seattle (WA)

On-site
USD 220,000 - 265,000
Medical, dental, vision
Flexible paid time off
Parental leave
+1