Senior Platform Reliability Engineer

Nscale

San Francisco (CA)

On-site

USD 150,000 - 210,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Nscale, the GPU cloud for AI, seeks a Senior Operational Engineer to lead cross-service improvements that enhance reliability, security, and cost awareness. You will design and deliver shared operational capabilities and partner with service teams to address systemic risks.

This hands-on role focuses on implementing readiness checks, runbooks, incident response, and automation, while mentoring engineers and driving meaningful improvements across Grafana, PagerDuty, Jira, Backstage, and cloud

Qualifications

  • 6–10 years in operations engineering or related production-focused role.
  • Experience designing and operating shared capabilities for monitoring, alerts, on-call, runbooks, service readiness, and operational reporting.
  • Hands-on experience building integrations across Grafana, PagerDuty, Jira, Backstage, and public cloud.
  • Experience using incident, change, service-health, continuity, patching, or cost data to drive operational improvement.
  • Ability to lead cross-service work, set practical standards, and support teams through incidents.

Responsibilities

  • Own design and delivery of shared operational capabilities and automation.
  • Set standards for ownership, on-call readiness, alerts, dashboards, recovery procedures, and operational evidence.
  • Partner with service teams to identify and resolve recurring operational issues.
  • Design and improve integrations across Grafana, PagerDuty, Jira, Backstage, and public-cloud platforms.
  • Turn incident findings into lasting engineering improvements.
  • Act as an incident commander and lead technical analysis for incidents; support corrective-action planning and delivery.
  • Build safe, observable, auditable automations that reduce manual work and improve execution.
  • Automate operational reporting, including SLA performance, health indicators, cost trends, patching status, and action closure.
  • Mentor engineers, review operational designs, and raise the technical standard of the wider team.
  • Contribute to continuity testing and focused failure-and-recovery experiments.

Skills

Operations engineering experience
SRE & cloud infra
Incident response leadership
Automation & integrations
Cross-team collaboration

Tools

Grafana
PagerDuty
Jira
Backstage
Public cloud platforms

Job description

Nscale, the GPU cloud for AI, seeks a Senior Operational Engineer to lead cross-service improvements that enhance reliability, security, and cost awareness. You will design and deliver shared operational capabilities and partner with service teams to address systemic risks.

This hands-on role focuses on implementing readiness checks, runbooks, incident response, and automation, while mentoring engineers and driving meaningful improvements across Grafana, PagerDuty, Jira, Backstage, and cloud

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE & Platform Reliability Engineer
Senior SRE & Platform Reliability Engineer

Nscale • Seattle (WA)

On-site
USD 140,000 - 180,000
Senior Operational Engineer: Reliability & Automation Lead
Senior Operational Engineer: Reliability & Automation Lead

Nscale • Houston (TX)

On-site
USD 120,000 - 180,000
Senior Operational Engineer
Senior Operational Engineer

Nscale • New York (NY)

On-site
USD 130,000 - 170,000
Senior Operational Engineer
Senior Operational Engineer

Nscale • Seattle (WA)

On-site
USD 140,000 - 180,000
Senior Operational Engineer
Senior Operational Engineer

Nscale • Houston (TX)

On-site
USD 120,000 - 180,000
Senior Operational Engineer
Senior Operational Engineer

Nscale • San Francisco (CA)

On-site
USD 150,000 - 210,000
Operational Engineer
Operational Engineer

Greenhouse Software, Inc. • San Francisco (CA)

On-site
USD 140,000 - 200,000
Senior Observability Platform Engineer – AI GPU Scale
Senior Observability Platform Engineer – AI GPU Scale

Nscale • United States

On-site
USD 160,000 - 230,000
Medical, dental, vision insurance
Flexible paid time off (PTO)
Parental leave
+1
Senior Observability Platform Engineer for AI GPU Cloud
Senior Observability Platform Engineer for AI GPU Cloud

Greenhouse Software, Inc. • Northern (KY)

Hybrid
USD 160,000 - 230,000
Medical
Dental
Vision
+3
Operational Engineer
Operational Engineer

Nscale • San Francisco (CA)

On-site
USD 110,000 - 170,000