Senior SRE: Automate Reliability for AI/GPU Platform

Nscale

Seattle (WA)

On-site

USD 130,000 - 200,000

Full time

20 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Nscale, the GPU cloud built for AI, seeks a career-level SRE to own the automation and tooling that keeps our platform running at scale. You’ll sit in the incident rotation, drive root cause analysis, and push improvements that quiet production systems over time.

In this role you’ll define SLOs/SLIs, build health dashboards, troubleshoot live issues, and collaborate with Engineering, Networking, and Infrastructure to raise reliability across the stack.

Qualifications

  • 3–6 years in SRE, systems engineering, or software engineering with production experience.
  • Strong programming skills (Python, Go, or similar) and automation bias.
  • Solid Linux, networking fundamentals, and distributed systems.
  • Proven ability to troubleshoot live production issues and own fixes.
  • Fluency with monitoring, observability, metrics, logs, dashboards, and alerts.
  • Comfort in fast-moving environments with shifting priorities.

Responsibilities

  • Build and own automation and tooling to keep the platform running.
  • Define and maintain SLOs, SLIs, and health dashboards.
  • Lead incident response; perform root cause analysis and post-incident reviews.
  • Investigate performance and reliability issues across Linux, networking, distributed services; fix at source.
  • Partner with Engineering, Networking, and Infrastructure to raise reliability across the stack.
  • Improve availability, scalability, and efficiency through code.

Skills

Python
Go
Automation
Linux
Networking
Observability
Distributed systems
Troubleshooting
Production systems

Tools

Kubernetes
Bare-metal environments
InfiniBand
RDMA

Job description

Nscale, the GPU cloud built for AI, seeks a career-level SRE to own the automation and tooling that keeps our platform running at scale. You’ll sit in the incident rotation, drive root cause analysis, and push improvements that quiet production systems over time.

In this role you’ll define SLOs/SLIs, build health dashboards, troubleshoot live issues, and collaborate with Engineering, Networking, and Infrastructure to raise reliability across the stack.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: Automate, Own, and Stabilize AI Platform
Senior SRE: Automate, Own, and Stabilize AI Platform

Nscale • San Francisco (CA)

On-site
USD 130,000 - 200,000
Equity
Career growth
Flexible schedule
Senior SRE – AI Infrastructure Reliability Leader
Senior SRE – AI Infrastructure Reliability Leader

Nscale • San Francisco (CA), Seattle (WA), Houston (TX)

On-site
USD 170,000 - 265,000
Equity
Ownership from start
Flexible schedule
Senior SRE Lead: Scale Reliability & AI Ops
Senior SRE Lead: Scale Reliability & AI Ops

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Senior SRE: GPU AI Platform Reliability & SLIs
Senior SRE: GPU AI Platform Reliability & SLIs

Mirantis • United States

Hybrid
USD 140,000 - 230,000
Competitive compensation package
Strong benefits plan
Professional development and training
+1
Principal SRE: AI Platform Reliability & Automation
Principal SRE: AI Platform Reliability & Automation

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
Equity
Benefits
Senior SRE — Scale AI Systems, Equity Eligible
Senior SRE — Scale AI Systems, Equity Eligible

NVIDIA AI • Santa Clara (CA)

On-site
USD 168,000 - 334,000
Senior Observability Platform Engineer – AI GPU Scale
Senior Observability Platform Engineer – AI GPU Scale

Nscale • United States

On-site
USD 160,000 - 230,000
Medical, dental, vision insurance
Flexible paid time off (PTO)
Parental leave
+1
Site Reliability Engineer
Site Reliability Engineer

Nscale • San Francisco (CA)

On-site
USD 130,000 - 200,000
Equity
Career growth
Flexible schedule
Site Reliability Engineer
Site Reliability Engineer

Nscale • Seattle (WA)

On-site
USD 130,000 - 200,000
Senior SRE & Automation Engineer — GPU Cloud Reliability
Senior SRE & Automation Engineer — GPU Cloud Reliability

Bitdeer Technologies Group • Austin (TX)

On-site
USD 150,000 - 230,000