Platform SRE: Own Automation for AI/GPU Ops

Nscale

Seattle (WA)

On-site

USD 130,000 - 200,000

Full time

37 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Competitive base + equity
Real scope early
Flexible work style

Job summary

Nscale, a GPU cloud platform, seeks a career-level SRE to own the automation and tooling that keep AI workloads reliable at scale. You will join incident rotations, drive root-cause analysis, and push improvements across production systems.

You will work with Linux, networking, and distributed services, building dashboards, defining SLOs/SLIs, and partnering with platform teams to raise the bar on reliability and efficiency.

Qualifications

  • 3–6 years in SRE, systems engineering, or software engineering with production experience.
  • Proficiency in Python or Go; strong automation bias.
  • Solid Linux, networking, and distributed systems knowledge.
  • Track record troubleshooting live production issues and owning the fix to retro.

Responsibilities

  • Build and own automation and tooling to keep the platform running; treat toil as a bug to fix.
  • Define and maintain SLOs, SLIs, and dashboards for service health.
  • Take point during incidents; troubleshoot, perform RCA, and drive post-incident reviews.
  • Investigate performance and reliability across Linux, networking, and distributed services; fix at source.
  • Partner with Engineering, Networking, and Infrastructure to raise reliability.
  • Improve availability, scalability, and efficiency through code.

Skills

Python/Go programming
Linux & Networking
Observability/Monitoring
Production troubleshooting
Automation mindset

Tools

Kubernetes
InfiniBand/RDMA
Bare-metal/Virt

Job description

Nscale, a GPU cloud platform, seeks a career-level SRE to own the automation and tooling that keep AI workloads reliable at scale. You will join incident rotations, drive root-cause analysis, and push improvements across production systems.

You will work with Linux, networking, and distributed services, building dashboards, defining SLOs/SLIs, and partnering with platform teams to raise the bar on reliability and efficiency.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: AI/GPU Scale & Automation
Senior SRE: AI/GPU Scale & Automation

Nscale • New York (NY), Northern (KY)

Hybrid
USD 130,000 - 200,000
Competitive base plus equity
Real scope early
Flexible work expectations
Senior SRE — AI Platform Reliability & Automation
Senior SRE — AI Platform Reliability & Automation

nscaleoperationsukltd • Houston (TX)

On-site
USD 130,000 - 200,000
Competitive base plus equity
Real scope early
Flexible work style
Senior SRE: Lead Reliable AI Platform & Mentor the Team
Senior SRE: Lead Reliable AI Platform & Mentor the Team

Nscale • Houston (TX)

On-site
USD 170,000 - 265,000
Competitive base + equity
Real ownership from the start
Flexible work culture
Site Reliability Engineer
Site Reliability Engineer

Nscale • Seattle (WA)

On-site
USD 130,000 - 200,000
Competitive base + equity
Real scope early
Flexible work style
Site Reliability Engineer New Houston; New York; San Francisco; Seattle
Site Reliability Engineer New Houston; New York; San Francisco; Seattle

Nscale • New York (NY), Northern (KY)

Hybrid
USD 130,000 - 200,000
Competitive base plus equity
Real scope early
Flexible work expectations
Site Reliability Engineer
Site Reliability Engineer

nscaleoperationsukltd • Houston (TX)

On-site
USD 130,000 - 200,000
Competitive base plus equity
Real scope early
Flexible work style
Senior Observability Platform Engineer – AI GPU Scale
Senior Observability Platform Engineer – AI GPU Scale

Nscale • United States

On-site
USD 160,000 - 230,000
Medical, dental, vision insurance
Flexible paid time off (PTO)
Parental leave
+1
Senior SRE: FinOps-Driven Infra, GPU & Scale
Senior SRE: FinOps-Driven Infra, GPU & Scale

Level AI • Mountain View (CA)

On-site
USD 180,000 - 240,000
Senior SRE: GPU AI Platform Reliability & SLIs
Senior SRE: GPU AI Platform Reliability & SLIs

Mirantis • United States

Hybrid
USD 140,000 - 230,000
Competitive compensation package
Strong benefits plan
Professional development and training
+1
Senior AI GPU Infra SRE - Scale, Automation & Equity
Senior AI GPU Infra SRE - Scale, Automation & Equity

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity