Senior SRE: Lead Reliable AI Platform & Mentor the Team

Nscale

Houston (TX)

On-site

USD 170,000 - 265,000

Full time

31 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Competitive base + equity
Real ownership from the start
Flexible work culture

Job summary

Nscale, a GPU cloud for AI, seeks a senior SRE to raise the reliability bar across the platform. You will own the hardest problems, influence architectural decisions, and mentor others while maintaining an on-call rotation that becomes lighter over time.

You’ll drive SLOs, observability, incident processes, and tooling that reduces toil. This is hands-on ownership with real impact on production at scale in a data center or cloud environment.

Qualifications

  • 6-10 years in SRE or related roles with ownership of production at scale.
  • Strong software engineering in Python or Go.
  • Deep Linux, networking and distributed systems knowledge.
  • Hands-on with Kubernetes and bare-metal or virtualization.
  • Experience running AI or GPU workloads or HPC.
  • SLOs, observability, incident process, on-call practices.
  • Senior voice in incidents and design reviews under pressure.
  • Able to raise colleagues and drive improvements.

Responsibilities

  • Own reliability for critical production services end to end; set direction.
  • Grow the team and mentor SREs through design reviews and debriefs.
  • Set standards: SLO framework, incident process, on-call practices.
  • Participate early in design reviews and architecture decisions.
  • Lead the hardest incidents and root cause analysis.
  • Build tooling and automation to remove toil for the team.

Skills

Python
Go
Linux
Networking
Distributed systems
Kubernetes
AI workloads
On-call management
Incident response
Mentoring

Job description

Nscale, a GPU cloud for AI, seeks a senior SRE to raise the reliability bar across the platform. You will own the hardest problems, influence architectural decisions, and mentor others while maintaining an on-call rotation that becomes lighter over time.

You’ll drive SLOs, observability, incident processes, and tooling that reduces toil. This is hands-on ownership with real impact on production at scale in a data center or cloud environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE — AI Platform Reliability & Automation
Senior SRE — AI Platform Reliability & Automation

nscaleoperationsukltd • Houston (TX)

On-site
USD 130,000 - 200,000
Competitive base plus equity
Real scope early
Flexible work style
Senior SRE: AI/GPU Scale & Automation
Senior SRE: AI/GPU Scale & Automation

Nscale • New York (NY), Northern (KY)

Hybrid
USD 130,000 - 200,000
Competitive base plus equity
Real scope early
Flexible work expectations
Senior SRE Lead - AI-Driven Reliability & Scale
Senior SRE Lead - AI-Driven Reliability & Scale

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 168,000 - 334,000
Equity
Benefits
Hybrid work model
Senior SRE Lead: Reliability, Automation & Incidents
Senior SRE Lead: Reliability, Automation & Incidents

Shield AI • San Diego (CA)

On-site
USD 183,000 - 275,000
Equity
Bonus
Benefits
Senior SRE: GPU AI Platform Reliability & SLIs
Senior SRE: GPU AI Platform Reliability & SLIs

Mirantis • United States

Hybrid
USD 140,000 - 230,000
Competitive compensation package
Strong benefits plan
Professional development and training
+1
Lead AI Platform Reliability Architect
Lead AI Platform Reliability Architect

NVIDIA • Santa Clara (CA)

Hybrid
USD 248,000 - 397,000
Senior SRE – AI Cloud Platform, Kubernetes Expert
Senior SRE – AI Cloud Platform, Kubernetes Expert

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health, dental, vision coverage for in
Wellness and commuter stipends
401k with 2% company match
+1
Senior SRE: AI-Driven Infra & Reliability
Senior SRE: AI-Driven Infra & Reliability

Jobless • Ann Arbor (MI)

Hybrid
USD 180,000 - 240,000
Health Care Coverage
Life Insurance
Health Savings Account
+3
Senior SRE: AI Cloud Platform & Kubernetes
Senior SRE: AI Cloud Platform & Kubernetes

Lambda Inc. • San Francisco (CA)

Hybrid
USD 190,000 - 270,000
Health insurance
401k with company match
Flexible PTO
+2
Senior Site Reliability Engineer -AI Infrastructure Operations
Senior Site Reliability Engineer -AI Infrastructure Operations

Nscale • Houston (TX)

On-site
USD 170,000 - 265,000
Competitive base + equity
Real ownership from the start
Flexible work culture