Senior Site Reliability Engineer -AI Infrastructure Operations

Nscale

Houston (TX)

On-site

USD 170,000 - 265,000

Full time

19 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Competitive base + equity
Real ownership from the start
Flexible work culture

Job summary

Nscale, a GPU cloud for AI, seeks a senior SRE to raise the reliability bar across the platform. You will own the hardest problems, influence architectural decisions, and mentor others while maintaining an on-call rotation that becomes lighter over time.

You’ll drive SLOs, observability, incident processes, and tooling that reduces toil. This is hands-on ownership with real impact on production at scale in a data center or cloud environment.

Qualifications

  • 6-10 years in SRE or related roles with ownership of production at scale.
  • Strong software engineering in Python or Go.
  • Deep Linux, networking and distributed systems knowledge.
  • Hands-on with Kubernetes and bare-metal or virtualization.
  • Experience running AI or GPU workloads or HPC.
  • SLOs, observability, incident process, on-call practices.
  • Senior voice in incidents and design reviews under pressure.
  • Able to raise colleagues and drive improvements.

Responsibilities

  • Own reliability for critical production services end to end; set direction.
  • Grow the team and mentor SREs through design reviews and debriefs.
  • Set standards: SLO framework, incident process, on-call practices.
  • Participate early in design reviews and architecture decisions.
  • Lead the hardest incidents and root cause analysis.
  • Build tooling and automation to remove toil for the team.

Skills

Python
Go
Linux
Networking
Distributed systems
Kubernetes
AI workloads
On-call management
Incident response
Mentoring

Job description

About Nscale

Nscale is the GPU cloud built for AI. We run high-performance, cost-efficient infrastructure for AI-native startups and global enterprises, from bare metal up through the platform services teams actually build on. Our culture runs on ownership, accountability, and speed. We move with urgency, we tell each other the truth, and everyone here stays close to the infrastructure that makes AI work.

The Role

This is a senior SRE role for someone who sets the reliability bar and then pulls the rest of the team up to it. You'll own the hardest problems on the platform: the automation other engineers build on, the services that can't go down, and the design decisions that determine whether either holds up at scale. You'll still carry a pager, but the real job is making sure it fires less, for everyone, over time.

What You'll Do
  • Own reliability for critical production services end to end; set the direction, not just respond to what breaks.
  • Grow the team, not just the systems; mentor other SREs through design review, pairing, and incident debriefs, and hold the bar that pulls everyone up to it.
  • Set the standards the rest of the team works to: the SLO framework, the incident process, and the on-call practices that keep it sustainable.
  • Get in early on design reviews and architecture decisions, so reliability is built in rather than bolted on after the first outage.
  • Lead the hardest incidents and the root causes nobody else can crack; turn each one into a change that keeps it from coming back.
  • Build the tooling and automation that removes toil for the whole team, not just your own surface area.
What You'll Bring
  • 6-10 years in SRE, systems engineering, or software engineering, with real ownership of production at scale in a data center or cloud environment.
  • Strong software engineering skills (Python, Go, or similar); you build tools other engineers adopt, not scripts that run once and rot.
  • Deep command of Linux, networking, and distributed systems, plus the judgment to know where the real failure modes hide.
  • Hands-on with Kubernetes and virtualized or bare-metal environments; comfortable close to the metal, not just the cloud console.
  • Experience running AI or GPU workloads, or high-performance computing (HPC); if not, the depth to get there fast.
  • Reliability practices you put in place that outlasted you: SLOs, observability and alerting at scale, incident process, on-call that people can actually live with.
  • A track record as the senior voice in incidents and design reviews, trusted to make the call under pressure.
  • A habit of raising the people around you without being asked to.
Nice to Have
  • Familiarity with high-performance networking (InfiniBand, RDMA).
On-Call and Pace

A quick note on the shape of the job. This role sits close to production, so there is an on-call rotation, and some weeks are busier than others. As a senior on the team, you help set how that rotation runs and, more to the point, how we make it lighter over time. We share the load fairly, and we treat every page as a signal worth acting on rather than just an interruption. The goal is to leave the systems quieter than you found them, so each rotation asks less of the person carrying it. If that's the kind of ownership you're drawn to, you'll do well here.

What We Offer
  • Competitive base plus equity, reviewed every 12 months.
  • Real ownership from the start, and a direct hand in how reliability works across the platform.
  • Flexibility that treats you as an adult; we care that the work gets done, and we trust you to shape your day.
Salary Range

$170,000 - $265,000 USD. Actual compensation varies with skill set, experience, and location, and the role may be eligible for bonus and equity.

Equal Opportunities Statement

At Nscale, we are committed to fostering an inclusive, diverse, and equitable workplace. We believe that a variety of perspectives enrich our work environment, and we encourage applications from candidates of all backgrounds, experiences, and abilities. We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.

If there’s anything we can do to accommodate your specific situation, please let us know.

The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.

Salary Range: $170,000 USD - $265,000 USD

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

nscaleoperationsukltd • Houston (TX)

On-site
USD 130,000 - 200,000
Competitive base plus equity
Real scope early
Flexible work style
Site Reliability Engineer
Site Reliability Engineer

Nscale • Seattle (WA)

On-site
USD 130,000 - 200,000
Competitive base + equity
Real scope early
Flexible work style
Site Reliability Engineer New Houston; New York; San Francisco; Seattle
Site Reliability Engineer New Houston; New York; San Francisco; Seattle

Nscale • New York (NY), Northern (KY)

Hybrid
USD 130,000 - 200,000
Competitive base plus equity
Real scope early
Flexible work expectations
Senior Infrastructure Support Engineer
Senior Infrastructure Support Engineer

Nscale • Seattle (WA)

On-site
USD 120,000 - 170,000
Remote-first culture
Equity plan
Flexible workplace
Senior Observability Platform Engineer
Senior Observability Platform Engineer

Nscale • United States

On-site
USD 160,000 - 230,000
Medical, dental, vision insurance
Flexible paid time off (PTO)
Parental leave
+1
Sr. Staff Security Engineer, Platform Security
Sr. Staff Security Engineer, Platform Security

Socket.dev • Houston (TX)

On-site
USD 210,000 - 250,000
Competitive salary + equity
Bonus + equity programs
Flexible workplace
Infrastructure Software Engineer, Fleet & Automation New Houston; New York; San Francisco; Seattle
Infrastructure Software Engineer, Fleet & Automation New Houston; New York; San Francisco; Seattle

Nscale • New York (NY)

On-site
USD 140,000 - 210,000
Competitive package
Equity
Growth opportunities
Senior Observability Platform Engineer
Senior Observability Platform Engineer

Socket.dev • United States

On-site
USD 160,000 - 230,000
Senior Staff Security Engineer, Incident Response
Senior Staff Security Engineer, Incident Response

Nscale • Houston (TX), New York (NY), San Francisco (CA), Seattle (WA)

On-site
USD 190,000 - 240,000
Base + bonus + equity
Flexible work environment
Dynamic progression plan
Sr. Staff Security Engineer, Platform Security
Sr. Staff Security Engineer, Platform Security

Nscale • Houston (TX), New York (NY), San Francisco (CA), Seattle (WA)

On-site
USD 210,000 - 250,000
Equity
Bonus
Healthcare