Site Reliability Engineer

nscaleoperationsukltd

Houston (TX)

On-site

USD 130,000 - 200,000

Full time

5 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Competitive base plus equity
Real scope early
Flexible work style

Job summary

Nscale is seeking a career-level SRE to own automated tooling and reliable systems for AI workloads. You will sit in rotation, reduce toil, and push for quieter, more stable production services across Linux and networking.

This role emphasizes ownership, fast iteration, and measurable reliability improvements. You’ll partner with Engineering and Infrastructure to raise the reliability bar while delivering scalable, efficient platforms for AI-native workloads.

Qualifications

  • 3–6 years in SRE or systems engineering with production exposure.
  • Strong programming skills (Python/Go) and automation bias.
  • Solid Linux, networking, and distributed systems knowledge.
  • Proven incident RCA and postmortem experience.
  • Fluency in monitoring and observability.

Responsibilities

  • Build and own automation and tooling to keep the platform running.
  • Define and maintain SLOs, SLIs, and dashboards.
  • Lead incidents, perform root cause analysis, and drive post-incident reviews.
  • Troubleshoot performance and reliability across Linux, networks, and distributed services.
  • Collaborate with Engineering, Networking, and Infrastructure to raise reliability.
  • Improve availability, scalability, and efficiency through code.

Skills

SRE experience
Python/Go programming
Linux networking
Incident response
Observability
Automation mindset

Tools

Kubernetes
Bare-metal / virtualization

Job description

About Nscale

Nscale is the GPU cloud built for AI. We run high-performance, cost-efficient infrastructure for AI-native
startups and global enterprises, from bare metal up through the platform services teams actually build
on. Our culture runs on ownership, accountability, and speed. We move with urgency, we tell each
other the truth, and everyone here stays close to the infrastructure that makes AI work.



The Role

This is a career-level SRE role for someone who wants to own systems, not just watch them. You'll take
real surface area: the automation and tooling other engineers depend on, and the reliability of
production services running AI and GPU workloads at scale. You'll sit in the incident rotation, and you'll
be expected to make the systems you touch quieter over time.



What You'll Do


  • Build and own the automation and tooling that keeps the platform running; treat operational toil as
    a bug to be fixed, not a fact of life.

  • Define and maintain SLOs, SLIs, and the dashboards that make service health obvious at a glance.

  • Take point during incidents; troubleshoot under pressure, drive root cause analysis, and run post-
    incident reviews that actually change the system.

  • Investigate performance and reliability problems across Linux, networking, and distributed services,
    then fix them at the source.

  • Partner with Engineering, Networking, and Infrastructure teams to raise the reliability bar across the
    stack.

  • Improve availability, scalability, and efficiency through code, not manual effort.


What You'll Bring


  • 3-6 years in SRE, systems engineering, or software engineering, including time running production in
    a data center or cloud environment.

  • Strong programming skills (Python, Go, or similar) and a genuine bias toward automating the work
    away.

  • Solid command of Linux, networking fundamentals, and distributed systems.

  • A track record of troubleshooting live production issues and owning the fix through to the retro.

  • Fluency with monitoring and observability; metrics, logs, dashboards, and alerting.

  • Comfort in a fast-moving environment where priorities shift and you fill gaps without waiting to be
    asked.


Nice to Have


  • Experience with AI or GPU workloads, or high-performance computing (HPC).

  • Familiarity with high-performance networking (InfiniBand, RDMA).

  • Kubernetes, plus virtualized or bare-metal environments.


On-Call and Pace

A quick note on the shape of the job. This role sits close to production, so there is an on-call rotation,
and some weeks are busier than others. We share it fairly, and we treat every page as a signal worth
acting on rather than just an interruption. The goal is to make the systems quieter over time, so each
rotation asks less of the person carrying it. If you take ownership of what you run and like leaving it in
better shape than you found it, you'll do well here.



What We Offer


  • Competitive base plus equity, reviewed every 12 months.

  • Real scope early, and a progression plan built around the skills you want to sharpen.

  • Flexibility that treats you as an adult; we care that the work gets done, and we trust you to shape
    your day.


Salary Range

$130,000 - $200,000 USD. Actual compensation varies with skill set, experience, and location, and the
role may be eligible for bonus and equity.



Equal Opportunities Statement

At Nscale, we are committed to fostering an inclusive, diverse, and equitable workplace. We believe that a variety of perspectives enrich our work environment, and we encourage applications from candidates of all backgrounds, experiences, and abilities. We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.



If there's anything we can do to accommodate your specific situation, please let us know.
For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Nscale • Seattle (WA)

On-site
USD 130,000 - 200,000
Competitive base + equity
Real scope early
Flexible work style
Site Reliability Engineer New Houston; New York; San Francisco; Seattle
Site Reliability Engineer New Houston; New York; San Francisco; Seattle

Nscale • New York (NY), Northern (KY)

Hybrid
USD 130,000 - 200,000
Competitive base plus equity
Real scope early
Flexible work expectations
Senior Site Reliability Engineer -AI Infrastructure Operations
Senior Site Reliability Engineer -AI Infrastructure Operations

Nscale • Houston (TX)

On-site
USD 170,000 - 265,000
Competitive base + equity
Real ownership from the start
Flexible work culture
Senior Infrastructure Support Engineer
Senior Infrastructure Support Engineer

Nscale • Seattle (WA)

On-site
USD 120,000 - 170,000
Remote-first culture
Equity plan
Flexible workplace
Sr. Staff Security Engineer, Platform Security
Sr. Staff Security Engineer, Platform Security

Socket.dev • Houston (TX)

On-site
USD 210,000 - 250,000
Competitive salary + equity
Bonus + equity programs
Flexible workplace
Sr. Staff Security Engineer, Platform Security
Sr. Staff Security Engineer, Platform Security

Nscale • Houston (TX), New York (NY), San Francisco (CA), Seattle (WA)

On-site
USD 210,000 - 250,000
Equity
Bonus
Healthcare
Support Desk Engineer
Support Desk Engineer

Nscale • United States

Remote
USD 50,000 - 100,000
Competitive package with reviews every 12 months
Dynamic progression plan tailored to ambitions
Human-first flexibility in a remote-first environment
Infrastructure Software Engineer, Fleet & Automation New Houston; New York; San Francisco; Seattle
Infrastructure Software Engineer, Fleet & Automation New Houston; New York; San Francisco; Seattle

Nscale • New York (NY)

On-site
USD 140,000 - 210,000
Competitive package
Equity
Growth opportunities
Principal Network Engineer
Principal Network Engineer

Nscale • Seattle (WA), New York (NY), San Francisco (CA), Houston (TX)

On-site
USD 180,000 - 240,000
Base salary + equity
Equity incentives
Dynamic progression plan
Director of Operations
Director of Operations

Socket.dev • North Dakota

On-site
USD 180,000 - 250,000
Base + equity
Dynamic progression plan
Human-First Flexibility