Member of Technical Staff, Site Reliability Engineer

Inferact

San Francisco (CA)

On-site

USD 200,000 - 400,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health benefits
401(k) match
Equity

Job summary

Inferact is hiring a Site Reliability Engineer in San Francisco to make vLLM-powered inference systems reliable, observable, and production-ready at scale.

You will define SLOs, improve monitoring, strengthen incident response, and drive post-mortems to reduce operational risk before it reaches users.

Qualifications

  • Bachelor's degree in CS, engineering, or equivalent
  • Experience operating production systems with meaningful traffic or criticality
  • Strong knowledge of SLOs, SLIs, error budgets, alerting, and post-mortem processes
  • Experience handling major production incidents with root cause analysis and prevention
  • Strong Linux, networking, observability, and distributed systems fundamentals
  • Ability to design simple, reliable systems and pre-empt failure modes
  • Proficiency in scripting languages for automation and reliability improvements

Responsibilities

  • Define SLOs and improve monitoring/alerting
  • Strengthen incident response and drive post-mortems
  • Reduce operational risk and improve production readiness at scale
  • Collaborate with engineering and infrastructure teams to improve service design and release safety
  • Lead incident reviews and prevention-oriented engineering efforts

Skills

SRE experience
Linux fundamentals
Observability
Python/Go
Automation scripting

Education

Bachelor's degree in CS or related

Tools

Kubernetes
Docker
Terraform
CI/CD
Cloud infrastructure

Job description

Overview

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.

About the Role

We're looking for a Site Reliability Engineer to help make vLLM-powered inference systems reliable, observable, and operationally simple at production scale. This role is for someone who thinks about failure before launch, designs systems that are easier to operate, and knows how to turn incidents into durable improvements rather than one-off fixes. You'll work across engineering and infrastructure to define SLOs, improve monitoring and alerting, strengthen incident response, drive post-mortems, and reduce operational risk before it reaches users. Your work will directly impact the reliability, availability, and production readiness of the systems powering AI inference at scale.

Skills and Qualifications

Minimum qualifications:

  • Bachelor's degree or equivalent experience in computer science, engineering, systems, infrastructure, or similar.
  • Strong experience operating production systems with meaningful traffic, user impact, or infrastructure criticality.
  • Deep understanding of SLOs, SLIs, error budgets, alerting, incident response, and post-mortem processes.
  • Experience live-fighting major production incidents, including mitigation, root cause analysis, escalation, and follow-through on prevention work.
  • Strong Linux, networking, systems debugging, observability, and distributed systems fundamentals.
  • Ability to design operationally simple systems and identify likely failure modes before launch.
  • Strong programming or scripting ability in Python, Go, Bash, or similar for automation, tooling, and reliability improvements.

Preferred qualifications:

  • Experience supporting ML infrastructure, inference systems, GPU workloads, Kubernetes-based platforms, or high-scale backend services.
  • Experience building or improving observability systems using metrics, logs, traces, dashboards, alerts, and runbooks.
  • Experience with Kubernetes, Docker, Terraform, cloud infrastructure, service meshes, CI/CD systems, or production deployment platforms.
  • Experience driving incident review culture, post-mortem processes, reliability reviews, and prevention-oriented engineering work.
  • Ability to partner with engineering teams to improve service design, release safety, capacity planning, and operational readiness.
Bonus points if you have:
  • Owned reliability for high-throughput, latency-sensitive, or mission-critical production systems.
  • Supported AI inference, model serving, GPU clusters, ML platforms, or distributed serving infrastructure.
  • Built automation that reduced toil, improved recovery time, or prevented repeat incidents.
  • Led incident response for severe outages with clear communication across engineering and leadership.
  • Created practical SLOs, dashboards, alerts, runbooks, or release gates that improved production reliability.
Logistics

Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.

Compensation

Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.

Visa sponsorship

We sponsor visas on a case-by-case basis.

Benefits

We offers generous health, dental, and vision benefits as well as 401(k) company match.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Site Reliability Engineer
Member of Technical Staff, Site Reliability Engineer

Linuxconfig • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Health benefits
Dental benefits
Vision benefits
+1
Head of Engineering
Head of Engineering

Inferact • San Francisco (CA)

On-site
USD 260,000 - 380,000
Health, dental, vision benefits
401(k) company match
Member of Technical Staff, Performance and Scale
Member of Technical Staff, Performance and Scale

Inferact • San Francisco (CA)

Hybrid
USD 200,000 - 400,000
Generous health, dental, and vision benefits
401(k) company match
Equity options
Member of Technical Staff, Cluster Administration
Member of Technical Staff, Cluster Administration

Inferact Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Health benefits
Dental benefits
Vision benefits
+1
Member of Technical Staff, Exceptional Generalist (Remote)
Member of Technical Staff, Exceptional Generalist (Remote)

Inferact • United States

Remote
USD 180,000 - 240,000
Competitive salary and equity
Visa sponsorship
Health coverage where applicable
Member of Technical Staff, Inference
Member of Technical Staff, Inference

Inferact • San Francisco (CA)

Hybrid
USD 200,000 - 400,000
Health, dental, and vision benefits
401(k) company match
Visa sponsorship on case-by-case basis
Site Reliability Engineer — AI Inference at Scale
Site Reliability Engineer — AI Inference at Scale

Linuxconfig • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Health benefits
Dental benefits
Vision benefits
+1
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Elorian • Palo Alto (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
Product Marketing Manager
Product Marketing Manager

Inferact • San Francisco (CA)

On-site
USD 120,000 - 160,000
Health benefits
Dental benefits
Vision benefits
+1