Site Reliability Engineer — AI Inference at Scale

Linuxconfig

San Francisco, Northern (CA, KY)

Hybrid

USD 200,000 - 400,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Health benefits
Dental benefits
Vision benefits
401(k) match

Job summary

Inferact is hiring a Site Reliability Engineer to make vLLM-powered inference systems reliable, observable, and operationally simple at production scale.

You will define SLOs, improve monitoring and alerting, strengthen incident response, drive post-mortems, and reduce operational risk before it reaches users. Your work will directly impact reliability and production readiness of AI inference systems.

Qualifications

  • Bachelor's degree or equivalent experience in computer science, engineering, systems, infrastructure, or similar.
  • Strong experience operating production systems with meaningful traffic, user impact, or infrastructure criticality.
  • Deep understanding of SLOs, SLIs, error budgets, alerting, incident response, and post-mortem processes.
  • Experience live-fighting major production incidents, including mitigation, root cause analysis, escalation, and follow-through on prevention work.
  • Strong Linux, networking, systems debugging, observability, and distributed systems fundamentals.
  • Ability to design operationally simple systems and identify likely failure modes before launch.
  • Strong programming or scripting ability in Python, Go, Bash, or similar for automation, tooling, and reliability improvements.

Responsibilities

  • Define SLOs and SLIs for production systems.
  • Improve monitoring, alerting, and incident response processes.
  • Drive post-mortems and prevention-oriented engineering work.
  • Reduce operational risk before it reaches users and improve system reliability at scale.
  • Collaborate with engineering to improve service design, release safety, and capacity planning.

Skills

Linux
Networking
Observability
Distributed systems
Python/Go/Bash
Incident response

Education

Bachelor's degree or equivalent

Tools

Kubernetes
Docker
Terraform
CI/CD

Job description

Inferact is hiring a Site Reliability Engineer to make vLLM-powered inference systems reliable, observable, and operationally simple at production scale.

You will define SLOs, improve monitoring and alerting, strengthen incident response, drive post-mortems, and reduce operational risk before it reaches users. Your work will directly impact reliability and production readiness of AI inference systems.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Site Reliability Engineer
Member of Technical Staff, Site Reliability Engineer

Linuxconfig • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Health benefits
Dental benefits
Vision benefits
+1
Member of Technical Staff, Site Reliability Engineer
Member of Technical Staff, Site Reliability Engineer

Inferact • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health benefits
401(k) match
Equity
Head of ML Systems & Inference
Head of ML Systems & Inference

Doist • San Francisco (CA)

On-site
USD 260,000 - 380,000
Health, dental, vision benefits
401(k) company match
Member of Technical Staff, Performance and Scale
Member of Technical Staff, Performance and Scale

Inferact • San Francisco (CA)

Hybrid
USD 200,000 - 400,000
Generous health, dental, and vision benefits
401(k) company match
Equity options
Site Reliability Engineer — ML Infra, Scale & Equity
Site Reliability Engineer — ML Infra, Scale & Equity

Baseten • New York (NY)

On-site
USD 165,000 - 330,000
Head of Engineering
Head of Engineering

Inferact • San Francisco (CA)

On-site
USD 260,000 - 380,000
Health, dental, vision benefits
401(k) company match
Staff Engineer - Customer-Facing AI Inference Infra
Staff Engineer - Customer-Facing AI Inference Infra

Simplify • San Francisco (CA)

On-site
USD 200,000 - 300,000
Housing stipend
Uber/Waymo rides
ML Inference Engineer — Scale & Reliability
ML Inference Engineer — Scale & Reliability

Clera • San Mateo (CA)

On-site
USD 180,000 - 240,000
Site Reliability Engineer, AI Infra & Observability
Site Reliability Engineer, AI Infra & Observability

Sierra • California (MO)

On-site
USD 150,000 - 210,000
Unlimited PTO
Medical, dental, vision
Retirement plan
+4
SRE for AI Platform & ML Inference Infra
SRE for AI Platform & ML Inference Infra

Cohere • New York (NY)

Hybrid
USD 140,000 - 200,000
Weekly lunch stipend
Health and dental benefits
RRSP matching
+3