Stand out for this role — generate a tailored resume and cover letter in about a minute.
Inferact is hiring a Site Reliability Engineer in San Francisco to make vLLM-powered inference systems reliable, observable, and production-ready at scale.
You will define SLOs, improve monitoring, strengthen incident response, and drive post-mortems to reduce operational risk before it reaches users.
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.
We're looking for a Site Reliability Engineer to help make vLLM-powered inference systems reliable, observable, and operationally simple at production scale. This role is for someone who thinks about failure before launch, designs systems that are easier to operate, and knows how to turn incidents into durable improvements rather than one-off fixes. You'll work across engineering and infrastructure to define SLOs, improve monitoring and alerting, strengthen incident response, drive post-mortems, and reduce operational risk before it reaches users. Your work will directly impact the reliability, availability, and production readiness of the systems powering AI inference at scale.
Minimum qualifications:
Preferred qualifications:
Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.
Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.
We sponsor visas on a case-by-case basis.
We offers generous health, dental, and vision benefits as well as 401(k) company match.