An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Zof AI in San Francisco is seeking a Site Reliability Engineer to own the execution layer of our control plane, including Kubernetes, containers, CI/CD, and isolation for untrusted code. You’ll manage scalable production infrastructure with a focus on security, reliability, and cost efficiency.
The role involves operating on-site in San Francisco, collaborating with engineers to ship safely and cost-effectively, and continuously improving observability and automation across large agent fleets.
Zof AI is seeking a Site Reliability Engineer to run the infrastructure that lets fleets of sandboxed agents execute customer code safely and cheaply. This role owns the execution layer of our control plane: Kubernetes and container orchestration, CI/CD pipelines, hard isolation for untrusted code, observability, and the cost controls that keep large agent fleets affordable. If you have worked as a Site Reliability Engineer, Platform Engineer, Cloud Engineer, or Infrastructure Engineer, this is that discipline at Zof AI. The ideal candidate has operated production infrastructure at scale and treats security, reliability, and cost per agent run as constraints they personally own.
Engineering · Mid to Senior · Full-time · On-site · San Francisco, CA
Must be able to run infrastructure for large agent workloads and use AI tools to automate operational work