Site Reliability Engineer

Zof AI

San Francisco, Northern (CA, KY)

Hybrid

USD 150,000 - 210,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Zof AI in San Francisco is seeking a Site Reliability Engineer to own the execution layer of our control plane, including Kubernetes, containers, CI/CD, and isolation for untrusted code. You’ll manage scalable production infrastructure with a focus on security, reliability, and cost efficiency.

The role involves operating on-site in San Francisco, collaborating with engineers to ship safely and cost-effectively, and continuously improving observability and automation across large agent fleets.

Qualifications

  • Experience running production infrastructure on major cloud platforms.
  • Hands-on with Kubernetes, containers, and orchestration.
  • Experience building CI/CD pipelines and infrastructure as code.
  • Familiarity with observability, monitoring, and on-call practices.
  • Judgment about security, reliability, and cost trade-offs.
  • Daily use of AI tools to automate operational and engineering work.

Responsibilities

  • Design and operate sandboxed environments for agents to reproduce defects and validate fixes.
  • Own Kubernetes, container, and compute infrastructure end to end.
  • Build CI/CD pipelines enabling rapid and safe ship cycles.
  • Harden isolation boundaries to keep untrusted code inside its sandbox.
  • Instrument fleet health with metrics, logs, tracing, and alerts to catch failures early.
  • Drive down cost per agent run through scheduling, autoscaling, and capacity work.

Skills

Cloud platforms
Kubernetes
CI/CD pipelines
Infrastructure as code
Observability
Security awareness

Tools

Terraform/Pulumi
Monitoring/Logging
CI/CD tooling
gVisor
Firecracker

Job description

Zof AI is seeking a Site Reliability Engineer to run the infrastructure that lets fleets of sandboxed agents execute customer code safely and cheaply. This role owns the execution layer of our control plane: Kubernetes and container orchestration, CI/CD pipelines, hard isolation for untrusted code, observability, and the cost controls that keep large agent fleets affordable. If you have worked as a Site Reliability Engineer, Platform Engineer, Cloud Engineer, or Infrastructure Engineer, this is that discipline at Zof AI. The ideal candidate has operated production infrastructure at scale and treats security, reliability, and cost per agent run as constraints they personally own.

Engineering · Mid to Senior · Full-time · On-site · San Francisco, CA

Responsibilities
  • Design and operate the sandboxed environments where agents reproduce defects and validate fixes.
  • Own Kubernetes, container, and compute infrastructure end to end.
  • Build CI/CD pipelines that let engineers ship safely many times a day.
  • Harden isolation boundaries so untrusted customer code stays inside its sandbox.
  • Instrument fleet health with metrics, logs, tracing, and alerting that catch failures early.
  • Drive down cost per agent run through scheduling, autoscaling, and capacity work.
  • Automate provisioning, deployment, rollback, and environment management.
  • Partner with engineers to make infrastructure fast and safe to build on.
Requirements
  • Experience running production infrastructure on a major cloud platform.
  • Working knowledge of Kubernetes, containers, and orchestration.
  • Experience building CI/CD pipelines, deployment automation, and infrastructure as code.
  • Familiarity with observability, monitoring, and on-call practice.
  • Judgment about security, reliability, and cost trade-offs.
  • Daily use of AI tools to automate operational and engineering work.
  • Clear written and verbal communication.
  • Comfort operating in a fast-moving environment.
Nice to have
  • Experience with sandboxing or multi-tenant isolation tooling such as gVisor or Firecracker.
  • Experience with Terraform, Pulumi, or similar infrastructure as code tooling.
  • Experience running large batch or job-based workloads cost efficiently.
  • Experience in early-stage infrastructure or platform teams.

Must be able to run infrastructure for large agent workloads and use AI tools to automate operational work

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Product Security Engineer
Product Security Engineer

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
Platform SRE: Scalable, Secure Infra for Agent Fleets
Platform SRE: Scalable, Secure Infra for Agent Fleets

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Backend Developer
Backend Developer

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
Deployment Engineer
Deployment Engineer

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Full Stack Developer
Full Stack Developer

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 110,000 - 170,000
Software Engineer in Test
Software Engineer in Test

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 120,000 - 180,000
AI Infrastructure Engineer, Sandbox Platform
AI Infrastructure Engineer, Sandbox Platform

Scale AI • Seattle (WA), New York (NY), San Francisco (CA)

On-site
USD 180,000 - 225,000
Senior AI Software Engineer
Senior AI Software Engineer

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Analytics Engineer
Analytics Engineer

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 120,000 - 180,000
MLOps Engineer
MLOps Engineer

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 240,000