Site Reliability Engineer — AI Infra & Equity

Zof AI, Inc.

San Francisco (CA)

On-site

USD 150,000 - 200,000

Full time

46 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Competitive salary
Meaningful equity

Job summary

Zof AI, Inc. in San Francisco, CA is hiring a Site Reliability Engineer to run the infrastructure for large agent workloads and safe, cost-effective operations.

This is a full-time on-site role reporting to the Infrastructure team, focused on Kubernetes, containers, CI/CD, and isolating untrusted code within sandboxed environments. The ideal candidate will have production-scale experience, strong security and reliability judgment, and a proactive approach to observability, cost control, and

Qualifications

  • Experience with production infrastructure
  • Strong knowledge of Kubernetes, containers, and orchestration
  • Experience building CI/CD pipelines, deployment automation, and infrastructure as code
  • Familiarity with observability, monitoring, and on-call practice
  • Judgment about security, reliability, and cost trade-offs
  • Daily use of AI tools to automate operational and engineering work
  • Clear written and verbal communication
  • Comfort operating in a fast-moving environment

Responsibilities

  • Design and operate the sandboxed environments where agents reproduce defects and validate fixes.
  • Own Kubernetes, container, and compute infrastructure end to end.
  • Build CI/CD pipelines that let engineers ship safely many times a day.
  • Harden isolation boundaries so untrusted customer code stays inside its sandbox.
  • Instrument fleet health with metrics, logs, tracing, and alerting that catch failures early.
  • Drive down cost per agent run through scheduling, autoscaling, and capacity work.
  • Automate provisioning, deployment, rollback, and environment management.
  • Partner with engineers to make infrastructure fast and safe to build on.

Skills

Communication
Security trade-offs
AI tooling automation
Fast-paced environment adaptation

Tools

Kubernetes
Containers
CI/CD pipelines
IaC (Terraform/Pulumi)
Observability tooling

Job description

Zof AI, Inc. in San Francisco, CA is hiring a Site Reliability Engineer to run the infrastructure for large agent workloads and safe, cost-effective operations.

This is a full-time on-site role reporting to the Infrastructure team, focused on Kubernetes, containers, CI/CD, and isolating untrusted code within sandboxed environments. The ideal candidate will have production-scale experience, strong security and reliability judgment, and a proactive approach to observability, cost control, and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Zof AI • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 210,000
Platform SRE: Scalable, Secure Infra for Agent Fleets
Platform SRE: Scalable, Secure Infra for Agent Fleets

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Site Reliability Engineer, AI Cloud Infrastructure
Site Reliability Engineer, AI Cloud Infrastructure

Anyscale • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Zof AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 200,000
Competitive salary
Meaningful equity
Site Reliability Engineer — ML Infra, Scale & Equity
Site Reliability Engineer — ML Infra, Scale & Equity

Baseten • New York (NY)

On-site
USD 165,000 - 330,000
Competitive compensation
100% coverage of medical, dental, and vision insurance
Generous PTO policy
+3
Site Reliability Engineer — ML Infra & Observability
Site Reliability Engineer — ML Infra & Observability

Baseten • San Francisco (CA)

On-site
USD 135,000 - 285,000
Competitive compensation including equity
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
+3
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineer — Scale AI Infra with Ownership
Site Reliability Engineer — Scale AI Infra with Ownership

Happyrobot Inc. • San Francisco (CA)

On-site
USD 100,000 - 140,000
Competitive salary + equity
Ownership & autonomy in projects
Opportunity to work with top-tier engineers
Site Reliability Engineer — Build Scalable AI Infra
Site Reliability Engineer — Build Scalable AI Infra

Future Secure AI Pty • Austin (TX)

On-site
USD 100,000 - 140,000
Flexible work environment
Competitive salary
Growth trajectory
Staff Site Reliability Engineer — AI-Driven Zero Trust
Staff Site Reliability Engineer — AI-Driven Zero Trust

Zscaler • San Jose (CA)

Hybrid
USD 119,000 - 170,000
Health plans
Vacation & sick time
Parental leave
+3