Senior Site Reliability Engineer — Kubernetes & AI-Driven Ops

Fal

San Francisco (CA)

On-site

USD 180,000 - 250,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Relocation assistance to San Francisco
Visa sponsorship
Competitive salary and equity
Health, dental, and vision insurance (

Job summary

fal is seeking a Software Engineer, Site Reliability in San Francisco. You will keep production infrastructure running at scale, owning the reliability of customer-facing systems—from Kubernetes clusters to deployment pipelines and the networking layer that connects it all.

You think in SLOs, automate ruthlessly, and treat every incident as a chance to improve the system. This is a full-time, onsite role with potential remote flexibility for senior levels.

Qualifications

  • 5+ years experience managing production systems and software development workflows.
  • Strong production experience setting up and operating Kubernetes at scale using IaC tools.
  • Deep knowledge of Linux networking, container networking, DNS.
  • Experience building CI/CD systems and GitOps workflows (FluxCD, ArgoCD).
  • Proficiency in Python and Go or Bash for tooling and automation.

Responsibilities

  • Own and operate our Kubernetes infrastructure: cluster lifecycle, upgrades, networking, and multi-tenant isolation for customer workloads.
  • Build and maintain CI/CD pipelines and deployment infrastructure.
  • Leverage AI to automate analysis and resolution of production issues, improving speed, reliability and maintainability.
  • Build dashboards, alerting, and anomaly detection across systems.
  • Define and enforce SLOs and build incident response processes.
  • Manage and improve networking, load balancing, and service mesh configurations.
  • Drive reliability improvements across the stack through automation, runbooks, and chaos engineering.

Skills

Kubernetes
CI/CD
Python
Go
Bash
Linux networking
Terraform
Ansible
Prometheus
Grafana
Loki
Datadog

Tools

FluxCD
ArgoCD
Terraform
Ansible

Job description

fal is seeking a Software Engineer, Site Reliability in San Francisco. You will keep production infrastructure running at scale, owning the reliability of customer-facing systems—from Kubernetes clusters to deployment pipelines and the networking layer that connects it all.

You think in SLOs, automate ruthlessly, and treat every incident as a chance to improve the system. This is a full-time, onsite role with potential remote flexibility for senior levels.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Site Reliability
Software Engineer, Site Reliability

fal • San Francisco (CA)

On-site
USD 180,000 - 250,000
Health, dental, and vision insurance
Relocation assistance
Learning and growth opportunities
+1
Site Reliability Engineer: AI-Driven Kubernetes Automation
Site Reliability Engineer: AI-Driven Kubernetes Automation

Kindredventures • United States

Remote
USD 8,000 - 15,000
Senior Kubernetes SRE — Scale AI Clusters & Automation
Senior Kubernetes SRE — Scale AI Clusters & Automation

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Equity compensation
Health, dental and vision coverage
Wellness and commuter stipends
+2
Software Engineer, Site Reliability
Software Engineer, Site Reliability

Fal • San Francisco (CA)

On-site
USD 180,000 - 250,000
Relocation assistance to San Francisco
Visa sponsorship
Competitive salary and equity
+1
Senior Site Reliability Engineer - AI Cloud Platform
Senior Site Reliability Engineer - AI Cloud Platform

Lambda • United States

Hybrid
USD 160,000 - 220,000
Senior SRE: AI Cloud Platform & Kubernetes Expert
Senior SRE: AI Cloud Platform & Kubernetes Expert

Lambda • Bellevue (WA)

On-site
USD 180,000 - 260,000
Health insurance
Dental insurance
Vision insurance
+3
Senior Site Reliability Engineer – Scalable AI Infra
Senior Site Reliability Engineer – Scalable AI Infra

Tavily Inc. • New York (NY)

Hybrid
USD 156,000 - 262,000
100% company-paid medical, dental, and vision coverage
Up to 4% company match 401(k) plan
20 weeks paid parental leave for primary caregivers
+2
Senior Site Reliability Engineer — AI Cloud Infra
Senior Site Reliability Engineer — AI Cloud Infra

Anyscale • San Francisco (CA)

On-site
USD 130,000 - 180,000
Senior Site Reliability Engineer – AI-First Infra
Senior Site Reliability Engineer – AI-First Infra

Evidently Ltd. • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 250,000
Senior SRE – AI Cloud Platform, Kubernetes Expert
Senior SRE – AI Cloud Platform, Kubernetes Expert

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health, dental, vision coverage for in
Wellness and commuter stipends
401k with 2% company match
+1