Software Engineer, Site Reliability

fal

San Francisco (CA)

On-site

USD 180,000 - 250,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health, dental, and vision insurance
Relocation assistance
Learning and growth opportunities
Regular team events

Job summary

fal is hiring an experienced Site Reliability Engineer (SRE) in San Francisco, CA, to manage Kubernetes infrastructure and improve production systems. The role requires 5+ years of experience and strong skills in automation using Python and infrastructure-as-code tools. Compensation ranges from $180,000 to $250,000, including equity and benefits, along with growth opportunities and relocation assistance. The company promotes a collaborative environment with regular team events.

Qualifications

  • 5+ years experience in managing production systems.
  • Strong experience with Kubernetes at scale.
  • Proficiency in Python and either Go or Bash.

Responsibilities

  • Own and operate Kubernetes infrastructure.
  • Build and maintain CI/CD pipelines.
  • Leverage AI for automating issue resolution.

Skills

Kubernetes management
Infrastructure-as-code (Terraform, Ansible)
Linux networking
CI/CD systems
Python
Monitoring and alerting tools (Prometheus, Grafana)
Communication skills

Tools

Prometheus
Grafana
Terraform
Ansible

Job description

You are a seasoned SRE who keeps production infrastructure running at scale. You own the reliability and availability of customer-facing systems — from Kubernetes clusters to deployment pipelines to the networking layer that connects it all. You think in SLOs, automate ruthlessly, and treat every incident as a chance to make the system better.

Key Responsibilities
  • Own and operate our Kubernetes infrastructure: cluster lifecycle, upgrades, networking, and multi-tenant isolation for customer workloads
  • Build and maintain CI/CD pipelines and deployment infrastructure
  • Leverage AI to an extreme level to automate analysis and resolution of production issues, and improve software development speed, reliability and maintainability
  • Build dashboards, alerting, and anomaly detection across our systems
  • Define and enforce SLOs and build out incident response processes
  • Manage and improve our networking, load balancing, and service mesh configurations
  • Drive reliability improvements across the stack through automation, runbooks, and chaos engineering
Requirements
  • 5+ years experience in managing critical production systems and software development workflows
  • Strong production experience setting up and operating Kubernetes at scale, using infrastructure-as-code (Terraform, Ansible)
  • Deep knowledge of Linux networking, container networking (CNI plugins, VXLAN, BGP), and DNS
  • Experience building CI/CD systems and GitOps workflows (FluxCD, ArgoCD)
  • Proficiency in Python and either Go or Bash for tooling and automation
  • Strong experience with logging, monitoring and alerting (Prometheus, Grafana, Loki, Thanos, VictoriaMetrics, Datadog)
  • Excellent communication and ability to drive technical decisions across teams
  • Self-starter who executes quickly, takes ownership, and constantly seeks improvement
Nice to have
  • Experience with managing GPU and AI/ML workloads
  • Experience with kernel-based monitoring and routing (eBPF, XDP)
  • Experience with security tooling (Falco, Coroot, SIEM)
  • Experience with bare metal Kubernetes networking (Calico, Cilium, MetalLB)
  • Experience with distributed storage systems (Ceph, Longhorn, etc.)
Compensation
  • $180,000-250,000 plus equity + benefits (Range is based across 3 levels MId, Senior and Staff)
Location

San Francisco, CA (willing to consider remote for Senior and Staff levels)

What we offer at fal
  • Interesting and challenging work
  • A lot of learning and growth opportunities
  • We are currently hiring in downtown San Francisco.
  • We offer relocation assistance to San Francisco.
  • Health, dental, and vision insurance (US)

Regular team events and offsites

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Site Reliability
Software Engineer, Site Reliability

Fal • San Francisco (CA)

On-site
USD 180,000 - 250,000
Relocation assistance to San Francisco
Visa sponsorship
Competitive salary and equity
+1
Member of Technical Staff, DevOps
Member of Technical Staff, DevOps

Reactor • San Francisco (CA)

On-site
USD 100,000 - 160,000
Competitive salary and early equity
Visa sponsorship
Generous health, dental, and vision coverage
Senior Site Reliability Engineer — Kubernetes & AI-Driven Ops
Senior Site Reliability Engineer — Kubernetes & AI-Driven Ops

Fal • San Francisco (CA)

On-site
USD 180,000 - 250,000
Relocation assistance to San Francisco
Visa sponsorship
Competitive salary and equity
+1
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • New York (NY)

Hybrid
USD 165,000 - 215,000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+1
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • New Jersey

On-site
USD 165,000 - 215,000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Senior Forward Deployed Engineer (DevOps/SRE)
Senior Forward Deployed Engineer (DevOps/SRE)

LeoForce • Pleasanton (CA)

On-site
USD 300,000 - 350,000
Medical benefits
401(k) plan
Equity
+1
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • North Carolina

On-site
USD 165,000 - 215,000
Pre‑IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Senior DevOps Engineer
Senior DevOps Engineer

Newton Research • Boston (MA)

On-site
USD 150,000 - 175,000
Equity
Competitive salary
Benefits
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The Recruiting Guy • Washington

On-site
USD 175,000 - 250,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000