Head of Site Reliability Engineering

Forge

San Francisco (CA)

On-site

USD 150,000 - 220,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Forge is seeking a Manager, Site Reliability Engineering to lead a growing SRE team, elevating reliability, observability, and incident response across our cloud-based platform. You will coach engineers, shape production operations, and partner with Platform, Security, Compliance, and Product teams to deliver secure, scalable services.

Ideal candidates have 5+ years leading SRE/DevOps, 10+ years in software or production operations, and strong communications across engineering and

Qualifications

  • 5+ years leading an SRE/DevOps or similar reliability function.
  • 10+ years total software engineering, infrastructure, or production operations.
  • Bachelor's degree in Computer Science, Engineering, or related field (or equivalent).
  • Experience building and operating large-scale cloud infrastructure and distributed systems.
  • Hands-on with observability, monitoring, alerting, incident response, troubleshooting, and production support.
  • Experience with CI/CD, infrastructure automation, cloud platforms, and ops tooling.
  • Strong technical judgment and ability to influence across teams.

Responsibilities

  • Lead Forge's Site Reliability Engineering team to maintain high availability for customers.
  • Drive incident management practices in partnership with engineering teams.
  • Build and manage observability infrastructure including monitoring, alerting, dashboards, and metrics.
  • Improve monitoring coverage and alert quality to reduce noise and speed detection.
  • Champion reliability best practices: service ownership, disaster recovery, production readiness.
  • Contribute to design, architecture, automation, and infrastructure improvements.
  • Collaborate with engineering to troubleshoot production issues and improve system reliability.
  • Hire, coach, and manage SRE team performance and career development.
  • Partner with Security, Compliance, and Risk to meet regulatory needs.

Skills

SRE leadership
Cloud infrastructure
Observability
CI/CD
Incident response

Education

Bachelor's degree in CS/Engineering

Tools

Kubernetes
Terraform
Ansible
Datadog
CloudWatch

Job description

Forge is seeking a Manager, Site Reliability Engineering to lead a growing SRE team, elevating reliability, observability, and incident response across our cloud-based platform. You will coach engineers, shape production operations, and partner with Platform, Security, Compliance, and Product teams to deliver secure, scalable services.

Ideal candidates have 5+ years leading SRE/DevOps, 10+ years in software or production operations, and strong communications across engineering and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineering Manager – Scale & Resilience
Senior Site Reliability Engineering Manager – Scale & Resilience

Socket.dev • New York (NY)

On-site
USD 150,000 - 220,000
SRE Engineering Manager: Reliability & Scale
SRE Engineering Manager: Reliability & Scale

Socket.dev • San Francisco (CA)

On-site
USD 150,000 - 220,000
Manager, Site Reliability Engineer
Manager, Site Reliability Engineer

Forge • San Francisco (CA)

On-site
USD 150,000 - 220,000
Manager, Site Reliability Engineer
Manager, Site Reliability Engineer

Socket.dev • San Francisco (CA)

On-site
USD 150,000 - 220,000
Manager, Site Reliability Engineer
Manager, Site Reliability Engineer

Socket.dev • New York (NY)

On-site
USD 150,000 - 220,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 240,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

New York Technology Partners • Chicago (IL)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

Hybrid
USD 140,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior SRE: Scale Infra, Automate, Elevate Reliability
Senior SRE: Scale Infra, Automate, Elevate Reliability

Fathom.ai • United States

On-site
USD 100,000 - 130,000
Competitive compensation
Supportive environment for personal growth
Dynamic and collaborative team