Foundational SRE: Build Reliability for SaaS Platform

Forward

Santa Clara (CA)

On-site

USD 230,000 - 250,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Forward is seeking its first dedicated SRE to build the reliability function for a distributed SaaS platform. You will define SLOs/SLIs, drive incident response, and partner with engineering, infrastructure, and product to meet enterprise reliability standards.

If you thrive when handed challenging problems rather than following a playbook, this foundational role offers a path to leadership as Forward scales its platform and observability capabilities across cloud environments.

Qualifications

  • 6+ years of site reliability engineering, DevOps, or infrastructure engineering in a SaaS or cloud environment.
  • Proven experience building or maturing an SRE function.
  • Strong networking fundamentals (TCP/IP, DNS, routing, load balancing).
  • Hands-on experience with Kubernetes and container orchestration.
  • Deep proficiency with observability tooling.
  • Scripting and automation skills in Python, Bash, or similar.
  • Experience with cloud platforms and infrastructure as code.

Responsibilities

  • Define and drive SRE practices from the ground up — SLOs, SLIs, error budgets, and the frameworks the engineering org will actually use.
  • Drive the reliability and operational excellence of the Forward SaaS platform.
  • Build and maintain observability infrastructure — logging, metrics, tracing, and alerting.
  • Lead incident response: on-call rotations, runbooks, post-mortems, and follow-through to prevent repeats.
  • Partner with engineering teams to embed reliability thinking into the SDLC — capacity planning and production readiness reviews.
  • Help define and build the SRE team as the company scales — foundational hire with a path to leadership.

Skills

SRE experience
Networking fundamentals
Scripting
Incident response
Communication

Tools

Kubernetes
Prometheus
Grafana
Datadog
Splunk
Python
Bash
Terraform
Ansible

Job description

Forward is seeking its first dedicated SRE to build the reliability function for a distributed SaaS platform. You will define SLOs/SLIs, drive incident response, and partner with engineering, infrastructure, and product to meet enterprise reliability standards.

If you thrive when handed challenging problems rather than following a playbook, this foundational role offers a path to leadership as Forward scales its platform and observability capabilities across cloud environments.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Forward • Santa Clara (CA)

On-site
USD 230,000 - 250,000
Founding SRE — Build Reliable Platform & Observability
Founding SRE — Build Reliable Platform & Observability

Legora • Town of Stockholm (NY)

On-site
USD 71,000 - 101,000
Founding SRE — Build Reliable, Scalable Systems
Founding SRE — Build Reliable, Scalable Systems

Incident IQ • Atlanta (GA)

On-site
USD 120,000 - 190,000
Medical benefits
Dental benefits
Vision benefits
+3
Remote Staff SRE: Reliability, Incident Recovery & Growth
Remote Staff SRE: Reliability, Incident Recovery & Growth

Fingerprint • United States

Remote
USD 150,000 - 210,000
Senior Platform SRE: Reliability Lead for Greenfield PaaS
Senior Platform SRE: Reliability Lead for Greenfield PaaS

Hidden Jobs • United States

Remote
USD 150,000 - 190,000
Site Reliability Engineering Lead
Site Reliability Engineering Lead

oneapp • United States

Remote
USD 180,000 - 260,000
Stock options
Health benefits from day one
401(k) with company match
+1
Founding SRE: Build Reliability from Day One
Founding SRE: Build Reliability from Day One

Incident IQ • Alpharetta (GA)

On-site
USD 120,000 - 180,000
Medical, Dental, Vision
401k Match
PTO
+1
Senior Site Reliability Engineer (Remote)
Senior Site Reliability Engineer (Remote)

Fathom.ai • United States

On-site
USD 100,000 - 130,000
Competitive compensation
Supportive environment for personal growth
Dynamic and collaborative team
Senior SRE Lead: Reliability & Cloud Platforms
Senior SRE Lead: Reliability & Cloud Platforms

Shield AI • San Mateo (CA)

On-site
USD 220,000 - 340,000
Bonus
Benefits
Equity
SRE Lead: Reliability & Cloud Observability Architect
SRE Lead: Reliability & Cloud Observability Architect

BlackCube Labs • San Diego (CA)

On-site
USD 190,000 - 280,000