Site Reliability Engineer

Forward

Santa Clara (CA)

On-site

USD 230,000 - 250,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Forward is seeking its first dedicated SRE to build the reliability function for a distributed SaaS platform. You will define SLOs/SLIs, drive incident response, and partner with engineering, infrastructure, and product to meet enterprise reliability standards.

If you thrive when handed challenging problems rather than following a playbook, this foundational role offers a path to leadership as Forward scales its platform and observability capabilities across cloud environments.

Qualifications

  • 6+ years of site reliability engineering, DevOps, or infrastructure engineering in a SaaS or cloud environment.
  • Proven experience building or maturing an SRE function.
  • Strong networking fundamentals (TCP/IP, DNS, routing, load balancing).
  • Hands-on experience with Kubernetes and container orchestration.
  • Deep proficiency with observability tooling.
  • Scripting and automation skills in Python, Bash, or similar.
  • Experience with cloud platforms and infrastructure as code.

Responsibilities

  • Define and drive SRE practices from the ground up — SLOs, SLIs, error budgets, and the frameworks the engineering org will actually use.
  • Drive the reliability and operational excellence of the Forward SaaS platform.
  • Build and maintain observability infrastructure — logging, metrics, tracing, and alerting.
  • Lead incident response: on-call rotations, runbooks, post-mortems, and follow-through to prevent repeats.
  • Partner with engineering teams to embed reliability thinking into the SDLC — capacity planning and production readiness reviews.
  • Help define and build the SRE team as the company scales — foundational hire with a path to leadership.

Skills

SRE experience
Networking fundamentals
Scripting
Incident response
Communication

Tools

Kubernetes
Prometheus
Grafana
Datadog
Splunk
Python
Bash
Terraform
Ansible

Job description

Forward is transforming how the world’s most complex networks are managed and secured. Founded in 2013 by four Stanford Ph.D.s, we built the industry’s first network digital twin — a mathematically precise model of the production network that gives IT teams unmatched visibility, verification, and agility across every major cloud and vendor environment.

Our customers include global leaders such as Goldman Sachs, PayPal, S&P Global, IBM, and Dell, as well as fast‑growing enterprises and government agencies. According to IDC, Forward customers realize an average of $14.2 million in annual benefits through improved efficiency and security.

Backed by world‑class investors including Andreessen Horowitz, Goldman Sachs, MSD Partners, and Threshold Ventures, Forward offers a people‑centric, innovative culture where brilliant minds are shaping the future of network reliability, security, and AI‑ready operations.

About the Role

This is not a "keep the lights on" SRE role. As our first or early SRE hire you will be building the reliability engineering function at Forward — defining how we think about availability, observability, incident response, and operational excellence across a complex, distributed SaaS platform. You will work closely with engineering, infrastructure, and product to ensure our platform meets the reliability bar our enterprise customers demand.

If you thrive in environments where you’re handed a problem rather than a playbook this role is for you.

What You'll Own
  • Define and drive SRE practices from the ground up — SLOs, SLIs, error budgets, and the frameworks the engineering org will actually use
  • Drive the reliability and operational excellence of the Forward SaaS platform
  • Build and maintain observability infrastructure — logging, metrics, tracing, and alerting — so the team always knows what's happening before customers do
  • Lead incident response: on‑call rotations, runbooks, post‑mortems, and the follow‑through to make sure the same incident doesn't happen twice
  • Partner with engineering teams to embed reliability thinking into the SDLC — capacity planning, load testing, chaos engineering, and production readiness reviews
  • Help define and build the SRE team as the company scales — this is a foundational hire with a path to leadership
What We're Looking For
  • 6+ years of experience in site reliability engineering, DevOps, or infrastructure engineering in a SaaS or cloud environment
  • Proven experience building or significantly maturing an SRE function — not just operating within one someone else built
  • Strong fundamentals in networking — TCP/IP, DNS, routing, switching, firewalls, and load balancing. Experience with network management or observability platforms is a significant plus
  • Hands‑on experience with Kubernetes and container orchestration in production environments
  • Deep proficiency with observability tooling — Prometheus, Grafana, Datadog, Splunk, or similar
  • Strong scripting and automation skills in Python, Bash, or similar
  • Experience with cloud platforms — AWS, GCP, or Azure — including infrastructure as code (Terraform, Ansible, or equivalent)
  • Track record of owning and improving incident response processes including blameless post‑mortems and SLO‑driven reliability improvements
  • Ability to communicate clearly with both engineering teams and non‑technical stakeholders — you can explain an outage to a customer‑facing team without jargon and explain an SLO to an executive without losing them
Nice to Have
  • Experience supporting enterprise or federal government customers with high availability requirements
  • Experience in a foundational or early SRE hire capacity at a growth stage company
What This Role Is Not
  • A pure ops or NOC role — you are building and engineering, not just monitoring
  • A siloed function — you will be deeply embedded with product and engineering teams
  • A ticket‑taker — you will be proactively identifying and solving reliability problems before they become incidents

The base pay range for this role is between $230,000 and $250,000. Base pay will depend on your skills, qualifications, experience, and location.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Foundational SRE: Build Reliability for SaaS Platform
Foundational SRE: Build Reliability for SaaS Platform

Forward • Santa Clara (CA)

On-site
USD 230,000 - 250,000
Site Reliability Engineering Lead
Site Reliability Engineering Lead

oneapp • United States

Remote
USD 180,000 - 260,000
Stock options
Health benefits from day one
401(k) with company match
+1
Site Reliability Engineer
Site Reliability Engineer

Stelvio Inc. • Town of Texas (WI)

On-site
USD 125,000 - 145,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

United States Digital Space LLC • Charlotte (TX)

On-site
USD 153,000 - 192,000
Discretionary incentive eligible
Benefits package
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Practice by Numbers • United States

On-site
USD 120,000 - 160,000
High ownership and autonomy
Strong engineering culture
Impactful work on healthcare infrastructure
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hobbsnews • Jersey City (NJ)

On-site
USD 153,000 - 192,000
Benefits eligible
Discretionary incentive plan
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
Senior SRE, Software Engineering (AWS / Scaling Infrastructure)
Senior SRE, Software Engineering (AWS / Scaling Infrastructure)

PulseRise Technologies • New York (NY)

On-site
USD 130,000 - 160,000
Site Reliability Engineer Engineer
Site Reliability Engineer Engineer

Modus Create • Aurora (IL)

Remote
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

National Black MBA Association • Jersey City (NJ)

On-site
USD 153,000 - 192,000
Benefits eligible
Annual discretionary plan