Site Reliability Engineer

Hidden Jobs

United States

On-site

USD 150,000 - 190,000

Full time

13 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Healthcare
Retirement matching
Paid family leave
Fertility support
Mental health resources
Learning opportunities

Job summary

myfitnesspal is seeking a Senior Site Reliability Engineer to build and operate reliability foundations that enable rapid, safe software releases. You will work within a productivity engineering and self-service platform team, owning production reliability, incident response, observability, infrastructure automation, Kubernetes operations, deployment safety, and security controls in CI/CD.

You will define SLI/SLO, lead incident reviews, implement policy-as-code, and mentor engineers on

Qualifications

  • 5+ years in site reliability, platform, or infrastructure engineering with senior ownership of production systems.
  • Strong programming ability in Go, Python, TypeScript, or a similar language for automation and production tooling.
  • Hands-on experience with a major cloud platform, Kubernetes, and Infrastructure as Code; AWS and Terraform experience are useful.
  • Proven experience leading incident response and implementing SLO-driven reliability practices.
  • Working knowledge of observability tooling; experience with Datadog is useful.
  • Practical experience securing CI/CD through scanning, dependency controls, or policy-as-code.

Responsibilities

  • Define and evolve SLI, SLO, and error-budget practices, using reliability data to influence priorities and product decisions.
  • Lead incident response, facilitate post-incident reviews, and convert findings into durable system improvements.
  • Build and maintain observability across metrics, logs, and traces while improving signal quality and reducing alert fatigue.
  • Design and operate resilient infrastructure with Infrastructure as Code, including capacity planning and cloud-cost optimization.
  • Manage production Kubernetes and container workloads and support safe deployment methods such as canary releases, progressive rollouts, and rapid rollback.
  • Integrate and tune SAST, DAST, SCA, dependency scanning, and related security controls in delivery pipelines.
  • Implement policy-as-code to prevent unsafe infrastructure and Kubernetes changes at admission time.
  • Maintain vulnerability triage and remediation service levels, improve on-call sustainability, and coach engineers on operational practices.

Skills

Go
Python
TypeScript
Automation
Incident response

Tools

Datadog
AWS
Terraform
Kubernetes
CI/CD security tooling

Job description

Role overview

Build and operate the reliability, delivery, automation, and security foundations that allow product teams to release software quickly and safely. Working within a productivity engineering and self-service platform team, this role owns production reliability practices, incident response, observability, infrastructure automation, Kubernetes operations, deployment safety, and security controls embedded in CI/CD.

Responsibilities
  • Define and evolve SLI, SLO, and error-budget practices, using reliability data to influence priorities and product decisions.
  • Lead incident response, facilitate post-incident reviews, and convert findings into durable system improvements.
  • Build and maintain observability across metrics, logs, and traces while improving signal quality and reducing alert fatigue.
  • Design and operate resilient infrastructure with Infrastructure as Code, including capacity planning and cloud-cost optimization.
  • Manage production Kubernetes and container workloads and support safe deployment methods such as canary releases, progressive rollouts, and rapid rollback.
  • Integrate and tune SAST, DAST, SCA, dependency scanning, and related security controls in delivery pipelines.
  • Implement policy-as-code to prevent unsafe infrastructure and Kubernetes changes at admission time.
  • Maintain vulnerability triage and remediation service levels, improve on-call sustainability, and coach engineers on operational practices.
Requirements
  • 5+ years in site reliability, platform, or infrastructure engineering with senior ownership of production systems.
  • Strong programming ability in Go, Python, TypeScript, or a similar language for automation and production tooling.
  • Hands-on experience with a major cloud platform, Kubernetes, and Infrastructure as Code; AWS and Terraform experience are useful.
  • Proven experience leading incident response and implementing SLO-driven reliability practices.
  • Working knowledge of observability tooling; experience with Datadog is useful.
  • Practical experience securing CI/CD through scanning, dependency controls, or policy-as-code.
  • Understanding of cloud security fundamentals, including IAM, least privilege, guardrails, and secrets management.
  • Strong judgment and communication skills when raising reliability or security issues across engineering teams.
Nice to have
  • Experience with policy-as-code frameworks such as OPA/Rego, Kyverno, or Conftest.
Benefits and work setup

The source describes comprehensive healthcare, retirement savings with employer matching, paid family leave, fertility support, mental health resources, wellness and technology allowances, performance-related rewards, learning opportunities, mentorship, and an inclusive workplace culture.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Site Reliability Engineer
Principal Site Reliability Engineer

Gen Digital Inc. • United States

Remote
USD 180,000 - 240,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Clearwater Analytics • Boise (ID)

On-site
USD 130,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 150,000 - 210,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Engg • Tempe (AZ)

On-site
USD 140,000 - 190,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink • United States

On-site
USD 140,000 - 190,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink, Inc. • United States

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Kovoro • Denver (CO), Northern (KY)

On-site
USD 150,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Ethos Group • Irving (TX)

On-site
USD 110,000 - 160,000
Platform Site Reliability Engineer
Platform Site Reliability Engineer

Specter • San Francisco (CA)

On-site
USD 180,000 - 230,000