Site Reliability Engineering Lead

oneapp

United States

Remote

USD 180,000 - 260,000

Full time

13 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Stock options
Health benefits from day one
401(k) with company match
Fully remote within the United States

Job summary

oneapp is seeking a Senior Site Reliability Engineer Leader to guide a small, senior SRE team while maintaining hands-on coding across Python or Go. You will define service ownership, observability, and incident practices to raise the reliability of customer-facing systems.

Responsibilities include leading the team, writing production code, building automation, and partnering across engineering to elevate reliability.

Qualifications

  • Strong software engineering foundation with production-quality coding in Python or Go.
  • Experience with large-scale, high-availability systems, preferably fintech or consumer platforms.
  • Recent people-leadership experience leading a small engineering/SRE team.

Responsibilities

  • Lead and grow a senior SRE team focused on high-leverage reliability work.
  • Write production code and contribute to the codebase; lead by example in design and reviews.
  • Define service ownership, alerting, dashboards, runbooks, and deployment safety.

Skills

Python
Go
Observability
Incident command
Leadership

Job description

Role overview

Lead a small, senior SRE team for a high-availability consumer fintech platform, combining people leadership with sustained hands-on coding and building. You will define service ownership, observability, and incident practices, set the technical bar by writing production code yourself, and partner across engineering to raise the reliability of customer-facing systems. It is a builder's leadership role measured by what the team ships as much as by the operating model put around it.

Responsibilities
  • Lead and grow a senior SRE team, focusing the group on high-leverage reliability work instead of manual, ticket-driven response.
  • Write production code, build automation and tooling, contribute to the team codebase, and lead by example in design and code review.
  • Define service ownership, alerting, dashboards, runbooks, deploy and rollback safety, and run regular service-health reviews.
  • Run incident command when needed, drive verified remediation, and prevent recurrence through self-service tooling and runbook automation.
  • Design durable on-call coverage through staffing, handoffs, and automation rather than open-ended volunteer hours.
  • Partner with service teams and engineering leaders to raise reliability standards while holding clear boundaries between SRE enablement and service-team ownership.
Requirements
  • Strong software engineering foundation with current production-quality coding in a modern language such as Python or Go.
  • Real reliability depth for production systems and services at scale, ideally in consumer, fintech, or other high-availability environments.
  • Background built at a high-bar, high-scale engineering organization with a strong engineering culture.
  • Recent people-leadership experience managing a small engineering team within roughly the last few years.
  • Hands-on experience with observability, alerting, safe deploys, automation, and incident command, building these rather than only specifying them.
  • Calm, credible communication with team members, partner teams, and during live incidents.
Benefits and work setup
  • Fully remote within the United States.
  • Competitive base salary with stock options and health benefits from day one.
  • 401(k) plan with company match.
  • Flexible time off and growth opportunities in a mission-driven, inclusive culture.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE, Software Engineering (AWS / Scaling Infrastructure)
Senior SRE, Software Engineering (AWS / Scaling Infrastructure)

PulseRise Technologies • New York (NY)

On-site
USD 130,000 - 160,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Gen Digital Inc. • United States

Remote
USD 180,000 - 240,000
Site Reliability Engineer
Site Reliability Engineer

Longbridge Securities • Town of Texas (WI)

On-site
USD 100,000 - 130,000
Competitive compensation package
Growth opportunities
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Engg • Tempe (AZ)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Stelvio Inc. • Town of Texas (WI)

On-site
USD 125,000 - 145,000
Remote SRE Lead — Hands-On Reliability Builder
Remote SRE Lead — Hands-On Reliability Builder

OnePay • United States

Remote
USD 180,000 - 240,000
Stock options
Health benefits from Day 1
401(k) plan with company match
+3
Senior Site Reliability Engineer (SRE
Senior Site Reliability Engineer (SRE

Govserviceshub • New York (NY)

On-site
USD 130,000 - 160,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Practice by Numbers • United States

On-site
USD 120,000 - 160,000
High ownership and autonomy
Strong engineering culture
Impactful work on healthcare infrastructure
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Kovoro • Denver (CO), Northern (KY)

On-site
USD 150,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000