Remote Lead Site Reliability Engineer: Performance & Scaling

Techholding

United States

Remote

USD 140,000 - 190,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Tech Holding is seeking a hands‑on Lead Site Reliability Engineer for a project based assignment. You will establish performance baselines, define SLOs, and drive remediation for scalability and reliability across services, databases, and infrastructure.

This is a hands‑on, cross‑functional role not a pure advisory position. You will partner with engineering, product, and leadership to forecast capacity, run resilience tests, and prepare for major launches and partnerships.

Qualifications

  • Extensive experience in Site Reliability Engineering and related disciplines.
  • Proven ability to support production systems at meaningful scale with latency and availability requirements.
  • Deep understanding of observability, capacity planning, and reliability engineering.
  • Hands-on experience with cloud infrastructure and production distributed systems.
  • Experience defining and operating against SLOs, SLIs, and reliability metrics.
  • Proficient in load, stress, soak, scalability, and resilience testing.
  • Ability to translate performance risks into business implications for leadership.

Responsibilities

  • Establish baselines for performance, throughput, latency, and capacity for core workloads.
  • Define and maintain SLOs, error budgets, dashboards, alerts, and thresholds.
  • Instrument full request paths across services, compute, storage, networks, and caches.
  • Identify bottlenecks and coordinate cross-functional remediation with engineering.
  • Build capacity models and cost estimates for scaling.
  • Lead load, stress, soak, spike, failure, and recovery testing in representative environments.
  • Develop demand scenarios for major customers and high-volume events.
  • Drive architecture hardening, graceful degradation, and resilience improvements.
  • Collaborate on automated performance testing and production release gates.
  • Own readiness assessments for pilots, partnerships, and launches.
  • Create runbooks for scale events, incidents, and dependency failures.
  • Lead incident investigations and incorporate lessons into future work.
  • Show business impact of capacity and reliability decisions to leadership.

Skills

Site Reliability
Performance engineering
Platform engineering
Distributed systems
Observability
Cloud infrastructure
SLOs/SLIs
Incident management
Load testing

Job description

Tech Holding is seeking a hands‑on Lead Site Reliability Engineer for a project based assignment. You will establish performance baselines, define SLOs, and drive remediation for scalability and reliability across services, databases, and infrastructure.

This is a hands‑on, cross‑functional role not a pure advisory position. You will partner with engineering, product, and leadership to forecast capacity, run resilience tests, and prepare for major launches and partnerships.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer - Scale & Reliability
Lead Site Reliability Engineer - Scale & Reliability

TQC Ltd • United States

Remote
USD 170,000 - 230,000
Senior Site Reliability Lead: Scale & Reliability Champion
Senior Site Reliability Lead: Scale & Reliability Champion

Twitter • San Francisco (CA)

On-site
USD 130,000 - 160,000
Lead Site Reliability Engineer - Architect & Own Production
Lead Site Reliability Engineer - Architect & Own Production

Optimal Market Technologies • New York (NY)

On-site
USD 175,000 - 200,000
Remote Site Reliability Engineer: Scale & Resilience
Remote Site Reliability Engineer: Scale & Resilience

Bright Vision Technologies • Nashua (NH)

On-site
USD 100,000 - 180,000
Remote Senior Site Reliability Engineer — Reliability Lead
Remote Senior Site Reliability Engineer — Reliability Lead

Priority Technology Holdings, Inc. • Alpharetta (GA)

On-site
USD 129,000 - 161,000
401(k) match
Employee Stock Purchase Program (ESPP)
Medical, dental, and vision coverage
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Axiom Pursuits • San Francisco (CA)

On-site
USD 150,000 - 180,000
Lead SRE - Remote/Hybrid, Enterprise-Scale Reliability
Lead SRE - Remote/Hybrid, Enterprise-Scale Reliability

Empower Retirement • Greenwood Village (CO)

Hybrid
USD 114,000 - 166,000
Medical insurance
401(k) with company match
Tuition reimbursement
+3
Lead Site Reliability Engineer: Reliability & Incidents
Lead Site Reliability Engineer: Reliability & Incidents

Referrals Only • Cincinnati (OH)

On-site
USD 120,000 - 180,000
Remote Senior SRE: Reliability, Performance & Security Lead
Remote Senior SRE: Reliability, Performance & Security Lead

TECEZE • United States

On-site
USD 120,000 - 190,000
Remote Lead Platform Engineer - Reliability & Scale
Remote Lead Platform Engineer - Reliability & Scale

Ladders • United States

On-site
USD 129,000 - 178,000
Medical benefits
401(k)
Paid time off
+3