Lead Site Reliability Engineer - Scale & Reliability

Doist

United States

Remote

USD 170,000 - 230,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Tech Holding is seeking a hands-on Lead Site Reliability Engineer to elevate the performance, reliability, and scalability of our platform. You will work across application services, infrastructure, databases, networking, caches, queues, and external dependencies to identify bottlenecks and lead remediation.

This is not a traditional DevOps role or an advisory architecture position. You will partner with engineering, product, and leadership to define SLOs, instrument the full request path, build

Qualifications

  • Extensive experience in Site Reliability Engineering or related discipline.
  • Experience supporting production systems with scale, latency, or availability requirements.
  • Strong understanding of observability, capacity planning, and reliability engineering.
  • Hands-on experience with cloud infrastructure and distributed systems.
  • Experience defining and operating against SLOs/SLIs and production reliability metrics.

Responsibilities

  • Establish baselines for throughput, latency, and capacity of critical platform journeys.
  • Define and maintain SLOs, error budgets, dashboards, alerts, and reliability thresholds.
  • Instrument the full request path across services, compute, storage, networking, and databases.
  • Identify bottlenecks and lead cross-functional remediation with engineering teams.
  • Build capacity models showing scale and cost implications for future growth.
  • Lead load, stress, soak, spike, and recovery testing in representative environments.
  • Develop realistic demand scenarios for major customers and partnerships.
  • Drive resilience improvements and graceful degradation planning.
  • Establish automated performance testing and production release gates with CI/CD.
  • Own readiness assessments for major pilots and launches.

Skills

SRE principles
Observability
Cloud infrastructure
Distributed systems
SLOs/SLIs
Load testing
Incident management
Performance tuning

Tools

Kubernetes
Cloud platforms

Job description

Tech Holding is seeking a hands-on Lead Site Reliability Engineer to elevate the performance, reliability, and scalability of our platform. You will work across application services, infrastructure, databases, networking, caches, queues, and external dependencies to identify bottlenecks and lead remediation.

This is not a traditional DevOps role or an advisory architecture position. You will partner with engineering, product, and leadership to define SLOs, instrument the full request path, build

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote Lead Site Reliability Engineer: Performance & Scaling
Remote Lead Site Reliability Engineer: Performance & Scaling

Techholding • United States

Remote
USD 140,000 - 190,000
Remote Lead SRE - Performance & Scalability
Remote Lead SRE - Performance & Scalability

Tech Holding • Northern (KY)

Hybrid
USD 150,000 - 210,000
Senior Site Reliability Lead: Scale & Reliability Champion
Senior Site Reliability Lead: Scale & Reliability Champion

Twitter • San Francisco (CA)

On-site
USD 130,000 - 160,000
Lead Site Reliability Engineer (Performance & Scalability) | Contract | Remote US New USA, Remote
Lead Site Reliability Engineer (Performance & Scalability) | Contract | Remote US New USA, Remote

Tech Holding • Northern (KY)

Hybrid
USD 150,000 - 210,000
Lead Site Reliability Engineer - Architect & Own Production
Lead Site Reliability Engineer - Architect & Own Production

Optimal Market Technologies • New York (NY)

On-site
USD 175,000 - 200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Axiom Pursuits • San Francisco (CA)

On-site
USD 150,000 - 180,000
Lead Site Reliability Engineer: Reliability & Incidents
Lead Site Reliability Engineer: Reliability & Incidents

Referrals Only • Cincinnati (OH)

On-site
USD 120,000 - 180,000
Senior Site Reliability Engineer: Reliability at Scale
Senior Site Reliability Engineer: Reliability at Scale

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000
Lead Site Reliability Engineer: AWS Cloud & Automation
Lead Site Reliability Engineer: AWS Cloud & Automation

Selby Jennings • Wilmington (NC)

On-site
USD 140,000 - 200,000
Remote Senior Site Reliability Engineer — Reliability Lead
Remote Senior Site Reliability Engineer — Reliability Lead

Priority Technology Holdings, Inc. • Alpharetta (GA)

On-site
USD 129,000 - 161,000
401(k) match
Employee Stock Purchase Program (ESPP)
Medical, dental, and vision coverage
+1