Sr. Site Reliability Engineer

Practice by Numbers

Bellevue (WA)

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Practice by Numbers is seeking a Senior SRE to own reliability of critical services in Bellevue, WA. You will design scalable distributed systems and define SLOs/SLIs while building tooling to reduce toil for engineers.

You will lead incident response, conduct blameless postmortems, and mentor peers while partnering with product engineering to bake reliability in from day one. This is a hands-on engineering leadership role requiring a strong CS/CE/EE background.

Qualifications

  • Engineering degree is mandatory: BS/MS in Computer Science, Computer Engineering, Electrical Engineering, or closely related engineering field.
  • 6+ years experience in software engineering, SRE, infrastructure/platform engineering, or related.
  • Strong programming skills in Go, Python, Java, or similar (production-quality code).
  • Proven experience building and operating production backend services or distributed systems.
  • Meaningful experience in on-call rotations, incident leadership, and post-incident improvement execution.

Responsibilities

  • Own reliability outcomes for critical services: availability, latency, incident rate, and recovery time.
  • Design and build reliable, scalable distributed systems that support mission-critical healthcare workflows.
  • Define and operationalize SLOs/SLIs and error budgets; drive adoption across teams and use them to prioritize work.
  • Lead incident response for high-severity issues; improve on-call effectiveness and reduce alert fatigue.
  • Run blameless postmortems and ensure follow-ups are implemented, measured, and stick.
  • Write software to eliminate operational toil: automation, self-service tooling, guardrails, and developer platforms.
  • Raise the bar on observability (metrics/logs/traces), alerting strategy, and operational readiness.
  • Improve resilience through capacity planning, load testing, performance tuning, and failure testing.
  • Mentor engineers (SRE and product engineers) on reliability practices, debugging, and production ownership.
  • Drive cross-team improvements like production readiness reviews, release safety (progressive delivery), and standard runbooks.

Skills

Go
Python
Java
SRE
Distributed systems

Education

BS/MS in CS/CE/EE or related engineering field

Tools

Docker
Kubernetes
Terraform
Prometheus
Grafana
OpenTelemetry
GitHub Actions

Job description

This is an engineering-first Senior SRE role.

We’re looking for Senior Engineers Who Have the following:

  • Built and shipped significant backend systems and/or distributed platforms
  • Owned services end-to-end in production (design → launch → on-call → reliability improvements)
  • Led incident response and drove durable follow-ups
  • Improved reliability by writing software and changing system design—not by adding manual process

You’ll partner closely with product engineering to ensure reliability is designed in from day one, while also building the tooling and platforms that make operating services safer and easier for every engineer.

Engineers here own services end-to-end—from design to production reliability.

Important: This is not a system administrator role. We are explicitly hiring an engineering leader in reliability. Engineering degree is an absolute requirement (BS/MS in CS/CE/EE or closely related engineering field).

What You’ll Do
  • Own reliability outcomes for critical services: availability, latency, incident rate, and recovery time.
  • Design and build reliable, scalable distributed systems that support mission-critical healthcare workflows.
  • Define and operationalize SLOs/SLIs and error budgets; drive adoption across teams and use them to prioritize work.
  • Lead incident response for high-severity issues; improve on-call effectiveness and reduce alert fatigue.
  • Run blameless postmortems and ensure follow-ups are implemented, measured, and stick.
  • Write software to eliminate operational toil: automation, self-service tooling, guardrails, and developer platforms.
  • Raise the bar on observability (metrics/logs/traces), alerting strategy, and operational readiness.
  • Improve resilience through capacity planning, load testing, performance tuning, and failure testing.
  • Mentor engineers (SRE and product engineers) on reliability practices, debugging, and production ownership.
  • Drive cross-team improvements like production readiness reviews, release safety (progressive delivery), and standard runbooks.
What We’re Looking For
  • Engineering degree is mandatory: BS/MS in Computer Science, Computer Engineering, Electrical Engineering, or a closely related engineering field.
  • 6+ years experience in software engineering, SRE, infrastructure/platform engineering, or related.
  • Strong programming skills in Go, Python, Java, or similar (production-quality code).
  • Proven experience building and operating production backend services or distributed systems.
  • Meaningful experience in on-call rotations, incident leadership, and post-incident improvement execution.
  • Strong debugging ability across complex systems: latency, saturation, cascading failures, dependency issues.
  • Experience with cloud infrastructure (AWS preferred, GCP/Azure acceptable).
Strong Signal
  • You’ve owned reliability for customer-facing services with clear, measurable improvements (e.g., higher availability, lower MTTR).
  • You’ve built internal platforms/tooling that made other engineers faster and reduced operational burden.
  • You’ve worked in an SRE culture with SLOs, error budgets, and blameless postmortems.
  • You’ve led multi-quarter reliability initiatives spanning multiple teams/services.
Technologies We Work With (Examples)
  • Cloud: AWS
  • Containers: Docker, Kubernetes
  • Infrastructure as Code: Terraform
  • Observability: Prometheus, Grafana, OpenTelemetry
  • Languages: Go, Python, TypeScript
  • CI/CD: GitHub Actions
This Role Is Not
  • System administration / IT ops / helpdesk
  • Manual server patching as a primary responsibility
  • A “click-ops” cloud operator role
Why Join PBN
  • Build and operate mission-critical healthcare infrastructure that supports real patient workflows.
  • High impact: reliability work directly improves customer trust and revenue-critical operations.
  • Small team with high ownership, autonomy, and ability to influence architecture.
  • Strong engineering culture focused on automation, simplicity, and measurable outcomes.
Compensation

The base pay range for this role is $120,000 – $150,000 per year.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Practice By Numbers, Inc. • Redmond (WA)

On-site
USD 130,000 - 160,000
High impact work on healthcare infrastructure
Strong engineering culture focused on automation
Small team with high ownership and autonomy
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Practice by Numbers • United States

On-site
USD 120,000 - 160,000
High ownership and autonomy
Strong engineering culture
Impactful work on healthcare infrastructure
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Kontakt.io • New York (NY)

Hybrid
USD 200,000 - 250,000
Hybrid work 3 days/week in NYC office.
Equity in a high-growth company
Health, dental, vision insurance
+1
SRE Leader
SRE Leader

Kontakt Micro-Location Sp. Z.o.o. • New York (NY)

Hybrid
USD 180,000 - 260,000
Equity in a high-growth company
Health, dental, and vision coverage
401k
+3
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Storm2 • Scottsdale (AZ)

Hybrid
USD 140,000 - 150,000
Competitive healthcare, dental, and vision coverage
401(k) with company match
Generous PTO and paid holidays
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Abbott • Sunnyvale (CA)

On-site
USD 90,000 - 180,000
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

Hybrid
USD 146,000 - 163,000
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

Hybrid
USD 140,000 - 200,000
Principal Site Reliability Engineer (SRE)
Principal Site Reliability Engineer (SRE)

Symmetrio • United States

Hybrid
USD 120,000 - 160,000
Health Care Plan (Medical, Dental & Vision)
Retirement Plan (401k, IRA)
Paid Time Off (Vacation, Sick & Public Holidays)