Senior Site Reliability Engineer — Remote Production Reliability

Fingerprint

Chicago (IL)

Remote

USD 152,000 - 205,000

Full time

6 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Fingerprint is seeking a Senior Site Reliability Engineer to own reliability for a platform handling millions of requests. This hands-on role requires writing code and infrastructure, defining SLIs/SLOs, and leading incident response with end-to-end ownership.

You’ll build secure, resilient systems, optimize capacity, and collaborate with product teams to ensure production readiness and scalable performance in a remote setting.

Qualifications

  • 6–10 years of experience in SRE, production engineering, infrastructure, or backend engineering within primarily cloud-based environments (AWS preferred).
  • Track record of owning a system end to end — designed, shipped, operated, and managed outcomes when issues occurred.
  • Hands-on experience defining and operating against SLIs, SLOs, and error budgets with targets.
  • Strong incident skills: led or served as primary responder on high-severity, customer-facing incidents.
  • Depth in distributed systems failure modes in high-throughput, low-latency environments.
  • Cloud infrastructure fundamentals: networking, load balancing, containerization (Kubernetes), and distributed systems.
  • Hands-on experience managing infrastructure through code and configuration (Terraform or equivalent).
  • Fluency with observability tooling and instrumenting systems (Datadog, Prometheus, Grafana, OpenTelemetry).
  • High ownership and autonomy; ability to work without clearly defined requirements.
  • Security-conscious, emphasizing safe deployment and risk awareness.

Responsibilities

  • Own the reliability of core production systems end to end — instrument, set targets, operate, and be accountable for behavior under real traffic.
  • Define and maintain SLIs/SLOs, wire into dashboards and alerts, use error budgets to drive fixes.
  • Drive alert quality by reducing noise and surfacing meaningful signals before customers notice issues.
  • Lead incident response: investigate across service boundaries, restore service, and publish actionable postmortems.
  • Build secure, resilient infrastructure with attention to failure modes: timeouts, backpressure, degraded performance, and blast radius containment.
  • Perform capacity and performance work with real data: load testing, profiling, and headroom planning.
  • Improve change safety: progressive delivery, automated rollback, production-like pre-production signals.
  • Manage infrastructure as code (Terraform) and align patterns with service architecture.
  • Develop developer-facing tooling to reduce toil and ease operation of services.
  • Run deliberate failure testing (game days, chaos engineering) to uncover gaps before customers do.
  • Partner with product teams on production readiness for new services, including capacity, failure modes, and runbooks.
  • Participate in on-call rotations with clearer runbooks and reduced pager fatigue.
  • Approach all work with security in mind and review code for vulnerabilities.

Skills

SRE / production engineering
Incident management
Distributed systems
Cloud fundamentals (AWS)
Go or Python programming
Observability tooling
On-call experience
Security mindset
Ownership / autonomy
AI-assisted debugging

Tools

Terraform
Kubernetes (EKS)
Redis/ElastiCache
Datadog/Prometheus/OpenTelemetry

Job description

Fingerprint is seeking a Senior Site Reliability Engineer to own reliability for a platform handling millions of requests. This hands-on role requires writing code and infrastructure, defining SLIs/SLOs, and leading incident response with end-to-end ownership.

You’ll build secure, resilient systems, optimize capacity, and collaborate with product teams to ensure production readiness and scalable performance in a remote setting.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE - Remote, Drive Production Reliability
Senior SRE - Remote, Drive Production Reliability

Real Work From Anywhere • United States

Remote
USD 152,000 - 205,000
Remote Senior Site Reliability Engineer — Reliability Lead
Remote Senior Site Reliability Engineer — Reliability Lead

Priority Technology Holdings, Inc. • Alpharetta (GA)

On-site
USD 129,000 - 161,000
401(k) match
Employee Stock Purchase Program (ESPP)
Medical, dental, and vision coverage
+1
Remote SRE Lead — Hands-On Reliability Builder
Remote SRE Lead — Hands-On Reliability Builder

OnePay • United States

Remote
USD 180,000 - 240,000
Stock options
Health benefits from Day 1
401(k) plan with company match
+3
Senior Production Engineer Platform Reliability (Remote)
Senior Production Engineer Platform Reliability (Remote)

GitHub • United States

Remote
USD 255,000 - 425,000
Senior Site Reliability Engineer - Remote
Senior Site Reliability Engineer - Remote

Bright-Vision-Technologies • United States

Remote
USD 100,000 - 150,000
Lead Site Reliability Engineer - Architect & Own Production
Lead Site Reliability Engineer - Architect & Own Production

Optimal Market Technologies • New York (NY)

On-site
USD 175,000 - 200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Axiom Pursuits • San Francisco (CA)

On-site
USD 150,000 - 180,000
Senior Site Reliability Engineer — Remote, AWS & Observability
Senior Site Reliability Engineer — Remote, AWS & Observability

Prove • United States

Hybrid
USD 140,000 - 190,000
Wellbeing reimbursement
401k Match
Parental Leave Policy
+5
Senior SRE: Platform Reliability & Incident Lead (Remote)
Senior SRE: Platform Reliability & Incident Lead (Remote)

Affirm, Inc. • Town of Poland (NY)

On-site
USD 32,000 - 48,000
Health insurance
Equity rewards
Flexible Spending Wallets
+1
Senior Site Reliability Lead: Scale & Reliability Champion
Senior Site Reliability Lead: Scale & Reliability Champion

Twitter • San Francisco (CA)

On-site
USD 130,000 - 160,000