Senior SRE - Remote, Drive Production Reliability

Real Work From Anywhere

United States

Remote

USD 152,000 - 205,000

Full time

9 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Fingerprint is seeking a Senior Site Reliability Engineer to join our fully remote Infrastructure team and own production reliability for our high‑volume identification platform. You will write code and infrastructure, define SLIs/SLOs, and lead incident responses to keep systems fast, available, and predictable at scale.

You will collaborate with product engineering, implement observability, manage outages, and optimize capacity with Terraform and containerized services.

Qualifications

  • 6–10 years of experience in SRE, production engineering or backend infra on cloud‑based environments (AWS preferred).
  • Owned a system end to end—designed, shipped, operated and managed consequences of issues.
  • Hands‑on experience with SLIs, SLOs and error budgets; targets adopted and acted on.
  • Strong incident skills leading high‑severity, customer‑facing incidents and postmortems.
  • Depth in distributed systems, high throughput/low latency, and failure modes.
  • Cloud fundamentals: networking, load balancing, containers (Kubernetes).
  • Proficient in Terraform or equivalent, with code‑driven infrastructure.
  • Go or Python production software development; shipping fixes, not just recommendations.
  • Fluent with observability tooling and instrumentation of systems.
  • Experience operating Redis/ElastiCache in production; depth is a differentiator.
  • Software engineering best practices: version control, code reviews, tests, safe deployments.
  • Strong ownership and autonomy; able to work without clearly defined requirements.
  • English communication skills; able to document and review effectively.
  • AI‑native by default; familiarity with AI tools to aid incident investigations and tooling.

Responsibilities

  • Own the reliability of core production systems end to end; instrument and operate them under real traffic.
  • Define and maintain SLIs/SLOs, dashboards and alerts; use error budgets to guide fixes.
  • Improve alert quality, reduce noise, and ensure issues are detected before customers notice.
  • Lead incident response; investigate across service boundaries and write actionable postmortems.
  • Build secure, resilient infra with attention to failure modes: timeouts, retries, and load shedding.
  • Perform capacity/performance work using real data: load testing, profiling, headroom planning.
  • Enhance change safety: progressive delivery, automated rollbacks, production‑readiness signals.
  • Manage infrastructure code/configuration (Terraform) and align with service architecture.
  • Develop developer tooling to reduce toil and improve operational efficiency.
  • Conduct deliberate failure testing (game days/chaos) to uncover gaps before customers.
  • Collaborate with product engineering for high‑risk services: runbooks, rollback plans, on‑call handoff.
  • Participate in on‑call rotation with improved runbooks and escalation clarity.
  • Approach engineering with a security lens; review peer work for vulnerabilities.

Skills

Go
Python
Incident response
SRE
Cloud (AWS)
Kubernetes
Terraform

Tools

Terraform
Datadog
Prometheus
Grafana
OpenTelemetry
Redis/ElastiCache

Job description

Fingerprint is seeking a Senior Site Reliability Engineer to join our fully remote Infrastructure team and own production reliability for our high‑volume identification platform. You will write code and infrastructure, define SLIs/SLOs, and lead incident responses to keep systems fast, available, and predictable at scale.

You will collaborate with product engineering, implement observability, manage outages, and optimize capacity with Terraform and containerized services.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer — Remote Production Reliability
Senior Site Reliability Engineer — Remote Production Reliability

Fingerprint • Chicago (IL)

Remote
USD 152,000 - 205,000
Senior SRE: Cloud Reliability, Terraform & Kubernetes (Remote)
Senior SRE: Cloud Reliability, Terraform & Kubernetes (Remote)

Motion Recruitment • Chicago (IL)

On-site
USD 140,000 - 190,000
Remote Staff SRE — Reliability & Incident Leader
Remote Staff SRE — Reliability & Incident Leader

Real Work From Anywhere • United States

Remote
USD 177,000 - 240,000
Senior SRE Manager: Remote Platform Reliability & Automation
Senior SRE Manager: Remote Platform Reliability & Automation

Ferguson Enterprises, Inc. • United States

Hybrid
USD 106,000 - 185,000
Senior SRE — Scale, Automation & Uptime
Senior SRE — Scale, Automation & Uptime

Hirebridge • Northern (KY)

Hybrid
USD 110,000 - 145,000
Bonus
Staff SRE - Remote-Optional, Scale & Reliability Leader
Staff SRE - Remote-Optional, Scale & Reliability Leader

Pivotal Health • New York (NY)

Hybrid
USD 180,000 - 240,000
Competitive compensation with equity
Full health, dental, vision coverage
401(k) retirement savings plan
+2
Remote SRE Lead — Hands-On Reliability Builder
Remote SRE Lead — Hands-On Reliability Builder

OnePay • United States

Remote
USD 180,000 - 240,000
Stock options
Health benefits from Day 1
401(k) plan with company match
+3
Site Reliability Engineer
Site Reliability Engineer

Motion Recruitment Partners LLC • Chicago (IL), Northern (KY)

Hybrid
USD 140,000 - 170,000
Remote Senior Site Reliability Engineer — Reliability Lead
Remote Senior Site Reliability Engineer — Reliability Lead

Priority Technology Holdings, Inc. • Alpharetta (GA)

On-site
USD 129,000 - 161,000
401(k) match
Employee Stock Purchase Program (ESPP)
Medical, dental, and vision coverage
+1
Remote Senior SRE: Platform Reliability & Incidents
Remote Senior SRE: Platform Reliability & Incidents

Tamarind Intelligence • United States

Remote
USD 99,000 - 140,000
Health coverage for you and dependents
Tech spending stipend
Employee stock purchase plan