Site Reliability Engineer

Talentify

Woonsocket (RI)

Hybrid

USD 150,000 - 190,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health benefits
Referral program
Growth opportunities

Job summary

Talentify is seeking an experienced Site Reliability Engineer to own end-to-end reliability of critical platforms across hybrid cloud and on-prem environments in Woonsocket, RI. You will drive SLI/SLO health, incident response, and observability, collaborating with product and operations to embed reliability in design and delivery.

You will lead postmortems, perform root cause analysis, and mentor teams in chaos engineering, fault injection, and SRE best practices, shaping organizational

Qualifications

  • 8+ years of Senior Software engineering experience in SRE, DevOps, platform engineering, or related production-systems roles in distributed systems at production scale with active on-call responsibility.
  • Experience tuning and validating timeseries anomaly detection models in a production observability context.
  • Strong programming proficiency in Python, React, and Java at production quality.
  • Hands-on experience designing SLIs, SLOs, and managing error budgets for customer-facing or business-critical services.

Responsibilities

  • Own and drive the end-to-end reliability, availability, and performance of critical retail and pharmacy technology platforms across hybrid cloud and on-premises environments.
  • Establish and maintain SLI/SLO health, alerting strategies, observability standards, and business-aligned monitoring for the assigned application domain.
  • Lead production incident response as Incident Commander, drive root cause analysis, postmortems, and continuous reliability improvements.
  • Partner with engineering, product, and operations teams to embed reliability, resiliency, scalability, and operational readiness into system design and delivery.
  • Build and optimize automation, self-service capabilities, and operational tooling to eliminate toil, improve efficiency, and reduce manual intervention.
  • Design and execute proactive reliability initiatives, including production readiness reviews, dependency risk assessments, fault injection, and chaos engineering exercises.
  • Mentor engineers, champion SRE best practices, and enable teams to independently detect, respond to, and learn from production issues with minimal SRE involvement.

Skills

Python
React
Java
Incident Commander
Observability
Anomaly detection
SRE
DevOps
Chaos engineering
LLM integration
Kafka
Istio
Envoy
Terraform
Ansible
Kubernetes
GCP
Rancher K3s
Airflow
Tidal
BigQuery
PostgreSQL

Tools

Prometheus
Grafana
Open Telemetry
Loki
Splunk
Elasticsearch

Job description

Our client, an IT Services and Consulting company, is looking for a Site Reliability Engineer for their Woonsocket, RI/ Hybrid location.

Responsibilities:

  • Own and drive the endtoend reliability, availability, and performance of critical retail and pharmacy technology platforms across hybrid cloud and onpremises environments.
  • Establish and maintain SLI/SLO health, alerting strategies, observability standards, and businessaligned monitoring for the assigned application domain.
  • Lead production incident response as Incident Commander, drive root cause analysis, postmortems, and continuous reliability improvements.
  • Partner with engineering, product, and operations teams to embed reliability, resiliency, scalability, and operational readiness into system design and delivery.
  • Build and optimize automation, selfservice capabilities, and operational tooling to eliminate toil, improve efficiency, and reduce manual intervention.
  • Design and execute proactive reliability initiatives, including production readiness reviews, dependency risk assessments, fault injection, and chaos engineering exercises.
  • Mentor engineers, champion SRE best practices, and enable teams to independently detect, respond to, and learn from production issues with minimal SRE involvement. Influence organizational adoption of SLOdriven engineering, observability, incident management, and reliability practices through collaboration, credibility, and measurable outcomes.

Requirements:

  • 8+ years of Senior Software engineering experience in SRE, DevOps, platform engineering, or related production-systems roles in distributed systems at production scale with active on-call responsibility.
  • Demonstrated experience as an oncall Incident Commander (IC) for P1 or P2 incidents — structured leadership updates, not just participant involvement.
  • Experience tuning and validating timeseries anomaly detection models in a production observability context — this is a Required qualification, not a preferred one; anomalybased detection is a core function of this role.
  • Strong programming proficiency in Python, React, and Java at production quality — capable of writing operational tooling that other engineers will rely on.
  • Handson experience designing SLIs, SLOs, and managing error budgets for customerfacing or businesscritical services.
  • Deep observability platform experience: Prometheus, Grafana, Open Telemetry, and at least two of the log aggregation solutions (Loki, Splunk, Elasticsearch).
  • Fleetscale deployment awareness: familiarity with progressive rollout strategies, blast radius management, and configuration drift as a reliability risk in large unattended node deployments.
  • Strong cloud platform expertise in Google Cloud Platform (GCP) and Rancher K3s.
  • Advanced Kubernetes operational experience: debugging, resource management, networking policies, and workload failure modes. Experience with AIassisted tooling and development.
  • Experience diagnosing and resolving workflow orchestration issues, batch processing failures, scheduler performance problems, and building observability on data pipelines: Apache Airflow and Tidal.
  • Experience owning Production Readiness Reviews or service launch gates.
  • Strong proficiency in transforming largescale operational and telemetry data into actionable business insights using SQLbased analytics and reporting frameworks: Google BigQuery, PostgreSQL.
  • Handson chaos or fault injection experience.
  • TIC (Technical Incident Commander) certification or equivalent structured incident command training.
  • Experience operating distributed systems in retail, pharmacy, healthcare, or other operationally sensitive environments where failures have direct patient or customer impact.
  • LLM integration for operational use cases (alert summarization, runbook suggestion, incident triage assistance) — design or implementation experience.
  • Experience with streaming data platforms: Kafka.
  • Experience with service mesh and traffic management: Istio, Envoy.
  • Infrastructureascode proficiency at production scale: Terraform or Ansible.
  • Years of Experience: 14.00 Years of Experience

Why Should You Apply?

  • Health Benefits
  • Referral Program
  • Excellent growth and advancement opportunities
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer -- SINDC5717546
Site Reliability Engineer -- SINDC5717546

Compunnel Inc. • Denton (TX)

On-site
USD 120,000 - 150,000
Site Reliability Engineer Engineer
Site Reliability Engineer Engineer

Modus Create • Aurora (IL)

Remote
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

On-site
USD 140,000 - 200,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 240,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

United States Digital Space LLC • Charlotte (TX)

On-site
USD 153,000 - 192,000
Discretionary incentive eligible
Benefits package
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Mindlance • Irving (TX)

On-site
USD 120,000 - 160,000
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

On-site
USD 130,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SDI International • Chicago (IL)

On-site
USD 130,000 - 180,000
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

On-site
USD 146,032 - 162,257
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1