SRE Lead

TechDigital Group

Woonsocket (RI)

On-site

USD 140,000 - 190,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

TechDigital Group is seeking a Senior Software Engineer specializing in SRE/DevOps to lead on-call incident response, build robust observability, and shape reliability across distributed production systems.

The role requires hands-on work with Python, React, Java, and a strong focus on SLIs/SLOs, with GCP and Kubernetes at scale. Expect collaboration across engineering teams and a focus on production readiness.

Qualifications

  • 8+ years of Senior Software engineering in SRE/DevOps or production systems.
  • On-call Incident Commander for P1/P2 incidents with structured updates.
  • Experience tuning time-series anomaly detection in production observability.
  • Proficient in Python, React, and Java for production tooling.
  • Designing SLIs/SLOs and managing error budgets for critical services.
  • Strong observability stack: Prometheus, Grafana, OpenTelemetry and 2+ log solutions.
  • Fleet-scale deployment awareness and drift management for large clusters.
  • Cloud experience with GCP and Rancher K3s; advanced Kubernetes ops.
  • Experience with Airflow and Tidal for data/workflow orchestration.

Skills

SRE/DevOps
On-call leadership
Anomaly detection
Python
React
Java
SLIs/SLOs
Observability
Prometheus
Grafana
OpenTelemetry
Log aggregation
Loki
Splunk
Elasticsearch
Kubernetes
GCP
Rancher K3s
Airflow
Tidal
Terraform
Ansible
Chaos engineering

Tools

Git
CI/CD

Job description

Job Description/ Responsibilities
  • 8+ years of Senior Software engineering experience in SRE, DevOps, platform engineering, or related production-systems roles in distributed systems at production scale with active on-call responsibility
  • Demonstrated experience as an on-call Incident Commander (IC) for P1 or P2 incidents — structured leadership updates, not just participant involvement
  • Experience tuning and validating time-series anomaly detection models in a production observability context — this is a Required qualification, not a preferred one; anomaly-based detection is a core function of this role
  • Strong programming proficiency in Python, React, and Java at production quality — capable of writing operational tooling that other engineers will rely on
  • Hands‑on experience designing SLIs, SLOs, and managing error budgets for customer‑facing or business‑critical services
  • Deep observability platform experience: Prometheus, Grafana, OpenTelemetry, and at least two of the log aggregation solution (Loki, Splunk, Elasticsearch)
  • Fleet‑scale deployment awareness: familiarity with progressive rollout strategies, blast radius management, and configuration drift as a reliability risk in large unattended node deployments
  • Strong cloud platform expertise in Google Cloud Platform (GCP) and Rancher K3s.
  • Advanced Kubernetes operational experience: debugging, resource management, networking policies, and workload failure modes. Experience with AI‑assisted tooling and development.
  • Experience diagnosing and resolving workflow orchestration issues, batch processing failures, scheduler performance problems, and building observability on data pipeline : Apache Airflow and Tidal.
Preferred Qualifications
  • Experience owning Production Readiness Reviews or service launch gates.
  • Strong proficiency in transforming large‑scale operational and telemetry data into actionable business insights using SQL-based analytics, and reporting frameworks: Google BigQuery, PostgreSQL.
  • Hands‑on chaos or fault injection experience.
  • TIC (Technical Incident Commander) certification or equivalent structured incident command training
  • Experience operating distributed systems in retail, pharmacy, healthcare, or other operationally sensitive environments where failures have direct patient or customer impact
  • LLM integration for operational use cases (alert summarization, runbook suggestion, incident triage assistance) — design or implementation experience
  • Experience with streaming data platforms: Kafka.
  • Experience with service mesh and traffic management: Istio, Envoy.
  • Infrastructure-as-code proficiency at production scale: Terraform or Ansible
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Site Reliability Engineer, SRE
Staff Site Reliability Engineer, SRE

Jobtailor • California (MO)

On-site
USD 120,000 - 210,000
Lead SRE
Lead SRE

JPMorgan Chase & Co. • Plano (TX)

On-site
USD 150,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Resolve Tech Solutions • Irving (TX)

On-site
USD 120,000 - 160,000
Senior Manager SRE
Senior Manager SRE

Expedite Talent Solutions • United States

Hybrid
USD 130,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Software Engineer Manager
Software Engineer Manager

Jobtailor • Alabama

On-site
USD 120,000 - 180,000
Observability Engineer / Site Reliability Engineer
Observability Engineer / Site Reliability Engineer

Ontrac Solutions • New York (NY)

On-site
USD 120,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
SRE Lead: Production Reliability & Observability Architect
SRE Lead: Production Reliability & Observability Architect

TechDigital Group • Woonsocket (RI)

On-site
USD 140,000 - 190,000
Java SRE Engineer
Java SRE Engineer

EITACIES Inc. • Santa Clara (CA)

On-site
USD 120,000 - 160,000