Senior Site Reliability Engineer: Scale, Automate, Observe

TalentDome Staffing

United States

On-site

USD 140,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

TalentDome Staffing is seeking a Senior Site Reliability Engineer to enhance the reliability, scalability, and performance of production systems. You will apply SRE principles to reduce toil, implement self-healing infrastructure, and drive capacity planning in high-scale environments.

Ideal candidates have 6+ years in SRE/DevOps with strong Python/Go coding, extensive AWS, and hands-on experience with Kubernetes.

Qualifications

  • 6+ years of hands-on SRE/DevOps or backend engineering with production reliability ownership.
  • Strong software engineering in Python or Go.
  • Deep knowledge of AWS core services (EC2, Lambda, RDS, VPC, IAM, S3) and cost/performance optimization.
  • Linux systems mastery including performance tuning, kernel/network troubleshooting, and security hardening.
  • Kubernetes/EKS production experience at scale.
  • Experience defining and operationalizing SLIs, SLOs, and error budgets.
  • Understanding of distributed systems, fault tolerance, and failover strategies.
  • Experience building full-stack observability pipelines and leading incident responses.

Responsibilities

  • Define and drive SLIs, SLOs, and error budgets across services.
  • Build automation to reduce toil and enable self-service for engineers.
  • Design and maintain IaC on AWS using Terraform/CloudFormation or Pulumi.
  • Architect observability with Prometheus, Grafana, Datadog, Splunk, and OpenTelemetry.
  • Act as incident commander during outages and lead blameless postmortems.
  • Forecast capacity, run load/chaos tests, and tune systems to prevent bottlenecks.
  • Partner with teams to deploy safe pipelines (canary, blue/green, feature flags) using ArgoCD/Jenkins/GitLab CI.
  • Embed security into pipelines, manage access, network segmentation, and vulnerability management.
  • Lead on-call rotations and runbooks; mentor engineers on reliability best practices.
  • Influence architecture decisions and champion SRE culture across the org.

Skills

Python
Go
Linux
CI/CD
Observability

Tools

AWS
Terraform/CloudFormation
Pulumi
Prometheus
Grafana
Datadog
Splunk
OpenTelemetry
Kubernetes
EKS
ArgoCD
Jenkins
GitLab CI

Job description

TalentDome Staffing is seeking a Senior Site Reliability Engineer to enhance the reliability, scalability, and performance of production systems. You will apply SRE principles to reduce toil, implement self-healing infrastructure, and drive capacity planning in high-scale environments.

Ideal candidates have 6+ years in SRE/DevOps with strong Python/Go coding, extensive AWS, and hands-on experience with Kubernetes.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

JobCubby • Barrington (RI), Northern (KY)

Hybrid
USD 110,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

myBridge Corporation • Austin (TX)

On-site
USD 120,000 - 160,000
Senior SRE: Scale, Reliability & Observability Leader
Senior SRE: Scale, Reliability & Observability Leader

Brez Technology Inc. • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Private Medical, Dental and Vision Benefits
Retirement Savings plan with matching contributions
Workspace benefits for your home office
+4
Senior Site Reliability Engineer - Automate & Scale
Senior Site Reliability Engineer - Automate & Scale

United States Digital Space LLC • New York (NY)

Hybrid
USD 179,000 - 226,000
Unlimited PTO
Employee stock options
Medical, dental, vision with HSA
+6
Senior SRE: Scale Reliability, Observability & CI/CD
Senior SRE: Scale Reliability, Observability & CI/CD

Breakout Tools • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior SRE: Scale, Reliability & Observability Lead
Senior SRE: Scale, Reliability & Observability Lead

Alien Blue • Chicago (IL)

On-site
USD 190,000 - 268,000
Comprehensive Healthcare Benefits
401k Matching
Flexible Vacation
+2
Senior SRE — Cloud Infra, Kubernetes & OpenShift
Senior SRE — Cloud Infra, Kubernetes & OpenShift

Jobtailor • North Carolina

Hybrid
USD 130,000 - 190,000
Senior Site Reliability Engineer – Scale & Observability
Senior Site Reliability Engineer – Scale & Observability

Inspire Brands, Inc. • Atlanta (GA)

On-site
USD 120,000 - 180,000
Senior Site Reliability Engineer: Observability & Cloud
Senior Site Reliability Engineer: Observability & Cloud

VBeyond Corporation • Jersey City (NJ)

On-site
USD 100,000 - 260,000