Site Reliability Engineer

TalentDome Staffing

United States

On-site

USD 140,000 - 210,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

TalentDome Staffing is seeking a Senior Site Reliability Engineer to enhance the reliability, scalability, and performance of production systems. You will apply SRE principles to reduce toil, implement self-healing infrastructure, and drive capacity planning in high-scale environments.

Ideal candidates have 6+ years in SRE/DevOps with strong Python/Go coding, extensive AWS, and hands-on experience with Kubernetes.

Qualifications

  • 6+ years of hands-on SRE/DevOps or backend engineering with production reliability ownership.
  • Strong software engineering in Python or Go.
  • Deep knowledge of AWS core services (EC2, Lambda, RDS, VPC, IAM, S3) and cost/performance optimization.
  • Linux systems mastery including performance tuning, kernel/network troubleshooting, and security hardening.
  • Kubernetes/EKS production experience at scale.
  • Experience defining and operationalizing SLIs, SLOs, and error budgets.
  • Understanding of distributed systems, fault tolerance, and failover strategies.
  • Experience building full-stack observability pipelines and leading incident responses.

Responsibilities

  • Define and drive SLIs, SLOs, and error budgets across services.
  • Build automation to reduce toil and enable self-service for engineers.
  • Design and maintain IaC on AWS using Terraform/CloudFormation or Pulumi.
  • Architect observability with Prometheus, Grafana, Datadog, Splunk, and OpenTelemetry.
  • Act as incident commander during outages and lead blameless postmortems.
  • Forecast capacity, run load/chaos tests, and tune systems to prevent bottlenecks.
  • Partner with teams to deploy safe pipelines (canary, blue/green, feature flags) using ArgoCD/Jenkins/GitLab CI.
  • Embed security into pipelines, manage access, network segmentation, and vulnerability management.
  • Lead on-call rotations and runbooks; mentor engineers on reliability best practices.
  • Influence architecture decisions and champion SRE culture across the org.

Skills

Python
Go
Linux
CI/CD
Observability

Tools

AWS
Terraform/CloudFormation
Pulumi
Prometheus
Grafana
Datadog
Splunk
OpenTelemetry
Kubernetes
EKS
ArgoCD
Jenkins
GitLab CI

Job description

We are seeking a highly skilled Senior Site Reliability Engineer to drive the reliability, scalability, and performance of our client's production systems. The ideal candidate combines deep software engineering ability with system-level expertise, applying core SRE principles to reduce toil, minimize downtime, and build self-healing infrastructure across complex, high-scale environments.

Key Responsibilities
  • Reliability Engineering: Define and drive adoption of SLIs, SLOs, and error budgets across services, using them to guide engineering priorities and release decisions.
  • Automation & Toil Reduction: Build tools and automation (Python, Go, Bash) to eliminate manual operational work and enable self-service capabilities for engineering teams.
  • Infrastructure as Code (IaC): Design and maintain scalable, resilient infrastructure on AWS using Terraform, CloudFormation, or Pulumi, ensuring consistency and repeatability.
  • Observability: Architect monitoring, logging, tracing, and alerting systems (Prometheus, Grafana, Datadog, Splunk, OpenTelemetry) that give clear, actionable signals into system health.
  • Incident Management: Act as an incident commander during major outages, lead blameless postmortems, and drive systemic fixes to prevent recurrence.
  • Capacity Planning & Performance: Forecast growth, execute load/chaos testing, and tune systems proactively to stay ahead of scaling bottlenecks.
  • CI/CD & Deployment Safety: Partner with engineering teams to build safe, progressive delivery pipelines (canary, blue/green, feature flags) using tools like ArgoCD, Jenkins, or GitLab CI.
  • Security & Compliance: Embed security best practices into infrastructure and deployment pipelines, including access control, network segmentation, and vulnerability management.
  • On-Call Leadership: Participate in and help evolve on-call rotations, escalation policies, and runbooks to reduce alert fatigue and improve response times.
  • Mentorship & Culture: Champion SRE best practices across the organization, mentor engineers on reliability thinking, and influence upstream architecture decisions.
Qualifications & Experience
  • 6+ years of hands-on experience in Site Reliability Engineering, DevOps, or Backend/Systems Engineering with a track record of owning production reliability at scale.
  • Strong Software Engineering Background: Proficiency in Python, Go, or similar languages—focused on building maintainable services and tooling, not just basic scripting.
  • AWS Expertise: Deep technical knowledge of AWS core services (EC2, Lambda, RDS, VPC, IAM, S3, Terraform/CloudFormation) alongside cost and performance optimization.
  • Linux Systems Mastery: Demonstrated proficiency in performance tuning, kernel/network troubleshooting, and security hardening.
  • Container Orchestration: Hands-on experience operating Kubernetes/EKS in production environments at scale.
  • SRE Frameworks: Proven experience defining and operationalizing SLIs, SLOs, and error budgets.
  • Distributed Systems: Strong understanding of consistency, fault tolerance, failover strategies, and graceful degradation.
  • Observability & Incident Command: Experience building full-stack observability pipelines and leading major incident response efforts.
Nice to Have
  • Experience with multi-cloud environments (AWS, GCP, Azure).
  • Chaos engineering experience (Gremlin, Chaos Mesh, or custom fault-injection tooling).
  • Active AWS Certifications (DevOps Engineer Professional, Solutions Architect).
  • Familiarity with compliance frameworks (SOC2, ISO 27001, HIPAA).
  • Experience with service mesh technologies (Istio, Linkerd).
  • Background in internal platform/developer experience (DevEx) teams.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer Lead
Site Reliability Engineer Lead

Good co India • United States

Remote
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Kovoro • Denver (CO), Northern (KY)

On-site
USD 150,000 - 190,000
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

On-site
USD 130,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Harvey Nash • United States

Remote
USD 120,000 - 150,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

On-site
USD 140,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Clearwater Analytics • Boise (ID)

On-site
USD 130,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Knack Solutions • Richmond (VA)

On-site
USD 100,000 - 130,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SDI International • Chicago (IL)

On-site
USD 130,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000