Site Reliability Engineer

TalentDome Staffing

United States

On-site

USD 140,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

TalentDome Staffing is seeking a Senior Site Reliability Engineer to enhance the reliability, scalability, and performance of production systems. You will apply SRE principles to reduce toil, implement self-healing infrastructure, and drive capacity planning in high-scale environments.

Ideal candidates have 6+ years in SRE/DevOps with strong Python/Go coding, extensive AWS, and hands-on experience with Kubernetes.

Qualifications

  • 6+ years of hands-on SRE/DevOps or backend engineering with production reliability ownership.
  • Strong software engineering in Python or Go.
  • Deep knowledge of AWS core services (EC2, Lambda, RDS, VPC, IAM, S3) and cost/performance optimization.
  • Linux systems mastery including performance tuning, kernel/network troubleshooting, and security hardening.
  • Kubernetes/EKS production experience at scale.
  • Experience defining and operationalizing SLIs, SLOs, and error budgets.
  • Understanding of distributed systems, fault tolerance, and failover strategies.
  • Experience building full-stack observability pipelines and leading incident responses.

Responsibilities

  • Define and drive SLIs, SLOs, and error budgets across services.
  • Build automation to reduce toil and enable self-service for engineers.
  • Design and maintain IaC on AWS using Terraform/CloudFormation or Pulumi.
  • Architect observability with Prometheus, Grafana, Datadog, Splunk, and OpenTelemetry.
  • Act as incident commander during outages and lead blameless postmortems.
  • Forecast capacity, run load/chaos tests, and tune systems to prevent bottlenecks.
  • Partner with teams to deploy safe pipelines (canary, blue/green, feature flags) using ArgoCD/Jenkins/GitLab CI.
  • Embed security into pipelines, manage access, network segmentation, and vulnerability management.
  • Lead on-call rotations and runbooks; mentor engineers on reliability best practices.
  • Influence architecture decisions and champion SRE culture across the org.

Skills

Python
Go
Linux
CI/CD
Observability

Tools

AWS
Terraform/CloudFormation
Pulumi
Prometheus
Grafana
Datadog
Splunk
OpenTelemetry
Kubernetes
EKS
ArgoCD
Jenkins
GitLab CI

Job description

We are seeking a highly skilled Senior Site Reliability Engineer to drive the reliability, scalability, and performance of our client's production systems. The ideal candidate combines deep software engineering ability with system-level expertise, applying core SRE principles to reduce toil, minimize downtime, and build self-healing infrastructure across complex, high-scale environments.

Key Responsibilities
  • Reliability Engineering: Define and drive adoption of SLIs, SLOs, and error budgets across services, using them to guide engineering priorities and release decisions.
  • Automation & Toil Reduction: Build tools and automation (Python, Go, Bash) to eliminate manual operational work and enable self-service capabilities for engineering teams.
  • Infrastructure as Code (IaC): Design and maintain scalable, resilient infrastructure on AWS using Terraform, CloudFormation, or Pulumi, ensuring consistency and repeatability.
  • Observability: Architect monitoring, logging, tracing, and alerting systems (Prometheus, Grafana, Datadog, Splunk, OpenTelemetry) that give clear, actionable signals into system health.
  • Incident Management: Act as an incident commander during major outages, lead blameless postmortems, and drive systemic fixes to prevent recurrence.
  • Capacity Planning & Performance: Forecast growth, execute load/chaos testing, and tune systems proactively to stay ahead of scaling bottlenecks.
  • CI/CD & Deployment Safety: Partner with engineering teams to build safe, progressive delivery pipelines (canary, blue/green, feature flags) using tools like ArgoCD, Jenkins, or GitLab CI.
  • Security & Compliance: Embed security best practices into infrastructure and deployment pipelines, including access control, network segmentation, and vulnerability management.
  • On-Call Leadership: Participate in and help evolve on-call rotations, escalation policies, and runbooks to reduce alert fatigue and improve response times.
  • Mentorship & Culture: Champion SRE best practices across the organization, mentor engineers on reliability thinking, and influence upstream architecture decisions.
Qualifications & Experience
  • 6+ years of hands-on experience in Site Reliability Engineering, DevOps, or Backend/Systems Engineering with a track record of owning production reliability at scale.
  • Strong Software Engineering Background: Proficiency in Python, Go, or similar languages—focused on building maintainable services and tooling, not just basic scripting.
  • AWS Expertise: Deep technical knowledge of AWS core services (EC2, Lambda, RDS, VPC, IAM, S3, Terraform/CloudFormation) alongside cost and performance optimization.
  • Linux Systems Mastery: Demonstrated proficiency in performance tuning, kernel/network troubleshooting, and security hardening.
  • Container Orchestration: Hands-on experience operating Kubernetes/EKS in production environments at scale.
  • SRE Frameworks: Proven experience defining and operationalizing SLIs, SLOs, and error budgets.
  • Distributed Systems: Strong understanding of consistency, fault tolerance, failover strategies, and graceful degradation.
  • Observability & Incident Command: Experience building full-stack observability pipelines and leading major incident response efforts.
Nice to Have
  • Experience with multi-cloud environments (AWS, GCP, Azure).
  • Chaos engineering experience (Gremlin, Chaos Mesh, or custom fault-injection tooling).
  • Active AWS Certifications (DevOps Engineer Professional, Solutions Architect).
  • Familiarity with compliance frameworks (SOC2, ISO 27001, HIPAA).
  • Experience with service mesh technologies (Istio, Linkerd).
  • Background in internal platform/developer experience (DevEx) teams.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

SRE • Puerto Rico

Hybrid
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Jobtailor • New Hampshire

On-site
USD 110,000 - 160,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Virtual Tech Gurus • Puerto Rico

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineer – Lead
Site Reliability Engineer – Lead

Jobtailor • Arizona

On-site
USD 140,000 - 230,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

Hybrid
USD 140,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Brooksource • Hapeville (GA)

On-site
USD 110,000 - 170,000
Director, Site Reliability Engineering
Director, Site Reliability Engineering

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000
Site Reliability Engineer
Site Reliability Engineer

Knack Solutions • Richmond (VA)

On-site
USD 100,000 - 130,000