Senior Site Reliability Engineer (AWS / EKS)

Salve.Lab

Seattle (WA)

On-site

USD 180,000 - 260,000

Full time

9 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Salve.Lab is seeking a Senior Site Reliability Engineer to own production reliability on AWS and Amazon EKS. This is a hands-on role focused on operating highly available services, participating in on-call rotations, troubleshooting complex Kubernetes issues, and driving permanent improvements after incidents.

You will work with customers during escalations, define SLIs/SLOs, build runbooks, and implement IaC with Terraform/Terragrunt and GitOps tooling (Argo CD or FluxCD).

Qualifications

  • Hands-on, production ownership of SRE/Production Engineering roles.
  • Experience operating production environments on AWS with EKS.

Responsibilities

  • Own the reliability, availability and operational health of production services on AWS and EKS.
  • Lead or participate in on-call rotations and incidents with customers.
  • Define, monitor and improve SLIs, SLOs and error budgets.
  • Develop runbooks and automated remediation to reduce toil.
  • Build and maintain infrastructure as code using Terraform/Terragrunt and GitOps tools.

Skills

AWS
Kubernetes
Incident management
Production systems
English communication

Tools

Terraform
Terragrunt
Argo CD
FluxCD
Prometheus
Grafana
OpenTelemetry
Datadog
ELK
Kafka
MSK
Karpenter
KEDA

Job description

Role Overview

B2B Contract | Fully remote

We are looking for a Senior Site Reliability Engineer with deep, hands-on experience operating highly available production environments on AWS and Amazon EKS. This is a true SRE position, not a cloud architecture, infrastructure design, or monitoring-focused role. You will take direct ownership of production reliability, participate in the on-call rotation, respond to critical incidents, troubleshoot complex Kubernetes and distributed-system failures, and drive permanent improvements following incidents. The role also requires strong technical communication. You will interact directly with customers during technical discussions and production escalations, clearly explaining issues, making sound technical decisions, and driving problems through to resolution. We are looking for someone who has spent significant time running production systems, not simply designing them.

Key Responsibilities
  • Own the reliability, availability and operational health of production services running on AWS and Amazon EKS.
  • Operate and troubleshoot Kubernetes clusters in production, including cluster lifecycle, upgrades, node management, networking, scaling, capacity and workload reliability.
  • Participate actively in on-call and pager rotations and take ownership of production incidents.
  • Lead or play a key technical role during P1/P2 and Sev1/Sev2 incidents, including diagnosis, mitigation, recovery and communication.
  • Coordinate technical incident bridges and communicate directly with customers during production escalations when required.
  • Perform root cause analysis and lead blameless postmortems, ensuring incidents result in concrete engineering improvements.
  • Define, monitor and improve SLIs, SLOs and error budgets for production services.
  • Develop and maintain actionable alerts, operational runbooks and automated remediation.
  • Build and improve infrastructure using Terraform/Terragrunt and Infrastructure as Code practices.
  • Operate GitOps-based delivery environments using tools such as Argo CD or FluxCD.
  • Improve Kubernetes scaling and efficiency using technologies such as Karpenter, KEDA and native Kubernetes autoscaling capabilities.
  • Build and improve observability using technologies such as Prometheus, Grafana, OpenTelemetry, Datadog and/or ELK.
  • Support highly available distributed and event-driven systems, including environments using technologies such as Kafka/MSK.
  • Design, implement and validate disaster recovery and business continuity mechanisms against measurable RTO and RPO objectives.
  • Identify recurring operational problems and eliminate toil through automation and engineering.
  • Improve AWS performance, scalability, security and cost efficiency across production environments.
  • Work closely with software, platform and engineering teams to build reliability into systems throughout the development lifecycle.
  • Contribute to continuous improvement of incident management, operational readiness and SRE engineering practices.
Requirements
  • Significant professional experience as a hands-on Site Reliability Engineer, Production Engineer or senior Platform Engineer with direct production ownership.
  • Several years of recent, hands-on experience operating production environments on AWS.
  • Strong, demonstrable experience operating Amazon EKS in production.
  • Deep Kubernetes operational knowledge beyond application deployment, including cluster administration, upgrades, nodes, autoscaling, networking, troubleshooting and production failure scenarios.
  • Proven participation in a production on-call/pager rotation.
  • Demonstrable ownership of significant production incidents, including troubleshooting, mitigation, recovery, RCA and post-incident improvements.
  • Practical experience with SLIs, SLOs, error budgets, alerting and runbooks.
  • Strong Infrastructure as Code experience with Terraform and/or Terragrunt.
  • Production experience with Kubernetes delivery and GitOps practices; Argo CD or FluxCD strongly preferred.
  • Strong production observability experience with technologies such as Prometheus, Grafana, OpenTelemetry, Datadog or ELK.
  • Experience operating highly available, distributed production systems.
  • Strong understanding of AWS networking, IAM, security, availability and resilience.
  • Experience implementing and testing disaster recovery strategies with measurable RTO/RPO objectives.
  • Proven external customer-facing technical experience, including technical discussions, production escalations, architecture/reliability conversations or incident communication.
  • Ability to explain complex technical problems clearly and make sound decisions during high-pressure production incidents.
  • Strong troubleshooting mindset and ability to work independently during complex production failures.
  • Strong professional English communication skills (min. C1) for regular interaction with clients,
  • A coherent track record demonstrating sustained hands-on production engineering ownership.
What's on Offer
  • Full-time permanent B2B cooperation.
  • Fully remote working environment.
  • Senior hands-on engineering position with meaningful ownership of business-critical production systems.
  • Opportunity to work on complex AWS, Kubernetes and distributed-system environments at scale.
  • Direct influence over reliability engineering, operational practices and platform improvements.
  • Modern engineering environment with strong emphasis on automation, observability and continuous improvement.
  • Collaboration with experienced engineering, platform and product teams.
  • Opportunity to introduce and use modern approaches, including AI-assisted engineering and operational automation.
  • Long-term opportunity for engineers who want to remain deeply technical and close to production.
Diversity and Inclusion Commitment

We are dedicated to creating and sustaining an inclusive, respectful workplace for all -regardless of gender, ethnicity, or background. We actively encourage applicants from all identities and experience levels to apply and bring your authentic self to our fast-paced, supportive team.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 150,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Motion Recruitment Partners LLC • Chicago (IL), Northern (KY)

Hybrid
USD 140,000 - 170,000
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 150,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Kovoro • Denver (CO), Northern (KY)

Hybrid
USD 150,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior Staff Site Reliability Engineer
Senior Staff Site Reliability Engineer

Archer • San Jose (CA)

On-site
USD 160,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Evlo AI • Chicago (IL)

On-site
USD 120,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veriipro • Charlotte (NC)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veriipro • Charlotte (NC)

On-site
USD 140,000 - 190,000