AWS SRE

Savvyan Technologies

Newark (NJ)

On-site

USD 140,000 - 210,000

Full time

16 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Savvyan Technologies seeks an experienced Senior AWS Site Reliability Engineer (SRE) to design, build, automate, and maintain highly available cloud infrastructure and applications in a dynamic environment.

Ideal candidates will be hands-on with AWS, Kubernetes, Terraform, CI/CD, and observability tools, and will collaborate across development, security, and operations teams to improve reliability and automate processes.

Qualifications

  • 6+ years of experience in SRE/DevOps, Cloud Engineering or related roles.
  • Strong hands-on experience with AWS.
  • Hands-on experience with Terraform/Terragrunt.
  • Hands-on experience with Kubernetes/EKS.
  • Strong Linux/Unix administration knowledge.
  • Experience with CI/CD pipelines and deployment automation.
  • Proficiency in Python, Bash, or Go.
  • Strong understanding of AWS networking and security.
  • Experience with monitoring, logging, and observability.
  • Incident management and RCA leadership experience.

Responsibilities

  • Design, implement, and maintain highly available and scalable infrastructure on AWS.
  • Manage AWS services including EC2, EKS, ECS, S3, RDS, Lambda, VPC, IAM, CloudWatch, Route 53, and Load Balancers.
  • Build and maintain infrastructure using IaC and Terraform/Terragrunt.
  • Design and administer Kubernetes/EKS environments.
  • Develop and maintain CI/CD pipelines using Jenkins, GitLab CI/CD, GitHub Actions, or similar tools.
  • Automate provisioning, deployments, monitoring, and support processes.
  • Develop scripts in Python, Bash, or Go.
  • Implement monitoring and observability with CloudWatch, Splunk, Datadog, Dynatrace, Prometheus, Grafana.
  • Participate in 24x7 on-call rotation and incident response.

Skills

SRE experience
DevOps
Cloud engineering
Linux administration
CI/CD pipelines
Python/Bash/Go
AWS networking
Incident management
Terraform/Terragrunt
Kubernetes/EKS
Monitoring/observability

Tools

Terraform/Terragrunt
Kubernetes/EKS
Jenkins
GitLab CI/CD
GitHub Actions
CloudWatch
Prometheus
Grafana
Splunk
Datadog
Dynatrace

Job description

Job Title: Senior AWS Site Reliability Engineer (SRE)
Location: New Jersey / Columbus, OH
Experience: 6–10+ years
Job Summary

We are seeking an experienced AWS Site Reliability Engineer (SRE) to design, build, automate, and maintain highly available, scalable, secure, and reliable cloud infrastructure and applications.

The ideal candidate will have strong hands-on experience with AWS, Kubernetes, Terraform, CI/CD, monitoring/observability, Linux, Python/Bash scripting, and production incident management. The engineer will work closely with development, infrastructure, security, and operations teams to improve system reliability and automate operational processes.

Key Responsibilities
  • Design, implement, and maintain highly available and scalable infrastructure on AWS.
  • Manage AWS services including EC2, EKS, ECS, S3, RDS, Lambda, VPC, IAM, CloudWatch, Route 53, and Load Balancers.
  • Build and maintain infrastructure using Terraform/Terragrunt and Infrastructure as Code (IaC) best practices.
  • Design, administer, and troubleshoot Kubernetes/EKS environments.
  • Develop and maintain CI/CD pipelines using Jenkins, GitLab CI/CD, GitHub Actions, or similar tools.
  • Automate infrastructure provisioning, application deployments, monitoring, and operational processes.
  • Develop scripts and automation using Python, Bash, or Go.
  • Implement monitoring, alerting, logging, and observability using tools such as CloudWatch, Splunk, Datadog, Dynatrace, Prometheus, and Grafana.
  • Participate in 24x7 on-call rotation and respond to production incidents.
  • Troubleshoot complex infrastructure, networking, application, and cloud-related issues.
  • Lead incident response, root-cause analysis (RCA), and post-incident remediation activities.
  • Define and improve SLOs, SLIs, SLAs, availability, latency, and reliability metrics.
  • Implement auto-scaling, fault tolerance, disaster recovery, backup, and high-availability solutions.
  • Identify opportunities to eliminate manual operational tasks through automation.
  • Work with development teams to ensure applications are designed and deployed with reliability, scalability, and performance in mind.
  • Implement AWS security best practices involving IAM, encryption, networking, secrets management, and least-privilege access.
  • Monitor AWS infrastructure for performance and cost optimization opportunities.
  • Create and maintain technical documentation, runbooks, operational procedures, and troubleshooting guides.
Required Skills
  • 6+ years of experience in SRE, DevOps, Cloud Engineering, Infrastructure Engineering, or related roles.
  • Strong hands-on experience with AWS.
  • Strong experience with Terraform/Terragrunt.
  • Hands-on experience with Kubernetes/EKS.
  • Strong understanding of Linux/Unix administration.
  • Experience with CI/CD pipelines and deployment automation.
  • Proficiency in Python, Bash, or Go.
  • Strong understanding of AWS networking, VPC, subnets, security groups, IAM, DNS, and load balancing.
  • Experience with monitoring, logging, alerting, and observability.
  • Strong production troubleshooting and incident-management experience.
  • Experience with distributed systems, microservices, and cloud-native architectures.
  • Strong understanding of high availability, scalability, resiliency, and disaster recovery.
Preferred Skills
  • AWS Certified DevOps Engineer or AWS Certified Solutions Architect.
  • Experience with Helm, ArgoCD, GitOps, Docker, Ansible, or CloudFormation.
  • Experience with Prometheus and Grafana.
  • Experience with Splunk, Datadog, Dynatrace, or New Relic.
  • Experience implementing blue/green, canary, and rolling deployments.
  • Experience with service mesh technologies such as Istio.
  • Experience with AWS cost optimization and FinOps.
  • Experience working in financial services, healthcare, telecommunications, or other highly regulated environments.
  • Strong knowledge of security and compliance requirements for cloud environments.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Brooksource • Hapeville (GA)

On-site
USD 110,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

JobCubby • Barrington (RI), Northern (KY)

On-site
USD 110,000 - 170,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Veriipro • Atlanta (GA)

On-site
USD 110,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Kovoro • Denver (CO), Northern (KY)

Hybrid
USD 150,000 - 190,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

Calance • United States

Hybrid
USD 150,000 - 200,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

myBridge Corporation • Austin (TX)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer – AWS/Datacenter
Senior Site Reliability Engineer – AWS/Datacenter

Jobtailor • Bellevue (WA)

On-site
USD 160,000 - 220,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000
Senior SRE Engineer
Senior SRE Engineer

Compunnel, Inc. • Alpharetta (GA)

On-site
USD 140,000 - 190,000