Lead Site Reliability Engineer: AWS Cloud & Automation

Selby Jennings

Wilmington (NC)

On-site

USD 140,000 - 200,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Selby Jennings is seeking a Lead Site Reliability Engineer to design, implement, and maintain reliable, scalable cloud infrastructure on AWS. You will lead a team of SREs, drive incident response, and improve platform reliability and efficiency across engineering and operations.

The role focuses on automation, monitoring, and cost optimization, with opportunities to shape disaster recovery and business continuity strategies in a data-driven, cloud-first environment.

Qualifications

  • Bachelor's degree in Computer Science or a related field, or equivalent professional experience.
  • Advanced degree in Computer Science or a related discipline is preferred.

Responsibilities

  • Lead and mentor a team of Site Reliability Engineers, including both full-time employees and contractors.
  • Design, implement, and maintain scalable, secure, and highly available cloud infrastructure in AWS.
  • Build and support monitoring, alerting, and observability solutions to ensure platform health and uptime.
  • Automate infrastructure provisioning and configuration management using Infrastructure-as-Code tools.
  • Develop and enhance CI/CD pipelines to improve deployment efficiency and software delivery.
  • Lead incident response efforts, conduct root cause analysis, and implement long-term solutions.
  • Partner with engineering teams to optimize performance, reliability, scalability, and cloud costs.
  • Promote operational best practices across infrastructure and application environments.
  • Develop and maintain disaster recovery and business continuity capabilities.

Skills

Scripting and programming experience
Git
SQL
Python
DevOps
Cloud
Site Reliability Engineer

Education

Bachelor's degree in Computer Science or related field
Advanced degree preferred

Tools

AWS
Kubernetes
Terraform
CI/CD
Git
Docker
Datadog

Job description

Selby Jennings is seeking a Lead Site Reliability Engineer to design, implement, and maintain reliable, scalable cloud infrastructure on AWS. You will lead a team of SREs, drive incident response, and improve platform reliability and efficiency across engineering and operations.

The role focuses on automation, monitoring, and cost optimization, with opportunities to shape disaster recovery and business continuity strategies in a data-driven, cloud-first environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer - Cloud Platform & Automation
Senior Site Reliability Engineer - Cloud Platform & Automation

Ridgeline, Inc. • San Ramon (CA)

Hybrid
USD 153,000 - 210,000
Unlimited vacation
Educational reimbursement
Comprehensive insurance plans
Remote Senior SRE – AWS Cloud Reliability & Automation
Remote Senior SRE – AWS Cloud Reliability & Automation

IMPACT Technology Recruiting • New York (NY)

On-site
USD 120,000 - 160,000
Lead SRE — AWS Cloud Reliability & Automation
Lead SRE — AWS Cloud Reliability & Automation

Anza Mortgage Insurance Company • Wilmington (NC)

On-site
USD 140,000 - 210,000
Competitive Compensation
Comprehensive Benefits
401(k) with company matching
+3
Lead SRE: Cloud Reliability & AWS Automation
Lead SRE: Cloud Reliability & AWS Automation

Anza Mortgage Insurance Company • McLean (VA)

On-site
USD 140,000 - 190,000
Competitive compensation
Comprehensive benefits
401(k) with company match
+2
Senior SRE: Cloud Reliability & Automation Lead
Senior SRE: Cloud Reliability & Automation Lead

Ridgeline • New York (NY)

Hybrid
USD 153,000 - 210,000
Unlimited vacation
Educational reimbursements
$0 cost employee insurance plans
Senior Cloud SRE: AWS, Serverless & Incident Response
Senior Cloud SRE: AWS, Serverless & Incident Response

Apply • Northern (KY)

Hybrid
USD 120,000 - 150,000
Lead SRE — Cloud Reliability & Automation Leader
Lead SRE — Cloud Reliability & Automation Leader

anza-mortgage-insurance-company • McLean (VA)

On-site
USD 150,000 - 210,000
Competitive Compensation
Comprehensive Benefits
401(k) with company matching
+3
Senior Cloud SRE: AWS Serverless & Reliability
Senior Cloud SRE: AWS Serverless & Reliability

MeridianLink • United States

Remote
USD 140,000 - 180,000
Site Reliability Engineer: Build Resilient, Automated Cloud
Site Reliability Engineer: Build Resilient, Automated Cloud

SRE • Puerto Rico

Hybrid
USD 120,000 - 180,000
Lead SRE - Remote/Hybrid, Enterprise-Scale Reliability
Lead SRE - Remote/Hybrid, Enterprise-Scale Reliability

Empower Retirement • Greenwood Village (CO)

Hybrid
USD 114,000 - 166,000
Medical insurance
401(k) with company match
Tuition reimbursement
+3