Lead Site Reliability Engineer

Selby Jennings

Wilmington (NC)

On-site

USD 140,000 - 200,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Selby Jennings is seeking a Lead Site Reliability Engineer to design, implement, and maintain reliable, scalable cloud infrastructure on AWS. You will lead a team of SREs, drive incident response, and improve platform reliability and efficiency across engineering and operations.

The role focuses on automation, monitoring, and cost optimization, with opportunities to shape disaster recovery and business continuity strategies in a data-driven, cloud-first environment.

Qualifications

  • Bachelor's degree in Computer Science or a related field, or equivalent professional experience.
  • Advanced degree in Computer Science or a related discipline is preferred.

Responsibilities

  • Lead and mentor a team of Site Reliability Engineers, including both full-time employees and contractors.
  • Design, implement, and maintain scalable, secure, and highly available cloud infrastructure in AWS.
  • Build and support monitoring, alerting, and observability solutions to ensure platform health and uptime.
  • Automate infrastructure provisioning and configuration management using Infrastructure-as-Code tools.
  • Develop and enhance CI/CD pipelines to improve deployment efficiency and software delivery.
  • Lead incident response efforts, conduct root cause analysis, and implement long-term solutions.
  • Partner with engineering teams to optimize performance, reliability, scalability, and cloud costs.
  • Promote operational best practices across infrastructure and application environments.
  • Develop and maintain disaster recovery and business continuity capabilities.

Skills

Scripting and programming experience
Git
SQL
Python
DevOps
Cloud
Site Reliability Engineer

Education

Bachelor's degree in Computer Science or related field
Advanced degree preferred

Tools

AWS
Kubernetes
Terraform
CI/CD
Git
Docker
Datadog

Job description

This organization is focused on modernizing the mortgage insurance industry through a technology-first approach. Rather than relying on legacy systems and manual processes, the company leverages software, automation, artificial intelligence, analytics, and scalable operating models to deliver better customer experiences and drive business efficiency. Its mission is to build a more modern, data-driven insurance platform designed for long-term growth and innovation.

About the Role

Lead Site Reliability Engineer

About the Company

This organization is focused on modernizing the mortgage insurance industry through a technology-first approach. Rather than relying on legacy systems and manual processes, the company leverages software, automation, artificial intelligence, analytics, and scalable operating models to deliver better customer experiences and drive business efficiency. Its mission is to build a more modern, data-driven insurance platform designed for long-term growth and innovation.

About the Role

The company is seeking a highly skilled and motivated Lead Site Reliability Engineer to play a key role in designing, implementing, and maintaining reliable, scalable, and high-performing cloud infrastructure within AWS. This individual will work closely with software engineering, operations, and cross-functional teams to improve platform reliability, enhance developer productivity, and drive operational excellence through automation, monitoring, and incident response practices.

Key Responsibilities
  • Lead and mentor a team of Site Reliability Engineers, including both full-time employees and contractors.
  • Prioritize, assign, and review technical work while providing guidance and feedback on code and infrastructure changes.
  • Design, implement, and maintain scalable, secure, and highly available cloud infrastructure in AWS.
  • Build and support monitoring, alerting, and observability solutions to ensure platform health and uptime.
  • Automate infrastructure provisioning and configuration management using Infrastructure-as-Code tools.
  • Develop and enhance CI/CD pipelines to improve deployment efficiency and software delivery.
  • Lead incident response efforts, conduct root cause analysis, and implement long-term solutions.
  • Partner with engineering teams to optimize performance, reliability, scalability, and cloud costs.
  • Promote operational best practices across infrastructure and application environments.
  • Develop and maintain disaster recovery and business continuity capabilities.
Qualifications
  • Bachelor's degree in Computer Science or a related field, or equivalent professional experience.
  • Advanced degree in Computer Science or a related discipline is preferred.
Required Technical Skills
  • AWS cloud infrastructure and services
  • Kubernetes and container orchestration platforms
  • Infrastructure as Code (Terraform or similar tools)
  • CI/CD and deployment automation
  • Git and modern version control practices
  • Containerization technologies (Docker)
  • Monitoring, logging, and observability platforms
  • Scripting and programming experience
  • Database administration and managementIncident and problem management
  • Security, compliance, and cloud governance
Preferred Experience
  • Experience with enterprise monitoring and observability platforms
  • Cloud networking and security best practices
  • Disaster recovery and resilience planning
  • Workflow automation and orchestration tools
  • Experience in insurance, financial services, or other regulated industries
  • AWS certifications or equivalent cloud certifications preferred
Desired Skills and Experience
  • Infrastructure as a Code
  • AWS
  • Git
  • SQL
  • Python
  • Kubernetes
  • Datadog
  • Scripting
  • DevOps
  • Site Reliability Engineer
  • Cloud
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobgether • United States

Hybrid
USD 120,000 - 160,000
Competitive compensation package
Flexible work arrangements
Professional development opportunities
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Jobtailor • Rockville (MD)

On-site
USD 120,000 - 180,000
Lead DevOps (Site Reliability Engineer)
Lead DevOps (Site Reliability Engineer)

Anza Mortgage Insurance Corporation • Wilmington (DE)

On-site
USD 140,000 - 200,000
Competitive pay
Comprehensive benefits
401(k) plan
+3
Lead DevOps (Site Reliability Engineer)
Lead DevOps (Site Reliability Engineer)

Anza Mortgage Insurance Company • Wilmington (NC)

On-site
USD 140,000 - 210,000
Competitive Compensation
Comprehensive Benefits
401(k) with company matching
+3
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Jobgether • United States

Remote
USD 150,000 - 200,000
Competitive salary
Comprehensive healthcare coverage
401(k) plan with company matching
+3
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Mike Albert Fleet Solutions • Cincinnati (OH)

Hybrid
USD 100,000 - 135,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Empower Retirement • Greenwood Village (CO)

Hybrid
USD 114,000 - 166,000
Medical insurance
401(k) with company match
Tuition reimbursement
+3
Lead DevOps (Site Reliability Engineer)
Lead DevOps (Site Reliability Engineer)

Anza Mortgage Insurance Company • McLean (VA)

On-site
USD 140,000 - 190,000
Competitive compensation
Comprehensive benefits
401(k) with company match
+2